Video and audio saliency prediction method based on continuous learning
By using continuous learning-based methods in video audio significance prediction tasks, replaying and mixing old task data to train models, the problem of catastrophic forgetting in multimodal continuous learning is solved, and the stability and performance of the model is improved.
Patent Information
- Application Number
- CN202510194621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art lacks effective methods for mitigating catastrophic forgetting in multimodal continuous learning and world modeling, especially in video and audio significance prediction tasks.
Using a continuous learning-based video audio significance prediction method, the video and audio samples of previous tasks are played back through the world model, mixed with the current task data and trained to achieve the mitigation of catastrophic forgetting of the video audio significance model.
Effectively alleviates the problem of catastrophic forgetting of video audio significance models in task flow, improves the stability and plasticity of the model, and ensures that new data can be maintained close to historical data training performance when new data appears.
Smart Images

Figure CN120126049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video and audio saliency prediction, and particularly to a video and audio saliency prediction method based on continuous learning. Background Art
[0002] Audio-visual saliency prediction methods use artificial intelligence techniques to identify and predict the parts of audio-visual content that are most likely to attract human attention. This method has a wide range of downstream applications in fields such as education, gaming, medicine, autonomous driving, and robotics, including video summarization, quality assessment, virtual human interaction, neuroscience research, and environmental perception. Given the wide variety of downstream tasks, it is very important to design a model that can effectively adapt to various tasks in a continuous learning manner and flexibly adapt to data changes. Unfortunately, few studies have addressed the scenario of continuous learning in saliency prediction.
[0003] Existing continuous learning methods can be roughly divided into three categories: regularization-based methods, generative replay-based methods, and architecture-based methods. Regularization-based methods mainly balance the learning and forgetting of old and new tasks by incorporating explicit regularization terms. EWC adds a quadratic penalty term to the loss function to penalize the change of each parameter according to its contribution or "importance" to performing the old task. LwF learns new training samples while using the predictions of the old task output heads to calculate the distillation loss. Replay-based methods mainly alleviate catastrophic forgetting in previous tasks by replaying the data distribution of old tasks. Early work selected samples according to fixed principles, randomly retaining a fixed number of old training samples from each training batch. iCaRL combines empirical replay with knowledge distillation to distill new and old samples. DER / DER++ and X-DER retain old training samples and logical outputs, performing logical matching throughout the optimization trajectory. Architecture-based methods mainly prevent catastrophic forgetting by creating new network modules or neurons for new tasks during the learning process while maintaining the performance on existing tasks. PNNs expand the network by adding new neural network layers or modules when learning each new task.
[0004] Humans can design a mental model of the world based on what their senses can perceive, and make decisions and take actions according to this model. Inspired by this, neural network-based world models aim to predict future sensory information by learning multi-modal feature representations. CPL establishes a hybrid world model based on visual control and prediction tasks, which models the physical dynamics in non-stationary environments and effectively avoids catastrophic forgetting. RWM considers the uncontrollable states in the autonomous driving scenario and separates the controllable and uncontrollable components in future scenarios from a visual perspective by improving the world model. However, considering the original intention of the brain-inspired world model, audio information, as an extremely important modality for humans to perceive the world, has not been considered to be added to the world model yet.
[0005] Modeling and predicting human attention using an audiovisual saliency model is more in line with the human visual attention mechanism, where vocal objects tend to attract more human attention. For example, the DAVS model is constructed by using a parametric neural network to map spatio-temporal coordinates to corresponding saliency values to build an effective mapping. STANet is a new multi-modal class activation mapping (CAM) model with a weak supervision training method to effectively alleviate the need for audiovisual saliency prediction on large-scale datasets.
[0006] In summary, although the above three research directions have received great attention from researchers in their respective fields, there are still many problems to be solved. Specifically, first, current research on continuous learning and world modeling still mainly focuses on single-modal tasks, and there is a completely new field to explore in the context of multi-modal continuous learning and world modeling. Second, in the context of multi-modal saliency prediction tasks, previous research has mainly focused on modality fusion. However, catastrophic forgetting that may occur in real-world scenarios also deserves high attention. Summary of the Invention
[0007] The objective of the technical solution of the present invention is to provide an audio-visual saliency prediction method based on a continuous learning method.
[0008] The technical solution of the present invention provides an audio-visual saliency prediction method based on continuous learning, including the following steps:
[0009] Set the basic parameters of the video-audio model, and the basic parameters of the video-audio model include the parameters of the world model and the saliency model, the embedding dimension of the task label, the number of encoder convolutional layers, and the number of LSTM modules;
[0010] Select a new task that has not been trained from the task sequence, and obtain the video, audio, and saliency map as the new task dataset for reading;
[0011] Determine whether there is an already trained old task currently. If so, the world model generates a complete temporally consistent video and audio based on the old task label. The saliency model outputs the predicted saliency map data based on the video and audio, and randomly inserts the generated video, audio, and saliency map data into the dataset of the new task as a new dataset. Use the new dataset as the model training data, calculate the loss value using the saliency loss function and the world loss function, and backpropagate to update the parameters of the model;
[0012] If not, use the new task dataset as the model training data, calculate the loss value using the saliency loss function and the world loss function, and backpropagate to update the parameters of the model;
[0013] Determine whether there is a next new task. If so, continue to select an untrained new task from the task sequence, and obtain the video, audio, and saliency map as the new task dataset for reading;
[0014] If not, use all the tasks in the task sequence to test the world model and the saliency model, and evaluate the ability of the world model and the saliency model to resist catastrophic forgetting.
[0015] Preferably, the world model samples through a Gaussian mixture distribution based on the task label. The generator generates the first frame of the video and audio, and continues to learn the dynamic distribution changes of video-audio multimodal information in the latent space and the input / output observation space. The task-separated multimodal spatio-temporal feature representation is further obtained from the first frame data, and a complete video and audio consistent with the real world and temporally continuous are generated.
[0016] Preferably, the formula for generating a complete video and audio consistent with the real world and temporally continuous is as follows:
[0017] X t+1:T ~World Model(X 1:t ,act t:T-1 )
[0018] Where X 1:t and X t+1:T represent the currently observed information and the future information predicted by the world model respectively, and act t:T-1 is the optional behavior influence sequence applied by the downstream model in the current environment, which is specifically the task label k of the task in the setting of continual learning.
[0019] Preferably, the world model constructs an encoding module, a representation module, and a dynamics module from the visual and auditory perspectives. The formula of the encoding module is as follows:
[0020]
[0021] The formula of the representation module is as follows:
[0022]
[0023] The formula of the dynamics module is as follows:
[0024]
[0025] Among them, V and A respectively represent video and audio information, k is the task label as a sign of the change in the task-related data distribution to distinguish the current task, t represents the time step, and z represents the latent variable compressed by the world model.
[0026] Preferably, the world model associates the video and audio in the new dataset, and uses the world loss function to calculate the loss value between the associated video and audio and the corresponding video and audio in the old task, so as to update the parameters of the world model.
[0027] Preferably, the formula of the world loss function is as follows:
[0028]
[0029] Among them, α M is a hyperparameter that adjusts the balance between the reconstruction loss and the KL loss within the same modality, and β is a hyperparameter that adjusts the weight update between the visual and auditory modalities.
[0030] Preferably, the saliency model uses the saliency loss function to calculate the loss value between the predicted saliency map data and the corresponding saliency map in the trained old task, so as to update the parameters of the saliency model.
[0031] Preferably, the formula of the saliency loss function is as follows:
[0032]
[0033] Loss = α * CC + (1 - α) * NSS
[0034] Among them, α is the weight coefficient, cov and ρ respectively represent covariance and standard deviation, ⊙ represents element-wise multiplication, y den , y fix and they represent the corresponding saliency map, annotation point map in the old task and the saliency map data predicted by the saliency model.
[0035] Preferably, the test is evaluated using the following formula:
[0036]
[0037] Among them, K is the total number of tasks, K i is the total number of tasks after learning the task, R j,j,R ki,j represent the performance metrics CC or NSS of the learning task and the new task, respectively.
[0038] The technical solution of the present invention proposes a method for predicting video and audio saliency based on continual learning. By using a world model to replay video and audio samples of previous tasks and mixing them with current task data for training, this method effectively alleviates catastrophic forgetting in the task stream of the video and audio saliency model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a schematic flowchart of a method for predicting video and audio saliency based on continual learning provided by an embodiment of the present invention;
[0040] Figure 2 is a structural diagram of the world model provided by an embodiment of the present invention;
[0041] Figure 3 is a structural diagram of the generative model provided by an embodiment of the present invention;
[0042] Figure 4 is a schematic diagram of the video and audio replay algorithm provided by an embodiment of the present invention;
[0043] Figure 5 is a schematic flowchart of the training process in a multi-task stream provided by an embodiment of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0045] As Figure 1 shown, an embodiment of the present invention provides a method for predicting video and audio saliency based on continual learning, including the following steps:
[0046] Set the basic parameters of the video and audio model. The basic parameters of the video and audio model include the parameters of the world model and the saliency model, the embedding dimension of the task label, the number of convolutional layers of the encoder, and the number of LSTM modules.
[0047] Select a new untrained task from the task sequence, and obtain video, audio, and saliency maps as the new task dataset for reading.
[0048] Determine whether there is an already trained old task currently. If so, the world model generates a complete and temporally consistent video and audio based on the old task label. The saliency model outputs the predicted saliency map data based on the video and audio, and randomly inserts the generated video, audio, and saliency map data into the dataset of the new task as a new dataset. Using the new dataset as the model training data, calculate the loss value using the saliency loss function and the world loss function and backpropagate to update the parameters of the model. Keeping the ratio of the total amount of generated old task data to the amount of new task data at 1:1 can enable the model to obtain a better trade-off between plasticity and stability.
[0049] The world model samples through a Gaussian mixture distribution based on the task label. The generator generates the first frame of the video and audio, and continues to learn the dynamic distribution changes of video-audio multimodal information in the latent space and the input / output observation space. Obtain the multimodal spatio-temporal feature representation separated by tasks from the first frame data, and generate a complete video and audio that is consistent with the real world and temporally continuous.
[0050] The generator is a conditional Gaussian mixture variational autoencoder. In the encoding stage, the embedding of the task label k jointly with the video frame / audio frame is used as the input to encode and obtain the latent variable z. In the decoding stage, the latent variable z combined with the embedding of the task label is input into the decoder to generate the original video frame / audio frame. That is, send the task label k to be played back into the generator to generate the first frame of the required video-audio dataset. The formula is as follows:
[0051] Sample the latent variable from the conditional Gaussian mixture distribution:
[0052]
[0053] Generate an image through the generator decoding:
[0054]
[0055] Send a frame of video-audio data into the world model to generate a spatio-temporally consistent video-audio dataset. The calculation formula is as follows:
[0056] X t+1:T ~World Model(X 1:t , act t:T-1 )
[0057] Among them, X 1:t and X t+1:T respectively represent the currently observed information and the future information predicted by the world model, and act t:T-1 is an optional sequence of behavior effects applied by the downstream model in the current environment, which is specifically the task label k of the task in the setting of continual learning.
[0058] The world model is a label-guided LSTM network (Long Short-Term Memory network). First, the input video frames / audio frames are encoded into corresponding features through the encoding part. Then, the LSTM predicts the feature representation of the current frame based on the features of the previous frame, the hidden state, and the embedding of the class label. The hidden state and cell state of the current time step are predicted from the feature representation f of the previous time step and the hidden state h of the previous time step. Finally, the video frames / audio frames of each time step are generated by decoding the hidden state.
[0059] The world model constructs an encoding module, a representation module, and a dynamics module from the visual and auditory perspectives to handle covariate shift, target shift, and dynamic shift that occur in a continuous learning environment respectively. Each module contains two parallel structures to process the audio and video information of the video respectively. Inputting a frame of data into the encoding module to obtain a latent variable containing information for predicting future moments, and then jointly inputting the data frame and the latent variable into the dynamics module to obtain the visible next frame of data. In the world model, the encoding module and the representation module use a stacked LSTM structure to learn the task-specific variations of the latent variable & from a Gaussian mixture distribution from the perspective of spatio-temporal dynamics, and jointly learn the prior and posterior distributions using KL divergence. The dynamics model then transforms the latent variable into the data space through a reconstruction loss. The specific representation is as follows:
[0060] Encoding module:
[0061] Representation module:
[0062] Dynamics module:
[0063] Among them, V and A represent video and audio information respectively, k is the task label as a sign of the change in the task-related data distribution to distinguish the current task, t represents the time step, and z represents the latent variable compressed by the world model.
[0064] The saliency model uses a saliency loss function to calculate the loss value between the predicted saliency map data and the corresponding saliency map in the trained old task to update the parameters of the saliency model. The world model associates the video and audio in the new dataset and uses a world loss function to calculate the loss value between the associated video and audio and the corresponding video and audio in the old task to update the parameters of the world model. Through multiple iterations, video and audio data containing complete temporal information can be obtained to realize the generation and playback of old task data. Finally, the generated video and audio dataset is input into the saliency model that has been trained on the old task to obtain the corresponding saliency map.
[0065] The saliency loss function includes CC (linear correlation coefficient) and NSS (normalized scanpath saliency), and the formula is as follows:
[0066]
[0067] Loss = α * CC+(1 - α) * NSS
[0068] where α is the weight coefficient, cov and ρ represent covariance and standard deviation respectively, ⊙ represents element-wise multiplication, y den , y fix and they represent the saliency map, annotation point map and saliency map data predicted by the saliency model corresponding to the old task.
[0069] The world loss function adopts the following formula:
[0070]
[0071] where α M is a hyperparameter that adjusts the balance between the reconstruction loss and the KL loss within the same modality, β is a hyperparameter that adjusts the weight update between the visual and auditory modalities, and α and β are set to 10 -4 and 0.1 respectively.
[0072] If not, use the new task dataset as the model training data, calculate the loss value using the loss function and backpropagate to update the model parameters.
[0073] Determine whether there is a next new task. If so, continue to select an untrained new task from the task sequence, and obtain the video, audio and saliency map as the new task dataset for reading.
[0074] If not, use all tasks in the task sequence to test the model and evaluate the model's ability to resist catastrophic forgetting.
[0075] During testing, use FA-CC (Final Average CC) and FA-NSS (Final Average NSS) to evaluate the average CC and NSS metrics of each task after completing all training tasks, and BF-CC (Backward Forgetting CC) and BF-NSS (Backward Forgetting NSS) to evaluate the impact of the learning task on the performance of the previous task to supplement backward forgetting. The specific formulas are as follows:
[0076]
[0077] where K is the total number of tasks, K i is the total number of tasks after the learning task, and R j,j , R ki,j represent the performance metrics CC or NSS of the learning task and the new task respectively.
[0078] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements any one of the above-mentioned video and audio saliency prediction methods based on continual learning. Among them, the memory includes various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0079] Compared with the prior art, the embodiment of the present invention reconsiders the catastrophic forgetting problem in video and audio saliency prediction from the perspective of continual learning for the first time, and alleviates the forgetting phenomenon of the model by means of generative replay.
[0080] A video and audio saliency prediction method based on continual learning provided by an embodiment of the present invention has the following beneficial effects compared with the prior art:
[0081] 1. The embodiment of the present invention explores the continual learning problem in the video and audio saliency prediction task. Compared with the classical classification tasks in continual learning, video and audio saliency prediction is more challenging in terms of the number of modalities and the complexity of the data distribution space. The embodiment of the present invention adopts a continual learning method based on generative replay, which can accurately model the video and audio modal information and completely align the temporal information between modalities, and more effectively capture the covariate shift of the data distribution to alleviate catastrophic forgetting.
[0082] When the embodiment of the present invention obtains a new batch of data sets, only using the current new data to fine-tune the model can obtain a performance close to that of fine-tuning with all historical training data. While effectively reducing the computational and storage loads, it alleviates the catastrophic forgetting phenomenon caused by only using new data for fine-tuning.
[0083] The world model in the embodiment of the present invention includes an image and audio generator and a video and audio LSTM module. When a new task appears, it effectively models the video and audio of the new task data and the corresponding temporal information through a mixture of Gaussian models. The world model not only retains the knowledge of the known tasks but also efficiently learns the data distribution of the new tasks. In the replay stage, the video and audio generated by the world model have good temporal consistency, so the generated old task data ensures the balance between the stability and plasticity of the video and audio saliency model.
[0084] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in this technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A video and audio saliency prediction method based on continuous learning, characterized in that: The following steps are involved: Set the basic parameters of the video and audio model, including the parameters of the world model and saliency model, the embedding dimension of the task label, the number of encoder convolution layers, and the number of LSTM modules; Select a new task that has not been trained from the task sequence, and obtain the video, audio, and saliency map as the new task dataset for reading; Determine whether there is an old task that has been trained. If so, the world model generates complete time-consistent video and audio based on the old task label. The saliency model outputs the predicted saliency map data based on the video and audio. The generated video, audio, and saliency map data are randomly inserted into the dataset of the new task as a new dataset. The new dataset is used as model training data. The saliency loss function and the world loss function are used to calculate the loss value and back-propagate to update the model parameters. If it does not exist, the new task dataset is used as the model training data, the loss value is calculated using the significant loss function and the world loss function, and the parameters of the model are updated by backpropagation; Determine whether there is a next new task. If so, continue to select a new task that has not been trained from the task sequence, and obtain the video, audio, and saliency map as the new task data set for reading; If it does not exist, all tasks in the task sequence are used to test the world model and the saliency model to evaluate their ability to resist catastrophic forgetting.
2. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The world model is sampled through a Gaussian mixture distribution based on the task label, and the generator generates the first frame of video and audio, and continues to learn the dynamic distribution changes of video and audio multimodal information in the latent space and input / output observation space. The task-separated multimodal spatiotemporal feature representation is further obtained from the first frame data to generate complete video and audio that is consistent with the real world and continuous in time.
3. A method for predicting video and audio saliency based on continuous learning as claimed in claim 2, characterized in that: The formula for generating complete and time-continuous video and audio that is consistent with the real world is as follows: X t+1:T ~World Model(X 1:t ,act t:T-1 ) Among them, X 1:t and X t+1:T Represents the currently observed information and the future information predicted by the world model, act t:T-1 It is the optional sequence of behavioral influences applied by the downstream model in the current environment, specifically the task label k of the task in the setting of continuous learning.
4. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The world model constructs the encoding module, representation module and dynamics module from the perspective of vision and hearing. The encoding module formula is as follows: The module formula is as follows: The dynamics module formula is as follows: Among them, V and A represent video and audio information respectively, k is the task label as a sign of the change in the distribution of task-related data to distinguish the current task, t represents the time step, and z represents the hidden variable compressed by the world model.
5. The video and audio saliency prediction method based on continuous learning as claimed in claim 1, characterized in that: The world model associates the video and audio in the new data set, and uses the world loss function to calculate the loss value of the associated video and audio with the corresponding video and audio in the old task, so as to update the parameters of the world model.
6. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The world loss function formula is as follows: Among them, α M is a hyperparameter that adjusts the balance between reconstruction loss and KL loss within the same modality, and β is a hyperparameter that adjusts the weight update between visual and auditory modalities.
7. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The saliency model uses a saliency loss function to calculate the loss value of the predicted saliency map data and the corresponding saliency map in the trained old task, so as to update the parameters of the saliency model.
8. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The significant loss function formula is as follows: Loss = α*CC+(1-α)*NSS Among them, α is the weight coefficient, cov and ρ represent covariance and standard deviation respectively, ⊙ represents element multiplication, and y den ,y fix and They represent the corresponding saliency maps in the old tasks, the annotation point maps, and the saliency map data predicted by the saliency model.
9. The method for predicting video and audio saliency based on continuous learning according to claim 1, characterized in that: The test is evaluated using the following formula: Where K is the total number of tasks, K i is the total number of tasks after learning the task, R j,j ,R ki,j represent the performance indicators CC or NSS of the learning task and the new task respectively.
Citation Information
Cited By
Cross-modal saliency prediction method and device for audio visual touch
CN121278634A