Text-guided visual saliency prediction method
By constructing the TDiffSal model, utilizing a multi-head fusion module and a saliency prediction diffusion module, and combining a text encoder and a U-Net structure, the problem of not utilizing complete text information in visual saliency prediction is solved, generating more accurate visual saliency images and improving the robustness and feature fusion effect of the model.
Patent Information
- Application Number
- CN202511049860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, visual saliency prediction tasks have failed to effectively utilize complete text information for guidance, and there is a lack of research on the saliency of complete text and images.
The TDiffSal model is constructed, including a saliency prediction diffusion module, a multi-head fusion module, and a combined loss function. The text description is converted into an embedding vector through a text encoder and fused with image features. A saliency image is generated using a multi-head cross-attention mechanism and a U-Net structure. The model is optimized by combining latent space and pixel space losses.
It improves the robustness and generalization ability of the model, effectively resolves the potential semantic associations between text and visual modalities, significantly improves the multimodal feature fusion effect, and generates more accurate visual saliency images.
Smart Images

Figure CN120953580A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a visual saliency prediction technique, and more particularly to a text-guided visual saliency prediction method. Background Technology
[0002] Existing research primarily focuses on exploring the relationship between images and text, particularly in object detection. For example, Liu et al. proposed an attribute-guided network that uses attributes derived from denotative representations as supervisory signals to locate corresponding objects. Furthermore, other text-based multimodal models have been applied to image captioning and visual question answering (VQA). For instance, Xu et al. studied a principled approach entirely based on the key idea of editing and realigning existing alternative text related to the image. To generate training data, they also performed manual annotation, where the annotator starts with existing alternative text and realigns it with the image content in multiple rounds, thus constructing titles with rich visual concepts. Moshiur et al. proposed a question-independent attention mechanism and successfully applied it to visual question answering tasks, achieving excellent results. Existing research has demonstrated the intrinsic correlation between text and images; however, in the field of visual saliency prediction, this research remains in the theoretical stage, and relevant datasets are relatively scarce.
[0003] While the attention relationship between text and image has been applied in the aforementioned tasks, no model currently implements a text-guided visual saliency prediction task based on complete text information. Existing tasks only focus on the relationships between objects within the text and image, without addressing the saliency of the complete text and image. Therefore, it is necessary to explore a model specifically designed for visual saliency prediction guided by complete text information. Summary of the Invention
[0004] To address the issue that existing tasks only focus on the relationships between objects within text and images, without studying the saliency of the complete text and image, a text-guided visual saliency prediction method is proposed.
[0005] The technical solution of this invention is as follows:
[0006] A text-guided method for visual saliency prediction includes the following steps:
[0007] Step 1: Data Preparation
[0008] Obtain the text dataset and the image saliency detection dataset; perform pairwise processing on the text dataset;
[0009] Step 2: Construction of the visual saliency prediction model:
[0010] Construct the TDiffSal model, including a significance prediction diffusion module, a multi-head fusion module, and a combined loss function;
[0011] Step 3: Initial training of the image saliency detection dataset:
[0012] The initial training of the model is initiated using an image saliency detection dataset: the original image and the true saliency image are mapped to the latent space, and the features of the original image are denoised to generate noisy image features; the corresponding text description is converted into a text embedding vector by a text encoder; the noisy image features and the text embedding vector are input into a multi-head fusion module to obtain a text-guided image feature vector, and then the saliency prediction diffusion module outputs the saliency prediction results in the latent space;
[0013] The dual-loss optimization model is calculated: the sum of the latent space loss and the pixel space loss is used as the total loss, and the model parameters are updated through backpropagation; the loss is evaluated after each training round, and if the current loss is the minimum value, it is saved as the optimal weight; otherwise, the iteration continues until the preset number of training rounds is reached to complete the initial training.
[0014] Step 4: Fine-tuning the text dataset:
[0015] Load the optimal weights saved from the initial training and fine-tune them using a text dataset: map the original images and true saliency images in the text dataset to the latent space and add noise, input the corresponding text description encoded into the embedding vector, and output the latent space prediction results through the multi-head fusion module and the saliency prediction diffusion module; similarly, strengthen the ability of the text to guide saliency prediction by calculating the dual loss optimization model.
[0016] Step 5: Test Prediction
[0017] Using the final saved optimal weights, the original images from the test dataset are input into the trained TDiffSal model, which outputs the final saliency image prediction results.
[0018] Furthermore, the specific steps include:
[0019] S1. Obtain a text dataset containing saliency images that match the corresponding video screenshots and corresponding text description information, and an image saliency detection dataset containing the original image and the corresponding saliency image. Perform pairwise processing on the data in the text dataset and match the paired text descriptions in the paired text dataset with the original image.
[0020] S2. Construct the visual saliency prediction model, also known as the TDiffSal model. The visual saliency prediction model includes three modules: a saliency prediction diffusion module, a multi-head fusion module, and a combined loss function. The multi-head fusion module is used for image-text feature fusion, the saliency prediction diffusion module is used to predict the saliency image after fusion, and the combined loss function is used to train the model. The visual saliency prediction model is trained using the image saliency detection dataset. It is determined whether the number of training rounds has been exceeded. If it has not been exceeded, proceed to step S3. If it has been exceeded, save the current model weights, end the training, and proceed to step S7.
[0021] S3. Input the original image and the corresponding saliency image from the image saliency detection dataset, and map them to the latent space; then add noise to the original image in the latent space to obtain the features of the noisy image; set the text input corresponding to the image, and process it into a text embedding vector through a text encoder;
[0022] S4. The noisy image and the text embedding vector are fused through a multi-head fusion module to obtain a text-guided image feature vector; this vector is then input into the saliency diffusion model for training to obtain the prediction results of the latent space.
[0023] S5. Calculate the mean squared error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean squared error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the model parameters, so that the final output saliency image is closer and closer to the true saliency map.
[0024] S6. Update the optimal training result weights. If the current calculated loss is the minimum, save the current training weights as the optimal weights. If there were previously optimal weights, they will be overwritten. After this step, return to step S2.
[0025] S7. The model reads the optimal weights saved from previous training, uses the text dataset to further train the visual saliency prediction model, and determines whether the number of training rounds has been exceeded. If it has not been exceeded, proceed to step S8. If it has been exceeded, save the current model weights, end the training, and proceed to step S12.
[0026] S8. Input the original image and the corresponding saliency image from the input text dataset, and map them to the latent space; then add noise to the original image in the latent space to obtain the features of the noisy image; input the corresponding descriptive text, and process it into a text embedding vector through a text encoder;
[0027] S9. The noisy image and the text embedding vector are fused through a multi-head fusion module to obtain a text-guided image feature vector; this vector is then input into the saliency diffusion model for training to obtain the prediction results of the latent space.
[0028] S10. Calculate the mean square error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean square error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the model parameters so that the final output saliency image is closer and closer to the true saliency map.
[0029] S11. Update the optimal training result weights. If the current calculated loss is the minimum, save the current training weights as the optimal weights. If there were previously optimal weights, they will be overwritten. After this step, return to step S7.
[0030] S12. The model uses the optimal weights to predict the test dataset, which yields the final saliency image of the predicted test dataset.
[0031] Furthermore, the specific implementation process of the multi-head fusion module is as follows:
[0032] The key elements for calculating attention are key K, query Q, and value V. The multi-head fusion module (MHF) uses visual features as Q, and text embedding vectors as K and V. The three vectors are obtained by multiplying the corresponding input by the corresponding projection matrix, as shown in the following formula:
[0033] Q = x t W q K = cW k V = cW k
[0034] Where x t represents the visual feature vector input at step t, c represents the current input text embedding vector, and W... q W k These represent the corresponding learnable projection matrices. After obtaining the K, Q, and V vectors needed to calculate attention, the dot product of Q and K yields the correlation score between the two features. Finally, the SoftMax function converts the score into attention weights for each key. After the above operations, the attention weights need to be multiplied by V to obtain the final output. The specific calculation formula is as follows:
[0035]
[0036] Where d represents the scaling factor, Attn represents the attention output, and T represents the transpose; this step is the final cross-attention output of single-head computation; multi-head computation requires dividing the input KQV into multiple subspaces, calculating the corresponding attention value for each subspace using the above method, and finally concatenating the outputs; the specific formula is shown below:
[0037]
[0038] head i This represents the attention output of the i-th head. This represents the image feature output containing textual meaning after fusion by the MHF module; after obtaining the MHF fused features, the model input will be changed from the originally set x t Updated to
[0039] Furthermore, the specific implementation process of the saliency prediction diffusion module is as follows:
[0040] The diffusion process of the saliency prediction diffusion module predicts saliency and generates a visual saliency image. Text-guided image features serve as input to the denoising neural network in the saliency prediction diffusion module, and the output is the final predicted saliency image in the latent space. The denoising neural network used in the denoising process is a U-Net structure network. In this saliency prediction diffusion module, the inputs are text features, text-guided image features, and a time step. This denoising neural network directly predicts the final generated image, and its calculation process can be formalized as follows:
[0041]
[0042] in This represents the predicted visual saliency map, F. Diff The saliency prediction diffusion module transforms the saliency prediction task into a saliency map generation task. First, the feature processing is performed by the MHF module, and then the visual saliency image is further generated by the saliency prediction diffusion module to obtain the text-guided visual saliency image in the latent space. Finally, the saliency map is restored from the latent space to the pixel space by the decoder.
[0043] Furthermore, the specific implementation process of the combined loss function is as follows:
[0044] The loss function is divided into loss calculation for the latent space and loss calculation for the true pixel space; in calculating the latent space loss, the true saliency image X is first calculated. sal By mapping the encoder to the latent space, we obtain its latent space representation z.sal Subsequently, the predicted latent space saliency map is calculated. With z sal The mean squared error between them yields the latent space training loss L. latent The mathematical formula for its calculation is as follows:
[0045]
[0046] Where N represents the number of samples calculated in this batch; the predicted latent space saliency map The pixel space representation X is obtained by mapping the pixel space to the decoder. pre Then, the pixel spatial loss L is calculated using the same loss calculation method as for the latent space. pixel The mathematical formula for its calculation is as follows:
[0047]
[0048] Finally, the final model loss is obtained by adding the loss function values of the two parts; its calculation formula is as follows:
[0049] L all =L latent +λL pixel
[0050] Where λ represents the weight value assigned to the pixel spatial loss result.
[0051] Optionally, the EGO4D or SALICON datasets can be selected for training.
[0052] The beneficial effects of this invention are as follows:
[0053] (1) To address the problem of limited datasets, a pre-training strategy is proposed to enhance the model’s saliency generation capability by utilizing existing datasets. This effectively solves the overfitting problem in training with small datasets and improves the robustness and generalization ability of the model.
[0054] (2) The powerful representation ability of the diffusion model in the text-visual modality and the iterative optimization form can be well adapted to the text-guided visual saliency prediction task, which plays a guiding role in the construction of the text-guided visual saliency prediction model.
[0055] (3) By introducing a multi-head fusion module, the potential semantic association between text and visual modalities is effectively analyzed, which significantly improves the multi-modal feature fusion effect. Attached Figure Description
[0056] Figure 1 This is a flowchart of the text-guided visual saliency prediction model of the present invention;
[0057] Figure 2 This is a framework diagram of the text-guided visual saliency prediction model of the present invention. Detailed Implementation
[0058] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0059] This embodiment provides a text-guided visual saliency prediction method. First, the dataset used in this embodiment is introduced:
[0060] The SALICON dataset is a large dataset specifically designed for human eye-tracking gaze prediction. It carefully selects 20,000 images from the Microsoft COCO dataset, each accompanied by a binary saliency map annotating the most salient regions. Notably, SALICON is the largest dataset in the field of image saliency detection. Our experiments are based on the latest version used in the LSUN 2017 challenge, which contains 10,000 training images, 5,000 validation images, and 5,000 test images.
[0061] The EGO4D dataset is currently the world's largest machine learning dataset and benchmark suite based on first-person (egocentric) video. It contains over 3,600 hours of densely narrated video content, covering various scenarios including family activities, outdoor scenes, work environments, and leisure activities. This extensive dataset covers annotations for five benchmark tasks and includes records from 926 unique camera wearers located in 74 different global locations across 9 countries. Some video clips in the EGO4D dataset also include audio, 3D environment meshes, eye-tracking information, stereoscopic views, and / or footage of the same event simultaneously captured by multiple egocentric cameras. To process the dataset into usable data for model training, all videos in the EGO4D dataset containing gaze information were captured using a frame-skipping method to obtain corresponding video screenshots. A saliency image matching the corresponding video screenshot was generated based on the given gaze information. For all matched video screenshots, corresponding textual descriptions were saved. After this processing, the dataset was divided into 1,452 training dataset images and 306 test training dataset images.
[0062] Experimental results show that the TDiffSal model outperforms other state-of-the-art methods.
[0063] like Figure 1 As shown, the method specifically includes the following steps:
[0064] S1. Obtain the EGO4D and SALICON datasets, perform pairwise processing on the text dataset data, and match the paired text descriptions in the EGO4D dataset with the original images.
[0065] S2. Construct a visual saliency prediction model (i.e., the TDiffSal model). The visual saliency prediction model includes three modules: a saliency prediction diffusion module, a multi-head fusion module, and a combined loss function. The multi-head fusion module is used for image-text feature fusion, the saliency prediction diffusion module is used to predict the saliency image after fusion, and the combined loss function is used to train the model. The SALICON dataset is used to train the saliency prediction of the TDiffSal model. It is determined whether the training rounds have been exceeded. If not, proceed to step S3. If they have been exceeded, save the current model weights, end the training, and proceed to step S7.
[0066] S3. Input the original image and its corresponding saliency image, and map both to the latent space. Then, add noise to the original image in the latent space to obtain the features of the noisy image. Set the text input corresponding to the image to "Saliency", and process it into a text embedding vector through a text encoder.
[0067] S4. The noisy image and text embedding vector are fused using a multi-head fusion module to obtain a text-guided image feature vector. This vector is then input into a saliency diffusion model for training to obtain prediction results for the latent space.
[0068] To more fully explore the potential semantic relationships between visual and textual modalities and more effectively utilize textual features to guide the visual saliency prediction process, the model introduces a multi-head fusion module (MHF module). The core of the MHF module lies in achieving dynamic alignment and complementary fusion of textual and visual features through a multi-head cross-attention mechanism. The advantage of multi-head cross-attention is that multiple heads can focus on different regions and are applicable across different modalities, ensuring that the final calculation result is based on comprehensive features and retains as much textual and visual attention-related information as possible. Generally, the key elements for calculating attention are the key, query, and value (KQV), all derived from the same feature vector. However, to align and fuse textual and visual features, the MHF module uses visual features as Q and textual embedding vectors as K and V. The three vectors are obtained by multiplying the corresponding input by the corresponding projection matrix, as shown in the following formula:
[0069] Q = x t W q K = cW k V = cW k
[0070] Where x t represents the visual feature vector input at step t, c represents the current input text embedding vector, and W... q W k These represent the corresponding learnable projection matrices. After obtaining the K, Q, and V vectors needed to calculate the attention, performing a dot product operation on Q and K yields the correlation score between the two features. Finally, the SoftMax function is used to convert the score into an attention weight for each key. After the above operations, the attention weights need to be multiplied by V to obtain the final output. The specific calculation formula is as follows:
[0071]
[0072] Where d represents the scaling factor, used in the formula to prevent the inner product from becoming too large, Attn represents the attention output, and T represents the transpose. This step represents the final cross-attention output of single-head computation. Multi-head computation requires dividing the input KQV into multiple subspaces, calculating the corresponding attention value for each subspace using the method described above, and finally concatenating the outputs. The specific formula is shown below:
[0073]
[0074] head i This represents the attention output of the i-th head. This represents the image feature output containing textual meaning after fusion by the MHF module. After obtaining the MHF fused features, the model input will be changed from the originally set x... t Updated to During iterative denoising training, these text-guided image features replace the original visual features, ensuring that the saliency map generation is driven by textual semantics and remains within the normal range of visual saliency prediction. This setting allows the final output saliency map to more accurately capture the guiding scale of textual features for visual saliency, avoiding both over-guidance and under-guidance.
[0075] Through the above operations, the TDiffSal model can use the multi-head fusion module and the saliency prediction diffusion module to jointly guide the image saliency prediction of text adaptation.
[0076] The MHF module addresses the problem of text-guided visual saliency. The subsequent diffusion process of the saliency prediction diffusion module then predicts saliency, generating a visual saliency image. Text-guided image features serve as input to the denoising neural network in the saliency prediction diffusion module, with the output being the final predicted saliency image in the latent space. The denoising neural network used in the denoising process is a U-Net structure. U-Net, through its encoder-decoder structure and skip connections, achieves multi-scale extraction and fusion of image features, thus efficiently completing the noise prediction task. In this saliency prediction diffusion module, the inputs are text features, text-guided image features, and a time step. This denoising neural network can generate the final result in two ways: first, the neural network predicts the noise value to be removed, and the final result is obtained by subtracting the predicted noise from the noisy image; second, the neural network directly predicts the final generated image. In this saliency prediction diffusion module, the second method is used, as it is more suitable for the current saliency prediction task than for noise prediction. Its calculation process can be formalized as follows:
[0077]
[0078] in This represents the predicted visual saliency map, F. Diff This is the saliency prediction diffusion module. The TDiffSal model transforms the saliency prediction task into a saliency map generation task. First, feature processing is performed by the MHF module, and then the saliency prediction diffusion module further generates the visual saliency image, obtaining a text-guided visual saliency image within the latent space. The final saliency map is then restored from the latent space to the pixel space by a decoder. The decoder used here is the same one used by the original stable diffusion model (no changes are made to the encoder-decoder part; it's the encoder-decoder trained on the original stable diffusion model).
[0079] S5. Calculate the mean squared error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean squared error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the TDiffSal model parameters, so that the final output saliency image is closer and closer to the true saliency map.
[0080] The joint loss function mainly consists of two parts: one is the loss calculation for the latent space, and the other is the loss calculation for the true pixel space. In calculating the latent space loss, the first step is to calculate the loss for the true saliency image X. salBy mapping the encoder to the latent space, we obtain its latent space representation z. sal Subsequently, the predicted latent space saliency map is calculated. With z sal The mean squared error between them yields the latent space training loss L. latent The mathematical formula for its calculation is as follows:
[0081]
[0082] Where N represents the number of samples calculated in this batch. To calculate the loss in the true pixel space to further constrain the model training to produce results that conform to the true saliency image, it is also necessary to calculate the predicted latent space saliency map. The pixel space representation X is obtained by mapping the pixel space to the decoder. pre Then, the pixel spatial loss L is calculated using the same loss calculation method as for the latent space. pixel The mathematical formula for its calculation is as follows:
[0083]
[0084] Finally, the final model loss is obtained by adding the loss function values from the two parts. The calculation formula is as follows:
[0085] L all =L latent +λL pixel
[0086] Here, λ represents the weight value assigned to the pixel space loss result. This method allows the model's prediction results to approximate the true saliency image in both the pixel space and the latent space. The weight parameter is introduced to facilitate dynamic adjustment of the contributions of the two loss functions. This approach eliminates the need to train the encoder-decoder part, allowing training only the diffusion and fusion parts, thus reducing computational complexity. The effectiveness of the joint loss function has also been verified through subsequent experiments.
[0087] S6. Update the optimal training result weights. If the current calculated loss is minimized, save the current training weights as the optimal weights. If an optimal weight already exists, it will be overwritten. After this step, return to step S2.
[0088] S7. The model reads the optimal weights saved from previous training, uses the EGO4D dataset, and divides it into training and test datasets. The TDiffSal model is further trained using the training dataset. It is determined whether the number of training rounds has been exceeded. If not, proceed to step S8. If it has been exceeded, the current model weights are saved, training ends, and the process proceeds to step S12.
[0089] S8. Input the original image and its corresponding saliency image, and map both to the latent space. Then, add noise to the original image in the latent space to obtain the features of the noisy image. Input the corresponding descriptive text, and process it into a text embedding vector through a text encoder.
[0090] S9. The noisy image and text embedding vector are fused using a multi-head fusion module to obtain a text-guided image feature vector. This vector is then input into a saliency diffusion model for training to obtain prediction results for the latent space.
[0091] S10. Calculate the mean squared error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean squared error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the model parameters, so that the final output saliency image is closer and closer to the true saliency map.
[0092] S11. Update the optimal training result weights. If the current calculated loss is minimized, save the current training weights as the optimal weights. If an optimal weight already exists, it will be overwritten. After this step, return to step S7.
[0093] S12. The model uses the optimal weights to predict the training dataset, and the saliency image of the final predicted test dataset can be obtained.
[0094] To verify whether the multi-head fusion module is more effective than the fusion method of the Stable Diffusion model itself, the model was quantitatively compared under two different conditions using the EGO4D dataset: feature fusion using the multi-head fusion module and feature fusion without MHF. Specifically, all 1758 images in the EGO4D dataset were randomly shuffled into three groups of 306 images each. The experiment was conducted three times, with one group of data used as the test set in each experiment, and the other two groups (1452 images in total) used for model training. Finally, the mean and standard deviation of the evaluation values calculated in the three experiments were calculated. Three evaluation metrics were compared: AUC, s-AUC, and NSS. The specific experimental values are shown in Table 1, where the baseline plus pre-training represents the result without the multi-head fusion module, and TDiffSal represents the result after using the multi-head fusion module. The baseline refers to the Stable Diffusion model (the baseline is the Stable Diffusion model, and this model adds a multi-head fusion module and a combined loss module on top of it). The experimental results in the table show that the data obtained using the MHF module are significantly better than those obtained without it, and the stability of the results across the three experiments is also better. Therefore, we can conclude that using the multi-head fusion module has a significant positive impact on the model's saliency prediction task.
[0095] Table 1. Ablation experiment results on EGO4D
[0096]
[0097] To verify the effectiveness of pre-training on the SALICON dataset, experiments compared the performance metrics of adapting the text dataset using pre-trained weights versus training directly on the text dataset without pre-trained weights. The comparison results used mean and variance. Notably, to ensure a fair comparison, the experiments performed the same number of training epochs in both cases. The experimental results are shown in Table 1, where the base framework plus the multi-head fusion module represents the results without using pre-trained weights, while TDiffSal represents the results after using pre-trained weights. It is clear from the table that the performance metrics calculated after training the model without pre-trained weights are significantly lower than those calculated using pre-trained weights. Therefore, we can conclude that pre-training on the SALICON dataset has a positive impact on the training of the TDiffSal model.
[0098] The above-described embodiments are merely one implementation of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention should be determined by the appended claims.
Claims
1. A text-guided visual saliency prediction method, characterized in that, Includes the following steps: Step 1: Data Preparation Obtain the text dataset and the image saliency detection dataset; perform pairwise processing on the text dataset; Step 2: Construction of the visual saliency prediction model: Construct the TDiffSal model, including a significance prediction diffusion module, a multi-head fusion module, and a combined loss function; Step 3: Initial training of the image saliency detection dataset: Initial training of the model was initiated using an image saliency detection dataset: the original image and the true saliency image were mapped to the latent space, and the features of the original image were denoised to generate noisy image features; The text encoder converts the corresponding text description into a text embedding vector; the noisy image features and the text embedding vector are input into the multi-head fusion module to obtain the text-guided image feature vector, and then the saliency prediction diffusion module outputs the saliency prediction result of the latent space. The dual-loss optimization model is calculated: the sum of the latent space loss and the pixel space loss is used as the total loss, and the model parameters are updated through backpropagation; the loss is evaluated after each training round, and if the current loss is the minimum value, it is saved as the optimal weight; otherwise, the iteration continues until the preset number of training rounds is reached to complete the initial training. Step 4: Fine-tuning the text dataset: Load the optimal weights saved from the initial training and fine-tune them using a text dataset: map the original images and true saliency images in the text dataset to the latent space and add noise, input the corresponding text description encoded into the embedding vector, and output the latent space prediction results through the multi-head fusion module and the saliency prediction diffusion module; similarly, strengthen the ability of the text to guide saliency prediction by calculating the dual loss optimization model. Step 5: Test Prediction Using the final saved optimal weights, the original images from the test dataset are input into the trained TDiffSal model, which outputs the final saliency image prediction results.
2. The text-guided visual saliency prediction method according to claim 1, characterized in that, Specifically, the following steps are included: S1. Obtain a text dataset containing saliency images that match the corresponding video screenshots and corresponding text description information, and an image saliency detection dataset containing the original image and the corresponding saliency image. Perform pairwise processing on the data in the text dataset and match the paired text descriptions in the paired text dataset with the original image. S2. Construct the visual saliency prediction model, also known as the TDiffSal model. The visual saliency prediction model includes three modules: a saliency prediction diffusion module, a multi-head fusion module, and a combined loss function. The multi-head fusion module is used for image-text feature fusion, the saliency prediction diffusion module is used to predict the saliency image after fusion, and the combined loss function is used to train the model. The visual saliency prediction model is trained using the image saliency detection dataset. It is determined whether the number of training rounds has been exceeded. If it has not been exceeded, proceed to step S3. If it has been exceeded, save the current model weights, end the training, and proceed to step S7. S3. Input the original image and the corresponding saliency image from the image saliency detection dataset, and map them to the latent space; then add noise to the original image in the latent space to obtain the features of the noisy image; set the text input corresponding to the image, and process it into a text embedding vector through a text encoder; S4. The noisy image and the text embedding vector are fused through a multi-head fusion module to obtain a text-guided image feature vector; this vector is then input into the saliency diffusion model for training to obtain the prediction results of the latent space. S5. Calculate the mean squared error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean squared error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the model parameters, so that the final output saliency image is closer and closer to the true saliency map. S6. Update the optimal training result weights. If the current calculated loss is the minimum, save the current training weights as the optimal weights. If there were previously optimal weights, they will be overwritten. After this step, return to step S2. S7. The model reads the optimal weights saved from previous training, uses the text dataset to further train the visual saliency prediction model, and determines whether the number of training rounds has been exceeded. If it has not been exceeded, proceed to step S8. If it has been exceeded, save the current model weights, end the training, and proceed to step S12. S8. Input the original image and the corresponding saliency image from the text dataset, and map them to the latent space; then add noise to the original image in the latent space to obtain the features of the noisy image; Input the corresponding descriptive text, which is then processed into a text embedding vector by a text encoder; S9. The noisy image and the text embedding vector are fused through a multi-head fusion module to obtain a text-guided image feature vector; this vector is then input into the saliency diffusion model for training to obtain the prediction results of the latent space. S10. Calculate the mean square error between the latent space prediction result and the latent space representation of the true saliency image as the latent space loss. Map the latent space prediction result to the pixel space and calculate the mean square error between it and the true saliency image as the pixel space loss. Use the sum of the two losses for backpropagation to optimize the model parameters so that the final output saliency image is closer and closer to the true saliency map. S11. Update the optimal training result weights. If the current calculated loss is the minimum, save the current training weights as the optimal weights. If there were previously optimal weights, they will be overwritten. After this step, return to step S7. S12. The model uses the optimal weights to predict the test dataset, which yields the final saliency image of the predicted test dataset.
3. The text-guided visual saliency prediction method according to claim 1, characterized in that, The specific implementation process of the multi-head fusion module is as follows: The key elements for calculating attention are key K, query Q, and value V. The multi-head fusion module (MHF) uses visual features as Q, and text embedding vectors as K and V. The three vectors are obtained by multiplying the corresponding input by the corresponding projection matrix, as shown in the following formula: Q=x t W q ,K=cW k ,V=cW k Where x t represents the visual feature vector input at step t, c represents the current input text embedding vector, and W... q W k These represent the corresponding learnable projection matrices. After obtaining the K, Q, and V vectors needed to calculate attention, the dot product of Q and K yields the correlation score between the two features. Finally, the SoftMax function converts the score into attention weights for each key. After the above operations, the attention weights need to be multiplied by V to obtain the final output. The specific calculation formula is as follows: Where d represents the scaling factor, Attn represents the attention output, and T represents the transpose; this step is the final cross-attention output of single-head computation; multi-head computation requires dividing the input KQV into multiple subspaces, calculating the corresponding attention value for each subspace using the above method, and finally concatenating the outputs; the specific formula is shown below: head i This represents the attention output of the i-th head. This represents the image feature output containing textual meaning after fusion by the MHF module; after obtaining the MHF fused features, the model input will be changed from the originally set x t Updated to 4. The text-guided visual saliency prediction method according to claim 1, characterized in that, The specific implementation process of the significance prediction diffusion module is as follows: The diffusion process of the saliency prediction diffusion module predicts saliency and generates a visual saliency image. Text-guided image features serve as input to the denoising neural network in the saliency prediction diffusion module, and the output is the final predicted saliency image in the latent space. The denoising neural network used in the denoising process is a U-Net structure network. In this saliency prediction diffusion module, the inputs are text features, text-guided image features, and time steps. This denoising neural network directly predicts the final generated image, and its calculation process can be formalized as follows: in This represents the predicted visual saliency map, F. Diff The saliency prediction diffusion module transforms the saliency prediction task into a saliency map generation task. First, the feature processing is performed by the MHF module, and then the visual saliency image is further generated by the saliency prediction diffusion module to obtain the text-guided visual saliency image in the latent space. Finally, the saliency map is restored from the latent space to the pixel space by the decoder.
5. The text-guided visual saliency prediction method according to claim 1, characterized in that, The specific implementation process of the combined loss function is as follows: The loss function is divided into loss calculation for the latent space and loss calculation for the true pixel space; in calculating the latent space loss, the true saliency image X is first calculated. sal By mapping the encoder to the latent space, we obtain its latent space representation z. sal Subsequently, the predicted latent space saliency map is calculated. With z sal The mean squared error between them yields the latent space training loss L. latent The mathematical formula for its calculation is as follows: Where N represents the number of samples calculated in this batch; the predicted latent space saliency map The pixel space representation X is obtained by mapping the pixel space to the decoder. pre Then, the pixel spatial loss L is calculated using the same loss calculation method as for the latent space. pixel The mathematical formula for its calculation is as follows: Finally, the final model loss is obtained by adding the loss function values of the two parts; its calculation formula is as follows: THE all =L latent +λL pixel Where λ represents the weight value assigned to the pixel spatial loss result.
6. The text-guided visual saliency prediction method according to claim 1, characterized in that, The EGO4D and SALICON datasets were selected for training.