A task-driven visual attention prediction method, device and system
By combining bottom-up multi-low-level visual feature fusion and top-down task guidance, a task-driven visual attention prediction method is constructed, which solves the problems of insufficient accuracy and application limitations of visual attention prediction in the prior art, and achieves more accurate visual attention prediction under the interactive tasks of ordinary people.
Patent Information
- Application Number
- CN202210751774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The existing visual attention prediction methods are mainly based on bottom-up stimulus drive, and fail to effectively combine top-down task and target information, resulting in insufficient accuracy of visual attention prediction under interactive tasks such as information browsing, navigation, and search, and their applications are limited to specific groups and scenarios.
By constructing a task-driven visual attention prediction method, combining bottom-up multi-low-level visual feature fusion and top-down task guidance, a training system is built, including feature fusion module, task guidance module, feature inference module and decoder module, and optimize parameters to generate a visual attention prediction model.
It realizes more accurate results for visual attention prediction of ordinary people under interactive tasks such as information browsing, navigation, and search, and improves prediction accuracy based on task state, and is suitable for a wider range of people and scenarios.
Smart Images

Figure CN115147677B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and specifically relates to a task-driven visual attention prediction method, device and system. Background Art
[0002] Visual attention prediction refers to the prediction of the probability of different positions in an image or video receiving visual attention, and is applied to various fields. Most of the existing visual attention predictions focus on the study of human physiological instinctive attention. For example, the Chinese invention patent CN101980248A discloses a natural scene target detection method based on an improved visual attention model, which predicts natural scene detection targets based on some low-level visual features such as brightness, color and direction. The Chinese invention patent CN110251076B discloses a saliency prediction method based on contrast that integrates visual attention, highlighting the guidance of color on human eye attention. The Chinese invention patent CN110827193A discloses a panoramic video saliency detection method based on multi-channel features. The low-level and high-level visual features used in these methods to predict visual attention are driven by bottom-up stimuli, and do not involve attention prediction combined with top-down information (such as tasks and goals). Subjective consciousness such as tasks is very important in guiding people's visual attention, and it is of great significance to study task-driven attention.
[0003] At the application level, the current application of visual attention prediction is only for specific groups of people, and visual attention prediction for ordinary people in interactive tasks such as information browsing, navigation, and search has not been achieved. The Chinese invention patent CN114092900A discloses a method for predicting driver visual attention. The preprocessed image data is input into the attention prediction network that combines convolutional neural networks and Transformer, and an attention prediction model is trained; the preprocessed image is input into the attention prediction model, and an attention prediction probability map is output. However, this technology does not mention the top-down cognitive level of the driver, the prediction of visual attention will produce deviations, and the application is limited to driving scenarios. The Chinese invention patent CN111951637A discloses a method for extracting the visual attention allocation pattern of drone pilots associated with task scenarios, which distinguishes the attention allocation situations of pilots under different task scenarios and different fatigue levels. However, the application of this technology is limited to drone flight driving scenarios. Summary of the invention
[0004] In view of the above, the purpose of the present invention is to provide a task-driven visual attention prediction method, device and system, which, by combining bottom-up stimulus drive with top-down task guidance, can more accurately predict human visual attention of ordinary people in interactive tasks such as information browsing, navigation, and search.
[0005] To achieve the above-mentioned object of the invention, an embodiment provides a task-driven visual attention prediction method, comprising the following steps:
[0006] Obtain an image sequence, and perform noise data cleaning and data enhancement on the image sequence to serve as sample data;
[0007] Constructing a training system, the training system includes a bottom-up feature fusion module, a top-down task guidance module, a feature reasoning module, and a decoder module, wherein the bottom-up feature fusion module is used to extract and fuse multiple low-level visual features of the input image sequence to obtain visual features; the top-down task guidance module is used to extract features of the input task information, fuse them with the visual features, and then reconstruct them to obtain reconstructed features, and perform task prediction based on the reconstructed features to obtain task prediction results; the feature reasoning module is used to re-extract features of the input visual features to obtain new features; the decoder module is used to perform visual attention prediction on the input new features and output an attention probability map;
[0008] Construct a loss function, which includes prediction loss based on the attention probability map, reconstruction constraint loss based on the reconstruction features, and task constraint loss based on the task prediction results.
[0009] The training system parameters are optimized according to the sample data and the loss function. After the parameter optimization is completed, the bottom-up feature fusion module, spatiotemporal reasoning module and decoder module determined by the parameters are extracted to form a visual attention prediction model;
[0010] Visual focus prediction using visual attention prediction model.
[0011] In one embodiment, the bottom-up feature fusion module extracts low-level visual features from three aspects of color, contrast, and direction features for the input image sequence, and then uses a self-attention mechanism to align the three aspects of the low-level visual features and then add the features to obtain visual features.
[0012] In the top-down task guidance module of one embodiment, task information is presented in the form of labels on a graph, wherein the task labels include task labels used as coarse-grained prompts and sub-task labels used as fine-grained prompts; the task labels and sub-task labels on the image are encoded using the BERT model and then fused to obtain task features; the visual features are pooled and then fused with the task features to obtain fused features; the AVE model is used to reconstruct the fused features to obtain reconstructed features; and a multi-classification model is used to perform task prediction on the reconstructed features to obtain task prediction results.
[0013] In the feature inference module of one embodiment, a VGG model is used to re-extract input visual features to obtain new features.
[0014] In one embodiment, the prediction loss constructed based on the attention probability map includes the prediction loss based on the entire image. and pixel-based prediction loss Specifically expressed as:
[0015]
[0016]
[0017] Among them, a represents the attention probability map, represents the attention truth label map, ||·||1 represents the 1-norm, a ij represents the attention probability of the jth pixel in the i-th image, represents the attention truth label of the jth pixel in the i-th image, and ω is the attention truth label map Note that the area ratio, ⊙ represents the dot product operation, and W and H represent the length and width of the image respectively.
[0018] In one embodiment, the reconstruction constraint loss constructed based on the reconstruction feature It is expressed as:
[0019]
[0020] Among them, f x represents the fusion feature obtained by fusion of the input task information with the visual feature after feature extraction, f x|z represents the reconstructed features, μ and σ represent the fusion features f x The mean and variance of the Gaussian distribution of the learned latent feature fz;
[0021] The task constraint loss constructed based on the task prediction result It is expressed as:
[0022]
[0023] Among them, y represents the task prediction result, represents the task truth value, F ce (·) represents the standard function of cross entropy loss.
[0024] In one embodiment, the method of using a visual attention prediction model to perform visual main force prediction includes:
[0025] The bottom-up feature fusion module is used to extract and fuse multiple low-level visual features of the input image to obtain visual features;
[0026] The feature inference module is used to extract the input visual features to obtain new features;
[0027] The decoder module is used to predict the visual attention of the new input features and output the attention probability map.
[0028] To achieve the above-mentioned purpose of the invention, an embodiment provides a task-driven visual main force prediction device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores a visual attention prediction model constructed by the above-mentioned task-driven visual main force prediction method; when the processor executes the computer program, the following steps are implemented:
[0029] receiving an image sequence;
[0030] Calling the visual attention prediction model to perform attention prediction on the image sequence, including: using the bottom-up feature fusion module to extract and fuse multiple low-level visual features of the input image to obtain visual features; using the feature inference module to re-extract the input visual features to obtain new features; using the decoder module to perform visual attention prediction on the input new features to obtain an attention probability map;
[0031] Output the attention probability map and visualize it in the form of a heat map.
[0032] To achieve the above-mentioned invention object, the embodiment provides a task-driven visual main force prediction system, including a client and a server, wherein the client is used to receive an input image sequence through a page interface and transmit the image sequence to the server; and is also used to visualize the attention probability map;
[0033] The server is mounted with a visual attention prediction model constructed by the above-mentioned task-driven visual main force prediction method, which is used to use the visual attention prediction model to predict the attention of the incoming image sequence and return the attention probability map to the client.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] By constructing a visual attention prediction model based on the bottom-up fusion of multiple low-level visual features and the guidance of task information, the model can achieve visual attention prediction for more general people in interactive tasks such as information browsing, navigation, and search, and improve the accuracy of prediction results based on task status. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 is a flowchart of a task-driven visual main force prediction method provided by an embodiment;
[0038] Figure 2 It is a schematic diagram of the structure of the training system provided in the embodiment. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0040] Figure 1 FIG. 4 is a flowchart of a task-driven visual force prediction method provided by an embodiment. Figure 1 As shown, the task-driven visual main force prediction method provided by the embodiment includes the following steps:
[0041] Step 1: Obtain an image sequence, and perform noise data cleaning and data enhancement on the image sequence to serve as sample data.
[0042] In the embodiment, an eye tracker is used to collect image sequence datasets driven by users in different tasks (including interactive tasks such as information browsing, navigation, and search). Each image is labeled with objects, tasks, and subtasks related to the task. In order to make the image sequence dataset usable for task-driven attention prediction, all images are additionally annotated with attention objects. The dataset includes images of people in a natural state without task drive and images driven by specific tasks.
[0043] In the embodiment, noise data cleaning and data enhancement are performed on the input images, and it is ensured that the subjects do not overlap in all tasks.
[0044] Step 2: Build a training system. The training system includes a bottom-up feature fusion module, a top-down task guidance module, a feature reasoning module, and a decoder module.
[0045] like Figure 2As shown, the training system provided by the embodiment includes a bottom-up feature fusion module (BU), a top-down task guidance module (TD), a feature inference module (FI), and a decoder module (DE). Among them, the feature fusion module (BU) is used to extract and fuse multiple low-level visual features of the input image sequence to obtain visual features; the task guidance module (TD) is used to extract features from the input task information, fuse them with the visual features, and then reconstruct them to obtain reconstructed features, and perform task prediction based on the reconstructed features to obtain task prediction results; the feature inference module (FI) is used to re-extract features from the input visual features to obtain new features; the decoder module (DE) is used to perform visual attention prediction on the input new features and output an attention probability map.
[0046] In the feature fusion module (BU), the color, contrast, and direction features in the image are fused together to provide bottom-up clues for attention prediction. Specifically, low-level visual features are extracted from the input image sequence from three aspects: color, contrast, and direction features. Then, in order to better fuse multiple low-level visual features, the self-attention mechanism is used to align the low-level visual features of the three aspects. The aligned low-level visual features f h 、f c 、f d Added together to form the bottom-up visual feature f bu :
[0047] f bu =f h +f c +f d
[0048] In the embodiment, task information is presented in the form of labels on the graph, embedded as features combined with bottom-up clues, and tasks are also formulated as constraints to guide the prediction of task-driven attention. A task usually consists of several subtasks, and images may have the same task label but different subtask labels. Task labels are coarse-grained cues, while subtask labels are fine-grained cues.
[0049] The BERT model allows exploring the semantic relationship between different words, which helps to build a powerful feature representation to encode task information. Therefore, in the task guidance module (TD), the BERT model is used to encode the task label and subtask label on the image and then fused to obtain the top-down task feature f td . Visual features f bu After the pooling operation, the features are obtained feature Then fused with task features td Get the fusion feature f x , expressed as:
[0050]
[0051] Among them, F c (·) indicates the fusion operation cat.
[0052] In order to obtain a better fusion feature representation, in the task guidance module (TD), the AVE (Variational Autoencoder) model is also used to reconstruct the fusion features to obtain the reconstructed features. That is, first, the fusion feature f x Based on learning a latent feature f z , then in the latent feature f z The reconstruction feature f is calculated based on x|z , the formula is:
[0053]
[0054] in, represents the AVE model, μ and σ represent the fusion feature f x The mean and variance of the Gaussian distribution of the learned latent features fz.
[0055] In the task guidance module (TD), a multi-classification model is used to reconstruct the feature f x|z Perform task prediction to obtain the task prediction result y, which is expressed as:
[0056]
[0057] in, Represents a multi-classification model, which can be a fully connected network.
[0058] In the feature inference module (FI), the visual features f corresponding to the image sequence are bu It is input into the VGG network, which learns more accurate new features f by exploring the association between image features and task features. b , expressed as:
[0059]
[0060] in, Represents the VGG network.
[0061] In the decoder module (DE), the new feature f b is sampled as a probability map and displayed as a black and white heat map to indicate the possible location of human visual attention, that is, the new feature f b Perform visual attention prediction and output pixel-level attention probability map a, the formula is expressed as:
[0062]
[0063] in, represents the decoding operation to predict the visual attention probability.
[0064] Step 3: Construct a loss function. The loss function includes the prediction loss based on the attention probability map, the reconstruction constraint loss based on the reconstruction features, and the task constraint loss based on the task prediction results.
[0065] In the embodiment, the constructed loss function is expressed as for:
[0066]
[0067] Among them, represents the prediction loss based on the entire image, Represents pixel-based prediction loss, prediction loss and Composed of prediction losses built on the attention probability map, represents the reconstruction constraint loss based on the reconstruction features, represents the task constraint loss constructed based on the task prediction results, and λ1, λ2, λ3, and λ4 represent the weights of a single loss.
[0068] In the embodiment, the prediction loss and Respectively expressed as:
[0069]
[0070]
[0071] Among them, a represents the attention probability map, represents the attention truth label map, ||·||1 represents the 1-norm, a ij represents the attention probability of the jth pixel in the i-th image, represents the attention truth label of the jth pixel in the i-th image. The truth label is represented by 1 and 0, where 1 represents the attention area and 0 represents the non-attention area. ω is the attention truth label map Note that the area ratio, ⊙ represents the dot product operation, W and H represent the length and width of the image respectively, and the prediction loss Indicates a and The overlap between a and The difference in predicted loss Represents the weighted cross entropy, which measures a and difference.
[0072] In this embodiment, the reconstruction constraint loss It is expressed as:
[0073]
[0074] Among them, f x represents the fusion feature obtained by fusion of the input task information with the visual feature after feature extraction, f x|z represents the reconstructed features, μ and σ represent the fusion features f x The mean and variance of the Gaussian distribution of the learned latent features fz.
[0075] The reconstruction constraint loss is obtained by minimizing f x and f x|z The difference between them, the reconstruction constraint helps to learn a stable representation of the fusion feature to weaken the top-down task feature f td and bottom-up visual features f bu The gap between.
[0076] In the embodiment, the task constraint loss It is expressed as:
[0077]
[0078] Among them, y represents the task prediction result, represents the task truth value, F ce (·) represents the standard function of cross entropy loss.
[0079] The task constraint loss The purpose is to encourage the model to predict task-driven attention by introducing top-down task information into the back-propagation of the network.
[0080] The above two constraints are used to strengthen feature representation and task guidance. The above two constraints are not directly added to the visual features f actually used for attention prediction. bu But the entire network parameters can be updated to update the visual features f bu To convey information from top to bottom and bottom to top.
[0081] Step 4: Optimize the parameters of the training system according to the sample data and the loss function. After the parameter optimization is completed, extract the bottom-up feature fusion module, spatiotemporal reasoning module and decoder module determined by the parameters to form a visual attention prediction model.
[0082] In the embodiment, the image sequence in the sample data of step 1 is input into the training system, and the visual features, task features, task prediction results, reconstruction features and attention probability maps are obtained through calculation. Then, the loss is calculated according to the loss function, and the parameters of the training system are updated using the loss. A cross-validation is performed. After the parameter optimization is completed, the bottom-up feature fusion module, spatiotemporal reasoning module and decoder module determined by the extraction parameters are used to form a visual attention prediction model. The visual attention prediction model can combine the bottom-up visual feature-driven and top-down task-guided visual attention prediction to generate an attention probability map.
[0083] Step 5: Use the visual attention prediction model to perform visual main force prediction.
[0084] In the embodiment, a visual attention prediction model is used to perform visual main force prediction, including: using a bottom-up feature fusion module to extract and fuse multiple low-level visual features of an input image to obtain visual features; using a feature inference module to re-extract features of the input visual features to obtain new features; using a decoder module to predict visual attention of the input new features and output an attention probability map, which is presented to the system application in the form of a heat map to visualize the user's visual attention focus. Since it is an input image sequence with time information, the attention probability map presented in the form of a heat map can display the visual attention area, which can obtain the visual attention duration, and support visual attention prediction under interactive tasks such as information browsing, navigation, and search.
[0085] Based on the same inventive concept, an embodiment further provides a task-driven visual main force prediction device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores a visual attention prediction model constructed by the above-mentioned task-driven visual main force prediction method; when the processor executes the computer program, the following steps are implemented:
[0086] Step 1, receiving an image sequence;
[0087] Step 2, calling the visual attention prediction model to perform attention prediction on the image sequence, including: using a bottom-up feature fusion module to extract and fuse multiple low-level visual features of the input image to obtain visual features; using a feature inference module to re-extract the input visual features to obtain new features; using a decoder module to perform visual attention prediction on the input new features to obtain an attention probability map;
[0088] Step 3: Output the attention probability map and visualize it in the form of a heat map.
[0089] In the application, the memory can be a volatile memory at the near end, such as RAM, or a non-volatile memory, such as ROM, FLASH, floppy disk, mechanical hard disk, etc., or a remote storage cloud. The processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), that is, the steps of attention prediction of image sequences can be implemented by these processors.
[0090] Based on the same inventive concept, an embodiment also provides a task-driven visual main force prediction system, including a client and a server, wherein the client is used to receive an input image sequence through a page interface and transmit the image sequence to the server; the server is mounted with a visual attention prediction model constructed by the above-mentioned task-driven visual main force prediction method, and is used to use the visual attention prediction model to predict the attention of the input image sequence and return an attention probability map to the client. The client is also used to visualize the attention probability map.
[0091] The task-driven visual attention prediction method, device and system provided in the embodiments construct a visual attention prediction model by integrating multiple low-level visual features from the bottom up and guiding task information, so that the model can realize the visual attention prediction of more general people in interactive tasks such as information browsing, navigation, and search, and improve the accuracy of prediction results based on task status.
[0092] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A task-driven visual force prediction method, characterized in that: The following steps are involved: Obtain an image sequence, and perform noise data cleaning and data enhancement on the image sequence to serve as sample data; Constructing a training system, the training system includes a bottom-up feature fusion module, a top-down task guidance module, a feature reasoning module, and a decoder module, wherein the bottom-up feature fusion module is used to extract and fuse multiple low-level visual features of the input image sequence to obtain visual features; the top-down task guidance module is used to extract features of the input task information, fuse them with the visual features, and then reconstruct them to obtain reconstructed features, and perform task prediction based on the reconstructed features to obtain task prediction results; the feature reasoning module is used to re-extract features of the input visual features to obtain new features; the decoder module is used to perform visual attention prediction on the input new features and output an attention probability map; Construct a loss function, which includes prediction loss based on the attention probability map, reconstruction constraint loss based on the reconstruction features, and task constraint loss based on the task prediction results. The training system parameters are optimized according to the sample data and the loss function. After the parameter optimization is completed, the bottom-up feature fusion module, spatiotemporal reasoning module and decoder module determined by the parameters are extracted to form a visual attention prediction model; Visual focus prediction using visual attention prediction model.
2. The task-driven visual force prediction method according to claim 1, characterized in that: The bottom-up feature fusion module extracts low-level visual features from the three aspects of color, contrast, and direction features of the input image sequence, and then uses a self-attention mechanism to align the three aspects of the low-level visual features and then add the features to obtain the visual features.
3. The task-driven visual force prediction method according to claim 1, characterized in that: In the top-down task guidance module, task information is presented in the form of labels on the graph, wherein the task labels include task labels used as coarse-grained prompts and subtask labels used as fine-grained prompts; the task labels and subtask labels on the image are encoded using the BERT model and then fused to obtain task features; the visual features are pooled and then fused with the task features to obtain fused features; the AVE model is used to reconstruct the fused features to obtain reconstructed features; and a multi-classification model is used to perform task prediction on the reconstructed features to obtain task prediction results.
4. The task-driven visual force prediction method according to claim 1, characterized in that: In the feature inference module, the VGG model is used to re-extract the input visual features to obtain new features.
5. The task-driven visual force prediction method according to claim 1, characterized in that: The prediction loss constructed based on the attention probability map includes the prediction loss based on the entire image and pixel-based prediction loss Specifically expressed as: Among them, a represents the attention probability map, represents the attention truth label map, ||·||1 represents the 1-norm, a ij represents the attention probability of the jth pixel in the i-th image, represents the attention truth label of the jth pixel in the i-th image, and ω is the attention truth label map Note that the area ratio, ⊙ represents the dot product operation, and W and H represent the length and width of the image respectively.
6. The task-driven visual force prediction method according to claim 1, characterized in that: The reconstruction constraint loss constructed based on the reconstruction feature It is expressed as: Among them, f x represents the fusion feature obtained by fusion of the input task information with the visual feature after feature extraction, f x|z represents the reconstructed features, μ and σ represent the fusion features f x The mean and variance of the Gaussian distribution of the learned latent feature fz; The task constraint loss constructed based on the task prediction result It is expressed as: Among them, y represents the task prediction result, represents the task truth value, F ce (·) represents the standard function of cross entropy loss.
7. The task-driven visual force prediction method according to claim 1, characterized in that: The method of using the visual attention prediction model to predict the visual main force includes: The bottom-up feature fusion module is used to extract and fuse multiple low-level visual features of the input image to obtain visual features; The feature inference module is used to extract the input visual features to obtain new features; The decoder module is used to predict the visual attention of the input new features and output the attention probability map.
8. A task-driven visual force prediction device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The memory stores a visual attention prediction model constructed by the task-driven visual main force prediction method according to any one of claims 1 to 7; when the processor executes the computer program, the following steps are implemented: receiving an image sequence; Calling the visual attention prediction model to perform attention prediction on the image sequence, including: using the bottom-up feature fusion module to extract and fuse multiple low-level visual features of the input image to obtain visual features; using the feature inference module to re-extract the input visual features to obtain new features; using the decoder module to perform visual attention prediction on the input new features to obtain an attention probability map; Output the attention probability map and visualize it in the form of a heat map.
9. A task-driven visual force prediction system, comprising a client and a server, characterized in that: The client is used to receive an input image sequence through a page interface and transmit the image sequence to a server; and is also used to visualize the attention probability map; The server is mounted with a visual attention prediction model constructed by the task-driven visual main force prediction method described in any one of claims 1-7, and is used to use the visual attention prediction model to predict the attention of the incoming image sequence and return the attention probability map to the client.
Citation Information
Patent Citations
Improved visual attention model-based method of natural scene object detection
CN101980248A
A method and apparatus for saliency detection based on contrast, incorporating visual attention.
CN110251076B
Panoramic video saliency detection method based on multi-channel features
CN110827193A
Task scene associated unmanned aerial vehicle pilot visual attention distribution mode extraction method
CN111951637A
Method and system for predicting visual attention of driver, equipment and medium
CN114092900A