Method, medium and device for predicting saliency of multi-view video with spatial audio
By fusing visual and audio features from multi-view videos, a sound source energy distribution estimation module is constructed and cross-view alignment is performed, which solves the problem of visual and audio information fusion in multi-view video saliency prediction and improves prediction accuracy and consistency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies struggle to effectively integrate spatial audio and visual information in multi-view video saliency prediction, leading to inconsistent prediction results from different perspectives and failing to accurately reflect attention distribution in real-world scenes.
By extracting visual texture and motion features from multi-view videos, performing time-frequency transformation and multi-scale feature encoding, and combining the spatial audio source energy distribution estimation module, sound source energy distribution features are constructed. Multimodal fusion and cross-view saliency alignment are then performed to generate consistent saliency feature representations.
It improves the accuracy and consistency of multi-view video saliency prediction, enhances the ability to locate salient regions, and provides an effective supplement to audio semantics, forming a good synergistic gain.
Smart Images

Figure CN122454489A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information fusion and intelligent sensing technology, specifically to a method, medium, and device for multi-view video saliency prediction that integrates spatial audio. Background Technology
[0002] With the development of virtual reality, augmented reality, free-viewpoint video, and immersive media technologies, multi-view video, capable of simultaneously recording the same scene from different viewing directions, has become an important data format in 3D scene perception, immersive content generation, and intelligent media processing. Compared to single-view video, multi-view video contains richer spatial structure and viewpoint correlation information, and can more comprehensively reflect the distribution of targets, motion states, and areas of visual attention in a scene. Therefore, conducting saliency prediction based on multi-view video is of great significance.
[0003] Saliency prediction aims to estimate more attention-grabbing regions in a scene and has wide applications in tasks such as video understanding, content encoding, resource allocation, object detection, and interactive display. Most existing saliency prediction methods are designed for single-view images or ordinary videos, primarily relying on visual cues such as color, texture, and motion to estimate salient regions. For multi-view videos, since different viewpoints correspond to different observations of the same scene, and these viewpoints have both content differences and strong geometric and semantic correspondences, using single-view methods to predict independently for each viewpoint easily leads to inconsistent prediction results across different viewpoints, making it difficult to accurately reflect the true attentional distribution in the scene.
[0004] On the other hand, in immersive scenarios, sound, especially spatial audio with directional and spatial location information, can significantly guide human visual attention. For example, when a voice, musical instrument, or other sudden sound source is present in a certain direction, users tend to prioritize the area associated with that sound. Therefore, relying solely on visual information for saliency prediction is insufficient to fully characterize human attention mechanisms in real-world scenarios. Jointly modeling spatial audio information with multi-view video information is an important direction for improving the accuracy of saliency prediction.
[0005] However, existing technologies still have shortcomings in multi-view video saliency prediction by fusing spatial audio. On the one hand, there are significant appearance variations between different viewpoints in multi-view videos. How to effectively model cross-view correlations and improve the consistency of multi-view saliency prediction while controlling computational complexity remains a technical challenge. On the other hand, spatial audio and visual content belong to different modalities of data, and there are complex correspondences between them at the temporal, spatial, and semantic levels. How to achieve effective alignment between sound cues and multi-view image regions, thereby accurately representing the guiding role of sound sources in visual attention, also lacks an effective solution. Summary of the Invention
[0006] The present invention proposes a multi-view video saliency prediction method that integrates spatial audio, which can at least solve one of the technical problems in the background art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A method for multi-view video saliency prediction that integrates spatial audio includes the following steps: S100: Acquire a multi-view video sequence of the same scene and an audio signal that is time-synchronized with the multi-view video sequence, and establish the sequential relationship between the views. S200. Extract the visual texture features and motion features of video frames from each viewpoint in the multi-view video sequence, and fuse the texture features and motion features to obtain the visual representation features corresponding to each viewpoint. S300. Perform time-frequency transformation and multi-scale feature encoding on the audio signals of the multi-view video sequence that are synchronized in time to obtain audio spectrum features, and select content spectrum features from the audio spectrum features as audio semantic representation; S400. Construct a sound source energy distribution estimation module, jointly model visual representation features and audio spectrum features to obtain sound source energy distribution features corresponding to the current visual scene. S500: Multimodal fusion of visual representation features, content spectrum features and sound source energy distribution features is performed to obtain single-view multimodal saliency features corresponding to each viewpoint; S600. Based on the order relationship between the viewpoints, perform cross-viewpoint saliency alignment on the single-viewpoint multimodal saliency features corresponding to each viewpoint to obtain a cross-viewpoint consistent saliency feature representation. S700. Based on the saliency feature representation after cross-view alignment, generate saliency distribution maps for each view.
[0008] Furthermore, the method for establishing the sequential relationship between viewpoints in step S100 of the present invention includes: Let the input multi-view video sequence be denoted as . The audio signal synchronized with the multi-view video sequence is denoted as ; in, This represents the entire multi-view video dataset. This represents the entire set of synchronized audio data. Indicates the total number of viewpoints. Indicates the total number of time frames. Indicates the first The perspective in the first Video frames captured at each moment. Indicates the relationship with the first The audio segment corresponding to each moment; The viewpoint order is determined by the linear order obtained by arranging the cameras by their numbers and the adjacency order constructed by the spatial distribution of the cameras.
[0009] Furthermore, the visual representation feature acquisition method in step S200 of the present invention includes: For the The perspective in the first Video frames at that moment First, extract the texture features, which are represented as follows:
[0010] in, Indicates the first The perspective in the first Texture features corresponding to video frames at each time point This represents the texture feature extraction function; To characterize temporal changes and target motion information in the video, motion features are extracted, which are represented as follows:
[0011] in, Indicates the first The perspective in the first Motion characteristics corresponding to each moment Represents the motion feature extraction function. Indicates the preset time interval, symbol This indicates a splicing operation along the channel dimension; After obtaining texture and motion features, the two are fused to obtain visual representation features. The fusion process is represented as follows:
[0012] in, Indicates the first Visual representation features corresponding to each viewpoint This represents the texture features from that viewpoint. This indicates the motion characteristics from that perspective. This represents a visual feature fusion function, which consists of convolutional mapping, normalization, and nonlinear activation.
[0013] Furthermore, the method for obtaining content spectral features in step S300 of the present invention includes: For the The audio segment corresponding to each moment Perform a short-time Fourier transform to obtain the time-frequency representation of the audio:
[0014] in, Indicates the first The audio clips are indexed by frequency. and short window index Spectral values at that location This represents the short-time Fourier transform operator. Represents a frequency-dimensional index. Indicates the time window index; The original time-domain audio signal is converted into a two-dimensional time-frequency domain representation through short-time Fourier transform, and spectrum modeling is performed. Mapping the time-frequency representation to the Mel spectrum, denoted as:
[0015] In the formula, Indicates the first Mel spectrograms corresponding to each audio segment This represents the Mel filter group mapping operator; The Mel spectrum is input into the audio coding network for multi-scale feature extraction. The multi-scale audio spectrum features are represented as follows:
[0016] in, Indicates the first Audio spectral features at each scale level Indicates the audio coding network in the 1st... Feature extraction mapping of layers, Indicates the total number of scale layers; One layer is selected from the multi-scale audio spectral features as the content spectral feature, denoted as:
[0017] in, Indicates content spectral characteristics, This indicates a pre-selected hierarchical index.
[0018] Furthermore, the method for obtaining the sound source energy distribution features corresponding to the current visual scene in step S400 of the present invention includes: Let the first Layered audio spectrum features, after pre-gated processing, are represented as follows:
[0019] in, Indicates the first The fusion features of the layers after pre-gated processing This represents a mapping function that generates spatial gating weights for visual features. This represents a mapping function that generates spatially gated weights for audio features. This represents element-wise multiplication. Indicates the first Visual representation features from each perspective Indicates the first Layer audio spectrum features; After pre-gating processing, the fusion result is recalibrated through contextual attention enhancement processing. The process is as follows:
[0020] In the formula, Indicates the first Output features of the layered audio-visual attention unit Represents the pooling function. This indicates a context-sensitive mapping function; The sound source energy distribution estimation module adopts an inter-layer cascading approach to achieve multi-scale progressive fusion from low-level local details to high-level global semantics. The process is represented as follows:
[0021]
[0022] in, This indicates the energy distribution characteristics of the initial layer sound source. Indicates the first The energy distribution characteristics of the sound source after layer fusion. Denotes the upsampling mapping function, Represents the convolution fusion function. This indicates feature splicing.
[0023] Furthermore, the method for obtaining single-view multimodal salient features corresponding to each viewpoint in step S500 of the present invention includes: The three modal features are mapped to a unified feature space to form a shared representation, which is:
[0024] in, Indicates the first Shared modal representation from multiple perspectives , and These represent the visual feature mapping parameters, content spectrum feature mapping parameters, and sound source energy distribution feature mapping parameters, respectively. This represents a convolutional mapping or linear mapping operation. Indicates the first Visual representation features from each perspective Indicates content spectral characteristics, Indicates the first Sound source energy distribution characteristics from multiple perspectives; In obtaining shared representation Then, the shared representation is concatenated with the original features of each modality to obtain the enhanced modal features, represented as:
[0025] in, Indicates enhanced visual features, This indicates the enhanced content spectral characteristics. This indicates the energy distribution characteristics of the enhanced sound source. , and These represent the enhancement mapping parameters for the corresponding modes; Finally, the three enhanced modal features are further fused to obtain single-view multimodal saliency features, the process of which is represented as follows:
[0026] in, Indicates the first Single-view multimodal saliency features corresponding to each viewpoint This represents a multimodal fusion function, implemented using convolutional mapping.
[0027] Furthermore, the method for obtaining cross-view consistent saliency feature representation in step S600 of the present invention includes: Let the single-view multimodal saliency features of all perspectives constitute the feature sequence. ,in, Represents multi-view feature sequences. Indicates the first Single-view multimodal saliency features from multiple perspectives; To propagate saliency information between adjacent viewpoints, the forward and backward feature propagation processes are represented as follows:
[0028]
[0029] in, Indicates the first Hidden state characteristics of each perspective during forward propagation Indicates the first Hidden state characteristics of each perspective during backpropagation. This represents the forward time propagation mapping function. This represents the backward time propagation mapping function.
[0030] After obtaining the forward hidden state and backward hidden state Then, the two are fused to obtain the cross-view aligned feature representation, specifically as follows:
[0031] in, Indicates the first A salient feature or salient response map of each viewpoint after cross-viewpoint alignment. Indicates the fusion mapping parameters, This represents the bias parameter.
[0032] Furthermore, the method for optimizing and training the entire network through supervised learning in step S700 of the present invention includes: Saliency features after cross-view alignment Activation mapping is performed to obtain the final predicted significance distribution map, represented as:
[0033] in, Indicates the first The final significance distribution map output from each perspective. Indicates the activation function; Let the first The saliency annotation map corresponding to each viewpoint is as follows: The significance of the supervised loss is expressed as:
[0034] in, Indicates significant predicted loss. Indicates the height of the saliency plot. Indicates the width of the saliency plot. Represents the pixel coordinates on the saliency map. The predicted significance plot is shown on the coordinates. The output value at that location, This indicates the saliency plot on the coordinate axis. The actual value at that point, This represents the pixel-level loss function.
[0035] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0036] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0037] In summary, this invention demonstrates that each component module positively impacts saliency prediction performance. The sound source energy distribution estimation module shows the most significant performance improvement, indicating that modeling spatial sound energy distribution effectively enhances salient region localization capabilities. Simultaneously, the cross-viewpoint saliency alignment module improves the consistency of saliency representations across different viewpoints, while content spectral features provide effective audio semantic supplementation for saliency prediction. The combined use of these three modules creates a strong synergistic gain, significantly improving the accuracy of multi-viewpoint video saliency prediction by fusing spatial audio. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the overall process of a multi-view video saliency prediction method that integrates spatial audio according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the structure of the AVA module according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the cross-viewpoint saliency alignment module according to Embodiment 1 of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0040] like Figure 1 As shown in this embodiment, the multi-view video saliency prediction method that integrates spatial audio includes the following steps: S100: Acquire a multi-view video sequence of the same scene and an audio signal that is time-synchronized with the multi-view video sequence, and establish the sequential relationship between the views. S200. Extract the visual texture features and motion features of video frames from each viewpoint in the multi-view video sequence, and fuse the texture features and motion features to obtain the visual representation features corresponding to each viewpoint. S300. Perform time-frequency transformation and multi-scale feature encoding on the audio signals of the multi-view video sequence that are synchronized in time to obtain audio spectrum features, and select content spectrum features from the audio spectrum features as audio semantic representation; S400. Construct a sound source energy distribution estimation module, jointly model visual representation features and audio spectrum features to obtain sound source energy distribution features corresponding to the current visual scene. S500: Multimodal fusion of visual representation features, content spectrum features and sound source energy distribution features is performed to obtain single-view multimodal saliency features corresponding to each viewpoint; S600. Based on the order relationship between the viewpoints, perform cross-viewpoint saliency alignment on the single-viewpoint multimodal saliency features corresponding to each viewpoint to obtain a cross-viewpoint consistent saliency feature representation. S700. Based on the saliency feature representation after cross-view alignment, generate saliency distribution maps for each view.
[0041] The following provides a detailed explanation of each step: S100: Acquire a multi-view video sequence of the same scene and an audio signal that is time-synchronized with the multi-view video sequence, and establish the sequential relationship between the views. Let the input multi-view video sequence be denoted as . The audio signal synchronized with the multi-view video sequence is denoted as ; in, This represents the entire multi-view video dataset. This represents the entire set of synchronized audio data. Indicates the total number of viewpoints. Indicates the total number of time frames. Indicates the first The perspective in the first Video frames captured at each moment. Indicates the relationship with the first The audio segment corresponding to each moment.
[0042] "Synchronization" refers to audio segments Corresponding video frames from all perspectives The sampling times are matched to each other, thus ensuring that the audio modality and the visual modality have a unified time reference in subsequent joint modeling.
[0043] Based on the actual arrangement of the camera array, establish the viewpoint sequence relationship for each viewpoint; The viewpoint order relationship can be a linear order relationship obtained by arranging the cameras according to their numbers, or an adjacency order relationship constructed based on the spatial distribution of the cameras. Its purpose is to provide an input order for the subsequent propagation of cross-viewpoint saliency information.
[0044] S200. Extract the visual texture features and motion features of video frames from each viewpoint in the multi-view video sequence, and fuse the texture features and motion features to obtain the visual representation features corresponding to each viewpoint. For the The perspective in the first Video frames at that moment First, extract the texture features, which can be represented as:
[0045] in, Indicates the first The perspective in the first Texture features corresponding to video frames at each time point This represents a texture feature extraction function, which can be implemented by a pre-trained image backbone network and is used to extract color, texture, contour, local structure, and semantic appearance information from an image.
[0046] Secondly, to characterize the temporal changes and target motion information in the video, motion features are extracted. These motion features can be represented as:
[0047] in, Indicates the first The perspective in the first Motion characteristics corresponding to each moment Represents the motion feature extraction function. Indicates the preset time interval, symbol This indicates that a splicing operation is performed along the channel dimension. In other words, the current frame is... Video frames at subsequent adjacent times After being stitched together, the data are input into a motion feature extraction network to obtain feature representations that reflect changes in scene motion.
[0048] After obtaining texture and motion features, the two are fused to obtain visual representation features. The fusion process can be represented as follows:
[0049] in, Indicates the first Visual representation features corresponding to each viewpoint This represents the texture features from that viewpoint. This indicates the motion characteristics from that perspective. The visual feature fusion function can be composed of convolutional mapping, normalization, and nonlinear activation. This approach allows for the simultaneous preservation of both static appearance saliency cues and dynamic motion saliency cues, thus forming visual fusion features suitable for saliency prediction.
[0050] S300. Perform time-frequency transformation and multi-scale feature encoding on the audio signals of the multi-view video sequence that are synchronized in time to obtain audio spectrum features, and select content spectrum features from the audio spectrum features as audio semantic representation; For the The audio segment corresponding to each moment First, a short-time Fourier transform is performed to obtain the time-frequency representation of the audio:
[0051] in, Indicates the first The audio clips are indexed by frequency. and short window index Spectral values at that location This represents the short-time Fourier transform operator. Represents a frequency-dimensional index. Indicates the time window index.
[0052] This transformation converts the original time-domain audio signal into a two-dimensional time-frequency domain representation for subsequent spectral modeling.
[0053] Mapping the time-frequency representation to the Mel spectrum, denoted as:
[0054] In the formula, Indicates the first Mel spectrograms corresponding to each audio segment This represents the Mel filter group mapping operator. The Mel spectrum can better simulate the differences in human hearing at different frequencies, and is therefore more conducive to extracting audio features related to human auditory attention.
[0055] After obtaining the Mel spectrum, it is input into the audio coding network for multi-scale feature extraction. The multi-scale audio spectrum features can be represented as:
[0056] in, Indicates the first Audio spectral features at each scale level Indicates the audio coding network in the 1st... Feature extraction mapping of layers, This indicates the total number of scale layers. The audio coding network can adopt a U-Net structure or other coding structures with multi-scale feature extraction capabilities.
[0057] One layer is selected from the multi-scale audio spectral features as the content spectral feature, denoted as:
[0058] in, Indicates content spectral characteristics, This represents a pre-selected hierarchical index. The content spectral features combine higher-level semantic representation capabilities with a certain spatial resolution, and are used to characterize the semantic information of the audio content itself.
[0059] S400. Construct a sound source energy distribution estimation module, jointly model visual representation features and audio spectrum features to obtain sound source energy distribution features corresponding to the current visual scene. To enable a fine-grained correspondence between audio information and potential sound-producing regions in visual images, this invention incorporates a sound source energy distribution estimation module within the audio branch. This module comprises multiple cascaded audiovisual attention units, each including two stages: pre-gating processing and contextual attention enhancement processing.
[0060] For the Layered audio spectral features, after pre-gated processing, can be represented as:
[0061] in, Indicates the first The fusion features of the layers after pre-gated processing This represents a mapping function that generates spatial gating weights for visual features. This represents a mapping function that generates spatially gated weights for audio features. This represents element-wise multiplication. Indicates the first Visual representation features from each perspective Indicates the first Layer audio spectrum characteristics.
[0062] The above formula indicates that, firstly, based on visual characteristics... Visual gating weights are generated, and then these weights are multiplied element-wise with the visual features themselves; simultaneously, based on audio features... An audio gating weight is generated, and then this weight is multiplied element-wise with the audio feature itself. Finally, the gated visual and audio features are additively fused. This step can highlight potential sound-producing regions and visual regions related to the sound source in the initial stage, while suppressing background components that are weakly related to saliency prediction.
[0063] After pre-gating, the fusion result is further recalibrated through contextual attention enhancement, which can be represented as follows:
[0064] In the formula, Indicates the first Output features of the layered audio-visual attention unit To represent the pooling function, you can use the global average pooling function. This represents the context attention mapping function, which can be constructed using a lightweight bottleneck network. The formula means: first, pooling operations are used to fuse features... Global context information is extracted, channel weights are generated by the context attention mapping function, and finally the channel weights are used to enhance the original fused features channel by channel, thereby improving the discriminative feature response related to the vocal region and the salient region.
[0065] To achieve multi-scale progressive fusion from low-level local details to high-level global semantics, the sound source energy distribution estimation module adopts an inter-layer cascade approach, the process of which can be represented as follows:
[0066]
[0067] in, This indicates the energy distribution characteristics of the initial layer sound source. Indicates the first The energy distribution characteristics of the sound source after layer fusion. Denotes the upsampling mapping function, Represents the convolution fusion function. This indicates feature splicing.
[0068] This step means: first, analyze the sound source energy distribution characteristics of the previous layer. Upsampling is performed, and then compared with the audiovisual attention features of the current layer. The layers are then stitched together, and finally, the fusion result of the current layer is obtained through convolutional mapping.
[0069] Finally, the output of the last layer is taken as the sound source energy distribution feature from the current perspective, denoted as:
[0070] in, Indicates the first The sound source energy distribution characteristics obtained from each perspective are used to characterize the potential spatial distribution of sound in the current visual scene, thereby reflecting the guiding relationship between spatial audio and visual attention.
[0071] S500: Multimodal fusion of visual representation features, content spectrum features and sound source energy distribution features is performed to obtain single-view multimodal saliency features corresponding to each viewpoint; First, the features of the three modalities are mapped to a unified feature space to form a shared representation, which can be expressed as:
[0072] in, Indicates the first Shared modal representation from multiple perspectives , and These represent the visual feature mapping parameters, content spectrum feature mapping parameters, and sound source energy distribution feature mapping parameters, respectively. This represents a convolutional mapping or linear mapping operation. Indicates the first Visual representation features from each perspective Indicates content spectral characteristics, Indicates the first The characteristics of sound source energy distribution from multiple perspectives.
[0073] In obtaining shared representation Then, the shared representation is concatenated with the original features of each modality to obtain the enhanced modal features, specifically represented as follows:
[0074] in, Indicates enhanced visual features, This indicates the enhanced content spectral characteristics. This indicates the energy distribution characteristics of the enhanced sound source. , and These represent the enhancement mapping parameters for the corresponding modes.
[0075] Finally, the three enhanced modal features are further fused to obtain single-view multimodal saliency features, the process of which can be represented as follows:
[0076] in, Indicates the first Single-view multimodal saliency features corresponding to each viewpoint The multimodal fusion function is represented by a convolutional mapping. Through the above-described shared representation construction and modality enhancement process, the complementary relationships between modalities can be strengthened while preserving the unique information of each modality, thereby improving the saliency modeling capability.
[0077] S600. Based on the order relationship between the viewpoints, perform cross-viewpoint saliency alignment on the single-viewpoint multimodal saliency features corresponding to each viewpoint to obtain a cross-viewpoint consistent saliency feature representation. Let the single-view multimodal saliency features of all perspectives constitute the feature sequence. ,in, Represents multi-view feature sequences. Indicates the first Single-view multimodal saliency features. To propagate saliency information between adjacent views, the forward and backward feature propagation processes can be represented as:
[0078]
[0079] in, Indicates the first Hidden state characteristics of each perspective during forward propagation Indicates the first Hidden state characteristics of each perspective during backpropagation. This represents the forward time propagation mapping function. This represents the backward temporal propagation mapping function. The forward and backward temporal propagation mapping functions can be implemented using a bidirectional convolutional long short-term memory network to transfer spatial attention information between different perspectives.
[0080] After obtaining the forward hidden state and backward hidden state Then, the two are fused to obtain the cross-viewpoint aligned feature representation, which can be specifically represented as:
[0081] in, Indicates the first A salient feature or salient response map of each viewpoint after cross-viewpoint alignment. Indicates the fusion mapping parameters, This represents the bias parameter.
[0082] The purpose of this step is to supplement the current perspective with both forward and backward perspective information, thereby enhancing the consistency of representations of the same salient target among different perspectives and reducing the fragmentation of results caused by independent predictions from a single perspective.
[0083] S700. Based on the saliency features after cross-view alignment, generate saliency distribution maps for each view and optimize the entire network through supervised learning. Saliency features after cross-view alignment After performing activation mapping, the final predicted significance distribution map is obtained, which can be represented as:
[0084] in, Indicates the first The final significance distribution map output from each perspective. The activation function can be the Softmax function, the Sigmoid function, or other mapping functions suitable for saliency output.
[0085] Let the first The saliency annotation map corresponding to each viewpoint is as follows: The significance monitoring loss can then be expressed as:
[0086] in, Indicates significant predicted loss. Indicates the height of the saliency plot. Indicates the width of the saliency plot. Represents the pixel coordinates on the saliency map. The predicted significance plot is shown on the coordinates. The output value at that location, This indicates the saliency plot on the coordinate axis. The actual value at that point, This represents a pixel-level loss function, which can be a cross-entropy loss function, etc.
[0087] Example 1 like Figure 1 As shown, the multi-view video saliency prediction method fused with spatial audio according to an embodiment of the present invention includes the following steps: Step 1: Acquire multi-view video sequences and synchronized audio signals First, a multi-view video sequence of the same scene, simultaneously captured from multiple camera perspectives, and an audio signal synchronized with the multi-view video sequence in time are acquired.
[0088] Let the multi-view video sequence be represented as The audio signal synchronized with the multi-view video sequence is denoted as ,in, This represents the entire multi-view video dataset. This represents the entire set of synchronized audio data. Indicates the total number of viewpoints. Indicates the total number of time frames. Indicates the first The perspective in the first Video frames captured at each moment. Indicates the relationship with the first The audio segment corresponding to each moment.
[0089] In this embodiment, "synchronization" refers to an audio segment. Video frames from all viewpoints at the corresponding time point Aligning them on the timeline ensures that the visual and audio modalities have a unified temporal reference during subsequent joint modeling.
[0090] Furthermore, a sequential relationship between viewpoints is established based on the physical arrangement of the cameras. This sequential relationship can be a relationship along the camera array arrangement direction, or a propagation sequence relationship constructed based on the proximity of the cameras. The purpose of establishing this sequential relationship is to provide feature propagation paths for subsequent cross-viewpoint saliency alignment.
[0091] Step 2: Extract visual features from each viewpoint For each video frame from a viewpoint, its visual features are extracted. These visual features include texture features and motion features. For the first... The viewpoint at the first Video frames at any moment First, texture features are extracted, which can be represented as:
[0092] in, Indicates the first The viewpoint at the first The texture feature tensor corresponding to the video frame at a given time. This represents the texture feature extraction function. In this embodiment, the RGB backbone network uses a pre-trained ResNet18 to extract texture features in order to capture visually salient regions.
[0093] Then, to characterize the temporal changes and local motion information in the video, motion features are extracted. These motion features can be represented as...
[0094] in, Indicates the first The perspective in the first Motion characteristics corresponding to each moment Represents the motion feature extraction function. Indicates the preset time interval, symbol This indicates that a splicing operation is performed along the channel dimension. In other words, the current frame is... Video frames at subsequent adjacent times After concatenation, the data is input into a motion feature extraction network to obtain feature representations reflecting scene motion changes. In this embodiment, the motion feature extraction network is implemented using an optical flow feature extraction network, which consists of a series of convolutional layers, pooling layers, deconvolutional layers, and normalization layers.
[0095] After obtaining texture and motion features, the two are fused to obtain the visual representation features of the viewpoint, which can be specifically represented as follows:
[0096] in, Indicates the first Visual representation features of a viewpoint This represents the texture features of the viewpoint. This indicates the motion characteristics of the viewpoint. This represents the visual fusion mapping function. It can be implemented by convolutional mapping, which is used to encode static appearance cues and dynamic motion cues into the same feature space to facilitate subsequent saliency prediction.
[0097] Step 3: Extract audio spectrum features For audio signals synchronized with video, audio spectral features are extracted to characterize the spectral content and semantic information of spatial audio.
[0098] For the The audio segment corresponding to each moment First, a short-time Fourier transform is performed to obtain the time-frequency representation of the audio:
[0099] in, Indicates the first The audio clips are indexed by frequency. and short window index Spectral values at that location This represents the short-time Fourier transform operator. Represents a frequency-dimensional index. This indicates the time window index. This transformation converts the original time-domain audio signal into a time-frequency domain signal to facilitate subsequent extraction of spectral features.
[0100] Furthermore, the time-frequency representation is mapped to the Mel spectrum, specifically as follows:
[0101] in, Indicates the first Mel spectrograms corresponding to each audio segment This represents the Mel filter group mapping operator. The Mel spectrum can better simulate the differences in human hearing at different frequencies, and is therefore more conducive to extracting audio features related to human auditory attention.
[0102] After obtaining the Mel spectrum, it is input into the audio coding network to extract multi-scale audio spectral features, which can be specifically represented as:
[0103] in, Indicates the first Audio spectral features at each scale level Indicates the audio coding network in the 1st... Feature extraction mapping of layers, This indicates the total number of scale layers. In this embodiment, the audio coding network uses a U-Net encoder to obtain multi-scale spectral features that combine local details with high-level semantic information.
[0104] Furthermore, one layer is selected from the multi-scale audio spectral features as the content spectral feature, denoted as:
[0105] in, Indicates content spectral characteristics, This represents a pre-selected hierarchical index. The content spectral features balance good semantic expressiveness and resolution, and are used as audio semantic representations in subsequent multimodal fusion.
[0106] Step 4: Estimate the sound source energy distribution using the AVA module. To establish a fine-grained correspondence between audio cues and visual regions, this embodiment further constructs a sound source energy distribution estimation module. This module consists of multiple cascaded audiovisual attention modules, i.e., AVA modules, such as... Figure 2 As shown. Each AVA module includes a pre-gated processing stage and a contextual attention enhancement stage, which is used to progressively align audio energy with visual scene structure and infer fine-grained sound source energy distribution.
[0107] For the Layered audio spectral features, after pre-gated processing, can be represented as:
[0108] in, Indicates the first The fusion features of the layers after pre-gated processing This represents a mapping function that generates spatial gating weights for visual features. This represents a mapping function that generates spatially gated weights for audio features. This represents element-wise multiplication. Indicates the first Visual representation features from each perspective Indicates the first Layer audio spectrum characteristics.
[0109] The physical meaning of the above formula is as follows: spatial attention weights are generated based on visual and audio features respectively, and these attention weights are used to modulate the original features element by element. The modulated visual and audio features are then added together and fused. In this way, irrelevant background areas can be suppressed at a coarse-grained level, while potential sound-producing areas and sound-related visual areas can be highlighted.
[0110] After pre-gated processing, to further enhance the channel response relevant to salient regions, contextual attention enhancement is applied to the pre-gated fusion features, which can be specifically represented as:
[0111] in, Indicates the first Output features of the layered audio-visual attention unit This represents the pooling function; in this embodiment, the global average pooling function is used. The context attention mapping function, in this embodiment, consists of a convolutional module.
[0112] This step indicates starting with the fusion features. Global context information is extracted, and then the context attention gate mapping function is used to generate channel weights. The channel weights are then applied to the fused features to enhance the channels with stronger discriminative significance.
[0113] To achieve a progressive fusion from low-level local details to high-level semantic information, multiple AVA modules are stacked according to scale hierarchy, and a pyramid-shaped fusion is formed through upsampling and concatenation. This can be specifically represented as follows:
[0114]
[0115] in, This indicates the energy distribution characteristics of the initial layer sound source. Indicates the first The energy distribution characteristics of the sound source after layer fusion. Denotes the upsampling mapping function, Represents the convolution fusion function. This indicates feature splicing.
[0116] The purpose of the aforementioned progressive fusion process is to: upsample the sound source energy information already obtained in the previous scale level, concatenate it with the audiovisual attention features of the current layer, and then perform convolutional mapping to obtain a more refined sound source energy distribution feature for the current layer. Finally, the output of the last layer is used as the sound source energy distribution feature of the current viewpoint, denoted as:
[0117] in, Indicates the first The final sound source energy distribution features obtained from each perspective. These features reflect the potential spatial distribution of sound in the visual scene and are an important input for subsequent saliency fusion.
[0118] Step 5: Perform multimodal feature fusion In obtaining visual representation features Content Spectrum Characteristics and the characteristics of sound source energy distribution Then, multimodal fusion is performed on the three types of features to construct single-viewpoint multimodal saliency features. First, the three types of features are mapped to a unified feature space, and a shared representation is constructed, which can be specifically represented as:
[0119] in, Indicates the first Shared modal representation from multiple perspectives , and These represent the visual feature mapping parameters, content spectrum feature mapping parameters, and sound source energy distribution feature mapping parameters, respectively. This represents a convolutional or linear mapping operation. This step is used to extract common information from the three modalities.
[0120] Then, the shared representation is concatenated with the original features of each modality to form modality-enhanced features, which can be specifically represented as:
[0121] in, Indicates enhanced visual features, This indicates the enhanced content spectral characteristics. This indicates the energy distribution characteristics of the enhanced sound source. , and These represent the enhancement mapping parameters for the corresponding modes.
[0122] Finally, the three enhanced modal features are further concatenated and fused by convolution to obtain the single-viewpoint multimodal saliency features, which can be specifically represented as:
[0123] in, Indicates the first Single-view multimodal saliency features corresponding to each viewpoint This represents the multimodal fusion function. Through the above fusion method, while preserving the unique information of different modalities, the complementary relationship between modalities is also strengthened. The corresponding formula in the paper also first forms a shared representation, then enhances each modality separately, and finally concatenates them to obtain the single-view multimodal features.
[0124] Step 6: Perform cross-viewpoint saliency alignment Since different viewpoints correspond to different observations of the same scene, performing saliency prediction for each viewpoint individually can easily lead to inconsistent results between different viewpoints. Therefore, this embodiment further introduces a cross-viewpoint saliency alignment module, such as... Figure 3 As shown. This module can preferably be implemented using a bidirectional convolutional long short-term memory network to propagate saliency information between adjacent viewpoints.
[0125] Let the single-viewpoint multimodal saliency feature sequence corresponding to all viewpoints be represented as follows: ,in, Represents multi-view feature sequences. Indicates the first Single-view multimodal saliency features. The forward and backward propagation processes can be represented as:
[0126]
[0127] in, Indicates the first Hidden state characteristics of each perspective during forward propagation Indicates the first Hidden state characteristics of each perspective during backpropagation. This represents the forward time propagation mapping function. This represents the backward time propagation mapping function. In this embodiment, the forward and backward propagation mapping functions are constructed using a bidirectional ConvLSTM network.
[0128] After obtaining the forward and backward hidden states, they are concatenated and linearly mapped to obtain the aligned salient features or salient distribution, which can be specifically represented as:
[0129] in, Indicates the first A salient feature or salient response map of each viewpoint after cross-viewpoint alignment. Indicates the fusion mapping parameters, This represents the bias parameter.
[0130] This step allows the current viewpoint to utilize both forward neighboring viewpoint information and backward neighboring viewpoint information simultaneously, thereby making the representation of the same salient target more consistent across different viewpoints.
[0131] Step 7: Generate a significance distribution map and perform training and optimization. Saliency response features after cross-viewpoint alignment The output is mapped to obtain the final significance distribution map, which can be represented as follows:
[0132] in, Indicates the first The final significance distribution map output from each perspective. This represents the activation function; in this embodiment, the Softmax activation function is used.
[0133] Let the first The saliency annotation map corresponding to each viewpoint is as follows: The significance monitoring loss can then be expressed as:
[0134] in, Indicates significant predicted loss. Indicates the height of the saliency plot. Indicates the width of the saliency plot. Represents the pixel coordinates on the saliency map. The predicted significance plot is shown on the coordinates. The output value at that location, This indicates the saliency plot on the coordinate axis. The actual value at that point, This represents the pixel-level loss function; in this embodiment, the cross-entropy loss function is used.
[0135] After the above steps, a saliency distribution map corresponding to each viewpoint is obtained. The saliency distribution map can be directly used for tasks such as salient region detection, resource priority allocation, key region encoding, and transmission scheduling in multi-view video; furthermore, the saliency distribution map can also serve as upstream saliency information input in volumetric video reconstruction, free viewpoint content generation, and immersive media processing.
[0136] To verify the effectiveness of the proposed multi-view video saliency prediction method that integrates spatial audio, an ablation experiment was conducted. The experiment focused on a saliency prediction task, using the mean intersection-union ratio (mIoU) as the primary evaluation metric. A higher mIoU value indicates a greater overlap between the predicted salient region and the actual salient region, signifying better saliency prediction performance of the proposed method.
[0137] Table 1. Significance prediction results
[0138] In the experiment, a baseline model was first constructed, consisting only of texture feature extraction and motion feature extraction. Then, based on this baseline model, a sound source energy distribution estimation module, content spectrum features, and a cross-viewpoint saliency alignment module were gradually introduced to analyze the impact of each component on the overall saliency prediction performance. Finally, all the above modules were integrated to obtain the complete method of this invention, and its saliency prediction effect was tested. The experimental results are shown in Table 1, where √ indicates that the module was used, and × indicates that it was not used.
[0139] Data analysis shows that, compared with the baseline model which only includes texture feature extraction and motion feature extraction modules, its mIoU is 72.04%. Based on this, the mIoU increases to 79.53% after introducing the sound source energy distribution estimation module, an improvement of 7.49%. The mIoU increases to 75.01% after introducing content spectral features, an improvement of 2.97%. The mIoU increases to 76.10% after introducing the cross-viewpoint saliency alignment module, an improvement of 4.06%. When all three modules are introduced simultaneously, the mIoU of the complete method reaches 85.01%, an improvement of 12.97 percentage points compared to the baseline model.
[0140] The above results demonstrate that each component module proposed in this invention can positively impact saliency prediction performance. Among them, the sound source energy distribution estimation module shows the most significant performance improvement, indicating that modeling spatial sound energy distribution can effectively enhance the ability to locate salient regions. Meanwhile, the cross-viewpoint saliency alignment module can improve the consistency of saliency representations between different viewpoints, while content spectral features provide effective audio semantic supplementation for saliency prediction. When used together, these three modules can form a good synergistic gain, thereby significantly improving the accuracy of multi-viewpoint video saliency prediction with fused spatial audio, verifying the effectiveness of the method of this invention.
[0141] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0142] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0143] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the multi-view video saliency prediction methods that fuse spatial audio in the above embodiments.
[0144] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0145] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0146] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0147] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0148] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the saliency of multi-view videos by fusing spatial audio, characterized in that, Includes the following steps: S100: Acquire a multi-view video sequence of the same scene and an audio signal that is time-synchronized with the multi-view video sequence, and establish the sequential relationship between the views. S200. Extract the visual texture features and motion features of video frames from each viewpoint in the multi-view video sequence, and fuse the texture features and motion features to obtain the visual representation features corresponding to each viewpoint. S300. Perform time-frequency transformation and multi-scale feature encoding on the audio signals of the multi-view video sequence that are synchronized in time to obtain audio spectrum features, and select content spectrum features from the audio spectrum features as audio semantic representation; S400. Construct a sound source energy distribution estimation module, jointly model visual representation features and audio spectrum features to obtain sound source energy distribution features corresponding to the current visual scene. S500: Multimodal fusion of visual representation features, content spectrum features and sound source energy distribution features is performed to obtain single-view multimodal saliency features corresponding to each viewpoint; S600. Based on the order relationship between the viewpoints, perform cross-viewpoint saliency alignment on the single-viewpoint multimodal saliency features corresponding to each viewpoint to obtain a cross-viewpoint consistent saliency feature representation. S700. Based on the saliency feature representation after cross-view alignment, generate saliency distribution maps for each view.
2. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The method for establishing the sequential relationship between the various viewpoints in step S100 includes: Let the input multi-view video sequence be denoted as . The audio signal synchronized with the multi-view video sequence is denoted as ; in, This represents the entire multi-view video dataset. This represents the entire set of synchronized audio data. Indicates the total number of viewpoints. Indicates the total number of time frames. Indicates the first The perspective in the first Video frames captured at each moment. Indicates the relationship with the first The audio segment corresponding to each moment; The viewpoint order is determined by the linear order obtained by arranging the cameras by their numbers and the adjacency order constructed by the spatial distribution of the cameras.
3. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The visual representation feature acquisition method in step S200 includes: For the The perspective in the first Video frames at a given moment First, extract the texture features, which are represented as follows: in, Indicates the first The perspective in the first Texture features corresponding to video frames at each time point This represents the texture feature extraction function; To characterize temporal changes and target motion information in the video, motion features are extracted, which are represented as follows: in, Indicates the first The perspective in the first Motion characteristics corresponding to each moment Represents the motion feature extraction function. Indicates the preset time interval, symbol This indicates a splicing operation along the channel dimension; After obtaining texture and motion features, the two are fused to obtain visual representation features. The fusion process is represented as follows: in, Indicates the first Visual representation features corresponding to each viewpoint This represents the texture features from that viewpoint. This indicates the motion characteristics from that perspective. This represents a visual feature fusion function, which consists of convolutional mapping, normalization, and nonlinear activation.
4. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The method for obtaining the content spectral features in step S300 includes: For the The audio segment corresponding to each moment Perform a short-time Fourier transform to obtain the time-frequency representation of the audio: in, Indicates the first The audio clips are indexed by frequency. and short window index Spectral values at that location This represents the short-time Fourier transform operator. Represents a frequency-dimensional index. Indicates the time window index; The original time-domain audio signal is converted into a two-dimensional time-frequency domain representation through short-time Fourier transform, and spectrum modeling is performed. Mapping the time-frequency representation to the Mel spectrum, denoted as: In the formula, Indicates the first Mel spectrograms corresponding to each audio segment This represents the Mel filter group mapping operator; The Mel spectrum is input into the audio coding network for multi-scale feature extraction. The multi-scale audio spectrum features are represented as follows: in, Indicates the first Audio spectral features at each scale level Indicates the audio coding network in the 1st... Feature extraction mapping of layers, Indicates the total number of scale layers; One layer is selected from the multi-scale audio spectral features as the content spectral feature, denoted as: in, Indicates content spectral characteristics, This indicates a pre-selected hierarchical index.
5. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The method for obtaining the sound source energy distribution features corresponding to the current viewpoint visual scene in step S400 includes: Let the first Layered audio spectrum features, after pre-gated processing, are represented as follows: in, Indicates the first The fusion features of the layers after pre-gated processing This represents a mapping function that generates spatial gating weights for visual features. This represents a mapping function that generates spatially gated weights for audio features. This represents element-wise multiplication. Indicates the first Visual representation features from each perspective Indicates the first Layer audio spectrum features; After pre-gating processing, the fusion result is recalibrated through contextual attention enhancement processing. The process is as follows: In the formula, Indicates the first Output features of the layered audio-visual attention unit Represents the pooling function. This indicates a context-sensitive mapping function; The sound source energy distribution estimation module adopts an inter-layer cascading approach to achieve multi-scale progressive fusion from low-level local details to high-level global semantics. The process is represented as follows: in, This indicates the energy distribution characteristics of the initial layer sound source. Indicates the first The energy distribution characteristics of the sound source after layer fusion. Represents the upsampling mapping function. Represents the convolution fusion function. This indicates feature splicing.
6. The multi-view video saliency prediction method fused with spatial audio according to claim 5, characterized in that, The methods for obtaining single-view multimodal salient features corresponding to each viewpoint in step S500 include: The three modal features are mapped to a unified feature space to form a shared representation, which is: in, Indicates the first Shared modal representation from multiple perspectives , and These represent the visual feature mapping parameters, content spectrum feature mapping parameters, and sound source energy distribution feature mapping parameters, respectively. This represents a convolutional mapping or linear mapping operation. Indicates the first Visual representation features from each perspective Indicates content spectral characteristics, Indicates the first Sound source energy distribution characteristics from multiple perspectives; In obtaining shared representation Then, the shared representation is concatenated with the original features of each modality to obtain the enhanced modal features, represented as: in, Indicates enhanced visual features, This indicates the enhanced content spectral characteristics. This indicates the energy distribution characteristics of the enhanced sound source. , and These represent the enhancement mapping parameters for the corresponding modes; Finally, the three enhanced modal features are further fused to obtain single-view multimodal saliency features, the process of which is represented as follows: in, Indicates the first Single-view multimodal saliency features corresponding to each viewpoint This represents a multimodal fusion function.
7. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The method for obtaining the cross-view consistent saliency feature representation in step S600 includes: Let the single-view multimodal saliency features of all perspectives constitute the feature sequence. ,in, Represents multi-view feature sequences. Indicates the first Single-view multimodal saliency features from multiple perspectives; To propagate saliency information between adjacent viewpoints, the forward and backward feature propagation processes are represented as follows: in, Indicates the first Hidden state characteristics of each perspective during forward propagation Indicates the first Hidden state characteristics of each perspective during backpropagation. This represents the forward time propagation mapping function. Represents the backward time propagation mapping function; After obtaining the forward hidden state and backward hidden state Then, the two are fused to obtain the cross-view aligned feature representation, specifically as follows: in, Indicates the first A salient feature or salient response map of each viewpoint after cross-viewpoint alignment. Indicates the fusion mapping parameters, This represents the bias parameter.
8. The multi-view video saliency prediction method fused with spatial audio according to claim 1, characterized in that, The method for optimizing and training the entire network through supervised learning in step S700 includes: Saliency features after cross-view alignment Activation mapping is performed to obtain the final predicted significance distribution map, represented as: in, Indicates the first The final significance distribution map output from each perspective. Indicates the activation function; Let the first The saliency annotation map corresponding to each viewpoint is as follows: The significance of the supervised loss is expressed as: in, Indicates significant predicted loss. Indicates the height of the saliency plot. Indicates the width of the saliency plot. Represents the pixel coordinates on the saliency map. The prediction significance plot is shown on the coordinate axis. The output value at that location, This indicates the saliency plot on the coordinate axis. The actual value at that point, This represents the pixel-level loss function.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the method as described in any one of claims 1 to 8.
10. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 8.