Omnidirectional image quality evaluation method and system based on significance guidance

Through the significance-guided method, the global and local saliency features were extracted in combination with SalBiNet360 and Mamba models, and the problem of taking into account both global and local features in omnidirectional image quality evaluation was solved, achieving higher evaluation accuracy and user experience optimization.

CN120355660APending Publication Date: 2025-07-22CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407491.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When processing omnidirectional image quality evaluation methods, it is difficult to take into account both global and local features, resulting in insufficient evaluation accuracy. In addition, traditional models are prone to loss of features when processing large data volumes, affecting the accuracy of image quality evaluation.

Method used

Using a significance-based guided method, global and local significance were extracted through the SalBiNet360 network, long-distance semantic association was captured in combination with the Mamba model, global feature maps were generated, and significance feature extraction was performed using the Swin-Transformer+VGG model, and finally the quality evaluation score was calculated through the linear regression model.

Benefits of technology

It improves the accuracy of omnidirectional image quality evaluation, can more comprehensively grasp the semantics of image significance, optimize the image processing process, and improve the performance and user experience of VR devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355660A_ABST
    Figure CN120355660A_ABST
Patent Text Reader

Abstract

The invention relates to an omni-directional image quality evaluation method and system based on significance guidance. The method comprises the following steps: acquiring an omni-directional image to be evaluated and preprocessing the omni-directional image; extracting global saliency and local saliency of the omni-directional image through a saliency extraction model, and fusing the global saliency and the local saliency into a saliency map of the omni-directional image; capturing long-distance semantic association among different areas in the omnidirectional image through a feature extraction model to obtain a global feature map of the omnidirectional image; inputting the saliency map of the omni-directional image into a saliency guiding model to extract saliency features of the omni-directional image; splicing the saliency feature of the omnidirectional image and the global feature map of the omnidirectional image to obtain a saliency enhancement feature of the omnidirectional image; and inputting the significantly enhanced features of the omnidirectional image into a linear regression model to calculate a quality evaluation score of the omnidirectional image. According to the method, the saliency map is input into the saliency guiding model to extract the saliency features, the saliency features are spliced with the global feature map to obtain the saliency enhancement features, then the saliency enhancement features are input into the linear regression model to calculate the quality evaluation score, and the whole process optimizes omnidirectional image quality evaluation from multiple links. The method is of great significance in optimizing the image processing process, improving the performance of the VR equipment, improving the user experience and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an omnidirectional image quality evaluation method and system based on saliency guidance. Background Art

[0002] In recent years, VR and AR technologies have developed rapidly. As a special three-dimensional information carrier, omnidirectional images can provide a 360-degree panoramic viewing experience, making users feel as if they are in a real scene. They are widely used in many fields such as education, medical treatment, entertainment, military, and vehicle driving, showing great potential. For example, Tesla uses this technology in its autopilot assistance system to provide drivers with more comprehensive road condition information. However, omnidirectional images are prone to distortion, blurring, artifacts and other distortion phenomena during the processes of acquisition, compression, transmission and rendering, which affect the user experience and even cause physiological discomforts such as dizziness.

[0003] Existing objective omnidirectional image quality evaluation methods are divided into full reference (FR) and no reference (NR). FR-type OIQA uses peak signal-to-noise ratio / structural similarity and traditional machine learning methods; NR-type OIQA is further divided into two categories based on different projection spaces and based on viewports. NR-type OIQA is favored by scholars because it does not require a reference image and has the advantages of low time and economic costs, strong stability and high practicability. However, although many current NR OIQA metrics show high performance on various datasets, there is still room for improvement. The models based on saliency guidance only consider the local viewport area of the image when extracting saliency semantics, ignoring the role of the overall saliency of the image in visual guidance, resulting in limited performance. Facing the large data volume of omnidirectional images, many models need to reduce the dimension first when processing long sequence features, which will cause feature loss and affect the accuracy of image quality evaluation. In addition, when extracting saliency semantics based on traditional Transformer or graph convolution, it is difficult to simultaneously consider the hybrid extraction of global and local features, and it is impossible to comprehensively and accurately grasp the saliency semantics of the image. Summary of the Invention

[0004] In order to solve the defects existing in the background art, one aspect of the present invention provides an omnidirectional image quality evaluation method based on saliency guidance, including:

[0005] S1: Obtain the omnidirectional image to be evaluated and perform preprocessing;

[0006] S2: Extract the global saliency and local saliency of the omnidirectional image through a saliency extraction model, and fuse them into a saliency map of the omnidirectional image;

[0007] S3: Capture the long-distance semantic associations between different regions in the omnidirectional image through a feature extraction model to obtain a global feature map of the omnidirectional image;

[0008] S4: Input the saliency map of the omnidirectional image into the saliency-guided model to extract the saliency features of the omnidirectional image;

[0009] S5: Concatenate the saliency features of the omnidirectional image and the global feature map of the omnidirectional image to obtain the significantly enhanced features of the omnidirectional image;

[0010] S6: Input the significantly enhanced features of the omnidirectional image into the linear regression model to calculate the quality evaluation score of the omnidirectional image.

[0011] Another aspect of the present invention provides an omnidirectional image quality evaluation system based on saliency guidance. The system includes a memory and a processor; the memory is used to store application programs; the processor is used to run the application programs and execute the above-mentioned omnidirectional image quality evaluation method based on saliency guidance.

[0012] Yet another aspect of the present invention provides a computer storage medium. A computer program is stored on the computer storage medium, and when the computer program is executed by a processor, the above-mentioned omnidirectional image quality evaluation method based on saliency guidance is implemented.

[0013] The present invention has at least the following beneficial effects

[0014] The omnidirectional image quality evaluation method based on saliency guidance proposed by the present invention has many advantages. In terms of saliency extraction, the global saliency and local saliency of the omnidirectional image are simultaneously extracted by the model and fused into a saliency map, overcoming the limitation of only considering the local viewport area in the past and being able to grasp the image saliency more comprehensively. In terms of feature extraction, a specific model is used to capture the long-distance semantic associations between different regions of the omnidirectional image to obtain the global feature map, avoiding feature loss caused by dimensionality reduction and improving the evaluation accuracy. In the saliency semantic extraction, this method can take into account the mixed extraction of global and local features and more accurately grasp the saliency semantics of the image. Finally, by inputting the saliency map into the saliency-guided model to extract saliency features, concatenating them with the global feature map to obtain significantly enhanced features, and then inputting them into the linear regression model to calculate the quality evaluation score, the entire process optimizes the omnidirectional image quality evaluation from multiple links, which is of great significance for optimizing the image processing process, improving the performance of VR devices, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the overall framework of the present invention;

[0016] Figure 2 It is a schematic diagram of the framework of the parallelized channel and spatial attention module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0018] Please refer to Figure 1 , the present invention provides an omnidirectional image quality evaluation method based on saliency guidance, including:

[0019] S1: Obtain the omnidirectional image to be evaluated and perform preprocessing;

[0020] In this embodiment, the preprocessing of the omnidirectional image to be evaluated includes: grayscale conversion, denoising, geometric correction, brightness and contrast adjustment, etc.

[0021] S2: Extract the global saliency and local saliency of the omnidirectional image through a saliency extraction model, and fuse them into a saliency map of the omnidirectional image;

[0022] Preferably, the saliency extraction model includes: SalBiNet360 network;

[0023] In this embodiment, after the omnidirectional image is input into the SalBiNet360 network, the front-end part of the network first performs a series of feature extraction operations on the image. It constructs a feature pyramid structure through multiple convolutional layers, and convolutional layers at different levels can capture image features at different scales. For example, shallower convolutional layers can capture details such as edges and textures in the image, while as the network depth increases, more global and abstract features can be gradually extracted. In this process, the SalBiNet360 network uses its unique 360-degree perspective processing mechanism to comprehensively analyze each part of the omnidirectional image. It will comprehensively consider information in all directions of the image and grasp the distribution of significant features of the image as a whole. Taking an omnidirectional image of an outdoor park as an example, by analyzing the entire image, the network can identify the square area in the park where people gather. Due to the activities and relatively concentrated situation of the people, this area has a high significance globally. The SalBiNet360 network will assign a high significance value to this square area in the corresponding global significance output, thereby generating a global significance map reflecting the global significant features of the omnidirectional image. For the extraction of local significance, the SalBiNet360 network divides the omnidirectional image into multiple local regions for separate processing. It adopts a sliding window method, moving a window of a certain size block by block on the image. Assuming the window size is set to 64×64 pixels, the network will sequentially input each local image block within the window into a specific sub-network structure. This sub-network structure also includes components such as convolutional layers and pooling layers, which are specifically used to analyze the features of local image blocks. When processing each local image block, the SalBiNet360 network focuses on the detail information within the local region. For example, in a local region of the omnidirectional image of the park, there may be a blooming flower. Compared with the surrounding grass and trees, this flower has unique colors and shapes. The SalBiNet360 network can keenly capture these significant features of the flower when processing this local image block and assign a high significance value to the position where the flower is located in the corresponding local significance output, thereby generating local significance maps for each local region. For the extraction of local significance, the SalBiNet360 network divides the omnidirectional image into multiple local regions for separate processing. It adopts a sliding window method, moving a window of a certain size block by block on the image. Assuming the window size is set to 64×64 pixels, the network will sequentially input each local image block within the window into a specific sub-network structure. This sub-network structure also includes components such as convolutional layers and pooling layers, which are specifically used to analyze the features of local image blocks. When processing each local image block, the SalBiNet360 network focuses on the detail information within the local region. For example, in a local region of the omnidirectional image of the park, there may be a blooming flower. Compared with the surrounding grass and trees, this flower has unique colors and shapes.When the SalBiNet360 network processes this local image patch, it can keenly capture these prominent features of the flowers and assign a relatively high saliency value to the position where the flowers are located in the corresponding local saliency output, thereby generating local saliency maps for each local region. The SalBiNet360 network adopts an adaptive fusion strategy to fuse the global saliency map and multiple local saliency maps. It will dynamically adjust the weights of the global and local saliency information according to the content and features of the image. For some regions with obvious global dominant features, such as large lakes in the park, the global saliency information will occupy a larger weight in the fusion process, making the lake maintain a prominent saliency position in the final saliency map. For those regions with rich local details and important significance, such as unique sculptures beside the park path, the weight of the local saliency information will be correspondingly increased to ensure that the sculpture can clearly show its saliency in the saliency map. The specific fusion process is that for each pixel, the network will perform a weighted sum of its value in the global saliency map and its value in the local saliency map at the corresponding position according to the pre-calculated weights. In this way, the generated final saliency map can not only reflect the overall saliency features of the omnidirectional image but also highlight the key salient objects in each local region, providing more comprehensive and accurate saliency information for subsequent image processing tasks based on the saliency map.

[0024] S3: Capture the long-range semantic associations between different regions in the omnidirectional image through a feature extraction model to obtain the global feature map of the omnidirectional image;

[0025] Preferably, the feature extraction model includes: the Mamba model;

[0026] In this embodiment, the Mamba model internally contains a series of processing modules, and some of these modules are specifically responsible for feature encoding of different regions of the image. It divides the omnidirectional image into multiple local regions, similar to splitting a jigsaw puzzle into many small pieces. For each small patch region, the model extracts its local features through convolution operations and other means, such as the shape and color features of a vehicle in a certain small patch region, or the pose features of pedestrians.

[0027] After completing the local feature extraction, the Mamba model utilizes its long-range dependence modeling ability to start capturing the long-range semantic associations between different regions. For example, the model can identify the relationship between tall buildings in the distance and vehicles driving on the street. Although their positions in the image are far apart, the urban built environment represented by the tall buildings and the urban traffic elements represented by the vehicles have semantic connections, that is, the vehicles drive on the roads in the city, and the roads rely on the environment constructed by the urban buildings. The Mamba model can pay attention to these elements at different positions in the image through its special attention mechanism and integrate the association information between them.

[0028] For another example, there is also a semantic association between the street-side stores and pedestrians. The pedestrians may be customers of the stores. The Mamba model can capture this relationship even when the stores and pedestrians are far apart in the image. By capturing and analyzing the rich long-range semantic associations between various regions in the omnidirectional image, the Mamba model incorporates this associated information into the overall understanding of the image.

[0029] Finally, based on the extraction of local features from different regions of the omnidirectional image and the capture of long-range semantic associations, the Mamba model generates a global feature map. In this global feature map, the semantic relationships between the various elements in the image are clearly reflected. Elements such as high-rise buildings in the distance, moving vehicles, pedestrians, and street-side stores are no longer isolated but are organically integrated through their semantic associations. This global feature map comprehensively reflects the overall characteristics of the omnidirectional image of the urban street scene, providing a key information basis for subsequent in-depth analysis and processing of the omnidirectional image.

[0030] S4: Input the saliency map of the omnidirectional image into the saliency-guided model to extract the saliency features of the omnidirectional image;

[0031] Preferably, the saliency-guided model includes: a convolutional layer, a max-pooling layer, three global-local feature extraction layers of Swin-Transformer + VGG, three parallel channel and spatial attention modules, and an average-pooling layer; among them, the convolutional layer, the max-pooling layer, and the three

[0032] global-local feature extraction layers of Swin-Transformer + VGG are cascaded in sequence; the output features of the three parallel channel and spatial attention modules are respectively input into the three parallel channel and spatial attention modules; the output features of the three parallel channel and spatial attention modules are concatenated and then input into the average-pooling layer for processing to obtain the saliency features of the omnidirectional image.

[0033] Preferably, the global-local feature extraction layer of Swin-Transformer + VGG includes: a Swin-Transformer layer and a VGG convolutional layer.

[0034] Please refer to Figure 2 , preferably, the parallel channel and spatial attention module includes: a max-pooling module, an average-pooling module, a multi-layer perceptron, and a sigmoid activation function;

[0035] Define the input feature of the parallel channel and spatial attention module as F; input the input feature F into the max-pooling module and the average-pooling module respectively for processing to obtain feature F1 and feature F2;

[0036] Feature F1 and feature F2 are added to obtain feature F12; feature F1 and feature F2 are concatenated to obtain feature F21;

[0037] Feature F12 is input into a multi-layer perceptron for processing to obtain feature F121; feature F121 is input into a sigmoid activation function to obtain feature F122;

[0038] Feature F21 is input into a sigmoid activation function to obtain feature F211;

[0039] Feature F122 and feature F211 are weighted and integrated to obtain feature F3; feature F3 is multiplied by feature F to obtain the output feature of the channel and spatial attention module with feature parallelization.

[0040] In this embodiment, it is assumed that we have an omnidirectional image showing a beautiful forest with a winding path and an ancient cabin beside the path. After obtaining the saliency map through step S2, this saliency map is input into a saliency-guided model composed of a convolutional layer, a max-pooling layer, 3 global-local feature extraction layers of Swin-Transformer + VGG, 3 parallel channel and spatial attention modules, and an average-pooling layer; First, the saliency map enters the convolutional layer. The convolutional kernels in the convolutional layer slide on the saliency map to perform convolutional operations and extract basic features such as edges and textures in the saliency map. For example, the basic features of significant regions such as the contour edges of trees in the forest and the unique structural edges of the cabin are identified. Then, these features extracted by convolution enter the max-pooling layer. The max-pooling layer selects the maximum feature value within a small area, which is equivalent to downsampling the features, retaining the most prominent features while reducing the data volume. For example, in an area containing tree branches and leaves, the max-pooling layer will select the maximum value that best represents the significant features of the area and ignore some relatively less significant details.

[0041] Subsequently, the features enter the global-local feature extraction layers of 3 Swin-Transformer + VGG. Taking the first global-local feature extraction layer of Swin-Transformer + VGG as an example, the Swin-Transformer layer starts to work. It divides the saliency map into multiple windows and performs self-attention calculations within each window to capture the relationships between features within the window. For example, in the window containing the log cabin, it can discover the spatial relationships between parts such as the doors, windows, and roof of the log cabin, as well as their associations with the surrounding forest environment. After that, the features enter the VGG convolutional layer, which further performs convolutional operations on the features to extract more refined local features, such as detailed features like the wood grain texture on the surface of the log cabin. The following two global-local feature extraction layers of Swin-Transformer + VGG also extract and integrate the features of the saliency map from different scales and angles in a similar manner.

[0042] Next, the features from the three global-local feature extraction layers of 3 Swin-Transformer + VGG enter the output features corresponding to three parallel channel and spatial attention modules. Taking one of the parallel channel and spatial attention modules as an example, the input feature F is respectively input into the max-pooling module and the average-pooling module. Suppose the feature F contains information about various aspects such as the forest, trail, and log cabin. After being processed by the max-pooling module, the feature F1 is obtained, which highlights the most prominent parts of the feature, such as the prominent position information of the log cabin in the picture. After being processed by the average-pooling module, the feature F2 is obtained, which more comprehensively reflects the average situation of the feature, such as the overall feature intensity of the forest area. Adding the feature F1 and F2 gives the feature F12, which makes the prominent feature and the average feature complement each other and strengthens the expression of the overall feature. Concatenating the feature F1 and F2 gives the feature F21, which can retain more dimensional information. The feature F12 enters the multi-layer perceptron, which performs a non-linear transformation on it to further explore the potential relationships between features, obtaining the feature F121, and then passing through the sigmoid activation function to get the feature F122. The sigmoid activation function maps the feature values to between 0 and 1, highlighting the importance of the features. The feature F21 passes through the sigmoid activation function to get the feature F211. Finally, the feature F122 and F211 are weighted and integrated to get the feature F3. Here, the weighted integration assigns weights according to the importance of the features. For example, if the features of the log cabin are considered more important in the current scene, then the weights of the features related to the log cabin will be larger. Multiplying the feature F3 and the feature F gives the output feature of this parallel channel and spatial attention module, which fuses the original feature F and the features processed by the attention mechanism, highlighting the prominent feature part. The other two parallel channel and spatial attention modules also process the features in the same way.

[0043] Finally, the output features of the 3 parallel channels and the spatial attention module are concatenated and then input into the average pooling layer. The average pooling layer calculates the average of the concatenated features to obtain a comprehensive feature vector, which is the saliency feature of the omnidirectional image. It synthesizes the saliency feature information of each region, each scale, and different attention mechanisms in the saliency map, providing a key feature basis for subsequent analysis and quality evaluation of the omnidirectional image.

[0044] S5: Concatenate the saliency feature of the omnidirectional image and the global feature map of the omnidirectional image to obtain the significantly enhanced feature of the omnidirectional image;

[0045] S6: Input the significantly enhanced feature of the omnidirectional image into the linear regression model to calculate the quality evaluation score of the omnidirectional image.

[0046] In this embodiment, after concatenation in S5, the significantly enhanced feature fuses saliency and global information, highlighting both the details of key regions and the overall structural associations of the image. Entering S6, the linear regression model, based on these features, comprehensively considers all aspects of the image's performance and accurately calculates the quality evaluation score, providing a quantitative basis for image quality assessment.

[0047] Preferably, the linear regression model includes a fully connected layer.

[0048] In this embodiment, the training method follows the standard training strategy of existing omnidirectional image quality evaluation algorithms. The model is trained using the Adam optimizer with a weight decay value of 1×10 -5 for 300 epochs, with each training batch size set to 4 and the initial learning rate set to 5×10 -5 ; Based on the above scheme, in order to verify the effectiveness of the method of the present invention and the actual results, in this embodiment, the method of the present invention and other omnidirectional image quality evaluation methods are applied in the omnidirectional image datasets CVIQD and OIQA. The specific comparison results are shown in Table 1. The experimental indicators are compared using PLCC and SRCC. The closer the index coefficient is to 1, the greater the correlation between the model evaluation result and the true label. Compared with some other methods, the test results of the present invention show good performance and can objectively evaluate the quality of omnidirectional images.

[0049] Table 1 Test results of different methods on OIQA and CVIQD datasets

[0050]

[0051] As shown in Table (1), compared with the full-reference (FR) model S-SSIM, the present invention has an improvement of approximately 6.2% in the SRCC index and approximately 2.4% in the PLCC index on the CVIQD dataset; an improvement of approximately 4.1% in the SRCC index and approximately 0.2% in the PLCC index on the OIQA dataset. Compared with the no-reference (NR) Assesor360 model, there is an improvement of approximately 0.4% in the SRCC index on the CVIQD dataset, and improvements of approximately 1.0% and approximately 0.2% respectively in the SRCC and PLCC indexes on the OIQA dataset. Generally speaking, the TransVGG model has varying degrees of improvement compared to other models in multiple indexes, demonstrating its performance advantages.

[0052] Another aspect of the present invention provides an omnidirectional image quality evaluation system based on saliency guidance, the system comprising a memory and a processor; the memory is used for storing application programs; the processor is used for running the application programs to execute the omnidirectional image quality evaluation method based on saliency guidance as described above.

[0053] Another aspect of the present invention provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the omnidirectional image quality evaluation method based on saliency guidance as described above.

[0054] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0055] In summary, the omnidirectional image quality evaluation method based on saliency guidance proposed by the present invention has many advantages. In terms of saliency extraction, the global saliency and local saliency of the omnidirectional image are simultaneously extracted by the model and fused into a saliency map, overcoming the limitation of only considering the local viewport area in the past and being able to comprehensively grasp the image saliency. In terms of feature extraction, a specific model is used to capture the long-distance semantic associations between different regions of the omnidirectional image to obtain a global feature map, avoiding feature loss caused by dimensionality reduction and improving the evaluation accuracy. In the saliency semantics extraction, this method can take into account the hybrid extraction of global and local features and more accurately grasp the saliency semantics of the image. Finally, by inputting the saliency map into the saliency guidance model to extract saliency features, splicing them with the global feature map to obtain significantly enhanced features, and then inputting them into a linear regression model to calculate the quality evaluation score, the entire process optimizes the omnidirectional image quality evaluation from multiple links, which is of great significance for optimizing the image processing process, improving the performance of VR devices, and enhancing the user experience.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An omnidirectional image quality evaluation method based on saliency guidance, characterized in that, Including: S1: Obtain the omnidirectional image to be evaluated and perform preprocessing; S2: Extract the global saliency and local saliency of the omnidirectional image through the saliency extraction model, and fuse them into the saliency map of the omnidirectional image; S3: Capture the long-range semantic associations between different regions in the omnidirectional image through the feature extraction model to obtain the global feature map of the omnidirectional image; S4: Input the saliency map of the omnidirectional image into the saliency guidance model to extract the saliency features of the omnidirectional image; S5: Concatenate the saliency features of the omnidirectional image and the global feature map of the omnidirectional image to obtain the significantly enhanced features of the omnidirectional image; S6: Input the significantly enhanced features of the omnidirectional image into the linear regression model to calculate the quality evaluation score of the omnidirectional image.

2. The omnidirectional image quality evaluation method based on saliency guidance according to claim 1, wherein The saliency extraction model includes: SalBiNet360 network; the feature extraction model includes: Mamba model.

3. A method for omnidirectional image quality evaluation based on saliency guidance according to claim 1, wherein, The linear regression model includes a fully connected layer.

4. The omnidirectional image quality evaluation method based on saliency guidance according to claim 1, wherein The saliency guidance model includes: a convolutional layer, a max pooling layer, 3 global-local feature extraction layers of Swin-Transformer + VGG, 3 parallel channel and spatial attention modules, and an average pooling layer; among them, the convolutional layer, the max pooling layer, and the 3 global-local feature extraction layers of Swin-Transformer + VGG are cascaded in sequence; the output features of the 3 parallel channel and spatial attention modules are respectively input into the 3 parallel channel and spatial attention modules; the output features of the 3 parallel channel and spatial attention modules are concatenated and then input into the average pooling layer for processing to obtain the saliency features of the omnidirectional image.

5. The omnidirectional image quality evaluation method based on saliency guidance according to claim 4, characterized in that The global-local feature extraction layer of Swin-Transformer + VGG includes: a Swin-Transformer layer and a VGG convolutional layer.

6. The omnidirectional image quality evaluation method based on saliency guidance according to claim 4, characterized in that, The parallel channel and spatial attention module includes: a max pooling module, an average pooling module, a multi-layer perceptron, and a sigmoid activation function; Define the input feature of the parallel channel and spatial attention module as F; input the input feature F into the max pooling module and the average pooling module respectively for processing to obtain feature F1 and feature F2; Add feature F1 and feature F2 to obtain feature F12; concatenate feature F1 and feature F2 to obtain feature F21; Input feature F12 into the multi-layer perceptron for processing to obtain feature F121; input feature F121 into the sigmoid activation function to obtain feature F122; Input feature F21 into the sigmoid activation function to obtain feature F211; Weightedly integrate feature F122 and feature F211 to obtain feature F3; multiply feature F3 and feature F to obtain the output feature of the parallel channel and spatial attention module.

7. An omnidirectional image quality evaluation system based on saliency guidance, characterized in that, The system includes a memory and a processor; the memory is used to store the application program; the processor is used to run the application program and execute a method for evaluating the quality of an omnidirectional image based on saliency guidance according to any one of claims 1 to 6.

8. A computer storage medium, characterized in that, A remote monitoring program is stored on the computer storage medium. When the remote monitoring program is executed by a processor, it implements an omnidirectional image quality evaluation method based on saliency guidance as described in any one of claims 1 to 6.