Panoramic image quality evaluation method based on semantic feature and visual feature fusion

By integrating semantic features and visual features in the panoramic image quality evaluation method, and using the big model and feature fusion module, the problem that existing methods fail to effectively combine semantic information is solved, and a higher precision panoramic image quality evaluation is achieved.

CN120070448AInactive Publication Date: 2025-05-30LIAONING BEIDOU SATELLITE NAVIGATION PLATFORM CO LTD +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510549967.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing panoramic image quality evaluation methods focus on the mining and analysis of visual features, and fail to effectively combine semantic information, resulting in a large room for performance improvement and the application needs of multimodal information interaction.

Method used

A panoramic image quality evaluation method based on the fusion of semantic features and visual features is adopted, text semantic information is obtained through a big model, combined with panoramic image visual features, and feature fusion is used to fusion of features to generate quality evaluation scores.

Benefits of technology

It realizes effective interaction between low-level semantic information and high-level semantic information, improves the overall accuracy of the model, and can more accurately evaluate the panoramic image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070448A_ABST
    Figure CN120070448A_ABST
Patent Text Reader

Abstract

The invention relates to the field of panoramic image quality evaluation, in particular to a panoramic image quality evaluation method based on semantic feature and visual feature fusion. The method comprises the following steps: acquiring a target distorted panorama, inputting the target distorted panorama into a text generation model, and outputting a text description statement of the target distorted panorama; then through an imagebind model, image semantic features and text semantic features which are in semantic alignment are obtained; inputting the features into a semantic fusion module for semantic fusion to obtain semantic features; performing viewport division processing on the target distortion panorama to obtain a plurality of viewport images, and performing visual feature extraction to obtain visual perception features; and carrying out feature fusion on the semantic features and the visual perception features to obtain a quality evaluation score of the target distorted panoramic image for carrying out quality evaluation on the target distorted panoramic image. In this way, the quality of the distorted panorama can be evaluated by combining the text semantic information and the visual features of the panorama, and the overall performance of the model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of panoramic image quality evaluation, and more specifically, to a panoramic image quality evaluation method based on the fusion of semantic features and visual features. Background Art

[0002] With the continuous maturity and development of information technology, panoramic images with a large field of view have attracted much attention and become one of the research hotspots in fields such as industry, education, and gaming. Different from traditional planar images, panoramic images cover a 360×180 viewing range, providing users with a full-view visual experience. Similar to traditional planar images, panoramic images may introduce distortions at each stage of image processing such as acquisition, storage, and compression, resulting in impaired visual perception quality and affecting the user experience. Although researchers have conducted in-depth research on image quality evaluation in recent years, the research on panoramic image quality evaluation is still in its infancy.

[0003] Currently, researchers usually use planar image quality evaluation methods to evaluate the quality of panoramic images in the early stage, or analyze them in combination with the large-view visual characteristics of panoramic images, with poor performance. With the support of deep learning technology, considering the important role of viewport images in the perception of large-scale panoramic images, researchers have proposed relevant panoramic image quality evaluation methods in combination with the viewport characteristics of panoramic images, and the performance has been improved.

[0004] Although the existing technology has a certain performance improvement compared with traditional evaluation methods, due to the lack of in-depth semantic understanding, its performance still has room for further improvement. The current panoramic image quality evaluation methods basically focus on the mining and analysis of visual features, do not consider the effective combination with other modal information for quality evaluation, and cannot meet the application requirements. In summary, in the panoramic image quality evaluation method, in the face of the continuous progress of large model technology, the influence and interaction relationship of multi-modal factors should be considered to construct a practical panoramic image quality evaluation method. Summary of the Invention

[0005] According to the present invention, a panoramic image quality evaluation scheme based on the fusion of semantic features and visual features is provided. This scheme uses a large model to obtain text semantic information and combines it with the visual features of panoramic images to effectively improve the overall performance of the model.

[0006] In a first aspect of the present invention, a panoramic image quality evaluation method based on the fusion of semantic features and visual features is provided. The method includes: Obtain a target distorted panoramic image; Input the target distorted panoramic image into the text generation model to output the text description statement of the target distorted panoramic image; input the target distorted panoramic image and the text description statement of the target distorted panoramic image into the imagebind model to obtain semantically aligned image semantic features and text semantic features; input the semantically aligned image semantic features and text semantic features into the semantic fusion module for semantic fusion to obtain semantic features; and Perform a split viewport process on the target distorted panoramic image to obtain a number of viewport images, and then perform visual feature extraction to obtain visual perception features; Fuse the semantic features and visual perception features to obtain the quality evaluation score of the target distorted panoramic image; the quality evaluation score is used to evaluate the quality of the target distorted panoramic image.

[0007] Further, the inputting the semantically aligned image semantic features and text semantic features into the semantic fusion module for semantic fusion to obtain semantic features includes: Input the semantically aligned image semantic features into the first softmax module for reweighting, and input them into the first multi-layer perceptron and the second multi-layer perceptron; Multiply the output of the first multi-layer perceptron by the semantically aligned text semantic features to obtain the first semantic fusion feature; Add the first semantic fusion feature to the output of the second multi-layer perceptron to obtain the second semantic fusion feature; Input the second semantic fusion feature into the first ReLu activation function and the first convolutional block in sequence, and output the semantic features.

[0008] Further, the visual feature extraction to obtain visual perception features includes: Input all the viewport images into the visual perception feature extraction module to obtain the local visual features of each viewport; Stitch all the local visual features together to obtain the visual perception features.

[0009] Further, the visual perception feature extraction module sequentially includes a spatial transformation embedding module, a first visual state space block, a first downsampling module, a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module, and a fourth visual state space block; Among them, the spatial transformation embedding module is used to obtain the position perception information of the viewport image; the local visual features output by the fourth visual state space block are used as the output of the visual perception feature extraction module.

[0010] Further, the first visual state space block sequentially includes a first LN layer, a first 2D selective scan block, a second LN layer, and a first FFN layer; The first LN layer obtains the output of the spatial transformation embedding module, performs normalization, and then outputs it to the first 2D selective scanning block for feature extraction, outputting the first extraction result; The output of the spatial transformation embedding module and the first extraction result are added together to obtain the first addition result; The first addition result is sequentially input into the second LN layer and the first FFN layer, and the first enhancement result is output; The first addition result and the first enhancement result are added together to obtain the output of the first visual state space block.

[0011] Further, the second visual state space block sequentially includes a third LN layer, a second 2D selective scanning block, a fourth LN layer, and a second FFN layer; The third LN layer obtains the output of the first downsampling module, performs normalization, and then outputs it to the second 2D selective scanning block for feature extraction, outputting the second extraction result; The output of the first downsampling module and the second extraction result are added together to obtain the second addition result; The second addition result is sequentially input into the fourth LN layer and the second FFN layer, and the second enhancement result is output; The second addition result and the second enhancement result are added together to obtain the output of the second visual state space block.

[0012] Further, the third visual state space block sequentially includes a fifth LN layer, a third 2D selective scanning block, a sixth LN layer, and a third FFN layer; The fifth LN layer obtains the output of the second downsampling module, performs normalization, and then outputs it to the third 2D selective scanning block for feature extraction, outputting the third extraction result; The output of the second downsampling module and the third extraction result are added together to obtain the third addition result; The third addition result is sequentially input into the sixth LN layer and the third FFN layer, and the third enhancement result is output; The third addition result and the third enhancement result are added together to obtain the output of the third visual state space block.

[0013] Further, the fourth visual state space block sequentially includes a seventh LN layer, a fourth 2D selective scanning block, an eighth LN layer, and a fourth FFN layer; The seventh LN layer obtains the output of the third downsampling module, performs normalization, and then outputs it to the fourth 2D selective scanning block for feature extraction, outputting the fourth extraction result; The output of the third downsampling module and the fourth extraction result are added together to obtain the fourth addition result; Input the fourth addition result into the eighth LN layer and the fourth FFN layer in sequence to output the fourth enhancement result; Add the fourth addition result and the fourth enhancement result to obtain the output of the fourth visual state space block.

[0014] Further, the feature fusion of the semantic feature and the visual perception feature to obtain the quality evaluation score of the target distorted panoramic image includes: Input the semantic feature into the first fully connected layer to obtain the first output result; and Input the semantic feature into the second fully connected layer to obtain the second output result; and Input the visual perception feature into the third fully connected layer to obtain the third output result; Perform self-attention fusion on the first output result, the second output result, and the third output result to obtain the fusion result; Add the fusion result and the third output result, and sequentially send the addition result to the ninth LN layer and the fifth FFN layer. Send the output of the fifth FFN layer to the fourth fully connected layer to output the quality evaluation score of the target distorted panoramic image.

[0015] In the second aspect of the present invention, an electronic device is provided. The electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the first aspect of the present invention.

[0016] Compared with the prior art, the present invention has the following beneficial technical effects: Through the fusion of semantic features and visual perception features, the effective interaction between low-level semantic information and high-level semantic information is realized, which can improve the overall accuracy of the model.

[0017] It should be understood that the content described in the invention content part is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Combined with the drawings and referring to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more obvious. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 Shows a flowchart of a panoramic image quality evaluation method based on the fusion of semantic features and visual features according to an embodiment of the present invention; Figure 2Shows a block diagram of a semantic fusion module according to an embodiment of the present invention; Figure 3 Shows a block diagram of visual feature extraction according to an embodiment of the present invention; Figure 4 Shows a block diagram of quality evaluation according to an embodiment of the present invention; Figure 5 Shows a schematic diagram of a semantic fusion module according to an embodiment of the present invention; Figure 6 Shows a schematic diagram of a visual perception feature extraction module according to an embodiment of the present invention; Figure 7 Shows a schematic diagram of quality evaluation according to an embodiment of the present invention; Figure 8 Shows an image semantic feature map according to an embodiment of the present invention; Figure 9 Shows a text semantic feature map according to an embodiment of the present invention; Figure 10 Shows a semantic feature map according to an embodiment of the present invention; Figure 11 Shows a visual perception feature map according to an embodiment of the present invention; Figure 12 Shows an output feature map of the first fully connected layer according to an embodiment of the present invention; Figure 13 Shows an output feature map of the second fully connected layer according to an embodiment of the present invention; Figure 14 Shows a feature map output by the third fully connected layer according to an embodiment of the present invention; Figure 15 Shows a self-attention fusion result map according to an embodiment of the present invention; Figure 16 Shows a block diagram of an exemplary electronic device capable of implementing an embodiment of the present invention.

[0019] Wherein, 1600 is an electronic device, 1601 is a computing unit, 1602 is a ROM, 1603 is a RAM, 1604 is a bus, 1605 is an I / O interface, 1606 is an input unit, 1607 is an output unit, 1608 is a storage unit, and 1609 is a communication unit. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] In addition, the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0022] In the present invention, text description statements are extracted from the target distorted panoramic image, and image semantic features and text semantic features that are semantically aligned are extracted according to the text description statements and the target distorted panoramic image. Then, the semantically aligned image semantic features and text semantic features are fused to obtain semantic features; feature extraction is performed on the viewport image of the target distorted panoramic image to obtain visual perception features; the semantic features and visual perception features are fused to obtain the quality evaluation score of the target distorted panoramic image. Through the fusion of semantic features and visual perception features, effective interaction between low-level semantic information and high-level semantic information is achieved, and the overall accuracy of the model can be improved.

[0023] Figure 1 The flowchart of a panoramic image quality evaluation method based on the fusion of semantic features and visual features according to an embodiment of the present invention is shown.

[0024] The method includes: S101. Obtain a target distorted panoramic image.

[0025] Specifically, in this embodiment, the publicly available panoramic image dataset CVIQ is used to evaluate the model proposed in the present invention. The dataset contains 528 distorted panoramic images, with a total of three different types of coding compression (JPEG, H.264 / AVC, and H.265 / HEVC). Each distorted image provides a subjective score MOS value, and the resolution of all images is 4096 × 2048.

[0026] S102. Input the target distorted panoramic image into a text generation model to output a text description statement of the target distorted panoramic image; input the target distorted panoramic image and the text description statement of the target distorted panoramic image into an imagebind model to obtain semantically aligned image semantic features and text semantic features; input the semantically aligned image semantic features and text semantic features into a semantic fusion module for semantic fusion to obtain semantic features; and perform a viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then perform visual feature extraction to obtain visual perception features.

[0027] Specifically, the text generation model is DepictQA (Depicted image Quality Assessment method, a multimodal large language model for image quality perception), which uses an MLLM (Multimodal Large Language Model) to make a human-like, language-based description of image quality.

[0028] As some optional implementation manners of this embodiment, assume that the target distorted panoramic image is input into the DepictQA model, and the output text description statement is "The quantization distortion is evident, resulting in a loss of detail and color depth, which affects the images's overall clarity. It has obvious quantization distortion, retains more detail and color accuracy, making it bad to reflect the scene".

[0029] In this embodiment, the imagebind model can embed text and image data into a unified embedding space, thereby realizing the understanding and conversion between different modal data.

[0030] As some optional implementation manners of this embodiment, assume that the text description statement output by the DepictQA model is input into the imagebind model, and the output semantically aligned image semantic features (as Figure 8 shown) and text semantic features (as Figure 9 shown).

[0031] In this embodiment, as Figure 2As shown in the figure, inputting the semantically aligned image semantic features and text semantic features into the semantic fusion module for semantic fusion to obtain semantic features includes: S201. Input the semantically aligned image semantic features into the first softmax module for reweighting, including the first multi-layer perceptron (MLP, Multilayer Perceptron) and the second multi-layer perceptron.

[0032] S202. Multiply the output of the first multi-layer perceptron by the semantically aligned text semantic features to obtain the first semantic fusion feature.

[0033] S203. Add the first semantic fusion feature to the output of the second multi-layer perceptron to obtain the second semantic fusion feature.

[0034] S204. Input the second semantic fusion feature into the first ReLu activation function and the first convolutional block in sequence, and output the semantic features.

[0035] Among them, the size of the first convolutional block is 3*3 and the stride is 8.

[0036] As some optional implementation manners of this embodiment, assume that Figure 8 the shown image semantic features and Figure 9 the shown text semantic features are input into the semantic fusion module for semantic fusion, and the output semantic features are as Figure 10 shown.

[0037] Modal alignment can ensure that the semantics of the image and text modalities are aligned consistently in the shared space, solve the heterogeneity between modalities, and prepare for feature interaction and fusion. Reweighting the image semantic information using the softmax module can improve the model's ability to understand image semantics; two-level fusion can capture richer detailed information, which helps to achieve smoother and more reliable semantic information interaction and fusion between text and images.

[0038] In this embodiment, as Figure 3 shown, the visual feature extraction to obtain visual perception features includes: S301. Input all the viewport maps into the visual perception feature extraction module to obtain the local visual features of each viewport.

[0039] In this embodiment, as Figure 5 shown, the visual perception feature extraction module (i.e., the VMamba model) sequentially includes a spatial transformation embedding module (Stem), a first visual state space block (VSS Block), a first downsampling module (DownSampling), a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module, and a fourth visual state space block; Among them, the spatial transformation embedding module is used to obtain the position perception information of the viewport map; the local visual features output by the fourth visual state space block are used as the output of the visual perception feature extraction module.

[0040] Specifically, the viewport acquisition method uses FoV Selection (Field of View Selection) to obtain the viewport map of the target distorted panoramic image.

[0041] Extracting the local visual features of each viewport through the feature extraction module conforms to the visual characteristics of the human eye when viewing panoramic images and can better reflect the quality loss information.

[0042] In this embodiment, as Figure 6 shown, the first visual state space block sequentially includes a first LN layer, a first 2D selective scanning block (SS2D Block), a second LN layer, and a first FFN layer.

[0043] Specifically, the first LN layer obtains the output of the spatial transformation embedding module, performs normalization, and then outputs it to the first 2D selective scanning block for feature extraction to obtain a first extraction result. The output of the spatial transformation embedding module and the first extraction result are added together to obtain a first addition result. The first addition result is sequentially input into the second LN layer and the first FFN layer to output a first enhancement result. The first addition result and the first enhancement result are added together to obtain the output of the first visual state space block.

[0044] The first visual state space block adopting the SS2D Block scanning method can simultaneously obtain the one-dimensional ordered features and the unsequenced context information of the two-dimensional visual data, expand the receptive field, obtain the dynamic weights with linear complexity, and help reduce the computational complexity of the visual processing task.

[0045] In this embodiment, the second visual state space block sequentially includes a third LN layer, a second 2D selective scanning block, a fourth LN layer, and a second FFN layer.

[0046] Specifically, the third LN layer obtains the output of the first downsampling module, performs normalization, and then outputs it to the second 2D selective scanning block for feature extraction to obtain a second extraction result. The output of the first downsampling module and the second extraction result are added together to obtain a second addition result. The second addition result is sequentially input into the fourth LN layer and the second FFN layer to output a second enhancement result. The second addition result and the second enhancement result are added together to obtain the output of the second visual state space block.

[0047] The second visual state space block undergoes downsampling processing to obtain a low-scale feature map. By combining with the 2D selective scanning module to obtain multi-level features, the accuracy of the model can be improved.

[0048] In this embodiment, the third visual state space block sequentially includes a fifth LN layer, a third 2D selective scanning block, a sixth LN layer, and a third FFN layer.

[0049] Specifically, the fifth LN layer obtains the output of the second downsampling module, performs normalization, and then outputs it to the third 2D selective scanning block for feature extraction to obtain a third extraction result. The output of the second downsampling module and the third extraction result are added to obtain a third addition result. The third addition result is sequentially input into the sixth LN layer and the third FFN layer to output a third enhancement result. The third addition result and the third enhancement result are added to obtain the output of the third visual state space block.

[0050] The third visual state space block undergoes downsampling processing to obtain a low-scale feature map. By combining with the 2D selective scanning module to obtain multi-level features, the accuracy of the model can be improved.

[0051] In this embodiment, the fourth visual state space block sequentially includes a seventh LN layer, a fourth 2D selective scanning block, an eighth LN layer, and a fourth FFN layer.

[0052] Specifically, the seventh LN layer obtains the output of the third downsampling module, performs normalization, and then outputs it to the fourth 2D selective scanning block for feature extraction to obtain a fourth extraction result. The output of the third downsampling module and the fourth extraction result are added to obtain a fourth addition result. The fourth addition result is sequentially input into the eighth LN layer and the fourth FFN layer to output a fourth enhancement result. The fourth addition result and the fourth enhancement result are added to obtain the output of the fourth visual state space block.

[0053] The fourth visual state space block undergoes downsampling processing to obtain a low-scale feature map. By combining with the 2D selective scanning module to obtain multi-level features, the accuracy of the model can be improved.

[0054] S302. Concatenate all local visual features to obtain visual perception features.

[0055] Performing viewport segmentation on the distorted panoramic image conforms to the characteristics of human visual perception, helps to extract local visual features, and obtain rich visual information; the VMamba model can exhibit higher computational efficiency and performance potential in high-resolution panoramic image processing tasks, and effectively mine visual perception features.

[0056] As some alternative embodiments of this embodiment, it is assumed that the target distorted panoramic image is divided into a plurality of viewport images, and one of the viewport images is selected and input into the VMamba model for visual feature extraction, and the output visual perception features are as Figure 11 shown.

[0057] S103. Perform feature fusion on the semantic feature and the visual perception feature to obtain a quality evaluation score of the target distorted panoramic image; the quality evaluation score is used to evaluate the quality of the target distorted panoramic image.

[0058] In this embodiment, as Figure 4 , 7 shown, the performing feature fusion on the semantic feature and the visual perception feature to obtain a quality evaluation score of the target distorted panoramic image includes: S401. Input the semantic feature into a first fully connected layer to obtain a first output result; and input the semantic feature into a second fully connected layer to obtain a second output result; and input the visual perception feature into a third fully connected layer to obtain a third output result.

[0059] Specifically, input the semantic feature into the first fully connected layer for channel feature processing; input the semantic feature into the second fully connected layer for offset feature processing to dynamically adjust the context feature response and improve the accuracy of feature description; input the visual perception feature into the third fully connected layer to adjust the visual feature dimension.

[0060] As some alternative embodiments of this embodiment, it is assumed that Figure 10 the shown semantic feature is input into the first fully connected layer, and the output feature map is as Figure 12 shown; Figure 10 the shown semantic feature is input into the second fully connected layer, and the output feature map is as Figure 13 shown; Figure 11 the shown visual perception feature is input into the third fully connected layer, and the output feature map is as Figure 14 shown.

[0061] S402. Perform self-attention fusion on the first output result, the second output result, and the third output result to obtain a fusion result.

[0062] Specifically, the self-attention fusion includes:

[0063] Among them, is the self-attention fusion result; is the visual perception feature; is the output result of the semantic feature input into the first fully connected layer; is the output result of the semantic feature input into the second fully connected layer; is a parameter to prevent the introduction of gradient messages; is the weight matrix of the key; is the weight matrix of the query; is the weight matrix of the value. Among them , , are learnable parameters.

[0064] As some alternative implementation manners of this embodiment, assume that Figure 12 , 13 , the features shown in 14 are respectively input into the above self-attention fusion formula, and the obtained fusion result is as Figure 15 shown.

[0065] S403. Add the fusion result to the third output result, successively send the added result into the ninth LN layer and the fifth FFN layer, and send the output of the fifth FFN layer into the fourth fully connected layer to output the quality evaluation score of the target distortion panoramic image.

[0066] As some alternative implementation manners of this embodiment, the evaluation metrics used are PLCC (Pearson linear correlation coefficient), SRCC (Spearman's Rank Correlation Coefficient), and RMSE (Root Mean Square Error). Among them, the higher the PLCC and SRCC values and the smaller the RMSE value, the better the model performance.

[0067] The calculation formula of PLCC is as follows:

[0068]

[0069]

[0070] Among them, is the number of test images; is the index value of the number of test images; is the mean value of the subjective quality scores of all images in the prediction dataset; is the mean value of the model prediction quality scores; is the th subjective quality score provided by the test image in the dataset; is the quality score predicted by the model; The linear consistency between the predicted image quality score and the quality score provided by the dataset.

[0071] The calculation formula is as follows:

[0072] Where, is the monotonicity between the predicted quality score and the subjective score provided by the dataset.

[0073] The RMSE calculation formula is as follows:

[0074] Where, RMSE is the error between the predicted score and the subjective score.

[0075] In this embodiment, the MOS value of the subjective score provided by the dataset and the objective quality score obtained by the present invention are input into the PLCC, SRCC, and RMSE calculation formulas to obtain the evaluation index values. Then, the above two modules are respectively replaced with a simple splicing and fusion module to calculate the evaluation index values. The results are shown in Table 1 below. From the above results, it can be seen that the PLCC and SRCC of the present invention are closest to 1, and the RMSE is the smallest, indicating that the fusion method of the present invention can effectively strengthen the semantic features and improve the feature representation ability, and can also achieve the effective fusion of features, improve the model accuracy, and prove that the present invention can accurately evaluate the panoramic image quality.

[0076] Table 1 PLCC SRCC RMSE the model of the present invention 0.8562 0.8662 7.6584 Replace the semantic fusion model with a splicing model 0.8122 0.8103 8.7745 Replace the feature fusion model with a splicing model 0.7844 0.7752 9.0325 Replace both the semantic fusion and feature fusion models with a splicing model 0.4861 0.4607 17.8314 According to the embodiments of the present invention, the present invention has the following advantages compared with the prior art: (1) By using the generation model to obtain the text description information of the distorted panoramic image, it can effectively supplement the visual perception information of the image, and can obtain more semantic information on quality loss than the existing methods that solely rely on image information, improving the overall performance of the model.

[0077] (2) Using the aligned visual and text semantic features for fusion to solve the heterogeneity between modalities, which helps to obtain rich semantic information and improve the model feature representation ability.

[0078] (3) The interactive fusion of semantic information and visual perception information strengthens the deep fusion of features and improves the overall performance of the model.

[0079] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0080] The above is the introduction of the method embodiments. The following is a further description of the solution of the present invention through apparatus embodiments having the same inventive concept as the methods in the foregoing embodiments.

[0081] According to an embodiment of the present invention, the present invention also provides an electronic device.

[0082] Figure 16 FIG. shows a schematic block diagram of an electronic device 1600 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0083] The electronic device 1600 includes a computing unit 1601, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1602 or a computer program loaded from a storage unit 1608 into a random access memory (RAM) 1603. In the RAM 1603, various programs and data required for the operation of the electronic device 1600 can also be stored. The computing unit 1601, the ROM 1602, and the RAM 1603 are connected to each other through a bus 1604. An input / output (I / O) interface 1605 is also connected to the bus 1604.

[0084] A plurality of components in the electronic device 1600 are connected to the I / O interface 1605, including: an input unit 1606, such as a keyboard, a mouse, etc.; an output unit 1607, such as various types of displays, speakers, etc.; a storage unit 1608, such as a magnetic disk, an optical disk, etc.; and a communication unit 1609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1609 allows the electronic device 1600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0085] The computing unit 1601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1601 executes the various methods and processes described above, such as methods S101 to S103. For example, in some embodiments, methods S101 to S103 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1600 via the ROM 1602 and / or the communication unit 1609. When the computer program is loaded into the RAM 1603 and executed by the computing unit 1601, one or more steps of methods S101 to S103 described above can be executed. Alternatively, in other embodiments, the computing unit 1601 can be configured to execute methods S101 to S103 in any other suitable manner (e.g., by means of firmware).

[0086] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0087] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0088] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0089] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A panoramic image quality assessment method based on the fusion of semantic features and visual features, characterized in that: include: Obtain a distorted panoramic image of the target; Inputting the target distorted panoramic image into a text generation model to output a text description sentence of the target distorted panoramic image; inputting the target distorted panoramic image and the text description sentence of the target distorted panoramic image into an imagebind model to obtain semantically aligned image semantic features and text semantic features; The semantically aligned image semantic features and text semantic features are input into the semantic fusion module for semantic fusion to obtain semantic features; as well as Performing viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then performing visual feature extraction to obtain visual perception features; The semantic features and the visual perception features are feature fused to obtain a quality evaluation score of the target distorted panoramic image; the quality evaluation score is used to perform quality evaluation on the target distorted panoramic image.

2. The method according to claim 1, characterized in that The semantically aligned image semantic features and text semantic features are input into the semantic fusion module for semantic fusion to obtain semantic features, including: The semantic features of the semantically aligned images are input into the first softmax module for reweighting, and then input into the first multi-layer perceptron and the second multi-layer perceptron; Multiplying the output of the first multi-layer perceptron by the semantically aligned text semantic feature to obtain a first semantic fusion feature; Adding the first semantic fusion feature to the output of the second multi-layer perceptron to obtain a second semantic fusion feature; The second semantic fusion feature is sequentially input into the first ReLu activation function and the first convolution block to output the semantic feature.

3. The method according to claim 1, characterized in that The visual feature extraction to obtain visual perception features includes: Input all viewport images into the visual perception feature extraction module to obtain the local visual features of each viewport; All local visual features are concatenated to obtain visual perception features.

4. The method according to claim 3, characterized in that The visual perception feature extraction module includes, in sequence, a spatial transformation embedding module, a first visual state space block, a first downsampling module, a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module and a fourth visual state space block; Among them, the spatial transformation embedding module is used to obtain the position perception information of the viewport graph; the local visual features output by the fourth visual state space block are used as the output of the visual perception feature extraction module.

5. The method according to claim 4, characterized in that The first visual state space block includes, in sequence, a first LN layer, a first 2D selective scanning block, a second LN layer, and a first FFN layer; The first LN layer obtains the output of the spatial transformation embedding module, normalizes it, and outputs it to the first 2D selective scanning block for feature extraction, and outputs a first extraction result; Adding the output of the spatial transformation embedding module and the first extraction result to obtain a first addition result; Inputting the first addition result to the second LN layer and the first FFN layer in sequence, and outputting a first enhancement result; The first addition result and the first enhancement result are added to obtain an output of the first visual state space block.

6. The method according to claim 4, characterized in that The second visual state space block includes, in sequence, a third LN layer, a second 2D selective scanning block, a fourth LN layer, and a second FFN layer; The third LN layer obtains the output of the first downsampling module, normalizes it, and outputs it to the second 2D selective scanning block for feature extraction, and outputs a second extraction result; Adding the output of the first down-sampling module and the second extraction result to obtain a second addition result; Inputting the second addition result to the fourth LN layer and the second FFN layer in sequence, and outputting a second enhancement result; The second addition result and the second enhancement result are added to obtain an output of the second visual state space block.

7. The method according to claim 4, characterized in that The third visual state space block includes, in sequence, a fifth LN layer, a third 2D selective scanning block, a sixth LN layer, and a third FFN layer; The fifth LN layer obtains the output of the second down-sampling module, normalizes it, and outputs it to the third 2D selective scanning block for feature extraction, and outputs a third extraction result; Adding the output of the second down-sampling module and the third extraction result to obtain a third addition result; inputting the third addition result to the sixth LN layer and the third FFN layer in sequence, and outputting a third enhancement result; The third addition result and the third enhancement result are added to obtain the output of the third visual state space block.

8. The method according to claim 4, characterized in that The fourth visual state space block includes, in sequence, a seventh LN layer, a fourth 2D selective scanning block, an eighth LN layer, and a fourth FFN layer; The seventh LN layer obtains the output of the third down-sampling module, normalizes it, and outputs it to the fourth 2D selective scanning block for feature extraction, and outputs a fourth extraction result; Adding the output of the third down-sampling module and the fourth extraction result to obtain a fourth addition result; inputting the fourth addition result to the eighth LN layer and the fourth FFN layer in sequence, and outputting a fourth enhancement result; The fourth addition result and the fourth enhancement result are added to obtain an output of the fourth visual state space block.

9. The method according to claim 1, characterized in that: The step of fusing the semantic features with the visual perception features to obtain a quality evaluation score of the target distorted panoramic image includes: Inputting the semantic feature into a first fully connected layer to obtain a first output result; and Inputting the semantic feature into a second fully connected layer to obtain a second output result; and Inputting the visual perception feature into a third fully connected layer to obtain a third output result; Performing self-attention fusion on the first output result, the second output result, and the third output result to obtain a fusion result; The fusion result is added to the third output result, and the addition result is sent to the ninth LN layer and the fifth FFN layer in sequence. The output of the fifth FFN layer is sent to the fourth fully connected layer, and the quality evaluation score of the target distorted panoramic image is output.

10. An electronic device comprising at least one processor; and a memory connected to the at least one processor in communication; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Visual question-answering method and system based on multiple attention

    CN110516791A

  • Aesthetics quality evaluation model and method based on multi-modal learning

    CN115601772A

  • Model training method and device, perception reasoning method, electronic equipment and medium

    CN117290459A

  • Image classification method, device and equipment based on multi-scale attention mechanism

    CN118823489A

  • Quality evaluation method and device for distorted panoramic image

    CN119006386A