Panoramic image quality evaluation method and device based on psychological memory mechanism

Through the panoramic image quality evaluation method based on psychological memory mechanism, using a large model to perform multi-level interactive fusion of visual and text semantic information, the problem of insufficient panoramic image quality evaluation in the prior art is solved, and higher evaluation accuracy and performance are achieved.

CN120070449AActive Publication Date: 2025-05-30LIAONING BEIDOU SATELLITE NAVIGATION PLATFORM CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510549968.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing panoramic image quality evaluation methods have shortcomings in visual perception modeling and algorithm innovation, and cannot effectively evaluate the quality of panoramic images.

Method used

A panoramic image quality evaluation method based on psychological memory mechanism is adopted to obtain visual and text semantic information through a large model, combine human brain memory mechanism to perform multi-level interactive fusion, extract rich semantic features, and improve the overall performance of the model.

Benefits of technology

It realizes deep mining of the interactive relationship between visual and text modes, extracts high-level semantic features, and improves the accuracy and performance of panoramic image quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070449A_ABST
    Figure CN120070449A_ABST
Patent Text Reader

Abstract

The invention relates to the field of panoramic image quality evaluation, in particular to a panoramic image quality evaluation method and device based on a psychological memory mechanism. The method comprises the steps that text semantic features are obtained through a target distortion panorama; multi-level image semantic features are obtained through the target distortion panorama; obtaining high-level semantic features through the text semantic features and the multi-level image semantic features; performing viewport splitting processing on the target distortion panorama to obtain a plurality of viewport images; obtaining a local low-level semantic feature of each viewport through each viewport graph and the text semantic feature, and splicing the local low-level semantic features of all viewports to obtain a low-level semantic feature; and carrying out feature aggregation on the high-level semantic features and the low-level semantic features to obtain a quality evaluation score of the target distorted panoramic image. In this way, visual and text semantic information can be obtained by means of a large model, and the overall performance of the model can be effectively improved in combination with a human brain memory mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of panoramic image quality assessment, and more particularly, to a panoramic image quality assessment method and apparatus based on a psychological memory mechanism. Background Art

[0002] With the rapid development of camera technology, panoramic images are increasingly widely used in fields such as entertainment, education, and real estate. Different from traditional planar images, panoramic images have visual characteristics of high resolution, wide viewing angle, and spherical projection, which can enhance the user experience and comfort. However, quality losses may be introduced in the processes of acquisition, compression, transmission, and rendering of panoramic images, and effective quality assessment methods are needed to guide optimization and improve the perceptual quality of images.

[0003] Traditional image quality assessment methods (such as PSNR, SSIM) are mainly for planar images and cannot be directly used for panoramic images. The research on panoramic image quality assessment methods has gradually attracted the attention of the academic and industrial communities. In the early stage of this field, researchers mainly manually extracted geometric features, texture features, etc. of panoramic images and evaluated them in combination with traditional quality assessment indicators, with general performance. Subsequently, many deep learning-based methods have been proposed, such as using convolutional neural networks (CNNs) or vision transformers to extract features and predict image quality, with improved performance. Considering that when a person views a panoramic image, only the range of one viewport can be viewed at a time, researchers have proposed a panoramic image quality assessment method based on viewport features, with further improved performance.

[0004] However, existing panoramic image quality assessment methods have deficiencies in aspects such as visual perception modeling and algorithm innovation. Summary of the Invention

[0005] According to an embodiment of the present invention, a panoramic image quality assessment solution based on a psychological memory mechanism is provided. This solution can obtain visual and text semantic information with the help of a large model and, combined with the human brain memory mechanism, can effectively improve the overall performance of the model.

[0006] In a first aspect of the present invention, a panoramic image quality assessment method based on a psychological memory mechanism is provided. The method includes: Obtain a target distorted panoramic image; Input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; Input the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; then input the text semantic features and the multi-level image semantic features into a global semantic interaction and fusion model to obtain high-level semantic features; The target distorted panoramic image is subjected to viewport processing to obtain a plurality of viewport images; each viewport image and the text semantic features are respectively input into a local semantic interaction fusion model, a local low-level semantic feature of each viewport is output, and the local low-level semantic features of all viewports are then spliced ​​to obtain a low-level semantic feature; The high-level semantic features and the low-level semantic features are feature aggregated to obtain a quality evaluation score of the target distorted panoramic image.

[0007] In a second aspect of the present invention, a panoramic image quality evaluation device based on a psychological memory mechanism is provided. The device comprises: An acquisition module, used for acquiring a distorted panoramic image of a target; A first feature extraction module, used for inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; A second feature extraction module is used to input the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; and then input the text semantic features and the multi-level image semantic features into a global semantic interaction fusion model to obtain high-level semantic features; The third feature extraction module is used to perform viewport processing on the target distorted panoramic image to obtain a plurality of viewport images; input each viewport image and the text semantic features into the local semantic interaction fusion model respectively, output the local low-level semantic features of each viewport, and then splice the local low-level semantic features of all viewports to obtain low-level semantic features; An evaluation module is used to perform feature aggregation on the high-level semantic features and the low-level semantic features to obtain a quality evaluation score of the target distorted panoramic image.

[0008] In a third aspect of the present invention, an electronic device is provided. The electronic device has at least one processor; and a memory connected to the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method of the first aspect of the present invention.

[0009] In a fourth aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the first aspect of the present invention.

[0010] Compared with the prior art, the present invention has the following beneficial technical effects: Based on the human brain's memory mechanism, it realizes the multi-level interactive fusion of text information and visual information, which can deeply explore the interactive relationship between the two modalities, extract rich semantic features, and improve the feature representation ability of the model.

[0011] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. Brief Description of the Drawings

[0012] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where: Figure 1 FIG. shows a flowchart of a panoramic image quality evaluation method based on a psychological memory mechanism according to an embodiment of the present invention; Figure 2 FIG. shows a flowchart of a text semantic feature extraction model according to an embodiment of the present invention; Figure 3 FIG. shows a flowchart of a global semantic interaction and fusion model according to an embodiment of the present invention; Figure 4 FIG. shows a flowchart of a semantic interaction and fusion module according to an embodiment of the present invention; Figure 5 FIG. shows a flowchart of feature aggregation according to an embodiment of the present invention; Figure 6 FIG. shows a schematic diagram of a global visual semantic feature extraction model according to an embodiment of the present invention; Figure 7 FIG. shows a schematic diagram of a visual state space block according to an embodiment of the present invention; Figure 8 FIG. shows a schematic diagram of a global semantic interaction and fusion model according to an embodiment of the present invention; Figure 9 FIG. shows a schematic diagram of a semantic interaction and fusion module according to an embodiment of the present invention; Figure 10 FIG. shows a schematic diagram of feature aggregation according to an embodiment of the present invention; Figure 11 FIG. shows a schematic diagram of text semantic features according to an embodiment of the present invention; Figure 12 FIG. shows a schematic diagram of the second-level image semantic features of a global visual semantic feature extraction model according to an embodiment of the present invention; Figure 13 FIG. shows a schematic diagram of the third-level image semantic features of a global visual semantic feature extraction model according to an embodiment of the present invention; Figure 14 FIG. shows a schematic diagram of the fourth-level image semantic features of a global visual semantic feature extraction model according to an embodiment of the present invention; Figure 15Schematic diagram showing the first - level high - level semantic features according to an embodiment of the present invention; Figure 16 Schematic diagram showing the second - level high - level semantic features according to an embodiment of the present invention; Figure 17 Schematic diagram showing the high - level semantic features according to an embodiment of the present invention; Figure 18 Schematic diagram showing the text - enhanced semantic features according to an embodiment of the present invention; Figure 19 Schematic diagram showing the fused semantic features according to an embodiment of the present invention; Figure 20 Schematic diagram showing the local low - level semantic features of the viewport according to an embodiment of the present invention; Figure 21 Schematic diagram showing the output K of the fourth convolutional layer according to an embodiment of the present invention; Figure 22 Schematic diagram showing the output V of the fifth convolutional layer according to an embodiment of the present invention; Figure 23 Schematic diagram showing the output Q of the seventh convolutional layer according to an embodiment of the present invention; Figure 24 Schematic diagram showing the output result of the second self - attention module according to an embodiment of the present invention; Figure 25 Block diagram showing a panoramic image quality evaluation device based on a psychological memory mechanism according to an embodiment of the present invention; Figure 26 Block diagram showing an exemplary electronic device capable of implementing an embodiment of the present invention.

[0013] Wherein, Stem is a spatial transformation embedding module, VSS Block is a visual state space block, Down Sampling is a down - sampling module, FFN is a feed - forward neural network layer, LN is a normalization layer, SS2D Block is a 2D selective scanning block, 2600 is an electronic device, 2601 is a computing unit, 2602 is a ROM, 2603 is a RAM, 2604 is a bus, 2605 is an I / O interface, 2606 is an input unit, 2607 is an output unit, 2608 is a storage unit, 2609 is a communication unit. Detailed implementation manners

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0015] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0016] In the present invention, text semantic features and multi-level image semantic features are obtained from the target distorted panoramic image, and high-level semantic features are obtained through the text semantic features and multi-level image semantic features; a viewport image is obtained from the target distorted panoramic image, and low-level semantic features are obtained through the viewport image and text semantic features; the high-level semantic features and low-level semantic features are aggregated to obtain a quality evaluation score for the target distorted panoramic image. In this way, the characteristics of human brain memory perception can be fully explored by combining large model technology; at the same time, the existing panoramic image quality evaluation algorithm is improved by combining multi-modal features, effectively improving the overall performance of the model.

[0017] Figure 1 The flowchart of the panoramic image quality evaluation method based on the psychological memory mechanism according to the embodiment of the present invention is shown.

[0018] The method includes: S101. Obtain a target distorted panoramic image.

[0019] Specifically, in this embodiment, the publicly available panoramic image dataset CVIQ is used to evaluate the model proposed in the present invention. The dataset contains 528 distorted panoramic images, with a total of three different types of coding compression (JPEG, H.264 / AVC, and H.265 / HEVC). A subjective score MOS value is provided for each distorted image, and the resolution of all images is 4096 × 2048.

[0020] S102. Input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features.

[0021] In this embodiment, as Figure 2 shown, the step of inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features includes: S201. Input the target distorted panoramic image into a text generation model to output a text description statement of the target distorted panoramic image.

[0022] Specifically, the text generation model is DepictQA (Depicted image Quality Assessment method, a multimodal large language model for image quality perception), which uses MLLM (Multimodal Large Language Model) to make a human-like, language-based description of the image quality.

[0023] As some optional implementation manners of this embodiment, assume that the target distorted panoramic image is input into the DepictQA model, and the output text description statement is “The quantization distortion is evident, resulting in a loss of detail and color depth, which affects the images's overall clarity. It has obvious quantization distortion, retains more detail and color accuracy, making it bad to reflect the scene”.

[0024] S202. Input the text description statement of the target distorted panoramic image into the Imagebind model to obtain text semantic features (as Figure 11 shown).

[0025] Specifically, in this embodiment, the Imagebind model can embed text data into a unified embedding space to achieve the understanding and conversion between different modal data.

[0026] Extracting the text description statement for the distorted panoramic image can provide non-visual semantic information for the image and improve the model accuracy; using a large model to extract text semantic information can reduce the computational complexity and obtain representative text semantic information at the same time.

[0027] S103. Input the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; then input the text semantic features and the multi-level image semantic features into a global semantic interaction and fusion model to obtain high-level semantic features.

[0028] In this embodiment, as Figure 6As shown, the global visual semantic feature extraction model (VMamba) sequentially includes a spatial transformation embedding module (Stem, Spatial Transformer Embedding Module), a first visual state space block (VSS Block, Visual State Space Block), a first downsampling module (Down Sampling), a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module, and a fourth visual state space block.

[0029] Among them, the output of the first visual state space block is used as the first-level image semantic feature; the output of the second visual state space block is used as the second-level image semantic feature (as Figure 12 shown); the output of the third visual state space block is used as the third-level image semantic feature (as Figure 13 shown); the output of the fourth visual state space block is used as the fourth-level image semantic feature (as Figure 14 shown); the first-level image semantic feature, the second-level image semantic feature, the third-level image semantic feature, and the fourth-level image semantic feature constitute multi-level image semantic features.

[0030] In this embodiment, as Figure 7 shown, the first visual state space block sequentially includes a third LN (normalization) layer, a first 2D selective scan block (SS2d Block, 2D-selective-scan Block), a fourth LN layer, and a second FFN (Feed Forward Network) layer.

[0031] Specifically, the third LN layer obtains the output of the spatial transformation embedding module, performs normalization, and then outputs it to the first 2D selective scan block for feature extraction to output a first extraction result; the output of the spatial transformation embedding module and the first extraction result are added together to obtain a first addition result; the first addition result is sequentially input into the fourth LN layer and the second FFN layer to output a first enhancement result; the first addition result and the first enhancement result are added together to obtain the output of the first visual state space block.

[0032] In this embodiment, the second visual state space block sequentially includes a fifth LN layer, a second 2D selective scan block, a sixth LN layer, and a third FFN layer.

[0033] Specifically, the fifth LN layer obtains the output of the first downsampling module, normalizes it, and then outputs it to the second 2D selective scanning block for feature extraction, and outputs a second extraction result; adds the output of the first downsampling module and the second extraction result to obtain a second addition result; sequentially inputs the second addition result into the sixth LN layer and the third FFN layer, and outputs a second enhancement result; adds the second addition result and the second enhancement result to obtain the output of the second visual state space block.

[0034] In this embodiment, the third visual state space block sequentially includes a seventh LN layer, a third 2D selective scanning block, an eighth LN layer, and a fourth FFN layer.

[0035] Specifically, the seventh LN layer obtains the output of the second downsampling module, normalizes it, and then outputs it to the third 2D selective scanning block for feature extraction, and outputs a third extraction result; adds the output of the second downsampling module and the third extraction result to obtain a third addition result; sequentially inputs the third addition result into the eighth LN layer and the fourth FFN layer, and outputs a third enhancement result; adds the third addition result and the third enhancement result to obtain the output of the third visual state space block.

[0036] In this embodiment, the fourth visual state space block sequentially includes a ninth LN layer, a fourth 2D selective scanning block, a tenth LN layer, and a fifth FFN layer.

[0037] Specifically, the ninth LN layer obtains the output of the third downsampling module, normalizes it, and then outputs it to the fourth 2D selective scanning block for feature extraction, and outputs a fourth extraction result; adds the output of the third downsampling module and the fourth extraction result to obtain a fourth addition result; sequentially inputs the fourth addition result into the tenth LN layer and the fifth FFN layer, and outputs a fourth enhancement result; adds the fourth addition result and the fourth enhancement result to obtain the output of the fourth visual state space block.

[0038] The VMamba model can effectively capture the long-range dependencies and local details of images. Its hierarchical structure can effectively capture rich multi-scale semantic information, and the computational complexity is low.

[0039] In this embodiment, as Figure 3 、 8 shown, the inputting the text semantic feature and the multi-level image semantic feature into the global semantic interaction and fusion model to obtain a high-level semantic feature includes: S301. Input the text semantic features into the first convolutional layer Conv and the first ReLu activation function in sequence; then use the output of the first ReLu activation function and the second-level image semantic feature S2 as the input of the first LSTM module, and output the first-level high-level semantic feature (as shown in Figure 15 ).

[0040] S302. Input the text semantic features into the second convolutional layer Conv and the second ReLu activation function in sequence; then use the output of the second ReLu activation function, the third-level image semantic feature S3, and the first-level high-level semantic feature as the input of the second LSTM module, and output the second-level high-level semantic feature (as shown in Figure 16 ).

[0041] S303. Input the text semantic features into the third convolutional layer Conv and the third ReLu activation function in sequence; then use the output of the third ReLu activation function, the fourth-level image semantic feature S4, and the second-level high-level semantic feature as the input of the third LSTM module, output the third-level high-level semantic feature, and use the third-level high-level semantic feature as the high-level semantic feature (as shown in Figure 17 ).

[0042] Specifically, the convolutional kernel sizes of the first convolutional layer, the second convolutional layer, and the third convolutional layer are all 6, and the strides are all 6.

[0043] Considering that the human brain contains descriptive and non-descriptive memory features, using LSTM to perform multi-level interaction and fusion of text semantics and multi-level image semantic information, effectively simulating the fusion strategy of the human brain memory mechanism, can obtain accurate and rich high-level semantic information.

[0044] S104. Perform a split viewport process on the target distorted panoramic image to obtain several viewport images; input each viewport image and the text semantic features into the local semantic interaction and fusion model respectively, output the local low-level semantic features of each viewport, and then splice the local low-level semantic features of all viewports to obtain the low-level semantic features.

[0045] Specifically, the viewport acquisition method is to use FoV Selection (Field of View Selection) to obtain the viewport images of the target distorted panoramic image.

[0046] In this embodiment, the local semantic interaction and fusion model includes several semantic interaction and fusion modules, and each semantic interaction and fusion module corresponds to a viewport, and is used to take the viewport image of the corresponding viewport and the text semantic features as inputs, and output the local low-level semantic features of each viewport; each semantic interaction and fusion module includes a first self-attention module, a first cross-attention module, a first FFN layer, and a viewport semantic feature extraction module.

[0047] In this embodiment, as Figure 4 , 9 shown, taking the viewport image corresponding to the viewport and the text semantic features as inputs, and outputting the local low-level semantic features of each viewport, including: S401. Input the text semantic features into the first self-attention module (SA, Self-Attention), add the output of the first self-attention module to the text semantic features to obtain text-enhanced semantic features (as Figure 18 shown).

[0048] S402. Input the viewport image of the corresponding viewport into the viewport semantic feature extraction module (ViT, Vision Transformer) to obtain viewport semantic features; then use the viewport semantic features and the text-enhanced semantic features as inputs to the first cross-attention module (CA, Cross-Attention) to output fused semantic features (as Figure 19 shown).

[0049] S403. Add the text-enhanced semantic features and the fused semantic features and input them into the first FFN layer to obtain enhanced fused semantic features; then add the enhanced fused semantic features to the input of the first FFN layer to obtain the local low-level semantic features of each viewport (as Figure 20 shown).

[0050] Performing viewport segmentation on the distorted panoramic image conforms to the characteristics of human visual perception and helps to extract local visual semantic features; fusing text information with each viewport separately can mine rich low-level semantic information, which is complementary to high-level semantic information and improves the overall performance of the model; using self-attention and cross-attention modules for feature fusion can capture intra-modal and inter-modal semantic interaction relationships and improve the overall performance of the model.

[0051] S105. As Figure 5 , 10 shown, perform feature aggregation on the high-level semantic features and the low-level semantic features to obtain the quality evaluation score of the target distorted panoramic image, including: S501. Input the high-level semantic features into the first LN layer; input the output of the first LN layer into the fourth convolutional layer Conv and the fifth convolutional layer Conv respectively; and input the low-level semantic features into the sixth convolutional layer Conv; input the output of the sixth convolutional layer Conv into the second LN layer; then use the output of the second LN layer as the input of the seventh convolutional layer Conv.

[0052] S502. The output K of the fourth convolutional layer (as Figure 21 shown), the output V of the fifth convolutional layer (asFigure 22 as shown) and the output Q of the seventh convolutional layer (such as Figure 23 as shown) are input into the second self-attention module, and then the output of the second self-attention module (such as Figure 24 as shown) is used as the input of the first fully connected layer (FC) to output the quality evaluation score (Score) of the target distorted panoramic image.

[0053] As some optional implementation manners of this embodiment, feature aggregation includes:

[0054] where is the output of the second self-attention module; is the low-level semantic feature; is the high-level semantic feature; is the query weight matrix; is the transpose of the key weight matrix; is the transpose of the high-level semantic feature; is to prevent gradient messages from introducing parameters; is the value weight matrix.

[0055] As some optional implementation manners of this embodiment, the evaluation metrics used are PLCC (Pearson linear correlation coefficient), SRCC (Spearman's Rank Correlation Coefficient), and RMSE (Root Mean Square Error). The higher the PLCC and SRCC values and the smaller the RMSE value, the better the model performance.

[0056] The calculation formula of PLCC is as follows:

[0057]

[0058]

[0059] where is the number of test images; is the index value of the number of test images; is the mean of the subjective quality scores of all images in the prediction dataset; is the mean of the model-predicted quality scores; is the th subjective quality score provided by the test image dataset; is the quality score predicted by the model; The linear consistency between the predicted image quality score and the quality score provided by the dataset.

[0060] The calculation formula is as follows:

[0061] Where, is the monotonicity between the predicted quality score and the subjective score provided by the dataset.

[0062] The RMSE calculation formula is as follows:

[0063] Where, RMSE is the error between the predicted score and the subjective score.

[0064] In this embodiment, the MOS value of the subjective score provided by the dataset and the objective quality score obtained by the present invention are input into the PLCC, SRCC, and RMSE calculation formulas to obtain the evaluation index values. Then, the above two modules are respectively replaced with a simple splicing and fusion module to calculate the evaluation index values. The results are shown in Table 1 below. From the above results, it can be seen that the PLCC and SRCC of the present invention are closest to 1, and the RMSE is the smallest, indicating that the fusion method of the present invention can effectively strengthen semantic features and improve the feature representation ability, and can also achieve effective fusion of features, improve the model accuracy, and prove that the present invention can accurately evaluate the panoramic image quality.

[0065] Table 1 PLCC SRCC RMSE The model of the present invention 0.9625 0.9639 2.7753 Replace the global semantic interaction fusion with a local semantic interaction fusion module 0.9517 0.9501 3.5124 Replace the feature aggregation model with a splicing model 0.9501 0.9579 3.0325 The interactive fusion of high-level semantics and low-level semantics can further refine semantic features and improve the feature representation ability of the model; using the self-attention model for feature aggregation can capture the semantic interaction relationship between modalities and improve the overall performance of the model.

[0066] According to the embodiments of the present invention, the present invention has the following advantages and effects compared with the prior art: (1) Realize the multi-level interactive fusion of text information and visual information based on the human brain memory mechanism, which can deeply mine the interactive relationship between the two modalities, extract rich high-level semantic features, and improve the feature representation ability of the model.

[0067] (2) Considering that different viewports contain different local semantic information and contribute differently to the overall perception quality, and fusing it with text information can mine the low-level semantic information contained in local viewports, effectively complement the high-level semantic information, and improve the overall performance of the model.

[0068] (3) The feature aggregation of high-level semantic features and low-level semantic features can further refine the semantic features extracted by the model, and can further improve the accuracy of the model compared with simple splicing.

[0069] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0070] The above is the introduction of the method embodiments. The following is a further description of the solution of the present invention through device embodiments having the same inventive concept as the methods in the foregoing embodiments.

[0071] As Figure 25 shown, the device 2500 includes: An acquisition module 2510, configured to acquire a target distorted panoramic image.

[0072] A first feature extraction module 2520, configured to input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features.

[0073] A second feature extraction module 2530, configured to input the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; and then input the text semantic features and the multi-level image semantic features into a global semantic interaction and fusion model to obtain high-level semantic features.

[0074] A third feature extraction module 2540, configured to perform a sub-viewport processing on the target distorted panoramic image to obtain a plurality of viewport images; respectively input each viewport image and the text semantic features into a local semantic interaction and fusion model, output local low-level semantic features of each viewport, and then splice the local low-level semantic features of all viewports to obtain low-level semantic features.

[0075] An evaluation module 2550, configured to perform feature aggregation on the high-level semantic features and the low-level semantic features to obtain a quality evaluation score of the target distorted panoramic image.

[0076] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0077] In the technical solution of the present invention, the acquisition, storage, and application of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0078] According to an embodiment of the present invention, the present invention also provides an electronic device and a readable storage medium.

[0079] Figure 26 FIG. shows a schematic block diagram of an electronic device 2600 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0080] The electronic device 2600 includes a computing unit 2601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 2602 or a computer program loaded from a storage unit 2608 into a random access memory (RAM) 2603. In the RAM 2603, various programs and data required for the operation of the electronic device 2600 can also be stored. The computing unit 2601, the ROM 2602, and the RAM 2603 are connected to each other via a bus 2604. An input / output (I / O) interface 2605 is also connected to the bus 2604.

[0081] A plurality of components in the electronic device 2600 are connected to the I / O interface 2605, including: an input unit 2606, such as a keyboard, a mouse, etc.; an output unit 2607, such as various types of displays, speakers, etc.; a storage unit 2608, such as a magnetic disk, an optical disk, etc.; and a communication unit 2609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 2609 allows the electronic device 2600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0082] The computing unit 2601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 2601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2601 executes the various methods and processes described above, such as methods S101 - S105. For example, in some embodiments, methods S101 - S105 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 2608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 2600 via the ROM 2602 and / or the communication unit 2609. When the computer program is loaded into the RAM 2603 and executed by the computing unit 2601, one or more steps of methods S101 - S105 described above can be executed. Alternatively, in other embodiments, the computing unit 2601 can be configured to execute methods S101 - S105 in any other suitable manner (e.g., by means of firmware).

[0083] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0084] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0085] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0086] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0087] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0088] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0089] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0090] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A panoramic image quality evaluation method based on psychological memory mechanism, characterized in that: include: Obtain a distorted panoramic image of the target; Inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; Inputting the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; then inputting the text semantic features and the multi-level image semantic features into a global semantic interactive fusion model to obtain high-level semantic features; The target distorted panoramic image is subjected to viewport processing to obtain a plurality of viewport images; each viewport image and the text semantic features are respectively input into a local semantic interaction fusion model, a local low-level semantic feature of each viewport is output, and the local low-level semantic features of all viewports are then spliced ​​to obtain a low-level semantic feature; The high-level semantic features and the low-level semantic features are feature aggregated to obtain a quality evaluation score of the target distorted panoramic image.

2. The method according to claim 1, characterized in that The step of inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features includes: Inputting the target distorted panoramic image into a text generation model, and outputting a text description sentence of the target distorted panoramic image; The text description sentence of the target distorted panorama is input into the Imagebind model to obtain the text semantic features.

3. The method according to claim 1, characterized in that The global visual semantic feature extraction model includes, in sequence, a spatial transformation embedding module, a first visual state space block, a first downsampling module, a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module and a fourth visual state space block; Among them, the output of the first visual state space block is used as the first-level image semantic feature; the output of the second visual state space block is used as the second-level image semantic feature; the output of the third visual state space block is used as the third-level image semantic feature; the output of the fourth visual state space block is used as the fourth-level image semantic feature; the first-level image semantic feature, the second-level image semantic feature, the third-level image semantic feature and the fourth-level image semantic feature constitute a multi-level image semantic feature.

4. The method according to claim 3, characterized in that The step of inputting the text semantic features and the multi-level image semantic features into a global semantic interaction fusion model to obtain high-level semantic features includes: The text semantic features are sequentially input into the first convolutional layer and the first ReLu activation function; the output of the first ReLu activation function and the second-level image semantic features are then used as inputs of the first LSTM module to output the first-level high-level semantic features; The text semantic features are sequentially input into the second convolutional layer and the second ReLu activation function; the output of the second ReLu activation function, the third-level image semantic features and the first-level high-level semantic features are then used as inputs of the second LSTM module, and the second-level high-level semantic features are output; The text semantic features are sequentially input into the third convolutional layer and the third ReLu activation function; then the output of the third ReLu activation function, the fourth-level image semantic features and the second-level high-level semantic features are used as inputs of the third LSTM module, the third-level high-level semantic features are output, and the third-level high-level semantic features are used as high-level semantic features.

5. The method according to claim 1, characterized in that The local semantic interaction fusion model includes a plurality of semantic interaction fusion modules, and each semantic interaction fusion module corresponds to a viewport, and is used to take the viewport map of the corresponding viewport and the text semantic features as input, and output the local low-level semantic features of each viewport; Each semantic interaction fusion module includes a first self-attention module, a first cross-attention module, a first FFN layer and a viewport semantic feature extraction module.

6. The method according to claim 5, characterized in that The method takes the viewport image of the corresponding viewport and the text semantic features as input and outputs the local low-level semantic features of each viewport, including: Inputting the text semantic feature into the first self-attention module, adding the output of the first self-attention module to the text semantic feature to obtain a text enhanced semantic feature; Inputting the viewport image of the corresponding viewport into the viewport semantic feature extraction module to obtain the viewport semantic feature; then using the viewport semantic feature and the text enhancement semantic feature as the input of the first cross attention module to output the fused semantic feature; The text enhanced semantic feature and the fused semantic feature are added and input into the first FFN layer to obtain an enhanced fused semantic feature; and then the enhanced fused semantic feature is added to the input of the first FFN layer to obtain a local low-level semantic feature of each viewport.

7. The method according to claim 1, characterized in that The step of performing feature aggregation on the high-level semantic features and the low-level semantic features to obtain a quality evaluation score of the target distorted panoramic image includes: Inputting the high-level semantic features into the first LN layer; inputting the output of the first LN layer into the fourth convolutional layer and the fifth convolutional layer respectively; and Input the low-level semantic features into the sixth convolutional layer; input the output of the sixth convolutional layer into the second LN layer; and then use the output of the second LN layer as the input of the seventh convolutional layer; The outputs of the fourth convolutional layer, the fifth convolutional layer, and the seventh convolutional layer are input into the second self-attention module, and then the output of the second self-attention module is used as the input of the first fully connected layer to output the quality evaluation score of the target distorted panoramic image.

8. A panoramic image quality assessment device based on psychological memory mechanism, characterized in that: include: An acquisition module, used for acquiring a distorted panoramic image of a target; A first feature extraction module, used for inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; A second feature extraction module is used to input the target distorted panoramic image into a global visual semantic feature extraction model to obtain multi-level image semantic features; and then input the text semantic features and the multi-level image semantic features into a global semantic interaction fusion model to obtain high-level semantic features; The third feature extraction module is used to perform viewport processing on the target distorted panoramic image to obtain a plurality of viewport images; input each viewport image and the text semantic features into the local semantic interaction fusion model respectively, output the local low-level semantic features of each viewport, and then splice the local low-level semantic features of all viewports to obtain low-level semantic features; An evaluation module is used to perform feature aggregation on the high-level semantic features and the low-level semantic features to obtain a quality evaluation score of the target distorted panoramic image.

9. An electronic device comprising at least one processor; and a memory connected in communication with the at least one processor; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Phrase-level text image generation method and system based on self-attention mechanism

    CN115587160A

  • Non-reference screen content image quality evaluation method based on multi-region feature fusion

    CN116403063A

  • Autism feature database system based on multi-modal fusion

    CN117079757A

  • Transform-based no-reference panoramic image quality evaluation method and system

    CN118014966A

  • Quality evaluation method and device for distorted panoramic image

    CN119006386A