Panoramic image quality evaluation method and device based on multi-modal semantic fusion
Through the multimodal semantic fusion method, combining global and local visual semantic features with text semantic features, the problem of lack of accuracy in the panoramic image quality evaluation method in the prior art is solved, and a more accurate and universal panoramic image quality evaluation is achieved.
Patent Information
- Application Number
- CN202510549970.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing panoramic image quality evaluation methods lack accuracy and cannot effectively extract key features in the viewport, resulting in inaccurate judgment of the local quality of panoramic images, affecting the universality and accuracy of the evaluation.
Using a multimodal semantic fusion method, the text semantic features and global visual semantic features are obtained through the text semantic feature extraction model and the global visual semantic feature extraction model, and input them into the multimodal global semantic fusion model for fusing to obtain multimodal global semantic features. At the same time, the target distortion panorama is divided into viewport processing, local visual semantic features are extracted, and multimodal local semantic fusion is performed with text semantic features to obtain multimodal local semantic features. Finally, the multimodal global semantic features and local semantic features are stitched and fused to obtain the quality evaluation score of the panoramic image.
Through cross-modal fusion, rich semantic interaction relationships are obtained, the model's feature representation ability and overall performance are improved, the accurate evaluation of panoramic image quality is enhanced, and the universality and accuracy of evaluation is improved.
Smart Images

Figure CN120070450A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of panoramic image quality evaluation, and more specifically, to a panoramic image quality evaluation method and device based on multimodal semantic fusion. Background Art
[0002] Panoramic images have attracted attention because they can display a spherical scene of 360×180 and provide an immersive visual experience, becoming one of the research hotspots in various application fields. However, in the processes of panoramic image acquisition, stitching, compression, etc., distortion or noise will inevitably be introduced, making it difficult to ensure the visual quality of panoramic images; at the same time, due to problems such as more projection transformation formats and larger field of view ranges in panoramic images, it makes objective quality evaluation more complex.
[0003] With the rapid development of visual language models, researchers have tried to describe image quality linguistically to make it consistent with human subjective feelings. Since the human eye has different attentions to different regions of panoramic images, existing evaluation methods do not accurately simulate the visual distribution of the human eye, resulting in inaccurate quality assessment of key regions of the image, and thus a viewport-based quality evaluation method has been derived.
[0004] However, the existing technology lacks precision when extracting perceptual characteristics based on viewports, cannot well extract key features in viewports, affects the judgment of the local quality of panoramic images, and thus leads to poor universality and low precision in panoramic image quality evaluation. Summary of the Invention
[0005] According to an embodiment of the present invention, a panoramic image quality evaluation scheme based on multimodal semantic fusion is provided. This scheme uses a large model to obtain text semantic information and efficiently fuses it with the visual semantic information of panoramic images, which can effectively improve the overall performance of the model.
[0006] In a first aspect of the present invention, a panoramic image quality evaluation method based on multimodal semantic fusion is provided. The method includes: Obtain a target distorted panoramic image; Input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; Input the target distorted panoramic image into a global visual semantic feature extraction model to obtain global visual semantic features; input the text semantic features and the global visual semantic features into a multimodal global semantic fusion model to obtain multimodal global semantic features; Perform viewport processing on the target distorted panoramic image to obtain a number of viewport images, and then input the number of viewport images into a local visual semantic extraction model for local visual feature extraction to obtain the local visual semantic features of each viewport; respectively input the local visual semantic features of each viewport and the text semantic features into a multimodal local semantic fusion model, output the local visual semantic features of a number of viewports, and then splice the local visual semantic features of all viewports to obtain multimodal local semantic features; Splice and fuse the multimodal global semantic features and the multimodal local semantic features to obtain the quality evaluation score of the target distorted panoramic image.
[0007] In a second aspect of the present invention, there is provided a panoramic image quality evaluation device based on multimodal semantic fusion. The device includes: An acquisition module for acquiring a target distorted panoramic image; A first feature extraction module for inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; A second feature extraction module for inputting the target distorted panoramic image into a global visual semantic extraction model to obtain global visual semantic features; inputting the text semantic features and the global visual semantic features into a multimodal global semantic fusion model to obtain multimodal global semantic features; A third feature extraction module for performing viewport processing on the target distorted panoramic image to obtain a number of viewport images, and then inputting the number of viewport images into a local visual semantic extraction model for local visual feature extraction to obtain the local visual semantic features of each viewport; respectively inputting the local visual semantic features of each viewport and the text semantic features into a multimodal local semantic fusion model, outputting the local visual semantic features of a number of viewports, and then splicing the local visual semantic features of a number of viewports to obtain multimodal local semantic features; An evaluation module for splicing and fusing the multimodal global semantic features and the multimodal local semantic features to obtain the quality evaluation score of the target distorted panoramic image.
[0008] In a third aspect of the present invention, there is provided an electronic device. The electronic device has at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention.
[0009] In a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the method of the first aspect of the present invention.
[0010] Compared with the prior art, the present invention has the following beneficial technical effects: By using visual semantic features and text semantic features for cross-modal fusion, rich semantic interaction relationships between modalities can be obtained, the model feature representation ability can be improved, and thus the universality and accuracy of the model can be improved.
[0011] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where: Figure 1 shows a flowchart of a panoramic image quality evaluation method based on multi-modal semantic fusion according to an embodiment of the present invention; Figure 2 shows a flowchart of a text semantic feature extraction model according to an embodiment of the present invention; Figure 3 shows a flowchart of a multi-modal global semantic fusion model according to an embodiment of the present invention; Figure 4 shows a flowchart of extracting local visual semantic features according to an embodiment of the present invention; Figure 5 shows a schematic diagram of a multi-modal global semantic fusion model according to an embodiment of the present invention; Figure 6 shows a schematic diagram of text semantic features according to an embodiment of the present invention; Figure 7 shows a schematic diagram of global visual semantic features according to an embodiment of the present invention; Figure 8 shows a schematic diagram of global text-enhanced semantic features according to an embodiment of the present invention; Figure 9 shows a schematic diagram of global cross-modal semantic fusion features according to an embodiment of the present invention; Figure 10 shows a schematic diagram of multi-modal global semantic features according to an embodiment of the present invention; Figure 11 shows a schematic diagram of local visual semantic features of each viewport according to an embodiment of the present invention; Figure 12 shows a schematic diagram of local text-enhanced semantic features according to an embodiment of the present invention; Figure 13Schematic diagram showing local cross-modal semantic fusion features according to an embodiment of the present invention; Figure 14 Schematic diagram showing local visual semantic features of each viewport according to an embodiment of the present invention; Figure 15 Schematic diagram showing multi-modal local semantic features according to an embodiment of the present invention; Figure 16 Schematic diagram showing the spliced fusion feature map of multi-modal global semantic features and multi-modal local semantic features according to an embodiment of the present invention; Figure 17 Block diagram showing a panoramic image quality evaluation device based on multi-modal semantic fusion according to an embodiment of the present invention; Figure 18 Block diagram showing an exemplary electronic device capable of implementing the embodiments of the present invention.
[0013] Wherein, 1800 is an electronic device, 1801 is a computing unit, 1802 is a ROM, 1803 is a RAM, 1804 is a bus, 1805 is an I / O interface, 1806 is an input unit, 1807 is an output unit, 1808 is a storage unit, and 1809 is a communication unit. Detailed implementation manners
[0014] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0016] In the present invention, through a text semantic feature extraction model, a global visual semantic feature extraction model, and a local visual semantic extraction model, text semantic features, global visual semantic features, and local visual semantic features are obtained; then, through a multimodal global semantic fusion model and a multimodal local semantic fusion model, multimodal global semantic features and multimodal local semantic features are obtained; finally, the multimodal global semantic features and multimodal local semantic features are concatenated to obtain the quality evaluation score of the target distorted panoramic image. In this way, key features in the viewport can be more comprehensively extracted, improving the universality and accuracy of panoramic image quality evaluation.
[0017] Figure 1 Fig. 4 shows a flowchart of a panoramic image quality evaluation method based on multimodal semantic fusion according to an embodiment of the present invention.
[0018] The method includes: S101. Obtain a target distorted panoramic image.
[0019] Specifically, in this embodiment, the publicly available panoramic image dataset CVIQ is used to evaluate the model proposed by the present invention. The dataset contains 528 distorted panoramic images, with a total of three different types of coding compression (JPEG, H.264 / AVC, and H.265 / HEVC). Each distorted image provides a subjective score MOS value, and the resolution of all images is 4096 × 2048.
[0020] S102. Input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features.
[0021] In this embodiment, as Figure 2 shown, the step of inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features includes: S201. Input the target distorted panoramic image into a text generation model to output a text description statement of the target distorted panoramic image.
[0022] Specifically, the text generation model is DepictQA (Depicted image Quality Assessment method, a multimodal large language model for image quality perception), which uses an MLLM (Multimodal Large Language Model) to provide a human-like, language-based description of the image quality.
[0023] As some alternative embodiments of this embodiment, it is assumed that the target distorted panoramic image is input into the DepictQA model, and the output text description statement is "The quantization distortion is evident, resulting in a loss of detail and color depth, which affects the images's overall clarity. It has obvious quantization distortion, retains more detail and color accuracy, making it bad to reflect the scene".
[0024] S202. Input the text description statement of the target distorted panoramic image into the Bert model to obtain text semantic features.
[0025] As some alternative embodiments of this embodiment, it is assumed that the text description statement output by the DepictQA model is input into the Bert model, and the output text semantic features are as Figure 6 shown.
[0026] Extracting the text description statement for the target distorted panoramic image can provide non-visual semantic information for the image and improve the model accuracy. The Bert model has powerful language understanding ability and context understanding ability, and can effectively capture long-distance dependency relationships in the text and obtain rich text semantic information.
[0027] S103. Input the target distorted panoramic image into the global visual semantic feature extraction model to obtain global visual semantic features; input the text semantic features and the global visual semantic features into the multi-modal global semantic fusion model to obtain multi-modal global semantic features.
[0028] In this embodiment, the global visual semantic feature extraction model (VMamba) sequentially includes a spatial transformation embedding module (Stem, Spatial Transformer Embedding Module), a first visual state space block (VSSBlock, Visual State Space Block), a first downsampling module (Down Sampling), a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module, and a fourth visual state space block.
[0029] Among them, the global visual semantic features output by the fourth visual state space block serve as the output of the global visual semantic feature extraction model, where the global visual semantic features are as Figure 7 shown.
[0030] Specifically, the first visual state space block successively includes a first LN layer, a first 2D selective scan block (SS2d Block, 2D-selective-scan Block), a second LN layer, and a third FFN (Feed Forward Network) layer; the first LN layer obtains the output of the spatial transformation embedding module, performs normalization, and then outputs it to the first 2D selective scan block for feature extraction to output a first extraction result; the output of the spatial transformation embedding module and the first extraction result are added together to obtain a first addition result; the first addition result is successively input into the second LN layer and the third FFN layer to output a first enhancement result; the first addition result and the first enhancement result are added together to obtain the output of the first visual state space block.
[0031] Specifically, the second visual state space block successively includes a third LN layer, a second 2D selective scan block, a fourth LN layer, and a fourth FFN layer; the third LN layer obtains the output of the first downsampling module, performs normalization, and then outputs it to the second 2D selective scan block for feature extraction to output a second extraction result; the output of the first downsampling module and the second extraction result are added together to obtain a second addition result; the second addition result is successively input into the fourth LN layer and the fourth FFN layer to output a second enhancement result; the second addition result and the second enhancement result are added together to obtain the output of the second visual state space block.
[0032] Specifically, the third visual state space block successively includes a fifth LN layer, a third 2D selective scan block, a sixth LN layer, and a fifth FFN layer; the fifth LN layer obtains the output of the second downsampling module, performs normalization, and then outputs it to the third 2D selective scan block for feature extraction to output a third extraction result; the output of the second downsampling module and the third extraction result are added together to obtain a third addition result; the third addition result is successively input into the sixth LN layer and the fifth FFN layer to output a third enhancement result; the third addition result and the third enhancement result are added together to obtain the output of the third visual state space block.
[0033] Specifically, the fourth visual state space block sequentially includes a seventh LN layer, a fourth 2D selective scanning block, an eighth LN layer, and a sixth FFN layer; the seventh LN layer obtains the output of the third downsampling module, performs normalization, and then outputs it to the fourth 2D selective scanning block for feature extraction, and outputs a fourth extraction result; the output of the third downsampling module and the fourth extraction result are added together to obtain a fourth addition result; the fourth addition result is sequentially input into the eighth LN layer and the sixth FFN layer, and a fourth enhancement result is output; the fourth addition result and the fourth enhancement result are added together to obtain the output of the fourth visual state space block.
[0034] VMamba can capture the global perception of an image for long-range dependence modeling, and at the same time can effectively capture the key features of a panoramic image and maintain a good linear complexity.
[0035] In this embodiment, the multi-modal global semantic fusion model sequentially includes a first self-attention module, a first cross-attention modal fusion module, and a first FFN layer.
[0036] In this embodiment, as Figure 3 、 5 shown, inputting the text semantic feature and the global visual semantic feature into the multi-modal global semantic fusion model to obtain a multi-modal global semantic feature includes: S301. Input the text semantic feature into the first self-attention module (SA, Self-Attention), and then add the output of the first self-attention module to the text semantic feature to obtain a global text enhanced semantic feature.
[0037] As some optional implementation manners of this embodiment, assume that the Figure 6 shown text semantic feature is used as the input of step S301, and the obtained global text enhanced semantic feature is as Figure 8 shown.
[0038] S302. Input the global text enhanced semantic feature and the global visual semantic feature into the first cross-attention modal fusion module (CA, Cross-Attention), and then add the output of the first cross-attention modal fusion module to the global text enhanced semantic feature to obtain a global cross-modal semantic fusion feature.
[0039] As some optional implementation manners of this embodiment, assume that the Figure 8 shown global text enhanced semantic feature is used as the input of step S302, and the obtained global cross-modal semantic fusion feature is as Figure 9 shown.
[0040] S303. Input the global cross-modal semantic fusion feature into the first FFN layer, and then add the output of the first FFN layer to the global cross-modal semantic fusion feature to obtain the multi-modal global semantic feature.
[0041] As some alternative implementation manners of this embodiment, assume that Figure 9 the global cross-modal semantic fusion feature shown is used as the input of step S303, and the obtained multi-modal global semantic feature is as Figure 10 shown.
[0042] The text semantics can further strengthen the representation ability of the text semantic features of the model through the self-attention mechanism, and obtain the interaction relationship within the modality; through cross-attention fusion with the visual semantic features, the interaction relationship between modalities can be obtained, realizing efficient semantic fusion; the mining of semantic information within and between modalities and the relationship learning can improve the accuracy of the model.
[0043] S104. Perform a viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then input the plurality of viewport images into a local visual semantic extraction model to extract local visual features, obtaining the local visual semantic features of each viewport (as Figure 11 shown); respectively input the local visual semantic features of each viewport and the text semantic features into a multi-modal local semantic fusion model, output the local visual semantic features of a plurality of viewports, and then splice the local visual semantic features of all viewports to obtain the multi-modal local semantic feature (as Figure 15 shown).
[0044] Specifically, the local visual semantic extraction model is ViT (Vision Transforme). The viewport acquisition method is to use FoV Selection (Field of View Selection) to obtain the viewport images of the target distorted panoramic image. Performing viewport segmentation on the distorted panoramic image conforms to the characteristics of human visual perception and helps to extract local visual semantic features; ViT performs excellently in traditional computer vision tasks and can effectively obtain the visual semantic features of the viewport.
[0045] In this embodiment, the multi-modal local semantic fusion model includes a plurality of multi-modal semantic fusion modules, and each multi-modal semantic fusion module corresponds to a viewport, and is used to take the local visual semantic feature of the corresponding viewport and the text semantic feature as inputs and output the local visual semantic feature of each viewport. Each multi-modal semantic fusion module sequentially includes a second self-attention module, a second cross-attention modality fusion module, and a second FFN layer.
[0046] In this embodiment, as Figure 4As shown, using the local visual semantic features of the corresponding viewport and the text semantic features as inputs, the local visual semantic features of each viewport are output, including: S401. Input the text semantic features into the second self-attention module, and then add the output of the second self-attention module to the text semantic features to obtain local text enhanced semantic features.
[0047] As some optional implementation manners of this embodiment, assume that Figure 6 the shown text semantic features are used as the input of step S401, and the obtained local text enhanced semantic features are as Figure 12 shown.
[0048] S402. Input the local text enhanced semantic features and the local visual semantic features of each viewport into the second cross-attention modality fusion module, and then add the output of the second cross-attention modality fusion module to the local text enhanced semantic features to obtain local cross-modal semantic fusion features.
[0049] As some optional implementation manners of this embodiment, assume that Figure 12 the shown local text enhanced semantic features are used as the input of step S402, and the obtained local cross-modal semantic fusion features are as Figure 13 shown.
[0050] S403. Input the local cross-modal semantic fusion features into the second FFN layer, and then add the output of the second FFN layer to the local cross-modal semantic fusion features to obtain the local visual semantic features of each viewport.
[0051] As some optional implementation manners of this embodiment, assume that Figure 13 the shown local cross-modal semantic fusion features are used as the input of step S403, and the obtained local visual semantic features of each viewport are as Figure 14 shown.
[0052] Performing cross-modal interaction and fusion on the text semantic features and the local visual semantic features can effectively mine the local semantic information within and between modalities, and effectively complement the multi-modal global semantic features, improving the overall performance of the model.
[0053] S105. Concatenate and fuse the multi-modal global semantic features and the multi-modal local semantic features to obtain the quality evaluation score of the target distorted panoramic image.
[0054] As some optional implementation manners of this embodiment, assume that Figure 10 the shown multi-modal global semantic features and the multi-modal local semantic features as Figure 15 described are concatenated and fused, and the output feature map is as shown in Figure 16.
[0055] The multi-modal global semantic features and multi-modal local semantic features are concatenated and fused, which not only fuses the semantic information of different modalities but also considers the complementary characteristics of global semantic information and local semantic information, and can improve the accuracy of the model.
[0056] As some alternative implementation manners of this embodiment, the evaluation metrics used are PLCC (Pearson linear correlation coefficient), SRCC (Spearman's Rank Correlation Coefficient), and RMSE (Root Mean Square Error). The higher the values of PLCC and SRCC and the smaller the value of RMSE, the better the performance of the model.
[0057] The calculation formula of PLCC is as follows:
[0058]
[0059]
[0060] Wherein, is the number of test images; is the index value of the number of test images; is the mean value of the subjective quality scores of all images in the prediction dataset; is the mean value of the quality scores predicted by the model; is the th subjective quality score provided by the dataset of the test image; is the quality score predicted by the model; is the linear consistency between the predicted image quality score and the quality score provided by the dataset.
[0061] The calculation formula of SRCC is as follows:
[0062] Wherein, is the monotonicity between the predicted quality score and the subjective score provided by the dataset.
[0063] The calculation formula of RMSE is as follows:
[0064] Wherein, RMSE is the error between the predicted score and the subjective score.
[0065] In this embodiment, the MOS value of the subjective score provided by the data set and the objective quality score obtained by the present invention are input into the PLCC, SRCC, and RMSE calculation formulas to obtain evaluation index values. To verify the effectiveness of the semantic fusion module and the feature fusion module of the present invention, the above two modules are respectively replaced with a simple splicing fusion module, and the evaluation index values are calculated. The results are shown below. It can be seen from the above results that the PLCC and SRCC of the present invention are closest to 1, and the RMSE is the smallest, indicating that the multi-modal feature fusion method of the present invention can effectively strengthen semantic features and improve feature representation ability, and can also achieve effective fusion of features, improve the model accuracy, and prove that the present invention can accurately evaluate the quality of panoramic images.
[0066] Table 1 PLCC SRCC RMSE The model of the present invention 0.9371 0.9351 3.9078 Replace the multimodal semantic fusion model used to obtain multimodal local semantic features with a splicing model 0.7921 0.7794 9.0412 Replace all multimodal semantic fusion models with splicing models 0.4861 0.4607 17.8314 According to the embodiments of the present invention, the present invention has the following advantages and effects compared with the prior art: (1) Cross-modal fusion is performed using global visual semantic features and text semantic features, which can obtain rich global semantic interaction relationships between modalities and improve the model feature representation ability.
[0067] (2) Cross-modal fusion is performed using local visual semantic features and text semantic features, which helps to strengthen local visual semantic features and obtain rich multi-modal local semantic features.
[0068] (3) The multi-modal global semantic information and multi-modal local semantic information complement each other, strengthening the deep fusion of semantic features and helping to improve the overall performance of the model.
[0069] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0070] The above is the introduction of the method embodiments. The following is a further description of the solution of the present invention through an apparatus embodiment having the same inventive concept as the method in the foregoing embodiments.
[0071] As Figure 17 shown, the apparatus 1700 includes: An acquisition module 1710, configured to acquire a target distorted panoramic image.
[0072] The first feature extraction module 1720 is configured to input the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features.
[0073] The second feature extraction module 1730 is configured to input the target distorted panoramic image into a global visual semantic feature extraction model to obtain global visual semantic features; input the text semantic features and the global visual semantic features into a multi-modal global semantic fusion model to obtain multi-modal global semantic features.
[0074] The third feature extraction module 1740 is configured to perform a viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then input the plurality of viewport images into a local visual semantic extraction model for local visual feature extraction to obtain local visual semantic features of each viewport; respectively input the local visual semantic features of each viewport and the text semantic features into a multi-modal local semantic fusion model, output local visual semantic features of a plurality of viewports, and then splice the local visual semantic features of the plurality of viewports to obtain multi-modal local semantic features.
[0075] The evaluation module 1750 is configured to splice and fuse the multi-modal global semantic features and the multi-modal local semantic features to obtain a quality evaluation score of the target distorted panoramic image.
[0076] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0077] In the technical solution of the present invention, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0078] According to an embodiment of the present invention, the present invention also provides an electronic device.
[0079] Figure 18 FIG. shows a schematic block diagram of an electronic device 1800 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0080] The electronic device 1800 includes a computing unit 1801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 1802 or computer programs loaded from a storage unit 1808 into a random access memory (RAM) 1803. In the RAM 1803, various programs and data required for the operation of the electronic device 1800 can also be stored. The computing unit 1801, the ROM 1802, and the RAM 1803 are connected to each other via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.
[0081] Multiple components in the electronic device 1800 are connected to the I / O interface 1805, including: an input unit 1806, such as a keyboard, a mouse, etc.; an output unit 1807, such as various types of displays, speakers, etc.; a storage unit 1808, such as a magnetic disk, an optical disk, etc.; and a communication unit 1809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1809 allows the electronic device 1800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0082] The computing unit 1801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1801 executes the various methods and processes described above, such as methods S101~S105. For example, in some embodiments, methods S101~S105 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1800 via the ROM 1802 and / or the communication unit 1809. When the computer program is loaded into the RAM 1803 and executed by the computing unit 1801, one or more steps of the methods S101~S105 described above can be executed. Alternatively, in other embodiments, the computing unit 1801 can be configured to execute methods S101~S105 in any other appropriate manner (e.g., by means of firmware).
[0083] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0085] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.
[0086] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A panoramic image quality assessment method based on multimodal semantic fusion, characterized in that: include: Obtain a distorted panoramic image of the target; Inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; Inputting the target distorted panoramic image into a global visual semantic feature extraction model to obtain a global visual semantic feature; inputting the text semantic feature and the global visual semantic feature into a multimodal global semantic fusion model to obtain a multimodal global semantic feature; Performing viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then inputting the plurality of viewport images into a local visual semantic extraction model to extract local visual features to obtain local visual semantic features of each viewport; The local visual semantic features and text semantic features of each viewport are input into the multimodal local semantic fusion model respectively, the local visual semantic features of several viewports are output, and then the local visual semantic features of all viewports are spliced to obtain the multimodal local semantic features; The multimodal global semantic features and the multimodal local semantic features are spliced and fused to obtain a quality evaluation score of the target distorted panoramic image.
2. The method according to claim 1, characterized in that The step of inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features includes: Inputting the target distorted panoramic image into a text generation model, and outputting a text description sentence of the target distorted panoramic image; The text description sentence of the target distorted panoramic image is input into the Bert model to obtain text semantic features.
3. The method according to claim 1, characterized in that The global visual semantic feature extraction model includes, in sequence, a spatial transformation embedding module, a first visual state space block, a first downsampling module, a second visual state space block, a second downsampling module, a third visual state space block, a third downsampling module and a fourth visual state space block; The global visual semantic features output by the fourth visual state space block serve as the output of the global visual semantic feature extraction model.
4. The method according to claim 1, characterized in that The multimodal global semantic fusion model includes the first self-attention module, the first cross-attention modality fusion module, and the first FFN layer in sequence.
5. The method according to claim 4, characterized in that The step of inputting the text semantic feature and the global visual semantic feature into a multimodal global semantic fusion model to obtain a multimodal global semantic feature comprises: Inputting the text semantic feature into a first self-attention module, and then adding the output of the first self-attention module to the text semantic feature to obtain a global text enhanced semantic feature; Inputting the global text enhancement semantic feature and the global visual semantic feature into a first cross-attention modality fusion module, and then adding the output of the first cross-attention modality fusion module to the global text enhancement semantic feature to obtain a global cross-modality semantic fusion feature; The global cross-modal semantic fusion feature is input into the first FFN layer, and then the output of the first FFN layer is added to the global cross-modal semantic fusion feature to obtain a multimodal global semantic feature.
6. The method according to claim 1, characterized in that The multimodal local semantic fusion model includes a plurality of multimodal semantic fusion modules, and each multimodal semantic fusion module corresponds to a viewport, and is used to take the local visual semantic features of the corresponding viewport and the text semantic features as input, and output the local visual semantic features of each viewport; Each multimodal semantic fusion module includes a second self-attention module, a second cross-attention modality fusion module, and a second FFN layer in sequence.
7. The method according to claim 6, characterized in that Taking the local visual semantic features of the corresponding viewport and the text semantic features as input, the local visual semantic features of each viewport are output, including: Inputting the text semantic feature into a second self-attention module, and then adding the output of the second self-attention module to the text semantic feature to obtain a local text enhanced semantic feature; Inputting the local text enhancement semantic feature and the local visual semantic feature of each viewport into a second cross-attention modality fusion module, and then adding the output of the second cross-attention modality fusion module to the local text enhancement semantic feature to obtain a local cross-modality semantic fusion feature; The local cross-modal semantic fusion feature is input into the second FFN layer, and then the output of the second FFN layer is added to the local cross-modal semantic fusion feature to obtain the local visual semantic feature of each viewport.
8. A panoramic image quality assessment device based on multimodal semantic fusion, characterized in that: include: An acquisition module, used for acquiring a distorted panoramic image of a target; A first feature extraction module, used for inputting the target distorted panoramic image into a text semantic feature extraction model to obtain text semantic features; A second feature extraction module is used to input the target distorted panoramic image into a global visual semantic feature extraction model to obtain a global visual semantic feature; input the text semantic feature and the global visual semantic feature into a multimodal global semantic fusion model to obtain a multimodal global semantic feature; A third feature extraction module is used to perform viewport processing on the target distorted panoramic image to obtain a plurality of viewport images, and then input the plurality of viewport images into a local visual semantic extraction model to extract local visual features to obtain local visual semantic features of each viewport; The local visual semantic features and text semantic features of each viewport are respectively input into the multimodal local semantic fusion model, and the local visual semantic features of several viewports are output. The local visual semantic features of several viewports are then spliced to obtain multimodal local semantic features. An evaluation module is used to splice and fuse the multimodal global semantic features and the multimodal local semantic features to obtain a quality evaluation score of the target distorted panoramic image.
9. An electronic device comprising at least one processor; and a memory connected in communication with the at least one processor; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Text-guided image compression noise removal method based on multi-modal feature fusion
CN114283080A
No-reference image quality evaluation method based on image features and semantic description
CN118608467A
Quality evaluation method and device for distorted panoramic image
CN119006386A
Video text cross-modal retrieval method and device
CN119166853A
Cited By
Multi-modal panoramic image blind quality evaluation method and system based on AI generation description
CN120356071A
AI-based multi-modal panoramic image blind quality evaluation method and system based on description generation
CN120356071B