Method and device for predicting visually salient regions of multi-view scenes, and electronic device
Patent Information
- Application Number
- CN202411498575.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2044-10-25
AI Technical Summary
[0004]本发明提供一种多视角场景的视觉显著区域预测方法、装置及电子设备,以解决现有技术中,视觉注意力模型在捕获多视角场景中眼动行为的准确性不足的问题
[0015] The present invention provides a method, device, and electronic device for predicting visual salient regions in multi-view scenes. By fusing multi-view images and depth information of the scene to be predicted, it realizes the prediction of human visual attention in multi-view scenes, providing more accurate prediction results for visual salient regions in 3D display technology.
Smart Images

Figure CN119672791B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, and electronic device for predicting visually salient regions in multi-view scenes. Background Technology
[0002] 3D light field display technology offers a highly immersive and realistic experience. As technology advances, people's demands for the comfort and realism of 3D content continue to rise. Accurately predicting salient areas in 3D displays is crucial for improving the viewing experience and enhancing the realism of display devices.
[0003] While progress has been made in the study of visually salient regions based on RGB-D, most existing methods are limited to single-view analysis. In multi-view scenes, traditional visual attention models struggle to effectively extract and fuse features from 3D scenes, limiting their performance in practical applications. Current research largely focuses on multi-view visual attention prediction models that use color images or color images combined with depth images as input, attempting to capture visual information within the scene. However, these approaches cannot adequately provide the multi-view feature information required for 3D scenes, limiting the accuracy of visual attention models in capturing key features of multi-view scenes. Summary of the Invention
[0004] This invention provides a method, device, and electronic device for predicting visual salient regions in multi-view scenes, in order to solve the problem that the accuracy of visual attention models in capturing eye movement behavior in multi-view scenes is insufficient in the prior art.
[0005] This invention provides a method for predicting visually salient regions in multi-view scenes, comprising the following steps: Obtain views of the scene to be predicted from different angles, including the left view, middle view, and right view; The left and right views are mapped to the viewpoints corresponding to the middle view to obtain the mapped left and right views; The mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view are used as inputs and fed into the human eye gaze point prediction model to obtain the visual salient region prediction result of the scene to be predicted.
[0006] According to the visual salient region prediction method for multi-view scenes provided by the present invention, the human eye gaze point prediction model consists of a parallel multi-view encoder and a fused multi-view decoder; the parallel multi-view encoder includes: A first parallel encoder is used to obtain the features of the mapped left view; The second parallel encoder is used to acquire features of the mid-view and the depth map corresponding to the mid-view; A third parallel encoder is used to obtain the features of the mapped right view; The fusion multi-view decoder is used to fuse various features and perform decoding operations to generate a visually salient region prediction result for the scene to be predicted.
[0007] According to the visual salient region prediction method for multi-view scenes provided by the present invention, the first parallel encoder, the second parallel encoder, and the third parallel encoder have the same weight.
[0008] According to the visual salient region prediction method for multi-view scenes provided by the present invention, the step of mapping the left view and the right view to the viewpoint corresponding to the middle view to obtain the mapped left view and the mapped right view includes: Based on the camera's internal parameters, the depth map and color map corresponding to the left view, and the depth map and color map corresponding to the right view, the three-dimensional point cloud map of the left view and the three-dimensional point cloud map of the right view are reconstructed. Based on the camera extrinsic parameter matrix, the 3D point cloud map of the left view and the 3D point cloud map of the right view are transformed into the camera coordinate system; In the camera coordinate system, the 3D point cloud map of the left view and the 3D point cloud map of the right view are projected onto a 2D plane using the camera intrinsic parameter matrix and the camera extrinsic parameter matrix, to obtain the mapped left view and the mapped right view.
[0009] The visual salient region prediction method for any multi-view scene provided by the present invention further includes: Using a multi-view scene salient region dataset, the initial human eye gaze prediction model was trained to obtain the human eye gaze prediction model.
[0010] The visual salient region prediction method for multi-view scenes provided by the present invention is characterized in that, the step of training an initial human eye fixation point prediction model using a multi-view scene salient region dataset to obtain the human eye fixation point prediction model includes: During the training process, a corresponding loss function is configured for each layer of the decoder in the initial human eye gaze prediction model, and the weight of the loss term of each decoder layer increases as the decoding layer is deepened.
[0011] The present invention also provides a visually salient region prediction device for multi-view scenes, comprising: The view acquisition module is used to acquire views of the scene to be predicted from different angles, including the left view, the middle view, and the right view. A view processing module is used to map the left view and the right view to the viewpoint corresponding to the middle view, so as to obtain the mapped left view and the mapped right view; The region prediction module is used to input the mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view into the human eye gaze point prediction model to obtain the visual salient region prediction result of the scene to be predicted.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the visual salient region prediction method for multi-view scenes as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the visual salient region prediction method for multi-view scenes as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the visual salient region prediction method for multi-view scenes as described above.
[0015] The present invention provides a method, device, and electronic device for predicting visual salient regions in multi-view scenes. By fusing multi-view images and depth information of the scene to be predicted, it realizes the prediction of human visual attention in multi-view scenes, providing more accurate prediction results for visual salient regions in 3D display technology. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is one of the flowcharts illustrating the visually salient region prediction method for multi-view scenes provided by the present invention.
[0018] Figure 2 This is the second flowchart of the visual salient region prediction method for multi-view scenes provided by the present invention.
[0019] Figure 3 This is a schematic diagram illustrating the principle of the human eye fixation prediction model provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the visual salient region prediction device for multi-view scenes provided by the present invention.
[0021] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such descriptions can be used interchangeably where appropriate to allow embodiments to be implemented in a sequence other than that illustrated or described in this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved. The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or discussed mutual coupling, direct coupling, or communication connection may be through some interface, and the indirect coupling or communication connection between units may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed in multiple circuit units. Some or all of the units can be selected to achieve the purpose of the solution in this application according to actual needs.
[0024] In embodiments of this invention, the multi-view scene salient region dataset uses an open-source dataset as the scene map. A comprehensive visual salient region dataset is constructed by collecting eye-tracking data from different subjects. The viewpoint mapper utilizes camera parameters for geometric transformation to achieve projection and alignment of multi-angle views. The visual saliency prediction framework employs a unique encoder-decoder network structure, extracting multiple features through parallel branches and fusing these features to generate accurate salient region predictions.
[0025] In this embodiment of the invention, the encoder section is specifically designed with three parallel branches to extract image features from the left, right, and middle viewpoints, respectively. The middle viewpoint branch simultaneously captures composite image features from both the color and depth images. The decoder section then fuses these image features, and during training, a deep supervision mechanism is introduced to improve prediction accuracy.
[0026] Furthermore, the visually salient region prediction results in this invention can be used for positioning in post-processing such as depth compression and super-resolution of 3D displays, achieving a high degree of personalization of visually salient regions, saving costs while ensuring display quality.
[0027] The following is combined with Figures 1-5 The specific contents of this invention are described below.
[0028] Figure 1 This is one of the flowcharts of a method for predicting visually salient regions in a multi-view scene provided by the present invention, including steps S101-S103.
[0029] Step S101: Obtain views of the scene to be predicted from different angles, including the left view, the middle view, and the right view.
[0030] Step S102: Map the left view and right view to the viewpoint corresponding to the middle view to obtain the mapped left view and the mapped right view.
[0031] Specifically, based on the camera's internal parameters and the depth maps of the left and right views, a 2D image from a certain perspective can be reconstructed into a 3D point cloud. This 3D point cloud uses the camera's extrinsic and internal parameters for projection transformation, mapping the original viewpoint to another viewpoint. In this process, not only is the viewpoint itself transformed, but the image features from the original viewpoint are also effectively mapped to the new viewpoint, ensuring the integrity and consistency of the information after the viewpoint transformation.
[0032] Step S103: The mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view are input into the human eye gaze point prediction model to obtain the prediction result of the visual salient region of the scene to be predicted.
[0033] Figure 2This is the second flowchart of a method for predicting visually salient regions in multi-view scenes provided by the present invention.
[0034] In one possible implementation, such as Figure 2 As shown, in step S102, the left view and the right view are mapped to the viewpoint corresponding to the middle view to obtain the mapped left view and the mapped right view, including steps S201-S203.
[0035] Step S201: Based on the camera's internal parameters, the depth map and color map corresponding to the left view, and the depth map and color map corresponding to the right view, reconstruct the 3D point cloud map of the left view and the 3D point cloud map of the right view.
[0036] Specifically, for any one of the left and right views, the two-dimensional pixel information on the image is first mapped to a three-dimensional point in the camera coordinate system. This mapping process uses the inverse of the camera intrinsic parameter matrix K, which contains intrinsic parameters such as the camera's focal length and optical center. For the normalized image coordinates (x, y) in the depth map corresponding to the image, we can calculate its position in the camera coordinate system using the depth information Zc. Inverse perspective projection then maps it to a three-dimensional point in the camera coordinate system. The mathematical expression for this process is:
[0037] in It is the inverse of the camera intrinsic matrix K. Next, the 3D points in the camera coordinate system are mapped to the world coordinate system using the camera extrinsic matrix (rotation matrix R and translation vector t).
[0038] Step S202: Based on the camera extrinsic parameter matrix, convert the 3D point cloud map of the left view and the 3D point cloud map of the right view into the camera coordinate system.
[0039] Specifically, this involves mapping 3D points in the camera coordinate system to the world coordinate system. The mathematical expression for this is:
[0040] The coordinates in the world coordinate system are: Through the above steps, we obtained the 3D coordinates of each pixel in the depth map in the world coordinate system. By combining the 3D points mapped from each pixel in the world coordinate system, we construct a 3D point cloud map.
[0041] Step S203: In the camera coordinate system, the 3D point cloud map of the left view and the 3D point cloud map of the right view are projected onto a 2D plane through the camera intrinsic parameter matrix and the camera extrinsic parameter matrix to obtain the mapped left view and the mapped right view.
[0042] Specifically, suppose we have a matrix P containing a 3D point cloud, where each column represents the coordinates of a 3D point. We also have a camera intrinsic matrix K and a camera extrinsic matrix, the latter containing a rotation matrix R and a translation vector t. The projection process can be expressed by the following formula:
[0043] Where p is the coordinate of the projected two-dimensional point. This means transforming the 3D point cloud into the camera coordinate system using the camera extrinsic matrix. Finally, through the projection operation of the camera intrinsic matrix K, we project the 3D point cloud onto a 2D plane, obtaining the mapped left and right views.
[0044] Since this process consists of rotation and translation, it is a rigid transformation, thus preserving the relative position and shape of the point cloud without geometrically changing its shape. Therefore, the newly generated 2D image retains the feature information from the previous viewpoint. The above steps convert the features of the 3D scene into a 2D image, which serves as the input for the next eye gaze prediction network.
[0045] In one possible implementation, the human eye gaze prediction model consists of a parallel multi-view encoder and a fused multi-view decoder; the parallel multi-view encoder includes: The first parallel encoder is used to obtain the features of the mapped left view; The second parallel encoder is used to obtain features of the mid-view and the corresponding depth map; The third parallel encoder is used to obtain the features of the mapped right view; A multi-view decoder is used to fuse various features and perform decoding operations to generate visually salient region prediction results for the scene to be predicted.
[0046] In embodiments of the present invention, the human eye gaze prediction model is an encoder-decoder network structure. This network consists of two parts: a parallel multi-view encoder and a fused multi-view decoder. The parallel multi-view encoder receives color images from three viewpoints and a depth image from the middle viewpoint as input, thus comprehensively considering the three-dimensional structure and depth information of the scene during feature extraction.
[0047] like Figure 3 As shown, this section consists of three parallel encoders, the first of which is a parallel encoder. , Third parallel encoder The components represent the feature extraction processes for the left, middle, and right views, respectively. Specifically, the encoder... and The inputs are the left view and the right view. and the view on the right , and encoder The input is the middle view. and its corresponding depth map The three encoders not only share the same network structure but also their weights. Finally, the outputs of these three encoders are combined into a single multi-feature vector. It can be obtained through the following methods: ; ]
[0048] Multiple feature vectors v can capture key information from different perspectives. Through this parallel multi-view encoding method, complementary information from different perspectives can be effectively extracted and fused, enabling more accurate perception of depth and understanding of the scene.
[0049] In the multi-view decoder fusion, the feature maps of the three encoded channels are concatenated to form a comprehensive feature representation. In this way, we achieve global integration of multi-view information, comprehensively consider the geometric structure and content features of the scene, and better understand the semantics and structure of the scene.
[0050] Furthermore, the decoder uses skip connections to concatenate the encoder's feature map with the upsampled feature map to enhance feature fusion and preserve spatial details, thereby improving the network's depth perception capability for multi-view scenes.
[0051] In one possible implementation, the first parallel encoder, the second parallel encoder, and the third parallel encoder have the same weight.
[0052] In one possible implementation, the method for predicting visual salient regions in a multi-view scene further includes: training an initial human eye gaze prediction model using a multi-view scene salient region dataset to obtain a human eye gaze prediction model.
[0053] The following describes in detail the process of training the initial human eye fixation prediction model in an embodiment of the present invention.
[0054] For example, using the DTU dataset, 200 scenes were selected. Each large scene in the original dataset can be divided into 2-3 smaller scenes, and each smaller scene was analyzed from three perspectives: left, middle, and right, ensuring the diversity and representativeness of the dataset. The eye gaze regions in each scene were recorded by 20 subjects clicking on the images with their mice, resulting in a multi-view scene salient region dataset. This process ensures that the constructed multi-view scene salient region dataset not only comprehensively covers the gaze points of different observers but also provides rich and high-quality visual information for visual salient region prediction through precise eye-tracking data.
[0055] In one possible implementation, an initial human gaze prediction model is trained using a multi-view scene salient region dataset to obtain a human gaze prediction model. This includes: during the training process, configuring a corresponding loss function for each layer of the decoder in the initial human gaze prediction model, with the weight of the loss term of each decoder layer increasing as the decoding layer progresses.
[0056] Specifically, a corresponding loss function is set for each layer of the decoder. The mean squared error function is used, combined with the correlation coefficient as the final loss function. The loss terms of different decoding layers have different weights, which gradually increase as the decoding layers deepen. This deep supervision mechanism aims to better guide the model to learn multi-level, multi-scale feature representations. This deep supervision mechanism helps alleviate the gradient vanishing problem and improves the stability of the human gaze prediction model. The specific loss function is shown in the following formula:
[0057] in, The loss function is set for each layer. P is the gaze point predicted by the model, Q is the actual human gaze point, and N is the number of samples.
[0058] After training the human gaze prediction model, 10% of the data in a multi-view scene salient region dataset was selected as the test set to verify its effectiveness. The results show that the method can accurately predict human gaze information in 3D scenes, especially when the depth of field is large. This method comprehensively considers depth and multi-view information to accurately predict salient regions in the scene. Furthermore, the method exhibits strong robustness and is rarely affected by high-contrast edges and complex backgrounds.
[0059] By employing the above method, the present invention can achieve the following beneficial effects: 1. By fusing multi-view images and depth information of the scene to be predicted, the prediction of human visual attention in multi-view scenes is realized, providing more accurate prediction results of visually salient regions for 3D display technology.
[0060] 2. A corresponding loss function is set for each layer of the decoder, with different weights for the loss terms in different decoding layers. The weights gradually increase as the decoding layers deepen. This deep supervision mechanism aims to better guide the human eye gaze prediction model to learn multi-level, multi-scale feature representations. Through this multi-level deep supervision mechanism, the gradient vanishing problem can be alleviated, and network stability can be improved.
[0061] The visual salient region prediction device for multi-view scenes provided by the present invention will be described below. The visual salient region prediction device for multi-view scenes described below can be referred to in correspondence with the visual salient region prediction method for multi-view scenes described above.
[0062] Figure 4 A schematic diagram of the structure of a visually salient region prediction device for multi-view scenes provided by the present invention includes the following parts: The view acquisition module 410 is used to acquire views of the scene to be predicted from different angles, including the left view, the middle view and the right view.
[0063] The view processing module 420 is used to map the left view and the right view to the viewpoint corresponding to the middle view, so as to obtain the mapped left view and the mapped right view.
[0064] The region prediction module 430 is used to input the mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view into the human eye gaze point prediction model to obtain the visual salient region prediction result of the scene to be predicted.
[0065] In one possible implementation, the human eye gaze prediction model consists of a parallel multi-view encoder and a fused multi-view decoder; the parallel multi-view encoder includes: The first parallel encoder is used to obtain the features of the mapped left view; The second parallel encoder is used to obtain features of the mid-view and the corresponding depth map; The third parallel encoder is used to obtain the features of the mapped right view; A multi-view decoder is used to fuse various features and perform decoding operations to generate visually salient region prediction results for the scene to be predicted.
[0066] In one possible implementation, the first parallel encoder, the second parallel encoder, and the third parallel encoder have the same weight.
[0067] In one possible implementation, the view processing module 420 is specifically used for: Based on the camera's internal parameters, the depth map and color map corresponding to the left view, and the depth map and color map corresponding to the right view, reconstruct the 3D point cloud map of the left view and the 3D point cloud map of the right view. Based on the camera extrinsic parameter matrix, the 3D point cloud map of the left view and the 3D point cloud map of the right view are transformed into the camera coordinate system; In the camera coordinate system, the 3D point cloud maps of the left and right views are projected onto a 2D plane using the camera intrinsic and extrinsic parameter matrices, resulting in the mapped left and right views.
[0068] In one possible implementation, the visual salient region prediction device for multi-view scenes further includes: a network training module for training an initial human eye gaze prediction model using a multi-view scene salient region dataset to obtain a human eye gaze prediction model.
[0069] In one possible implementation, the network training module is specifically configured during the training process to configure a corresponding loss function for each layer of the decoder in the initial human eye gaze prediction model, with the weight of the loss term of each decoder layer increasing as the decoding layer progresses.
[0070] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a method for predicting visually salient regions in a multi-view scene, the method including: Obtain views of the scene to be predicted from different angles, including the left view, middle view, and right view; Map the left and right views to the viewpoints corresponding to the middle view to obtain the mapped left and right views; The mapped left view, mapped right view, middle view, and the corresponding depth map of the middle view are used as inputs and fed into the human eye gaze point prediction model to obtain the prediction results of the visual salient region of the scene to be predicted.
[0071] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0072] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a method for predicting visually salient regions in a multi-view scene provided by the methods described above. The method includes: Obtain views of the scene to be predicted from different angles, including the left view, middle view, and right view; Map the left and right views to the viewpoints corresponding to the middle view to obtain the mapped left and right views; The mapped left view, mapped right view, middle view, and the corresponding depth map of the middle view are used as inputs and fed into the human eye gaze point prediction model to obtain the prediction results of the visual salient region of the scene to be predicted.
[0073] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for predicting visually salient regions in a multi-view scene provided by the methods described above, the method comprising: Obtain views of the scene to be predicted from different angles, including the left view, middle view, and right view; Map the left and right views to the viewpoints corresponding to the middle view to obtain the mapped left and right views; The mapped left view, mapped right view, middle view, and the corresponding depth map of the middle view are used as inputs and fed into the human eye gaze point prediction model to obtain the prediction results of the visual salient region of the scene to be predicted.
[0074] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0075] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting visually salient regions in multi-view scenes, characterized in that, include: Obtain views of the scene to be predicted from different angles, including the left view, middle view, and right view; Mapping the left and right views onto the viewpoint corresponding to the middle view to obtain the mapped left and right views specifically includes: reconstructing the 3D point cloud map of the left view and the 3D point cloud map of the right view based on camera intrinsic parameters, the depth map and color map corresponding to the left view, and the depth map and color map corresponding to the right view; transforming the 3D point cloud map of the left view and the 3D point cloud map of the right view into the camera coordinate system based on the camera extrinsic parameter matrix; and projecting the 3D point cloud map of the left view and the 3D point cloud map of the right view onto a 2D plane in the camera coordinate system using the camera intrinsic parameter matrix and the camera extrinsic parameter matrix to obtain the mapped left and right views. The mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view are used as inputs and fed into the human eye gaze point prediction model to obtain the visual salient region prediction result of the scene to be predicted. The human eye gaze prediction model consists of a parallel multi-view encoder and a fused multi-view decoder; the parallel multi-view encoder includes: A first parallel encoder is used to obtain the features of the mapped left view; The second parallel encoder is used to obtain the stitched features of the middle view and the depth map corresponding to the middle view; A third parallel encoder is used to obtain the features of the mapped right view; The fusion multi-view decoder is used to fuse various features and perform decoding operations to generate a visually salient region prediction result for the scene to be predicted.
2. The method for predicting visually salient regions in multi-view scenes according to claim 1, characterized in that, The first parallel encoder, the second parallel encoder, and the third parallel encoder have the same weight.
3. The method for predicting visually salient regions in multi-view scenes according to claim 1 or 2, characterized in that, Also includes: Using a multi-view scene salient region dataset, the initial human eye gaze prediction model was trained to obtain the human eye gaze prediction model.
4. The method for predicting visually salient regions in multi-view scenes according to claim 3, characterized in that, The process of training an initial human gaze point prediction model using a multi-view scene salient region dataset to obtain the human gaze point prediction model includes: During the training process, a corresponding loss function is configured for each layer of the decoder in the initial human eye gaze prediction model, and the weight of the loss term of each decoder layer increases as the decoding layer is deepened.
5. A device for predicting visually salient regions in a multi-view scene, characterized in that, include: The view acquisition module is used to acquire views of the scene to be predicted from different angles, including the left view, the middle view, and the right view. The view processing module is used to map the left and right views onto the viewpoint corresponding to the middle view to obtain the mapped left and right views. Specifically, it includes: reconstructing the 3D point cloud map of the left view and the 3D point cloud map of the right view based on camera intrinsic parameters, the depth map and color map corresponding to the left view, and the depth map and color map corresponding to the right view; transforming the 3D point cloud map of the left view and the 3D point cloud map of the right view into the camera coordinate system based on the camera extrinsic parameter matrix; and projecting the 3D point cloud map of the left view and the 3D point cloud map of the right view onto a two-dimensional plane in the camera coordinate system through the camera intrinsic parameter matrix and the camera extrinsic parameter matrix to obtain the mapped left and right views. The region prediction module is used to input the mapped left view, the mapped right view, the middle view, and the depth map corresponding to the middle view into the human eye gaze point prediction model to obtain the visual salient region prediction result of the scene to be predicted. The human eye gaze prediction model consists of a parallel multi-view encoder and a fused multi-view decoder; the parallel multi-view encoder includes: A first parallel encoder is used to obtain the features of the mapped left view; The second parallel encoder is used to obtain the stitched features of the middle view and the depth map corresponding to the middle view; A third parallel encoder is used to obtain the features of the mapped right view; The fusion multi-view decoder is used to fuse various features and perform decoding operations to generate a visually salient region prediction result for the scene to be predicted.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the visual salient region prediction method for multi-view scenes as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the visual salient region prediction method for multi-view scenes as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the visual salient region prediction method for multi-view scenes as described in any one of claims 1 to 4.