360-degree panoramic image saliency target detection method and device based on icosahedron subdivision
Through the panoramic image significance object detection method based on icosahedral subdivision, the multi-level feature fusion and attention mechanism are used to solve the problems of low accuracy and distortion in panoramic image detection, and more efficient panoramic image feature extraction and detection are achieved.
Patent Information
- Application Number
- CN202510277684.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has low accuracy in panoramic image significance object detection, and there are problems of spherical plane mapping distortion and boundary discontinuity.
The 360-degree panoramic image significance object detection method based on icosahedral subdivision is adopted, and multi-level global and local features are obtained through equiregular columnar projection and icosahedral subdivision projection, and feature fusion is performed through adaptive fusion and attention mechanism.
It improves the accuracy of the detection of the significance target of panoramic images, avoids the problems of image distortion and boundary discontinuity, and can capture the characteristics of panoramic images more comprehensively and effectively.
Smart Images

Figure CN120125845A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a 360-degree panoramic image salient object detection method and device based on icosahedron subdivision. Background Art
[0002] Traditional image salient object detection methods extract spatial salient object information from the image plane. These methods consider factors such as image texture, frequency, and color. And with the development of deep learning, network models based on deep learning can fully extract salient features in the image spatial domain, and some salient object detection algorithms for the plane image spatial domain have achieved excellent results. In some application scenarios of panoramic images or videos, due to the characteristics of high resolution and large data volume of panoramic content, there are huge challenges in its real-time processing, encoding, and transmission. Some algorithms have been designed to address such challenges based on human visual characteristics, including the design of salient object detection algorithms for panoramic images. However, there are inherent problems when directly applying a two-dimensional image salient object detection model to panoramic image content, such as spherical-to-planar mapping distortion and discontinuous boundaries. Therefore, there is room for improvement in the salient object detection solution for panoramic content.
[0003] Regarding the problems faced by salient object detection in panoramic images, although there are already some related technologies, the accuracy of salient detection of some existing related technologies is relatively low. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a 360-degree panoramic image salient object detection method and device based on icosahedron subdivision.
[0005] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] The present invention provides a 360-degree panoramic image salient object detection method based on icosahedron subdivision, including:
[0007] Obtain a 360-degree panoramic image to be detected;
[0008] Perform equirectangular cylindrical projection and icosahedron subdivision projection on the to-be-detected 360-degree panoramic image respectively to obtain a panoramic image in the form of ERP projection and a panoramic image in the form of icosahedron projection respectively;
[0009] Input the panoramic image based on the ERP projection form and the panoramic image based on the icosahedron projection form into the trained saliency object detection network model. The trained saliency object detection network model obtains multi-level global features based on the panoramic image based on the ERP projection form, and obtains multi-level local features based on the panoramic image based on the icosahedron projection form. Perform adaptive fusion on the multi-level global features and the multi-level local features to obtain multi-level fusion features, and perform feature fusion on the multi-level fusion features based on the attention mechanism to obtain a saliency object result map.
[0010] The present invention also provides a 360-degree panoramic image saliency object detection device based on icosahedron subdivision, including a processor, a communication interface, a memory, and a communication bus. It is characterized in that the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0011] The memory is used to store computer programs;
[0012] When the processor is used to execute the program stored in the memory, it realizes the steps of the above-mentioned 360-degree panoramic image saliency object detection method based on icosahedron subdivision.
[0013] Compared with the prior art, the beneficial effects of the present invention are:
[0014] 1) In view of the rich scene content characteristics of panoramic images, the present invention extracts saliency information in the scene from different spatial scales, avoiding the problem of ignoring possible local detail features in the scene;
[0015] 2) The present invention adopts a planar projection method based on icosahedron subdivision, which makes up for the information loss and boundary discontinuity problems caused by traditional equidistant cylindrical projection and cube projection methods, and will not cause image distortion;
[0016] 3) The present invention performs adaptive fusion on the global features and local features obtained according to different projection methods, fully considering the advantages and disadvantages between different projection features. It can capture the global object information in the panoramic image while reducing the impact of the distortion caused by equidistant projection on the algorithm, and pay attention to the feature details of the panoramic content based on icosahedron subdivision projection, so as to obtain more comprehensive and effective panoramic image features, thereby improving the accuracy of subsequent saliency object detection;
[0017] 4) By performing feature fusion on multi-level fusion features based on the attention mechanism, the present invention can introduce spatial attention and channel attention mechanisms to perform refined processing on multi-level fusion features, highlight the most significant features of the feature channels while suppressing noise, adaptively select the advantages of each level of features for fusion, effectively improve the expression ability of multi-level features, and thus improve the accuracy of subsequent salient object detection.
[0018] The following will further elaborate on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0019] Figure 1 is a schematic flowchart of a 360-degree panoramic image salient object detection method based on icosahedron subdivision provided by an embodiment of the present invention;
[0020] Figure 2 is a schematic diagram of the principle of equirectangular projection of a 360° panoramic image;
[0021] Figure 3 is a planar representation of a 360° panoramic image based on icosahedron subdivision projection;
[0022] Figure 4 is a schematic diagram of the structure and processing flow of a salient object detection network model provided by an embodiment of the present invention;
[0023] Figure 5 is a schematic diagram of the structure of a feature pyramid module provided by an embodiment of the present invention;
[0024] Figure 6 is an exemplary schematic diagram of the structure of a multi-projection feature fusion module provided by an embodiment of the present invention;
[0025] Figure 7 is an exemplary schematic diagram of the structure of a multi-level feature fusion module provided by an embodiment of the present invention. Detailed Embodiments
[0026] The following further describes the present invention in detail in conjunction with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0027] The ideas of existing solutions can be roughly divided into two categories. One is to use a projection method with less distortion in the process of realizing the projection of spherical images onto planes, but the inventors found that this solution still leads to the problem of discontinuity of image boundaries. Specifically, since panoramic content usually presents the full scene perspective in a spherical manner, it is necessary to convert the panoramic image from a spherical projection format to a two-dimensional plane format before performing the processing task. The panoramic projection methods recommended by the international standard panoramic media format (OmnidirectionalMedia Format, OMAF) include equidistant cylindrical projection (ERP) and cube projection (Cube). Among them, equidistant cylindrical projection will produce projection distortion in the area far away from the equator of the spherical surface, which will cause semantic distortion in the subsequent image processing process. Compared with equidistant cylindrical projection, although the cube projection reduces the projection distortion of the polar regions of the spherical surface, the faces of the cube after projection are discontinuous, and the salient targets located in the boundary area of the plane image after projection are split, which affects the probability of being correctly detected. Another idea is to obtain salient targets by combining the characteristics of multiple different projection methods through feature fusion. For example, the global features of the panoramic image are extracted by ERP projection, and the local features of the panoramic image are obtained by using the Cube projection image. The detection of significant targets is achieved by fusing the global features and the local features. However, the inventors found that the key to this solution lies in the reasonable selection of the fusion strategy of different projection features. When the fusion strategy is inappropriate, it is difficult to obtain ideal performance. When realizing feature fusion, it is necessary to comprehensively consider the characteristics of features at different levels. How to reasonably balance the feature fusion strategy to link the contradictions in the feature fusion stage is also one of the problems that the existing significant target detection algorithm needs to solve. In view of the above problems, the present invention provides a 360-degree panoramic image significant target detection method based on icosahedral subdivision, which can use convolutional neural networks to extract multi-level features from equidistant cylindrical projections and icosahedral subdivision projections, and hierarchically process feature information through feature pyramids and feature fusion modules to efficiently realize the detection of significant targets in panoramic images. First, in view of the rich content of panoramic image scenes, the present invention extracts significant information in scenes from different spatial scales to avoid ignoring local detail features that may exist in the scene. Secondly, in order to solve the problems of image distortion and boundary discontinuity caused by the traditional panoramic image plane projection method, the present invention adopts a plane projection method based on icosahedron subdivision to make up for the information loss and boundary discontinuity that may be caused by the traditional equidistant cylindrical projection method and cubic projection method. Thirdly, in view of the contradictions and conflicts between multi-projection features, the present invention designs a multi-projection feature fusion module, which can consider the advantages and disadvantages of different projection features, and combines the feature fusion strategy of the attention mechanism with equidistant cylindrical projection features as the main and icosahedron features as the auxiliary, which is suitable for the task of salient target detection in panoramic images.In this way, the accuracy of predicting the salient object features of the panoramic image can be improved, and a global salient feature map reflecting details can be generated. Finally, to achieve the efficient fusion of multi-level features, the present invention designs a multi-scale feature fusion module based on the attention mechanism, which optimizes the integration process of global and local salient object features, ensuring the improvement of the accuracy and detail performance of the finally generated panoramic image salient object map.
[0028] Exemplarily, Figure 1 is a schematic flowchart of a 360-degree panoramic image salient object detection method based on icosahedron subdivision provided by an embodiment of the present invention. As Figure 1 shown, the method includes:
[0029] S101. Obtain a 360-degree panoramic image to be detected.
[0030] S102. Perform equirectangular cylindrical projection and icosahedron subdivision projection on the 360-degree panoramic image to be detected respectively, and obtain a panoramic image in the form of ERP projection and a panoramic image in the form of icosahedron projection respectively.
[0031] S103. Input the panoramic image in the form of ERP projection and the panoramic image in the form of icosahedron projection into a trained salient object detection network model together. The trained salient object detection network model obtains multi-level global features according to the panoramic image in the form of ERP projection, and obtains multi-level local features according to the panoramic image in the form of icosahedron projection, adaptively fuses the multi-level global features and the multi-level local features to obtain multi-level fusion features, and performs feature fusion on the multi-level fusion features based on the attention mechanism to obtain a salient object result map.
[0032] In the present invention, the above S102 is implemented through the following steps:
[0033] S1021. Perform equirectangular cylindrical projection on the 360-degree panoramic image to be detected to obtain a panoramic image in the form of ERP projection.
[0034] The equidistant cylindrical projection method is a common method for projecting a 360° panoramic image onto a two-dimensional plane. Through equirectangular cylindrical projection, the position relationship between each pixel position on the two-dimensional plane image and the corresponding pixel on the 360° panoramic image (i.e., the 360° spherical image) can be obtained. The process of the equidistant cylindrical projection method is to map the sampling points on the two-dimensional plane to the corresponding pixel positions on the sphere. Briefly speaking, first construct an outer circumscribed cylindrical surface of the 360° spherical image, | then project the 360° spherical image onto the outer circumscribed cylindrical surface, and then unfold the outer circumscribed cylindrical surface into a rectangle to obtain the corresponding two-dimensional plane image. Exemplarily, Figure 2 is a schematic diagram of the equirectangular cylindrical projection of a 360° panoramic image. AsFigure 2 As shown in Figure 2 , a two-dimensional plane coordinate system is defined, called the uv coordinate system. Among them, the sampling points on the plane are identified by the coordinates (m, n). Here, m and n represent the coordinates of the sampling points in the horizontal and vertical directions respectively. During the mapping process, first, the points on these two-dimensional planes are mapped into the uv coordinate system through a certain mathematical transformation. The specific formula is: Among them, w is the length of the circumscribed cylindrical surface of a 360° spherical image (i.e., the length of the plane), and h is the height of the circumscribed cylindrical surface of a 360° spherical image (i.e., the width of the plane). Next, the (u, v) coordinates of the sampling points are further converted into longitude and latitude coordinates on the sphere This conversion can be completed through the following formula: θ = (0.5 - v) × π, where The range of is [-π, π], and the range of θ is [-π / 2, π / 2]. Then, use the following formula to map the longitude and latitude coordinate points in the spherical coordinate system to the points (x, y, z) in three-dimensional space, and then obtain the pixel values on the uniform grid in the uv plane through the pixel interpolation method: y = sin(θ), In this way, the projection is completed, and a panoramic image in the form of ERP projection is obtained.
[0035] S1022. Perform icosahedron subdivision projection on the 360-degree panoramic image to be detected to obtain a panoramic image in the form of icosahedron projection.
[0036] The present invention uses the method of polyhedron subdivision to approximately represent the sphere. In three-dimensional space, a sphere can be approximately approximated by a tetrahedron, an octahedron, a dodecahedron, an icosahedron, etc. As the number of faces of the polyhedron increases, the distortion of the polyhedron relative to the sphere continuously decreases, but the number of parameters expressing the subdivided polyhedron will also increase significantly. Considering comprehensively the distortion relative to the sphere after subdivision and the amount of parameter data generated, the present invention selects to use a regular icosahedron as the subdivision unit to realize the planar representation of a 360° panoramic image based on icosahedron subdivision projection. Exemplarily, Figure 3 is the planar representation of a 360° panoramic image based on icosahedron subdivision projection, where Figure 3 in (a) represents the panoramic image in the form of ERP projection obtained after performing equirectangular cylindrical projection on the 360° panoramic image, Figure 3 in (b) represents Figure 3 the circumscribed regular icosahedron of the spherical image obtained by inverse-projecting the (a) figure in into a spherical image and then constructing its circumscribed icosahedron, Figure 3Among them, (c) represents the panoramic image in the form of an icosahedron projection obtained after performing icosahedron subdivision projection on the spherical image based on the regular icosahedron circumscribing the spherical image.
[0037] When the panoramic image in the spherical form is projected into the two-dimensional tangent plane form, it is necessary to determine the points on the projection plane and their positions on the sphere at these points. With the help of spatial geometric relationships, the coordinate conversion relationship from spherical longitude and latitude coordinates to three-dimensional rectangular coordinates can be established. The coordinate conversion relationship specifically adopts the following formula: Where R is the radius of the sphere, and the longitude and latitude coordinates of the pixel coordinates (i, j) are (φ (i,j) , θ (i,j) ), and the corresponding coordinates in the three-dimensional rectangular coordinate system are (X (i,j) , Y (i,j) , Z (i,j) ). In three-dimensional space, each face of the icosahedron is a triangle, which is defined by three vertices, and each vertex contains three coordinate values. Based on this vertex information, the specific position and orientation of each icosahedron face in space can be determined. Specifically, the projection of spherical data onto the icosahedron can be achieved through the following steps:
[0038] 1) Calculate the tangent plane: For each face of the icosahedron, determine its tangent plane through the normal line perpendicular to the face and passing through the center of the sphere. This plane serves as the projection plane, and record the coordinates of the center point of this plane.
[0039] 2) Determine the icosahedron projection plane: Define its boundary using the vertex coordinates of each face of the icosahedron. Each face is a triangle composed of adjacent vertices, and its boundary is uniquely determined by the vertex coordinates.
[0040] 3) Projection of spherical points onto the tangent plane: Adopt the stereographic projection method to map the points on the sphere onto the corresponding tangent plane. Specifically, with the center point of the tangent plane as the center, construct a square tangent to the triangular patch, and then project the inscribed spherical image onto the square tangent plane. This mapping process converts the spherical coordinates into two-dimensional coordinates on the icosahedron projection plane.
[0041] 4) Create the projection image: After the projection is completed, generate a projection image on each tangent plane. Each point in the image corresponds to a specific point in the original panoramic image, so that part of the original panoramic image is presented in each projection image.
[0042] In the present invention, the global feature extraction branch uses the panoramic image in the form of ERP projection as the input, aiming to obtain the macroscopic features and overall structural information of the image; the local feature extraction branch uses the panoramic image in the form of icosahedron projection as the input, aiming to extract the local details and fine-grained features in the image.
[0043] In the present invention, the trained saliency object detection network model includes: a trained global feature extraction module, a trained local feature extraction module, a trained first feature pyramid module, a trained second feature pyramid module, a trained multi-projection feature fusion module, and a trained multi-level feature fusion module. Exemplarily, Figure 4 is a schematic diagram of the structure and processing flow of the saliency object detection network model, as Figure 4 shown, the input of the global feature extraction module is a panoramic image in the form of an ERP projection, the output of the global feature extraction module is connected to the input of the first feature pyramid module, the output of the first feature pyramid module is connected to the input of the multi-projection feature fusion module, and, the input of the local feature extraction module is a panoramic image in the form of an icosahedron projection, the output of the local feature extraction module is connected to the input of the second feature pyramid module, the output of the second feature pyramid module is connected to the input of the multi-projection feature fusion module, the output of the multi-projection feature fusion module is connected to the input of the multi-level feature fusion module, and the output of the multi-level feature fusion module is the saliency object result map. In the present invention, the global feature extraction module combines with the first feature pyramid module to obtain multi-level global features according to the panoramic image in the form of an ERP projection; the local feature extraction module combines with the second feature pyramid module to obtain multi-level local features according to the panoramic image in the form of an icosahedron projection; the multi-projection feature fusion module adaptively fuses the multi-level global features and the multi-level local features to obtain multi-level fusion features; the multi-level feature fusion module performs feature fusion on the multi-level fusion features based on the attention mechanism to obtain the saliency object result map.
[0044] In some embodiments, the number of levels of the multi-level global features is L-1, where L is a positive integer greater than 1; based on this, "obtaining multi-level global features according to the panoramic image in the form of an ERP projection" in the above S103 can be implemented as: performing feature extraction on the panoramic image in the form of an ERP projection to obtain L levels of preliminary global features; performing feature pyramid processing on the L levels of preliminary global features to obtain L-1 levels of global features with the same number of channels.
[0045] In some embodiments, the number of levels of the multi-level local features is L-1; based on this, "obtaining multi-level local features according to the panoramic image in the form of an icosahedron projection" in the above S103 can be implemented as: performing feature extraction on the panoramic image in the form of an icosahedron projection to obtain L levels of preliminary local features; performing feature pyramid processing on the L levels of preliminary local features to obtain L-1 levels of local features with the same number of channels.
[0046] Exemplarily, both the global feature extraction module and the local feature extraction module can be ResNet50 networks, so L = 4. Exemplarily, assuming that the size of the panoramic image based on the ERP projection form or the panoramic image based on the icosahedron projection form is 3×H×W, then after the panoramic image based on the ERP projection form or the panoramic image based on the icosahedron projection form is input into the ResNet50 network, after passing through the initial convolutional layer and pooling layer of ResNe50, the size of the feature map obtained is After that, it successively passes through four convolutional blocks conv2_x, conv3_x, conv4_x, and conv5_x. Among them, the convolutional layers inside each block do not change the resolution of the feature map, but downsampling is performed between blocks. That is to say, from the first convolutional block, the second convolutional block, the third convolutional block to the fourth convolutional block, the resolution change of the feature map is as follows: Thus, preliminary features at four levels are obtained and where when n = 0, it represents the preliminary global feature, and when n = 1, it represents the preliminary local feature, as described above specifically Figure 4 . After obtaining the preliminary features at four levels and , and are used as the input of the Feature Pyramid Networks (FPN) module, so that the FPN module processes the features extracted by the ResNet50 network into features with a unified number of channels of 64 l' = 1, 2, 3, 4, as described above specifically Figure 4 . Exemplarily, Figure 5 is the structural schematic diagram of the feature pyramid module provided by the present invention. Among them, Scale represents the upsampling operation, 1×1Conv represents the convolutional layer with a convolutional kernel of 1×1, BN represents the batch normalization layer, represents the addition operation. Exemplarily, the processing principle of the feature pyramid module can be expressed by the following formula: l = 3, 2, 1, n = 0, 1, where represents the input feature of the l-th layer of the FPN, represents the input feature of the (l + 1)-th layer, Up represents the upsampling operation, represents the output feature of the FPN module. Since after being projected into subdivision sections, the resolution of the sections is too small, and for the deep-level features, due to the small feature size and large number of channels, the gain in detection accuracy for the image with small resolution is small, and at the same time, the computational complexity is additionally increased. In view of this situation, as Figure 5As shown, when the present invention processes multi-level features by FPN, the processing of the fourth-level features from ResNet50 is omitted, that is, one output is reduced. Thus, the feature sizes after being processed by FPN are respectively
[0047] In some embodiments, "adaptive fusion of multi-level global features and multi-level local features to obtain multi-level fusion features" in the above S103 can be implemented as the following steps:
[0048] S1031: Respectively extract the channel attention and spatial attention of the l-th level global feature to correspondingly obtain the l-th level channel attention feature and the l-th level spatial attention feature; the number of levels of multi-level global features, multi-level local features, and multi-level fusion features is all L - 1, where L is a positive integer greater than 1, and the value of l is a positive integer in the range of 1 to L - 1.
[0049] S1032: Fuse the l-th level global feature the l-th level channel attention feature and the l-th level spatial attention feature to obtain the l-th level preliminary fusion feature
[0050] Specifically, multiply the l-th level global feature the l-th level channel attention feature and the l-th level spatial attention feature to obtain the l-th level multiplied feature; concatenate the l-th level global feature with the l-th level multiplied feature and then perform convolution processing to obtain the l-th level preliminary fusion feature
[0051] S1033: Based on the l-th level local feature f 1 l , perform multi-scale context processing on the l-th level preliminary fusion feature to obtain the l-th level multi-projection fusion feature
[0052] Specifically, inverse-project the l-th level local feature f 1 l into the ERP form to obtain the l-th level local feature P T2E (f 1 l ); after performing convolution processing on the l-th level local feature P T2E (f 1 l ), then combine it with the l-th level preliminary fusion feature Perform cascading, then perform convolution processing and batch normalization processing on the cascaded features, and perform multi-scale fusion (Multi-Scale Fusion, MSF) processing on the features after batch normalization processing to obtain the multi-projection fusion features at the l-th level.
[0053] S1034. For the multi-projection fusion features at the l-th level Perform self-refinement (SR) processing to obtain the fused features at the l-th level.
[0054] Specifically, Figure 6 is an exemplary structural schematic diagram of the multi-projection feature fusion module. CA represents the channel attention extraction operation, SA represents the spatial attention extraction operation, T2E represents the projection format transformation from the icosahedron to ERP, SR represents the SR module, MSF represents the MSF module, Cat represents the cascading operation, Conv represents the convolutional layer, input represents the input, output represents the output, Conv represents the convolutional layer, and Relu represents the Relu activation function. The following combines Figure 6 to specifically illustrate the above steps S1031 to S1034. As Figure 6 shown, when and f 1 l are input into the multi-projection feature fusion module, first extract the channel attention of , then perform a max pooling operation on the extracted channel attention to obtain the channel attention features at the l-th level, and first extract the spatial attention of , then perform an average pooling operation on the extracted spatial attention to obtain the spatial attention features at the l-th level. After that, multiply the channel attention features at the l-th level and the spatial attention features at the l-th level, and then cascade them with , and perform 3×3 convolution processing after cascading to obtain the preliminary fusion features at the l-th level At the same time, for f Regarding 1 l , since f 1 l carries rich local detail information, if a similar attention extraction operation is performed on it, some key local information will be lost and local detail noise will be amplified, thereby affecting the expression of global information. Therefore, in the present invention, first project f 1 l inversely into the ERP form to obtain the local feature P at the l-th level in the ERP form T2E (f 1 l ), and then perform operations on P T2E (f 1l ) Perform 3×3 convolution processing to obtain the convolved feature Conv 3×3 (P T2E (f 1 l )); After that, combine with by concatenation, and then perform 3×3 convolution processing and batch normalization processing on the concatenated feature in sequence. Then, use the MSF module to perform MSF processing on the feature after batch normalization to obtain the multi-projection fusion feature at the l-th level That is l = 1, 2, 3, where Conv 3×3 represents a 3x3 convolution block, Concat represents the concatenation operation, and the structural diagram of the MSF module is as shown in Figure 5 shown. Next, perform 3×3 convolution processing on to obtain Then, use the SR module to perform SR processing on Specifically, refer to the above Figure 5 , and the specific processing logic of the SR module can be expressed as a formula: l = 1, 2, 3, where represents the multiplication operation. Here, the advantage of the SR module is that it further refines and enhances the feature map by using multiplication and addition operations, combines the complementary features between different levels of features to obtain a comprehensive feature expression, and avoids some defects in directly fusing two features through splicing.
[0055] In some embodiments, the step of "performing feature fusion on the multi-level fusion feature based on the attention mechanism to obtain the saliency target result map" in S103 above can be implemented as the following steps:
[0056] S1035: Perform upsampling processing on the i-th level fusion feature to obtain the i-th level upsampled feature; i takes positive integers in the range of 2 to L - 1.
[0057] S1036: Concatenate the first level fusion feature and the upsampled features from the second level to the L - 1 level, and perform convolution operation on the concatenated feature to obtain the concatenated convolution feature F′ PFA .
[0058] S1037: Extract the feature attention weight W and the spatial attention feature A of the concatenated convolution feature F′ PFA respectively s .
[0059] S1038: Based on the first level fusion feature the upsampled features from the second level to the L - 1 level, and the concatenated convolution feature F′ PFA, the feature attention weight W and the spatial attention feature A s , to obtain the feature F after fusing multi-level feature attention MLFA .
[0060] Specifically, multiply the first channel W of the attention weight W 1 , the first-level fused feature and the spatial attention feature A s to obtain the first-level multiplied feature; multiply the i-th channel W of the attention weight W i , the feature after the i-th level of upsampling, and the spatial attention feature A s to obtain the i-th level of multiplied feature; concatenate the first-level multiplied feature and the multiplied features from the 2nd to the L-1th levels, and perform convolution processing on the concatenated features to obtain the feature F after fusing multi-level feature attention MLFA .
[0061] S1039. Perform concatenation on the feature F after fusing multi-level feature attention MLFA and the concatenated convolutional feature F' PFA , and then perform convolution processing and non-linear activation processing to obtain the saliency target result map F fused .
[0062] Specifically, perform concatenation on the feature F after fusing multi-level feature attention MLFA and the concatenated convolutional feature F' PFA to obtain the concatenated feature Concat(F MLFA , F' PFA ); perform 3×3 convolution, 1×1 convolution processing, and non-linear activation processing on the concatenated feature Concat(F MLFA , F' PFA ) in sequence to obtain the saliency target result map F fused .
[0063] Exemplarily, Figure 7 is an exemplary structural schematic diagram of the multi-level feature fusion module. As Figure 7 shown, Up represents the upsampling operation, and Cat represents the concatenation operation. As Figure 7 shown, when obtaining the three-level fused features and , perform upsampling processing on and respectively, and then concatenate and the upsampled and , and perform a 3×3 convolution operation on the concatenated features to obtain F' PFA ; then, on the one hand, for F' PFAPerform average pooling and max pooling respectively, then concatenate the features after average pooling and max pooling, and then perform 1×1 convolution operation and non-linear activation processing on the concatenated features in sequence to obtain F′ PFA The feature attention weight W of 1×1 , that is, W = δ(Conv PFA (Concat(AvgPool(F′ PFA ), MaxPool(F′ 1×1 )), where AvgPool represents the average pooling operation, MaxPool represents the max pooling operation, Concat represents the concatenation operation, Conv PFA represents the 1×1 convolution block, δ represents non-linear activation, and the shared spatial attention map can be learned through the Conv1×1 convolution. Among them, Conv1×1 reduces the number of channels from 128 to 1; on the other hand, perform average pooling on F′ s to obtain the spatial attention feature A 1 . Then, multiply the first channel W of W and A s to obtain the first-level multiplication feature. Multiply W 2 , the feature after the second-level upsampling and A s to obtain the second-level multiplication feature. Multiply W 3 , the feature after the third-level upsampling and A s to obtain the third-level multiplication feature; then, concatenate these three-level multiplication features and perform 1×1 convolution processing on the concatenated features to obtain the feature F MLFA after fusing multi-level feature attention, that is , where U represents the upsampling operation. Here, the number of channels can be reduced from 512 to 128 through concatenation and Conv1×1. The weights of each level of features and the shared spatial attention are multiplied by the corresponding features at different levels to obtain multi-level adaptive features. By adaptively increasing or decreasing these learnable weights, the importance of different levels of features in the fusion process can be reflected. Then, concatenate F MLFA and the previously obtained F′ PFA , and perform 3×3 convolution, 1×1 convolution processing and non-linear activation processing on the concatenated feature Concat(F MLFA , F′ PFA ) in sequence to obtain the saliency target result map F fused in ERP format, that is F fused = Conv 3×3 (Conv 1×1 (Concat(F MLFA , F′PFA ))) Here, through the Conv3×3 and Conv1×1 convolutional layers, the number of feature channels can be reduced from 256 to 128.
[0064] In the present invention, the above-mentioned trained saliency object detection network model is trained using a specific training set and a specific loss function. Exemplarily, the training set can be the SOD360 dataset. Exemplarily, the present invention uses Binary Cross-Entropy Loss (BCE) and Mean Absolute Error (MAE) as the main loss functions to supervise the training process of the model. The Binary Cross-Entropy Loss is used to measure the difference in probability distribution between the saliency map predicted by the model and the true saliency map, and its specific form is: where y i represents the label of the i-th pixel in the true saliency map, and p i represents the saliency probability of the i-th pixel predicted by the model, and N is the total number of pixels in the panoramic image based on the ERP projection form or the panoramic image based on the icosahedron projection form. By minimizing the loss function, the model can better learn the boundary and detail information of the saliency object. At the same time, in order to further enhance the localization ability of the model for the saliency region, the present invention introduces the MAE loss. The MAE loss directly measures the absolute difference between the predicted value and the true value, and its specific form is: Finally, to constrain the difference in probability distribution and absolute difference of the saliency object between the true value and the predicted value, combining BCE and MAE as the loss function, the loss function of the present invention is designed as the sum of L BCE and L MAE , and its specific form is: L = L BCE + L MAE .
[0065] The present invention also provides a 360-degree panoramic image saliency object detection device based on icosahedron subdivision, including a processor, a communication interface, a memory, and a communication bus. It is characterized in that the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0066] The memory is used to store computer programs;
[0067] The processor is used to implement the steps of the above-mentioned 360-degree panoramic image saliency object detection method based on icosahedron subdivision when executing the program stored in the memory.
[0068] The present invention has the following advantages:
[0069] 1. The present invention explores a panoramic image projection representation method based on icosahedron subdivision and proposes a panoramic image saliency object detection network based on multi-projection feature fusion. The detection network includes panoramic image projection processing, feature extraction and processing, multi-projection feature fusion, and multi-level feature fusion modules.
[0070] 2. The present invention uses the equirectangular projection image and the projection image based on icosahedron subdivision as the input of the saliency object detection network. These two projection methods have their respective advantages. On the one hand, the equirectangular projection feature can better perceive global information, but it also has inevitable serious distortion and poor ability to perceive local details of saliency objects. On the other hand, the icosahedron-based feature has less distortion and strong ability to perceive local features of the image. However, due to the disconnection between its sections, it cannot guarantee the integrity and continuity of saliency objects in the panoramic image. Therefore, the present invention fuses the features of these two representation methods before feature decoding, can obtain more comprehensive and effective panoramic image features, thereby achieving a better decoding prediction effect. Finally, combining the advantages of the former's globality and continuity and the latter's locality and small projection distortion, their characteristics are organically fused through the network.
[0071] 3. In order to capture global object information in the panoramic image, reduce the influence of the distortion caused by the equidistant projection on the algorithm, and pay attention to the feature details of the panoramic content based on icosahedron subdivision projection, the present invention designs an adaptive multi-projection feature fusion module. This module adaptively selects and fuses features relying on the attention mechanism, improving the accuracy of panoramic image saliency object detection.
[0072] 4. The present invention designs a multi-level feature fusion module based on the attention mechanism. By introducing the spatial attention and channel attention mechanisms, this module refines the multi-projection features, suppresses noise while highlighting the most significant features of the feature channels, adaptively selects the advantages of each level of features for fusion, effectively improving the expression ability of multi-level features, and further improving the accuracy of the model's saliency object detection.
[0073] 5. The present invention constructs a joint loss function based on binary cross-entropy (BCE) and mean absolute error loss (MAE) for the training supervision of the saliency object detection network model. This loss function can effectively measure the probability distribution difference and absolute difference between the predicted saliency object map and the true saliency object map of the model, and realizes the fine adjustment of network parameters through a collaborative optimization strategy, further improving the accuracy and robustness of saliency object detection in panoramic images.
[0074] To verify the effectiveness of the present invention, the present invention will be compared with other existing excellent solutions on the publicly available panoramic saliency object detection datasets SOD360 and F-360iSOD below.
[0075] To verify the effectiveness of the method proposed by the present invention, some saliency object detection algorithms are selected for comparison, including some classic or excellent-performing image saliency object detection algorithms: DDS, FANet, ZoomNet, EGNet, GCPANet, CPDNet, R3Net. DDS is the first saliency object detection algorithm for panoramic content. This algorithm constructs a distortion adaptive module to handle the distortion caused by equidistant cylindrical projection, and introduces a multi-scale context integration block to perceive and distinguish rich scenes and salient objects in the global scene. ZoomNet is a 2D image saliency object detection algorithm based on a hybrid-scale triplet network model. This algorithm learns discriminative hybrid-scale semantics through the designed scale integration unit and hierarchical hybrid-scale unit, and fully focuses on the imperceptible cues between salient objects and background environments. GCPANet is an efficient algorithm for 2D image saliency object detection. This algorithm effectively integrates low-level appearance features, high-level semantic features, and global context features through some progressive context-aware feature interleaving aggregation modules. R3Net is a new cyclic residual refinement network equipped with residual refinement blocks to more accurately detect the salient objects in the input 2D image. The following Table 1 gives the evaluation results of the seven models under different metrics on dataset 360SOD and dataset F-360iSOD.
[0076] Table 1
[0077]
[0078]
[0079] Obviously, the present invention has better performance compared with other existing excellent solutions.
[0080] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0081] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0082] In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. Certain measures are recited in mutually different embodiments, but this does not mean that these measures cannot be combined to produce good effects.
[0083] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A 360-degree panoramic image salient object detection method based on icosahedron subdivision, characterized in that: include: Obtain a 360-degree panoramic image to be detected; Performing equirectangular cylindrical projection and icosahedral subdivision projection on the 360-degree panoramic image to be detected, respectively, to obtain a panoramic image based on an ERP projection form and a panoramic image based on an icosahedral projection form; The panoramic image based on the ERP projection form and the panoramic image based on the icosahedron projection form are input together into a trained salient target detection network model, the trained salient target detection network model obtains multi-level global features according to the panoramic image based on the ERP projection form, and obtains multi-level local features according to the panoramic image based on the icosahedron projection form, the multi-level global features and the multi-level local features are adaptively fused to obtain multi-level fused features, the multi-level fused features are feature fused based on the attention mechanism, and a salient target result map is obtained.
2. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 1, characterized in that: The number of levels of the multi-level global features is L-1, where L is a positive integer greater than 1; the multi-level global features are obtained according to the panoramic image based on the ERP projection form, including: Extracting features from the panoramic image based on the ERP projection form to obtain L-level preliminary global features; The L-level preliminary global features are processed by feature pyramid to obtain L-1-level global features with the same number of channels.
3. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 1, characterized in that: The number of levels of the multi-level local features is L-1, where L is a positive integer greater than 1; the multi-level local features are obtained according to the panoramic image based on the icosahedron projection form, including: Extracting features from the panoramic image based on the icosahedron projection to obtain L-level preliminary local features; The L-level preliminary local features are processed by feature pyramid to obtain L-1-level local features with the same number of channels.
4. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 1, characterized in that: The adaptively fusing the multi-level global features and the multi-level local features to obtain the multi-level fused features includes: Extract the l-th level global features respectively The channel attention and spatial attention of the multi-level global feature and the multi-level local feature are obtained, and the number of levels of the multi-level global feature, the multi-level local feature and the multi-level fusion feature is L-1, L is a positive integer greater than 1, and the value of l is a positive integer between 1 and L-1; For the first level global feature The l-th level channel attention feature and the l-th level spatial attention feature are fused to obtain the l-th level preliminary fusion feature Based on the l-th level local feature f1 l , for the first level of preliminary fusion features Perform multi-scale context processing to obtain the l-th level multi-projection fusion feature The first level multi-projection fusion feature Perform adaptive correction processing to obtain the l-th level fusion feature 5. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 4, characterized in that: The first level global feature The l-th level channel attention feature and the l-th level spatial attention feature are fused to obtain the l-th level preliminary fusion feature include: For the first level global feature The l-th level channel attention feature and the l-th level spatial attention feature are multiplied to obtain the l-th level multiplied feature; The l-th level global feature After cascading with the features after multiplication at the first level, convolution processing is performed to obtain the first level preliminary fusion features 6. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 4, characterized in that: Based on the l-th level local feature f1 l , for the first level of preliminary fusion features Perform multi-scale context processing to obtain the l-th level multi-projection fusion feature include: The l-th level local feature f1 l Back-projection is in ERP form, and the l-th level local feature P in ERP form is obtained T2E (f1 l ); The l-th level local feature P in the ERP form T2E (f1 l ) is convolved and then combined with the first level preliminary fusion feature Cascade, perform convolution and batch normalization on the cascaded features, perform multi-scale fusion on the features after batch normalization, and obtain the l-th level multi-projection fusion feature 7. The method for detecting salient objects in a 360-degree panoramic image based on icosahedron subdivision according to claim 1, characterized in that: The number of levels of the multi-level fusion feature is L-1, where L is a positive integer greater than 1; The method of performing feature fusion on the multi-level fusion features based on the attention mechanism to obtain a significant target result map includes: Perform upsampling on the i-th level fusion features to obtain the i-th level upsampled features; The value of i is a positive integer between 2 and L-1; The first level fusion features The features after upsampling from level 2 to level L-1 are cascaded, and the cascaded features are convolved to obtain the cascaded convolution feature F P ' FA ; Extract the cascade convolution features F respectively P ' FA The feature attention weight W and spatial attention feature A s ; Based on the first-level fusion features The features after upsampling from the 2nd level to the L-1th level, the cascade convolution features F P ' FA , the feature attention weight W and the spatial attention feature A s , and obtain the feature F after fusion of multi-level feature attention MLFA ; The feature F after fusion of multi-level feature attention MLFA And the cascaded convolutional feature F P ' FA After cascading, convolution processing and nonlinear activation processing are performed in sequence to obtain the significant target result map F fused .
8. The method for detecting salient objects in a 360-degree panoramic image based on icosahedral subdivision according to claim 7, characterized in that: Based on the first level fusion feature The features after upsampling from the 2nd level to the L-1th level, the cascade convolution features F P ' FA , the feature attention weight W and the spatial attention feature A s , and obtain the feature F after fusion of multi-level feature attention MLFA ,include: The first channel W of the attention weight W 1 , the first level fusion feature and the spatial attention feature A s Multiply to obtain the first-level multiplication feature; The i-th channel W of the attention weight W i , the i-th level up-sampled features and the spatial attention features A s Multiply them to get the i-th level multiplication feature; The first-level multiplication feature and the second-level to L-1-level multiplication features are cascaded, and the cascaded features are convolved to obtain the feature F after fusion of multi-level feature attention. MLFA .
9. The method for detecting salient objects in 360-degree panoramic images based on icosahedron subdivision according to claim 7, characterized in that: The feature F after fusion of multi-level feature attention MLFA And the cascaded convolutional feature F P ' FA After cascading, convolution processing and nonlinear activation processing are performed to obtain the significant target result map F fused ,include: The feature F after fusion of multi-level feature attention MLFA And the cascaded convolutional feature F P ' FA Cascade to obtain the cascaded feature Concat(F MLFA ,F P ' FA ); The cascaded feature Concat(F MLFA ,F P ' FA ) performs 3×3 convolution, 1×1 convolution and nonlinear activation processing in sequence to obtain the salient target result map F fused .
10. A 360-degree panoramic image salient object detection device based on icosahedral subdivision, comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1-9 when executing the program stored in the memory.
Citation Information
Cited By
Panoramic depth estimation method integrating geometric and semantic optimization and related device
CN121095311A
Panoramic video saliency detection method and system
CN121438186A
A panoramic video saliency detection method and system
CN121438186B