Semantic scene completion method, electronic device, and storage medium
Patent Information
- Application Number
- CN202410593387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-14
AI Technical Summary
[0003]但,语义场景补全方法通常仅依赖于当前帧,它只能提供非常有限的观察来恢复3D几何和语义信息,损害了语义占用预测的稳定性
[0018]本发明的有益效果:区别于现有技术,本发明公开了一种语义场景补全方法、电子设备以及存储介质;其中语义场景补全方法包括:从深度网络和上下文网络获取当前帧相关特征,基于所述深度网络和所述上下文网络构建当前帧三维体积特征;姿态网络获取历史帧相关特征和所述当前帧相关特征,所述姿态网络基于所述历史帧相关特征和所述当前帧相关特征构建跨帧特征亲和度;基于所述跨帧特征亲和度对所述历史帧相关特征进行动态优化;基于交叉注意力机制融合所述历史帧相关特征和所述当前帧相关特征得出语义场景补全信息。通过结合历史帧相关特征得出语义场景补全信息提升语义场景补全的生成质量,得到精细准确的语义场景补全结果。
Smart Images

Figure CN118505990B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic scene completion technology, and in particular to a semantic scene completion method, electronic device, and storage medium. Background Technology
[0002] Existing semantic scene completion methods infer geometric and semantic information of a scene jointly from incomplete observations.
[0003] However, semantic scene completion methods typically rely only on the current frame, which can only provide very limited observations to recover 3D geometric and semantic information, thus compromising the stability of semantic occupancy prediction. Summary of the Invention
[0004] To address the above problems, this invention provides a semantic scene completion method, comprising:
[0005] The current frame-related features are obtained from the deep network and the context network, and the three-dimensional volume features of the current frame are constructed based on the deep network and the context network.
[0006] The pose network acquires historical frame-related features and current frame-related features, and the pose network constructs cross-frame feature affinity based on the historical frame-related features and current frame-related features;
[0007] The relevant features of the historical frames are dynamically optimized based on the cross-frame feature affinity.
[0008] Semantic scene completion information is derived by fusing the relevant features of the historical frames and the relevant features of the current frame based on the cross-attention mechanism.
[0009] The deep network and the context network are constructed based on the Lift Splat Shoot algorithm.
[0010] The pose network fuses and aligns the historical frame-related features and the current frame-related features.
[0011] The pose network acquires the historical frame-related features and the current frame-related features; it constructs multiple sets of contexts through dilated convolutions with different dilation rates; the multiple sets of context features include historical multiple sets of contexts and current multiple sets of contexts.
[0012] The historical frame-related features and the current frame-related features are processed by dilated convolution with the same dilation rate to obtain the same group context; the group affinity is calculated based on the same group context.
[0013] Among them, the affinity of the same group is connected along the channel dimension to obtain the cross-frame feature affinity.
[0014] Specifically, dynamic feature aggregation is performed on the cross-frame feature affinity and the historical frame related features based on 3D deformable convolution.
[0015] Specifically, a multi-level deformable block with multi-level 3D deformable convolution is constructed based on the multiple historical frames.
[0016] To address the above problems, the present invention also provides an electronic device comprising: a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the semantic scene method described above.
[0017] To address the above problems, the present invention also provides a computer-readable storage medium storing program data that can be executed by a processor to implement the image compression method described in any of the preceding claims.
[0018] The beneficial effects of this invention are as follows: Unlike existing technologies, this invention discloses a semantic scene completion method, an electronic device, and a storage medium. The semantic scene completion method includes: obtaining current frame-related features from a deep network and a context network; constructing a three-dimensional volumetric feature of the current frame based on the deep network and the context network; obtaining historical frame-related features and the current frame-related features from a pose network; constructing a cross-frame feature affinity based on the historical frame-related features and the current frame-related features; dynamically optimizing the historical frame-related features based on the cross-frame feature affinity; and fusing the historical frame-related features and the current frame-related features based on a cross-attention mechanism to obtain semantic scene completion information. By combining historical frame-related features to obtain semantic scene completion information, the generation quality of semantic scene completion is improved, resulting in a refined and accurate semantic scene completion result. Attached Figure Description
[0019] Figure 1 This is a schematic diagram illustrating the steps of an embodiment of the semantic scene completion method provided by the present invention;
[0020] Figure 2 This is a schematic diagram of the structure of an embodiment of the semantic scene completion method provided by the present invention;
[0021] Figure 3 A schematic diagram of the structure of an embodiment of the electronic device provided by the present invention;
[0022] Figure 4 This is a schematic diagram of an embodiment of a computer-readable storage medium provided by the present invention. Detailed Implementation
[0023] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings to further illustrate the technical solutions of the present invention. However, the present invention is not limited to these embodiments.
[0024] The task of a robot acquiring the localization and category of an object is called semantic segmentation, while the task of obtaining the complete geometry of an object from a single viewpoint is called shape completion. Combining these two tasks, using deep learning-based computer vision techniques, allows for simultaneous 3D shape completion and semantic segmentation of a scene from a single viewpoint. This joint task is called 3D scene shape completion and semantic segmentation, or simply semantic scene completion.
[0025] Please see Figure 1 — Figure 2 As shown, Figure 1 This is a schematic diagram illustrating the steps of an embodiment of the semantic scene completion method provided by the present invention; Figure 2 This is a schematic diagram of the structure of an embodiment of the semantic scene completion method provided by the present invention.
[0026] Step S11: Obtain relevant features of the current frame from the deep network and the context network, and construct the three-dimensional volume features of the current frame based on the deep network and the context network; that is, input the relevant features of the current frame into the deep network and the context network to obtain the three-dimensional volume features of the current frame;
[0027] Making full use of contextual information is crucial in semantic scene completion methods. Since there are significant differences in scale, lighting, and viewpoint between objects in an image scene, the features extracted from the same semantic label pixel will also have different values, which will affect the recognition accuracy. By combining deep networks and context networks to construct the three-dimensional volume features of the current frame, the semantic scene completion method provided by this invention has better semantic segmentation performance.
[0028] Step S12: The pose network obtains relevant features of historical frames and relevant features of the current frame, and constructs cross-frame feature affinity based on the relevant features of historical frames and relevant features of the current frame;
[0029] Analyzing only the features of the current frame to obtain semantic scene completion is insufficient to fully reproduce the actual positional relationships. For example, during human movement, there are varying degrees of interaction between joints, making it necessary to convey the interaction information between joints. Since the changes in posture and shape between adjacent frames are very small, the posture information of the previous frame can help understand the positional information of the current frame. Historical frames include those with multiple timestamps, such as those with the three timestamps preceding the current frame.
[0030] The existence of differences between multiple historical frames and the current frame is used to determine the affinity between the historical frames and the current frame. Since the affinity between multiple historical frames and the current frame is different, the pose network constructs cross-frame feature affinity through multiple affinity values.
[0031] Step S13: Dynamically optimize historical frames based on cross-frame feature affinity; that is, after the relevant features of historical frames and the relevant features of the current frame enter the pose network, the cross-frame feature affinity is obtained. At the same time, the relevant features of historical frames are dynamically optimized based on the cross-frame feature affinity. The optimized relevant features of historical frames are then combined with the relevant features of the current frame after being processed by the deep network and the context network.
[0032] Step S14: Based on the cross-attention mechanism, fuse relevant features from historical frames and relevant features from the current frame to obtain semantic scene completion information; the specific calculation formula is as follows:
[0033] V ret =α·CrossAtt(Q,K,V)+V vox
[0034] α is a learnable coefficient, and CrossAtt represents a typical cross-attention mechanism in computer vision. V vox For features relevant to the current frame, the cross-attention mechanism queries Q by V. vox The key K and value V are generated from features related to historical frames, V ret This represents the final aggregated features used to generate semantic scene completion results.
[0035] In summary, the semantic scene completion method of this invention includes: obtaining relevant features of the current frame from a deep network and a context network; constructing 3D volumetric features of the current frame based on the deep network and the context network; obtaining relevant features of historical frames and the current frame from a pose network; constructing cross-frame feature affinity based on the relevant features of historical frames and the current frame from the pose network; dynamically optimizing the relevant features of historical frames based on the cross-frame feature affinity; and fusing the relevant features of historical frames and the relevant features of the current frame based on a cross-attention mechanism to obtain semantic scene completion information. By combining the relevant features of historical frames to obtain semantic scene completion information, the generation quality of semantic scene completion is improved, resulting in a refined and accurate semantic scene completion result.
[0036] Optionally, the deep network and context network are built based on the Lift Splat Shoot algorithm; the Lift Splat Shoot algorithm is an algorithm for autonomous driving perception; this algorithm transforms multi-view camera images into feature representations in 3D space, and its operation is as follows:
[0037] 1. Generate 3D features from the images of each camera by "lifting" them;
[0038] 2. Project these 3D features onto the rasterized bird's-eye view grid by "splatting" them;
[0039] 3. Interpretable feature learning is achieved by "shooting" the template motion trajectory into the network output.
[0040] Pose-based networks fuse and align features from historical frames and the current frame. Frame fusion improves the smoothness between historical and current frames, reducing stuttering during slow motion. Frame alignment prevents ghosting when current and historical frames are superimposed, improving efficiency for tasks such as object detection and motion estimation. Frame alignment includes explicit and implicit frame alignment.
[0041] Display frame alignment includes optical flow estimation and motion compensation. Optical flow is the motion vector between two consecutive frames, which is obtained by finding the motion vector of each pixel. The motion vector refers to the velocity.
[0042] The implicit frame alignment process is as follows: the input features are processed by a convolution to obtain a 2x channel offset field (in both x and y directions), and then the offset field is unfolded to obtain the offset amount; the input features and the offset field are simultaneously input into a variable convolution to obtain the output features.
[0043] Alternatively, the pose network can be formed based on the PoseNet paradigm in mainstream video depth prediction.
[0044] A pose network acquires relevant features from historical frames and the current frame; multiple sets of context features are constructed using dilated convolutions with different dilation rates; these multiple sets of context features include historical context features and current context features. Given relevant features from historical frames and the current frame, multiple sets of context features are constructed using 3D dilated convolutions with different dilation rates. Specifically, historical frame relevant features... After a series of dilated convolutions, multiple historical contexts are generated:
[0045]
[0046] Where GN and δ represent group normalization and GELU activation, respectively, and atrous convolution is... i is the dilation rate of spatial convolution.
[0047] In one embodiment, the historical frame related features include three sets of historical frames, at which point H i In this context, i is 3, the timestamp of the first historical frame is closest to the current frame, the timestamp of the third historical frame is furthest from the current frame, and the other historical frame is the second historical frame; the inflation rate corresponding to the first historical frame is 1, the inflation rate corresponding to the second historical frame is 2, and the inflation rate corresponding to the third historical frame is 4.
[0048] Current frame related features They were handled in the same way:
[0049]
[0050] Where GN} and δ represent grouped normalization and GELU activation, respectively. Atrous i Let be the dilation rate of the dilated convolution. The dilation rates are 1, 2, and 4, respectively.
[0051] In other embodiments, the historical frames may include 1, 2, 4, 5, 6 or more sets of historical frames, and the inflation rate may also be modified adaptively, which will not be elaborated here.
[0052] Historical frame-related features and current frame-related features are processed by dilated convolution with the same dilation rate to obtain the same group context; group affinity is calculated based on the same group context. Group affinity A i The calculation formula is as follows:
[0053]
[0054] Input context matrix C i and H i It is considered as a high-dimensional vector with different groups. and This represents the average context matrix of the same group of contexts.
[0055] The affinity of the same group is connected along the channel dimension to obtain the cross-frame feature affinity. The calculation formula is as follows:
[0056]
[0057] The value of n is determined by multiple sets of historical frames. For example, with 3 sets of historical frames, n is 3.
[0058]
[0059] Group affinity A i Further aggregation along the channel dimension into cross-frame feature affinity Affinity A in each group i During the calculation process, their respective average values are subtracted to achieve scale-aware isolation.
[0060] Dynamic feature aggregation based on cross-frame feature affinity and historical frame correlation features is performed using 3D deformable convolution. Cross-frame affinity For features related to historical frames, we use general 3D deformable convolution to perform dynamic feature aggregation. The features sampled by deformable convolution have cross-frame affinity. Multiplication is performed to aggregate relevant features. To utilize these aggregated features and further reason about dynamic modeling through hierarchical context, multi-layered historical frame aggregated features are output. for:
[0061]
[0062] in This represents different 3D deformable convolutions. Concat means concatenated along the channel dimension, and W represents a general 3D convolutional layer. The value of n is determined by multiple sets of historical frames. For example, with 3 sets of historical frames, n is 3.
[0063]
[0064] A multi-level deformable block with multiple concatenated 3D deformable convolutions is constructed based on multiple historical frames. The multiple concatenated 3D deformable convolutions correspond to multiple historical frames. For example, if there are 3 sets of historical frames, a multi-level deformable block with three concatenated 3D deformable convolutions is constructed.
[0065] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device includes a memory and a processor, wherein the memory and the processor are coupled to each other, the memory stores program data, and the processor executes the program data to implement the steps of any embodiment of the semantic scene completion method described above.
[0066] In this embodiment, the processor may also be referred to as a CPU (Central Processing Unit). The processor may be an integrated circuit chip with signal processing capabilities. The processor may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0067] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium stores program data that can be executed by a processor to implement the steps of any embodiment of the semantic scene completion method described above.
[0068] In this embodiment, the computer-readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data. Alternatively, it can be a server that stores the program data. The server can send the stored program data to other devices for execution, or it can run the stored program data itself.
[0069] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A semantic scene completion method, characterized in that, include: The current frame-related features are obtained from the deep network and the context network, and the three-dimensional volume features of the current frame are constructed based on the deep network and the context network. The pose network acquires historical frame-related features and current frame-related features, and constructs multiple sets of contexts through dilated convolutions with different dilation rates; the multiple sets of context features include historical multiple sets of contexts and current multiple sets of contexts; The attitude network constructs cross-frame feature affinity based on the historical frame-related features and the current frame-related features; The historical frame-related features and the current frame-related features are processed by dilated convolution with the same dilation rate to obtain the same group of context; the same group affinity is calculated based on the same group of context; the same group affinity is concatenated along the channel dimension to obtain the cross-frame feature affinity; The relevant features of the historical frames are dynamically optimized based on the cross-frame feature affinity. Semantic scene completion information is derived by fusing the relevant features of the historical frames and the relevant features of the current frame based on the cross-attention mechanism. The calculation formula is as follows: ; Learnable coefficient, This represents a typical cross-attention mechanism in computer vision. For features relevant to the current frame, use a cross-attention mechanism for querying. Depend on Generate, key Sum Generated from historical frame-related features, This represents the final aggregated features used to generate semantic scene completion results.
2. The semantic scene completion method according to claim 1, characterized in that, The deep network and the context network are constructed based on the Lift Splat Shoot algorithm.
3. The semantic scene completion method according to claim 2, characterized in that, The pose network fuses and aligns the historical frame-related features and the current frame-related features.
4. The semantic scene completion method according to claim 3, characterized in that, Dynamic feature aggregation is performed on the cross-frame feature affinity and the historical frame related features based on 3D deformable convolution.
5. The semantic scene completion method according to claim 4, characterized in that, Construct a multi-level deformable block with multiple cascaded 3D deformable convolutions based on multiple historical frames.
6. An electronic device, characterized in that, The electronic device includes a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the semantic scene completion method as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data that can be executed by a processor to implement the semantic scene completion method as described in any one of claims 1-5.