Adaptive scene analysis and target generation method and system based on deep learning
Scene features and semantic information are extracted through deep learning technology, combined with variational autoencoder and probability graph inference network, the problem of insufficient target generation accuracy and adaptability in complex environments is solved, and efficient and robust scenario understanding and target generation are achieved.
Patent Information
- Application Number
- CN202510487435.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
It is difficult for the prior art to efficiently extract scene features in complex environments, optimize the target generation process, and improve the adaptability of the model, resulting in limited accuracy and generalization capabilities of target generation.
Adaptive scenario analysis and target generation methods based on deep learning are adopted. By extracting the depth information and semantic labels of scene image sequences, a three-dimensional point cloud data and semantic correlation matrix is constructed, and input it into a variational autoencoder for feature distribution modeling. It is combined with a probability graph inference network to achieve target generation constraints, improving the accuracy and adaptability of target generation.
It realizes efficient extraction of scene features in complex environments, optimizes the target generation process, improves the adaptability of the model and the accuracy and generalization ability of target generation, and meets the needs of intelligent vision systems for efficient and robust scenario understanding and target generation.
Smart Images

Figure CN120014525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision technology, and in particular to a method and system for adaptive scene analysis and target generation based on deep learning. Background Art
[0002] With the development of computer vision and artificial intelligence technology, scene analysis and target generation are playing an increasingly important role in the fields of autonomous driving, intelligent security, robot navigation, etc. Traditional scene analysis methods mainly rely on manually designed features or pattern matching based on shallow models, which are difficult to adapt to complex and changing environments. In recent years, the rise of deep learning technology has made end-to-end feature extraction and target recognition possible, but existing methods often focus on the analysis of a single data source, such as RGB images or point cloud data, and fail to fully integrate the spatial geometric information and semantic information of the scene, thereby limiting the accuracy and applicability of target generation.
[0003] Existing methods usually only apply this information in specific tasks and lack a unified feature expression framework, resulting in insufficient utilization of information complementarity between different data modalities. In addition, the accuracy of current target generation technology in complex scenarios depends on a large amount of labeled data, with high training costs and limited generalization capabilities. How to efficiently extract scene features in complex environments, optimize the target generation process, and improve the adaptability of the model remains a key issue that needs to be addressed.
[0004] Therefore, there is an urgent need for an adaptive scene analysis and target generation method based on deep learning, which can effectively extract the spatial and semantic features of the scene, and model the feature distribution through variational autoencoders, thereby generating a more accurate target description vector. At the same time, the probabilistic graph reasoning network is combined to realize the fusion of scene information and target generation constraints, improve the accuracy and adaptability of target generation, and meet the needs of intelligent vision systems for efficient and robust scene understanding and target generation. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for adaptive scene analysis and target generation based on deep learning, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention,
[0007] A method for adaptive scene analysis and target generation based on deep learning is provided, comprising:
[0008] Extract the depth information of each frame in the scene image sequence, construct 3D point cloud data based on the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the 3D point cloud data, and combine the distance matrix and direction matrix to obtain the scene space feature vector;
[0009] The scene image sequence is semantically segmented to obtain a semantic label map, the labels in the semantic label map are used to construct a semantic association matrix according to the label co-occurrence frequency, and the scene semantic vector is calculated based on the semantic association matrix;
[0010] The scene spatial feature vector and the scene semantic vector are input into the variational autoencoder, the scene feature distribution is obtained through the encoding process of the variational autoencoder, the target feature vector is sampled from the scene feature distribution, the attention score is calculated based on the target feature vector, the weight of the target feature vector is updated according to the attention score, and a weighted target description vector is generated;
[0011] Construct a probabilistic graph reasoning network, map the scene spatial feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour feature according to the target generation constraint;
[0012] A regional attention map is constructed based on the target contour features, and the regional attention map is combined with the target description vector to generate the final target generation result.
[0013] In an optional embodiment,
[0014] Extract the depth information of each frame in the scene image sequence, construct 3D point cloud data based on the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the 3D point cloud data, and combine the distance matrix and direction matrix to obtain the scene space feature vector including:
[0015] Calculate the pixel mapping relationship of adjacent frames to obtain the displacement field, perform depth information restoration based on the displacement field to obtain a 3D point cloud, perform dual feature extraction of normal vector and curvature field on the 3D point cloud to achieve object segmentation, calculate the direction matrix between objects and the distance matrix considering the overlap of convex hulls, and fuse the direction matrix and distance matrix through the attention mechanism and residual connection to obtain the scene space features, including:
[0016] Acquire a scene image sequence, perform time-series completion on the depth information of each frame in the scene image sequence, calculate the corresponding relationship between the pixel positions of the current frame and the previous frame to obtain a forward mapping matrix, calculate the corresponding relationship between the pixel positions of the current frame and the next frame to obtain a backward mapping matrix, and calculate the displacement field of the pixel position based on the forward mapping matrix and the backward mapping matrix;
[0017] Calculating a motion consistency score according to the displacement field, using the motion consistency score as a weight coefficient to perform weighted combination on depth values of adjacent frames to obtain a depth completion function, using the depth completion function to repair the missing depth value of the current frame to obtain a repaired depth map, and multiplying the pixel coordinates and depth values in the repaired depth map with the camera intrinsic parameter matrix to obtain a three-dimensional point cloud;
[0018] A search radius is constructed for each point in the three-dimensional point cloud, a neighborhood point set is extracted within the search radius, a mean vector of the neighborhood point set is calculated, a covariance matrix is constructed using the neighborhood point set and the mean vector, an eigenvalue decomposition is performed on the covariance matrix to obtain a normal vector, an angle between the normal vector and the normal vector of the neighborhood point is calculated to obtain a normal deviation value, and points whose normal deviation values are greater than a preset deviation threshold are removed to obtain a filtered point cloud;
[0019] Calculate the normal change rate of each point in the filtered point cloud to obtain a normal vector field, use the normal vector field to calculate the local curvature to obtain a curvature field, input the normal vector field and the curvature field into a region growing algorithm to segment and obtain a plane point set, calculate the spatial distance of points outside the plane point set to obtain cluster labels, and segment the point cloud into multiple object point sets according to the cluster labels;
[0020] The centroid position and the point cloud covariance matrix of the object point set are respectively calculated, the point cloud covariance matrix is decomposed by eigenvalue to obtain a main direction vector, the main direction vector is projected on three coordinate planes to obtain a projection vector, the angle between objects is calculated based on the projection vector to obtain a direction matrix, the distance value between objects is calculated using the centroid position, the convex hull of the object point set is constructed and the convex hull overlap is calculated, the distance value is combined with the convex hull overlap to obtain a distance matrix, the direction matrix and the distance matrix are input into a multi-layer perceptron, the attention score is calculated on each layer of the feature map of the multi-layer perceptron, the features are selectively fused according to the attention score, and the fused features are transferred through residual connections to obtain a scene space feature vector.
[0021] In an optional embodiment,
[0022] The scene image sequence is semantically segmented to obtain a semantic label map, and the labels in the semantic label map are used to construct a semantic association matrix according to the label co-occurrence frequency. The scene semantic vector is calculated based on the semantic association matrix, including:
[0023] Construct a temporal consistency loss optimized segmentation network to obtain a semantic label sequence. Calculate the conditional entropy of the semantic label sequence to construct an information entropy weight matrix. Multiply the information entropy weight matrix with the co-occurrence frequency matrix to obtain a semantic association matrix. Construct a graph structure with semantic categories as nodes and semantic association matrices as edge weights. Use a multi-head attention mechanism and jump connections to fuse node features to obtain a scene semantic vector. Specifically, it includes:
[0024] A deep convolutional neural network is used to perform semantic segmentation on a scene image sequence to obtain a semantic label sequence, a forward mapping matrix and a backward mapping matrix are calculated according to the corresponding relationship between pixel positions between adjacent frames, the semantic prediction result of the current frame is aligned with the adjacent frames using the forward mapping matrix and the backward mapping matrix, the difference in the aligned prediction results is calculated to construct a temporal consistency loss function, and the semantic segmentation network parameters are optimized by minimizing the temporal consistency loss function to obtain a semantic label sequence with temporal consistency;
[0025] Extracting semantic categories from the semantic labels of each frame in the semantic label sequence, counting the co-occurrence frequencies of each pair of semantic categories in the same scene to construct a co-occurrence frequency matrix, calculating the joint probability and edge probability of the semantic category pairs in the co-occurrence frequency matrix to obtain conditional entropy, using the conditional entropy to construct an information entropy weight matrix, multiplying the information entropy weight matrix with the co-occurrence frequency matrix and normalizing them to obtain a semantic association matrix;
[0026] The semantic categories and semantic association matrix are respectively constructed as node sets and edge weights of a graph structure, and neighborhood node features are extracted for each node in the graph structure. The attention coefficient between nodes is calculated using a multi-head attention mechanism, and the neighborhood node features are weightedly aggregated according to the attention coefficient to obtain updated features of the nodes. The updated features of multiple layers of nodes are fused through jump connections, and the fused node features are concatenated to obtain a scene semantic vector containing object category information and semantic association patterns.
[0027] In an optional embodiment,
[0028] The scene space feature vector and the scene semantic vector are input into the variational autoencoder, the scene feature distribution is obtained through the encoding process of the variational autoencoder, the target feature vector is sampled from the scene feature distribution, the attention score is calculated based on the target feature vector, the target feature vector is weighted and updated according to the attention score, and the target description vector with weight is generated, including:
[0029] Input the scene space feature vector and the scene semantic vector into the feature decomposition unit of the variational autoencoder, obtain the prior probability distribution parameters through Bayesian inference, construct the conditional prior distribution based on the prior probability distribution parameters, extract the shared information and independent information in the scene space feature vector and the scene semantic vector according to the conditional prior distribution, respectively, obtain the feature shared subspace and the feature independent subspace, calculate the weights of the feature shared subspace and the feature independent subspace based on the mutual information maximization criterion, and perform weighted combination of the feature subspaces according to the weights to obtain the scene feature distribution;
[0030] The scene feature distribution is respectively input into multiple probability distribution branches of the variational autoencoder, distribution parameters are calculated for each branch based on the statistical characteristics of the scene feature distribution, information entropy values of each branch are calculated using the distribution parameters, and a target feature vector is obtained by sampling from the scene feature distribution according to the distribution parameters of the branch with the smallest information entropy value;
[0031] Each dimension of the target feature vector is constructed as a node in a conditional random field, the feature similarity between adjacent nodes is calculated to obtain a pairwise potential function, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, the pairwise potential function and the combined potential function are constructed as an energy function, and the energy function is minimized by iterative optimization to obtain an attention score;
[0032] The attention score is multiplied by each dimension of the target feature vector to obtain a weighted feature vector, the statistical value of the weighted feature vector is calculated to obtain a gating threshold, and the feature dimensions in the weighted feature vector that are greater than the gating threshold are retained to obtain a weighted target description vector.
[0033] In an optional embodiment,
[0034] Each dimension of the target feature vector is constructed as a node in a conditional random field, the feature similarity between adjacent nodes is calculated to obtain a pairwise potential function, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, and the pairwise potential function and the combined potential function are constructed as an energy function, including:
[0035] Constructing each dimension feature value of the target feature vector as a node in a conditional random field, calculating the correlation coefficient between the nodes, setting a connection threshold based on the correlation coefficient, establishing a connection relationship between node pairs whose correlation coefficient is greater than the connection threshold to obtain an adaptive node connection graph, wherein the node connection graph includes a node set and an edge set;
[0036] Calculate the Euclidean distances of the corresponding feature values of adjacent nodes in the node connection graph, and calculate the feature similarity between the adjacent nodes based on the Euclidean distances to obtain a pairwise potential function;
[0037] Calculating the mean of all node feature values in the node group in the node connection graph, calculating the distance between each node feature value and the mean based on the mean, and calculating the combined similarity of the node group according to the distance to obtain a combined potential function;
[0038] The first energy term is obtained by calculating the logarithmic value of the paired potential function, the second energy term is obtained by calculating the logarithmic value of the combined potential function, and the energy function is obtained by summing up the first energy term and the second energy term after respectively assigning weight coefficients.
[0039] In an optional embodiment,
[0040] Construct a probabilistic graph reasoning network, map the scene space feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour features according to the target generation constraint, including:
[0041] Construct a two-stream probability graph network to extract position and semantic features respectively, form target generation constraints by calculating the divergence between position probability distribution and uniform distribution, and the divergence between semantic probability distribution and preset distribution, concatenate the target generation constraints with the target description vector in the feature dimension, and output the target contour features through the generation network, specifically including:
[0042] Constructing a probabilistic graph reasoning network, the probabilistic graph reasoning network includes a position feature stream and a semantic feature stream, the position feature stream processes the scene space feature vector through multi-layer graph convolution to obtain a position feature map, and the semantic feature stream weights the scene semantic vector by calculating the attention weight to obtain a semantic feature;
[0043] Performing a normalization operation on the position feature map to obtain a position probability distribution, and performing a nonlinear transformation on the semantic feature to obtain a semantic probability distribution;
[0044] Calculate the divergence value between the position probability distribution and the uniform distribution to obtain the position constraint item, calculate the divergence value between the semantic probability distribution and the preset semantic distribution to obtain the semantic constraint item, and perform weighted summation of the position constraint item and the semantic constraint item to obtain the target generation constraint;
[0045] The target generation constraint and the target description vector are spliced in the feature dimension, and the spliced features are input into the generation network. The generation network amplifies the spatial scale of the input features through an upsampling convolution layer and outputs the target contour features.
[0046] In an optional embodiment,
[0047] Based on the target contour features, a regional attention map is constructed, and the regional attention map is combined with the target description vector to generate the final target generation results including:
[0048] The sliding window is used to extract the maximum response and average response features at the same time to construct spatial attention, and combined with the channel attention to form a regional attention mechanism to obtain the regional attention map. The residual connection is used to enhance the feature expression, and the regional attention map is multi-scale fused with the target description features and then decoded to generate the target result. Specifically, it includes:
[0049] Scanning the target contour features in a sliding window manner, calculating the statistical values of the features in the window based on each sliding window position, including calculating the maximum value to obtain a maximum response feature map, calculating the average value to obtain an average response feature map, performing convolution operations on the maximum response feature map and the average response feature map, adding them together and performing normalization processing to obtain a spatial attention map;
[0050] Average pooling is performed on the target contour feature in the spatial dimension and a channel attention map is obtained through a fully connected layer, and the spatial attention map, the channel attention map and the target contour feature are multiplied to obtain a regional attention map;
[0051] Performing convolution processing on the regional attention map to obtain local features, and performing residual connection between the local features and the regional attention map to obtain enhanced features;
[0052] Upsampling and expanding the target description vector, concatenating the expanded features with the enhanced features in the feature dimension, and performing multi-scale processing on the concatenated features through convolution layers with different convolution kernel sizes;
[0053] The multi-scale processed features are input into the decoder, which decodes the features through deconvolution layers and skip connections to generate target results.
[0054] According to a second aspect of the embodiments of the present invention,
[0055] Provided is a deep learning-based adaptive scene analysis and target generation system, comprising:
[0056] The first unit is used to extract the depth information of each frame image in the scene image sequence, construct three-dimensional point cloud data according to the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the three-dimensional point cloud data, and combine the distance matrix and the direction matrix to obtain the scene space feature vector;
[0057] The second unit is used to perform semantic segmentation on the scene image sequence to obtain a semantic label map, construct a semantic association matrix based on the labels in the semantic label map according to the label co-occurrence frequency, and calculate the scene semantic vector based on the semantic association matrix;
[0058] The third unit is used to input the scene space feature vector and the scene semantic vector into the variational autoencoder, obtain the scene feature distribution through the encoding process of the variational autoencoder, sample the target feature vector from the scene feature distribution, calculate the attention score based on the target feature vector, and update the weight of the target feature vector according to the attention score to generate a weighted target description vector;
[0059] Unit 4: Construct a probabilistic graph reasoning network, map the scene spatial feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour feature according to the target generation constraint;
[0060] In the fifth unit, a regional attention map is constructed based on the target contour features, and the regional attention map is combined with the target description vector to generate the final target generation result.
[0061] According to a third aspect of the embodiments of the present invention,
[0062] An electronic device is provided, comprising:
[0063] processor;
[0064] a memory for storing processor-executable instructions;
[0065] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0066] According to a fourth aspect of the embodiments of the present invention,
[0067] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0068] In this embodiment, by fusing depth information, three-dimensional point cloud and semantic labels, the spatial structure and semantic relationship of the scene are fully extracted, making target recognition and generation more accurate. At the same time, the variational autoencoder is used to construct the scene feature distribution, and the target description vector is optimized in combination with the attention mechanism, so that the model can adapt to different scenes and improve the stability and generalization ability of target feature extraction. In addition, the probabilistic graph reasoning network is used to model the scene spatial features and semantic features, so that the target generation process not only depends on local information, but also integrates global constraints to enhance the rationality of target generation. On this basis, the target contour features are optimized in combination with the regional attention map, so that the target generation results are clearer and more accurate, and the background interference is effectively reduced, and the adaptability of the model in complex environments is improved. It can effectively improve the overall accuracy of scene understanding and target generation, while reducing the dependence on large-scale labeled data and improving training and reasoning efficiency. It is widely used in fields such as autonomous driving, intelligent monitoring, robot navigation, etc., and can enhance the environmental perception ability of intelligent systems and provide more robust solutions for intelligent visual applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1A schematic diagram of a process of an adaptive scene analysis and target generation method based on deep learning according to an embodiment of the present invention;
[0070] Figure 2 This is an analysis diagram of the application effect of the scene semantic vector in downstream tasks according to an embodiment of the present invention;
[0071] Figure 3 This is a probability distribution divergence analysis diagram of an embodiment of the present invention;
[0072] Figure 4 A schematic diagram of the structure of a decoder according to an embodiment of the present invention;
[0073] Figure 5 Schematic diagram of the structure of an adaptive scene analysis and target generation system based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0075] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0076] Figure 1 FIG. 1 is a flow chart of a method for adaptive scene analysis and target generation based on deep learning according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0077] S101. Extracting the depth information of each frame in the scene image sequence, constructing three-dimensional point cloud data according to the depth information, calculating the distance matrix and direction matrix between objects in the scene based on the three-dimensional point cloud data, and combining the distance matrix and the direction matrix to obtain the scene space feature vector;
[0078] S102. Perform semantic segmentation on the scene image sequence to obtain a semantic label map, construct a semantic association matrix based on the label co-occurrence frequency in the semantic label map, and calculate the scene semantic vector based on the semantic association matrix;
[0079] S103. Input the scene spatial feature vector and the scene semantic vector into the variational autoencoder, obtain the scene feature distribution through the encoding process of the variational autoencoder, sample the target feature vector from the scene feature distribution, calculate the attention score based on the target feature vector, update the weight of the target feature vector according to the attention score, and generate a weighted target description vector;
[0080] S104. Construct a probabilistic graph reasoning network, map the scene space feature vector to a position probability distribution, map the scene semantic vector to a semantic probability distribution, construct a target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector combination into the generation network, and the generation network outputs the target contour feature according to the target generation constraint;
[0081] S105. Construct a regional attention map based on the target contour features, and combine the regional attention map with the target description vector to generate the final target generation result.
[0082] In an optional implementation, extracting depth information of each frame image in a scene image sequence, constructing three-dimensional point cloud data according to the depth information, calculating a distance matrix and a direction matrix between objects in the scene based on the three-dimensional point cloud data, and combining the distance matrix and the direction matrix to obtain a scene space feature vector includes:
[0083] Calculate the pixel mapping relationship of adjacent frames to obtain the displacement field, perform depth information restoration based on the displacement field to obtain a 3D point cloud, perform dual feature extraction of normal vector and curvature field on the 3D point cloud to achieve object segmentation, calculate the direction matrix between objects and the distance matrix considering the overlap of convex hulls, and fuse the direction matrix and distance matrix through the attention mechanism and residual connection to obtain the scene space features, including:
[0084] Acquire a scene image sequence, perform time-series completion on the depth information of each frame in the scene image sequence, calculate the corresponding relationship between the pixel positions of the current frame and the previous frame to obtain a forward mapping matrix, calculate the corresponding relationship between the pixel positions of the current frame and the next frame to obtain a backward mapping matrix, and calculate the displacement field of the pixel position based on the forward mapping matrix and the backward mapping matrix;
[0085] Calculating a motion consistency score according to the displacement field, using the motion consistency score as a weight coefficient to perform weighted combination on depth values of adjacent frames to obtain a depth completion function, using the depth completion function to repair the missing depth value of the current frame to obtain a repaired depth map, and multiplying the pixel coordinates and depth values in the repaired depth map with the camera intrinsic parameter matrix to obtain a three-dimensional point cloud;
[0086] A search radius is constructed for each point in the three-dimensional point cloud, a neighborhood point set is extracted within the search radius, a mean vector of the neighborhood point set is calculated, a covariance matrix is constructed using the neighborhood point set and the mean vector, an eigenvalue decomposition is performed on the covariance matrix to obtain a normal vector, an angle between the normal vector and the normal vector of the neighborhood point is calculated to obtain a normal deviation value, and points whose normal deviation values are greater than a preset deviation threshold are removed to obtain a filtered point cloud;
[0087] Calculate the normal change rate of each point in the filtered point cloud to obtain a normal vector field, use the normal vector field to calculate the local curvature to obtain a curvature field, input the normal vector field and the curvature field into a region growing algorithm to segment and obtain a plane point set, calculate the spatial distance of points outside the plane point set to obtain cluster labels, and segment the point cloud into multiple object point sets according to the cluster labels;
[0088] The centroid position and the point cloud covariance matrix of the object point set are respectively calculated, the point cloud covariance matrix is decomposed by eigenvalue to obtain a main direction vector, the main direction vector is projected on three coordinate planes to obtain a projection vector, the angle between objects is calculated based on the projection vector to obtain a direction matrix, the distance value between objects is calculated using the centroid position, the convex hull of the object point set is constructed and the convex hull overlap is calculated, the distance value is combined with the convex hull overlap to obtain a distance matrix, the direction matrix and the distance matrix are input into a multi-layer perceptron, the attention score is calculated on each layer of the feature map of the multi-layer perceptron, the features are selectively fused according to the attention score, and the fused features are transferred through residual connections to obtain a scene space feature vector.
[0089] For example, first, a scene image sequence is collected. The target scene is photographed by a high-resolution camera to ensure the clarity and details of each frame. Each frame will be processed to extract its depth information, usually using a depth sensor or through a computer vision algorithm for depth estimation.
[0090] Next, the depth information of each frame is time-completion. By analyzing the depth information between adjacent frames, the corresponding relationship between the pixel positions of the current frame and the previous frame is calculated to obtain the forward mapping matrix. At the same time, the corresponding relationship between the pixel positions of the current frame and the next frame is calculated to obtain the backward mapping matrix. Based on these two mapping matrices, the displacement field of the pixel position is calculated so that the depth information can be repaired in the subsequent steps.
[0091] According to the calculated displacement field, the motion consistency score is evaluated. The motion consistency score reflects the motion similarity between adjacent frames, and is used as a weight coefficient to weight the depth values of adjacent frames to form a depth completion function. Using this depth completion function, the missing depth values in the current frame are repaired to generate a repaired depth map. The pixel coordinates and depth values in the repaired depth map are multiplied with the camera intrinsic parameter matrix to obtain 3D point cloud data.
[0092] In the 3D point cloud, a search radius is constructed for each point to extract the neighborhood point set. The mean vector of the neighborhood point set is calculated, and the covariance matrix is constructed using the mean vector. The covariance matrix is subjected to eigenvalue decomposition to obtain the normal vector. The angle between the normal vector and the normal vector of the neighborhood point is calculated to obtain the normal deviation value. Points with normal deviation values greater than the preset threshold are removed to obtain the filtered point cloud.
[0093] Next, the normal change rate of each point in the filtered point cloud is calculated to obtain the normal vector field. The local curvature is calculated using the normal vector field to obtain the curvature field. The normal vector field and the curvature field are input into the region growing algorithm to perform object segmentation and obtain a plane point set. The spatial distance is calculated for points outside the plane point set to obtain cluster labels. According to the cluster labels, the point cloud is segmented into multiple object point sets.
[0094] For each object point set, calculate its centroid position and point cloud covariance matrix. Perform eigenvalue decomposition on the covariance matrix to obtain the main direction vector. Project the main direction vector on three coordinate planes to obtain the projection vector. Based on the projection vector, calculate the angle between objects to form a direction matrix. At the same time, use the centroid position to calculate the distance value between objects, construct the convex hull of the object point set and calculate the convex hull overlap. Combine the distance value with the convex hull overlap to obtain the distance matrix.
[0095] Finally, the direction matrix and distance matrix are input into the multi-layer perceptron. The attention score is calculated on each layer of the feature map of the multi-layer perceptron, and the features are selectively fused according to the attention score. The fused features are transferred through the residual connection, and finally the scene space feature vector is obtained.
[0096] In this embodiment, by calculating the pixel mapping relationship of adjacent frames, the displacement field information is obtained, and the depth information is repaired based on the motion consistency score, so that the point cloud data is more complete and the reconstruction error caused by the lack of depth is reduced. The dual feature extraction method of the normal vector and the curvature field is combined to refine the point cloud data and effectively improve the accuracy of object segmentation. In terms of object relationship calculation, the scheme constructs the direction matrix between objects by calculating the centroid position, covariance matrix and main direction vector, and calculates the distance matrix in combination with the convex hull overlap, so as to more accurately characterize the spatial relationship between objects. The attention mechanism and residual connection are used to fuse the direction matrix and the distance matrix, which effectively enhances the feature expression ability, so that the scene space feature vector can more fully describe the geometric structure and spatial distribution relationship between objects. It can improve the accuracy of scene understanding and object segmentation, making the point cloud data more applicable in complex environments. At the same time, the method based on depth information completion reduces the dependence on high-quality sensing equipment and improves the stability and robustness of three-dimensional point clouds.
[0097] In an optional implementation, semantic segmentation is performed on a scene image sequence to obtain a semantic label map, labels in the semantic label map are constructed into a semantic association matrix according to label co-occurrence frequencies, and the scene semantic vector is calculated according to the semantic association matrix, including:
[0098] Construct a temporal consistency loss optimized segmentation network to obtain a semantic label sequence. Calculate the conditional entropy of the semantic label sequence to construct an information entropy weight matrix. Multiply the information entropy weight matrix with the co-occurrence frequency matrix to obtain a semantic association matrix. Construct a graph structure with semantic categories as nodes and semantic association matrices as edge weights. Use a multi-head attention mechanism and jump connections to fuse node features to obtain a scene semantic vector. Specifically, it includes:
[0099] A deep convolutional neural network is used to perform semantic segmentation on a scene image sequence to obtain a semantic label sequence, a forward mapping matrix and a backward mapping matrix are calculated according to the corresponding relationship between pixel positions between adjacent frames, the semantic prediction result of the current frame is aligned with the adjacent frames using the forward mapping matrix and the backward mapping matrix, the difference in the aligned prediction results is calculated to construct a temporal consistency loss function, and the semantic segmentation network parameters are optimized by minimizing the temporal consistency loss function to obtain a semantic label sequence with temporal consistency;
[0100] Extracting semantic categories from the semantic labels of each frame in the semantic label sequence, counting the co-occurrence frequencies of each pair of semantic categories in the same scene to construct a co-occurrence frequency matrix, calculating the joint probability and edge probability of the semantic category pairs in the co-occurrence frequency matrix to obtain conditional entropy, using the conditional entropy to construct an information entropy weight matrix, multiplying the information entropy weight matrix with the co-occurrence frequency matrix and normalizing them to obtain a semantic association matrix;
[0101] The semantic categories and semantic association matrix are respectively constructed as node sets and edge weights of a graph structure, and neighborhood node features are extracted for each node in the graph structure. The attention coefficient between nodes is calculated using a multi-head attention mechanism, and the neighborhood node features are weightedly aggregated according to the attention coefficient to obtain updated features of the nodes. The updated features of multiple layers of nodes are fused through jump connections, and the fused node features are concatenated to obtain a scene semantic vector containing object category information and semantic association patterns.
[0102] For example, when deep convolutional neural networks process scene image sequences, they need to ensure the temporal stability of semantic prediction results. Traditional semantic segmentation methods usually process each frame of the image independently, ignoring the temporal correlation between adjacent frames, resulting in inconsistent semantic labels for the same object in adjacent frames. To solve this problem, it is necessary to construct a temporal consistency loss and optimize the parameters of the semantic segmentation network by minimizing the loss, so that the semantic prediction results remain stable in time.
[0103] In the image sequence, a deep convolutional neural network is first used to perform semantic segmentation on each frame of the image to obtain a preliminary semantic label sequence. In order to establish the correspondence between adjacent frames, it is necessary to calculate the forward mapping matrix and the backward mapping matrix based on the motion information of the pixel position. The function of these two matrices is to describe the pixel mapping relationship between the current frame and the adjacent frame. For example, in a video sequence, a slight shake of the camera or the movement of objects in the scene will cause the pixels between adjacent frames to shift. The forward mapping matrix can predict the position of the pixels of the current frame in the next frame, while the backward mapping matrix can predict the position of the pixels of the previous frame in the current frame.
[0104] Using the forward and backward mapping matrices, the semantic prediction results of the current frame can be aligned with the prediction results of the adjacent frames, and the difference between the aligned prediction results can be calculated. The size of the difference can measure the semantic consistency between adjacent frames. If the same position of adjacent frames has the same semantic label, it means that the network has good temporal consistency. On the contrary, if there is a large deviation in the semantic labels between adjacent frames, it means that the temporal consistency of the network is poor. By constructing a temporal consistency loss function and minimizing the loss during training, the parameters of the network can be optimized to maintain a high stability in the time dimension.
[0105] For example, in an autonomous driving scenario, suppose a car is driving on the street and there is a red car in front of it. In the first frame, the car is correctly segmented into the "vehicle" category, but in the next frame, due to lighting changes or motion blur, part of the car is misclassified as "road". At this time, the forward mapping matrix and the backward mapping matrix can help align the vehicle areas in the two frames and calculate the degree of label inconsistency. By optimizing the temporal consistency loss, this cross-frame label jitter can be reduced and the stability of semantic segmentation can be improved.
[0106] After obtaining a temporally consistent sequence of semantic labels, it is necessary to further analyze the co-occurrence relationship between different semantic categories in the scene to construct a semantic association matrix. The semantic association matrix is used to describe the regularity of the co-occurrence of objects of different categories in the same scene, providing support for subsequent scene understanding.
[0107] It is necessary to extract the semantic category of each frame from the semantic label sequence and count the co-occurrence frequency between different categories. For example, in an urban road scene, "vehicle" and "road" may co-occur, and "pedestrian" and "sidewalk" may co-occur. By counting the number of co-occurrences of each pair of semantic categories in multiple scenes, a co-occurrence frequency matrix can be constructed. Each element of the co-occurrence frequency matrix represents the frequency of two semantic categories co-appearing in the same frame of the image. For example, "vehicle-road" may have a higher co-occurrence frequency, while "vehicle-sky" may have a lower co-occurrence frequency.
[0108] In order to further measure the strength of the association between semantic categories, it is necessary to calculate conditional entropy. Conditional entropy can be used to measure the predictive power of the occurrence of a certain category on another category. For example, if in most cases, there must be a "vehicle" when there is a "road" in the scene, then the conditional entropy of "road" for "vehicle" is low, indicating that "road" has a strong predictive power for "vehicle". Conversely, if the co-occurrence relationship between two categories is relatively random, the conditional entropy is high.
[0109] Based on conditional entropy, an information entropy weight matrix can be constructed. The role of the information entropy weight matrix is to weight the co-occurrence frequency matrix so that category pairs with lower information entropy (i.e., stronger correlation) have higher weights in the semantic association matrix. By multiplying the information entropy weight matrix with the co-occurrence frequency matrix and performing normalization, the final semantic association matrix can be obtained, which can more accurately reflect the true association between semantic categories.
[0110] For example, in a smart security monitoring scenario, if a camera monitors an area, the area may include two categories, "table" and "chair". Since tables and chairs appear together in most cases, their co-occurrence frequency is high and the conditional entropy is low. Therefore, in the semantic association matrix, the weight of the category pair "table-chair" is high. However, "table-window" may only appear occasionally, so its weight is low.
[0111] Based on the semantic categories and semantic association matrix, a graph structure can be constructed to further extract the high-level semantic representation of the scene. In this graph structure, the semantic categories are regarded as the nodes of the graph, and the weights of the semantic association matrix are regarded as the weights of the edges. In order to learn richer node features, it is necessary to extract features from the neighborhood information of each node and fuse them through a multi-head attention mechanism.
[0112] In this process, the features of each node can be updated through the features of its neighboring nodes. For example, in a traffic scene, the features of the node "road" can be weighted and aggregated through the features of neighboring nodes such as "vehicles", "pedestrians" and "traffic lights". In order to better capture the importance of different neighboring nodes, a multi-head attention mechanism can be used to calculate the attention coefficient between nodes. The multi-head attention mechanism can calculate the attention weights separately in multiple different subspaces, making the final feature representation more comprehensive. For example, in an autonomous driving system, the attention of "road" to "vehicles" may be higher, but the attention to "buildings" may be lower, so the attention mechanism can dynamically adjust the feature weights of different categories.
[0113] After calculating the updated node features, it is necessary to perform feature fusion through skip connections. The role of skip connections is to prevent excessive smoothing of features during layered transmission, so that high-level features still retain the original fine-grained information. For example, in video surveillance analysis, the features of the "person" category may lose details due to excessive information aggregation after multi-layer propagation, while skip connections can maintain the uniqueness of the "person" category.
[0114] The features of all nodes are concatenated to form the overall semantic vector of the scene. This semantic vector contains not only the information of a single category, but also the association pattern between categories, which can be used for subsequent tasks such as scene classification and object detection.
[0115] Figure 2 This is an analysis diagram of the application effect of the scene semantic vector of the embodiment of the present invention in downstream tasks. The figure comprehensively shows the performance comparison between the scene semantic vector generated by this technical solution and classic methods such as Word2Vec and BERT in four types of downstream application tasks. As can be seen from the data in the figure, this technical solution shows significant performance advantages in all tasks. In the scene classification task, this technical solution achieved an average accuracy (mAP) of 0.94, while BERT and Word2Vec were 0.87 and 0.79, respectively, with relative improvements of 8.0% and 19.0%, respectively; in the behavior recognition task, the mAP of this technical solution was 0.92, while BERT and Word2Vec were 0.85 and 0.76, respectively, and the advantages were still obvious; in the complex anomaly detection task, this technical solution maintained a high accuracy of 0.90, while BERT and Word2Vec dropped to 0.83 and 0.74, respectively; in the image retrieval task, the mAP of this technical solution was 0.88, which was still significantly better than BERT's 0.81 and Word2Vec's 0.72. These results fully prove that the scene semantic vectors generated by this technical solution have richer semantic expression capabilities and stronger discriminability through innovative designs such as semantic segmentation under temporal consistency constraints, semantic association matrix optimized by information entropy weights, multi-head attention mechanism, and jump connection feature fusion. In particular, in scene classification and behavior recognition, two tasks closely related to scene understanding, this technical solution performs particularly well, which verifies that this solution can effectively capture the core semantic structure and key semantic relationships in the scene. At the same time, in anomaly detection and image retrieval, two tasks that require high generalization capabilities of semantic vectors, this solution also maintains high performance, reflecting that the scene semantic vectors it generates have good generalization and robustness, and can adapt to various complex downstream application scenarios.
[0116] When performing semantic analysis on scene image sequences, the prior art usually uses deep convolutional neural networks to segment independent frames, which lacks temporal information constraints, resulting in instability of semantic labels between adjacent frames. In addition, the correlation between semantic categories is usually based on simple co-occurrence statistics, and the dynamic characteristics of category relationships are not fully considered, which affects the accuracy and generalization ability of scene semantic representation. This application optimizes the semantic segmentation network by introducing temporal consistency loss, making the semantic label prediction results between adjacent frames more stable, effectively reducing the semantic drift caused by inter-frame noise, and improving segmentation accuracy. At the same time, the information entropy weight matrix is constructed using conditional entropy, so that the semantic association matrix not only depends on the category co-occurrence frequency, but also reflects the stability and amount of information between categories, enhancing the rationality of semantic relationship modeling. In addition, through the multi-head attention mechanism and the feature fusion method of jump connection, the semantic nodes in the graph structure are updated, so that the scene semantic vector can integrate local and global information and improve the ability to understand complex scenes. Compared with the existing technology, the present application has made improvements in semantic consistency, category association modeling and feature fusion methods, avoiding the limitations of single-frame analysis, making the final scene semantic vector more stable and accurate, and can better express the semantic association pattern of objects in the scene, thereby improving adaptability in complex scenes.
[0117] In an optional implementation, the scene space feature vector and the scene semantic vector are input into a variational autoencoder, a scene feature distribution is obtained through the encoding process of the variational autoencoder, a target feature vector is obtained by sampling from the scene feature distribution, an attention score is calculated based on the target feature vector, and the target feature vector is weighted and updated according to the attention score. Generating a weighted target description vector includes:
[0118] Input the scene space feature vector and the scene semantic vector into the feature decomposition unit of the variational autoencoder, obtain the prior probability distribution parameters through Bayesian inference, construct the conditional prior distribution based on the prior probability distribution parameters, extract the shared information and independent information in the scene space feature vector and the scene semantic vector according to the conditional prior distribution, respectively, obtain the feature shared subspace and the feature independent subspace, calculate the weights of the feature shared subspace and the feature independent subspace based on the mutual information maximization criterion, and perform weighted combination of the feature subspaces according to the weights to obtain the scene feature distribution;
[0119] The scene feature distribution is respectively input into multiple probability distribution branches of the variational autoencoder, distribution parameters are calculated for each branch based on the statistical characteristics of the scene feature distribution, information entropy values of each branch are calculated using the distribution parameters, and a target feature vector is obtained by sampling from the scene feature distribution according to the distribution parameters of the branch with the smallest information entropy value;
[0120] Each dimension of the target feature vector is constructed as a node in a conditional random field, the feature similarity between adjacent nodes is calculated to obtain a pairwise potential function, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, the pairwise potential function and the combined potential function are constructed as an energy function, and the energy function is minimized by iterative optimization to obtain an attention score;
[0121] The attention score is multiplied by each dimension of the target feature vector to obtain a weighted feature vector, the statistical value of the weighted feature vector is calculated to obtain a gating threshold, and the feature dimensions in the weighted feature vector that are greater than the gating threshold are retained to obtain a weighted target description vector.
[0122] The scene space feature vector and scene semantic vector are input into the variational autoencoder to achieve feature decomposition and optimization. The scene space feature vector includes the geometric structure, texture information and depth information of the scene, which is usually extracted by a deep convolutional neural network or a point cloud processing algorithm, and is mainly used to characterize the physical form of the scene. The scene semantic vector is used to represent the object categories and their relationships in the scene. It is generally extracted by a pre-trained semantic segmentation network or an object detection network, and contains category labels as well as information such as position and size in the scene. These two types of features are usually processed independently in traditional methods, but through the variational autoencoder, the correlation between the two can be established in the latent variable space, and shared patterns and independent patterns can be extracted to obtain a more complete scene representation.
[0123] The feature decomposition unit uses the Bayesian inference method to calculate the prior probability distribution parameters of the input features. The prior probability distribution is used to constrain the distribution form of the features in the latent space to make it more consistent with the actual physical laws. Based on the prior probability distribution, a conditional prior distribution is constructed, that is, given a certain scene feature, the distribution range of another type of feature is inferred. Subsequently, the scene space feature vector and the scene semantic vector are decomposed respectively to extract the shared information and independent information. The shared information contains the common patterns of the two, such as the combination relationship between the spatial structure and the object category, while the independent information corresponds to the geometric details of the scene space or the separate semantic category information. The proportion of shared information and independent information is calculated by the mutual information maximization criterion, which is used to evaluate the degree of information complementarity between the two variables, and the features are weighted and combined according to the calculated weights to form a scene feature distribution, so that the final feature representation is both comprehensive and has a strong ability to distinguish.
[0124] After being processed, the scene feature distribution is input into multiple probability distribution branches of the variational autoencoder. Each branch corresponds to a different scene feature pattern, such as scene types of different categories, object relationships of different scales, etc. In order to ensure that the final sampled features can stably express the target information, it is necessary to calculate the statistical characteristics of each branch, that is, to estimate the distribution parameters of each branch to measure whether the scene features it represents are stable. For each branch, the information entropy value is calculated. The information entropy measures the uncertainty of the distribution. A lower information entropy indicates that the distribution stability of the branch is higher and it is suitable as the source of the target feature vector. Therefore, the final target feature vector is obtained by sampling from the distribution parameters of the branch with the smallest information entropy value. This process ensures the stability and representativeness of the target feature.
[0125] The optimization of the target feature vector is achieved by modeling it through conditional random fields. Conditional random fields are a probabilistic graphical model for modeling sequential or structured data, in which each feature dimension is mapped to a node in the conditional random field. The feature similarity between adjacent nodes is calculated and a pairwise potential function is constructed. The pairwise potential function is used to describe the local correlation between features, for example, spatially close features usually have similar semantic information. In addition, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, which is used to characterize global patterns, such as the combination relationship of objects in the entire scene. Based on these two potential functions, an energy function is constructed, and the minimization process of the energy function is completed through an iterative optimization algorithm, and finally an attention score is calculated. The attention score is used to measure the importance of each feature dimension in the target feature vector. A higher attention score indicates that the dimension is particularly critical in the current scene, while a lower attention score indicates that the dimension may be redundant information or background noise.
[0126] The attention score is used to update the weight of the target feature vector to enhance the influence of important features while weakening the role of unimportant features. The specific implementation method is to multiply the attention score with each dimension of the target feature vector one by one, so that the dimensions with higher weights are enhanced and the dimensions with lower weights are suppressed. Subsequently, the feature vector after weight update is statistically calculated to determine the gating threshold. The gating threshold is used to screen the most representative feature dimensions. The feature dimensions below the threshold are discarded, and only the feature dimensions exceeding the threshold are retained. The final weighted target description vector can efficiently extract key features and provide more accurate semantic information in subsequent tasks.
[0127] In this embodiment, through the feature decomposition process of the variational autoencoder, spatial information and semantic information can be effectively integrated, making the scene representation more comprehensive, while avoiding the feature redundancy problem that may exist in the traditional method. By screening the optimal distribution through information entropy, the stability of the target feature vector is improved, so that it can maintain a strong generalization ability in different environments. The target features are modeled using conditional random fields and optimized through the attention mechanism, so that the feature representation is more in line with the task requirements, and the reliability and robustness of the final target representation are improved. The final target description vector is not only compact but also has strong discrimination, and can adapt to different application requirements.
[0128] In an optional implementation, each dimension of the target feature vector is constructed as a node in a conditional random field, feature similarities between adjacent nodes are calculated to obtain a pairwise potential function, combined similarities of all nodes in a node group are calculated to obtain a combined potential function, and constructing the pairwise potential function and the combined potential function as an energy function comprises:
[0129] Constructing each dimension feature value of the target feature vector as a node in a conditional random field, calculating the correlation coefficient between the nodes, setting a connection threshold based on the correlation coefficient, establishing a connection relationship between node pairs whose correlation coefficient is greater than the connection threshold to obtain an adaptive node connection graph, wherein the node connection graph includes a node set and an edge set;
[0130] Calculate the Euclidean distances of the corresponding feature values of adjacent nodes in the node connection graph, and calculate the feature similarity between the adjacent nodes based on the Euclidean distances to obtain a pairwise potential function;
[0131] Calculating the mean of all node feature values in the node group in the node connection graph, calculating the distance between each node feature value and the mean based on the mean, and calculating the combined similarity of the node group according to the distance to obtain a combined potential function;
[0132] The first energy term is obtained by calculating the logarithmic value of the paired potential function, the second energy term is obtained by calculating the logarithmic value of the combined potential function, and the energy function is obtained by summing up the first energy term and the second energy term after respectively assigning weight coefficients.
[0133] Exemplarily, first, each dimension of the target feature vector is constructed as a node in a conditional random field, and each node represents the value of the target feature vector in a specific dimension. The conditional random field is a probabilistic graph model for modeling complex relationships, which can effectively capture the mutual influence between features. The correlation coefficient between these nodes is calculated. The correlation coefficient is used to measure the linear relationship between two feature dimensions. When the correlation coefficient of the two feature dimensions is high, it means that they may share similar information. Based on the calculated correlation coefficient, a connection threshold is set. When the correlation coefficient of the two feature dimensions exceeds this threshold, it is considered that there is a strong correlation between them, and a connection relationship is established. The establishment of the connection relationship forms an adaptive node connection graph, which consists of a node set and an edge set, where the node set contains all the target feature dimensions, and the edge set represents the connection between those node pairs that have a high correlation and meet the threshold condition.
[0134] On the constructed adaptive node connection graph, the Euclidean distance of the corresponding eigenvalues of adjacent nodes is calculated. The Euclidean distance measures the spatial difference between two eigenvalues. A smaller Euclidean distance indicates that the eigenvalues of the two nodes are more similar. The feature similarity between adjacent nodes is calculated based on the Euclidean distance. The feature similarity is used to describe the consistency of adjacent features in the target feature space. The higher the feature similarity, the more representative the feature is in the entire target vector. This calculation process obtains a pairwise potential function, which describes the similarity relationship between node pairs and ensures that the conditional random field model can effectively focus on the correlation between features.
[0135] The mean of the feature values of all nodes in the node group is further calculated. The mean is used to characterize the central trend of the node group and reflect the overall feature performance of the nodes in the group. Based on the calculated mean, the distance between the feature value of each node and the mean is calculated. The distance is used to measure the degree of deviation of a single feature from the overall feature. The combined similarity of the node group is calculated based on the distance. The combined similarity is used to characterize the stability of the overall feature. If the feature value of a node deviates greatly from the mean, it means that the feature may be an outlier or affected by external noise, and the combined similarity is low; conversely, if the feature values of all nodes are close to the mean, it means that the group of features is relatively stable and the combined similarity is high. The calculation of the combined similarity obtains the combined potential function, which is used to describe the overall consistency of the target feature vector to ensure that the target feature has stable semantic expression capabilities in a global range.
[0136] Calculate the logarithm of the paired potential function to obtain the first energy term. The logarithm is used to stabilize the calculation results, making the calculation more robust and reducing the impact of extreme values. Similarly, calculate the logarithm of the combined potential function to obtain the second energy term. Assign weight coefficients to the first energy term and the second energy term respectively. The weight coefficients are used to control the contribution of the paired potential function and the combined potential function in the final energy function to ensure the balance between local feature similarity and global feature stability. Sum the weighted first energy term and the second energy term to obtain the final energy function, which is used to measure the overall quality of the target feature vector. In the subsequent optimization process, the energy function is minimized by an iterative solution method to ensure the optimal expression of the target feature vector.
[0137] In this embodiment, by constructing an adaptive node connection graph, feature pairs with higher correlation are adaptively selected for connection, so that feature relationship modeling is more flexible and not restricted by a fixed topological structure. By calculating the paired potential function, adjacent features are kept consistent in the target feature vector, thereby improving the separability of the feature space. By calculating the combined potential function, the stability of the overall feature is enhanced, ensuring that the final target feature vector has a stronger semantic expression ability. In the subsequent feature selection or task application process, the optimized target feature can effectively improve the accuracy of classification, recognition or matching, while reducing feature redundancy, improving computational efficiency, and making feature optimization more efficient and stable.
[0138] In an optional implementation, a probabilistic graph reasoning network is constructed, the scene space feature vector is mapped to a position probability distribution, the scene semantic vector is mapped to a semantic probability distribution, a target generation constraint is constructed based on the position probability distribution and the semantic probability distribution, the target generation constraint and the target description vector are combined and input into the generation network, and the generation network outputs the target contour features according to the target generation constraint, including:
[0139] Construct a two-stream probability graph network to extract position and semantic features respectively, form target generation constraints by calculating the divergence between position probability distribution and uniform distribution, and the divergence between semantic probability distribution and preset distribution, concatenate the target generation constraints with the target description vector in the feature dimension, and output the target contour features through the generation network, specifically including:
[0140] Constructing a probabilistic graph reasoning network, the probabilistic graph reasoning network includes a position feature stream and a semantic feature stream, the position feature stream processes the scene space feature vector through multi-layer graph convolution to obtain a position feature map, and the semantic feature stream weights the scene semantic vector by calculating the attention weight to obtain a semantic feature;
[0141] Performing a normalization operation on the position feature map to obtain a position probability distribution, and performing a nonlinear transformation on the semantic feature to obtain a semantic probability distribution;
[0142] Calculate the divergence value between the position probability distribution and the uniform distribution to obtain the position constraint item, calculate the divergence value between the semantic probability distribution and the preset semantic distribution to obtain the semantic constraint item, and perform weighted summation of the position constraint item and the semantic constraint item to obtain the target generation constraint;
[0143] The target generation constraint and the target description vector are spliced in the feature dimension, and the spliced features are input into the generation network. The generation network amplifies the spatial scale of the input features through an upsampling convolution layer and outputs the target contour features.
[0144] Exemplarily, the scene space feature vector and the scene semantic vector are first obtained. The scene space feature vector represents the spatial layout information in the scene, and the dimension is C×H×W, where C is the number of feature channels, and H and W are the height and width of the feature map, respectively. For example, for an indoor scene, C can be set to 64, and H and W can be set to 32. The scene semantic vector represents the semantic information in the scene, and the dimension is N×D, where N is the number of semantic entities in the scene, and D is the semantic feature dimension. For example, for a scene containing 5 objects, N can be set to 5 and D can be set to 128.
[0145] Next, a probabilistic graph reasoning network is constructed, which includes two branches: the position feature stream and the semantic feature stream. The position feature stream processes the scene space feature vector through multi-layer graph convolution to obtain a position feature map. Specifically, the scene space feature vector is first converted into a graph structure, in which each pixel is a node in the graph and there are edge connections between adjacent pixels. Then, a three-layer graph convolution network is applied to process the graph structure. The number of input channels of the first layer of graph convolution is C, and the number of output channels is 128; the number of input channels of the second layer of graph convolution is 128, and the number of output channels is 256; the number of input channels of the third layer of graph convolution is 256, and the number of output channels is 128. Each layer of graph convolution is followed by a batch normalization layer and a ReLU activation function. After three layers of graph convolution processing, a position feature map with a dimension of 128×H×W is obtained.
[0146] The semantic feature flow obtains semantic features by weighting the scene semantic vector by calculating the attention weight. Specifically, the scene semantic vector is first mapped to the query vector and the key vector through the fully connected layer, both with dimensions of N×64. Then, the dot product between the query vector and the key vector is calculated, and the attention weight matrix is obtained by normalizing it through the Softmax function, with a dimension of N×N. Finally, the attention weight matrix is multiplied by the scene semantic vector to obtain the weighted semantic features with a dimension of N×D. Next, the semantic features are aggregated into a D-dimensional vector through a pooling operation.
[0147] Perform normalization on the position feature map to obtain the position probability distribution. Specifically, perform Softmax normalization on the position feature map so that the sum of the probability values at all pixels is 1, and obtain the position probability distribution with a dimension of H×W. Perform nonlinear transformation on the semantic features to obtain the semantic probability distribution. Specifically, a fully connected layer is used to map the D-dimensional semantic features to a K-dimensional vector (K is the number of predefined semantic categories, such as 10), and then normalized by the Softmax function to obtain the semantic probability distribution with a dimension of K.
[0148] The position constraint term is obtained by calculating the divergence value between the position probability distribution and the uniform distribution. The uniform distribution U_pos means that each position has an equal probability, that is, the probability value of each pixel is 1 / (H×W). The KL divergence is used to calculate the difference between the position probability distribution and the uniform distribution to obtain the position constraint term. The semantic constraint term is obtained by calculating the divergence value between the semantic probability distribution and the preset semantic distribution. The preset semantic distribution is set according to prior knowledge. For example, for generating furniture targets in indoor scenes, the probability of the "table" category can be set to 0.3, the probability of the "chair" category to 0.4, and the probability of other categories to 0.3. The difference between the semantic probability distribution and the preset semantic distribution is calculated using the KL divergence to obtain the semantic constraint term.
[0149] The target generation constraint is obtained by weighted summing of the position constraint and the semantic constraint. The target generation constraint is concatenated with the target description vector in the feature dimension. The target description vector represents the text description of the target to be generated. The text description is converted into a vector representation with a dimension of M (e.g., 256) through a text encoder (such as BERT). The target generation constraint L with a dimension of 1 is expanded to the same dimension as the target description vector, and then concatenated with the target description vector in the feature dimension to obtain a feature vector with a dimension of M+1.
[0150] The concatenated features are input into the generative network, which amplifies the spatial scale of the input features through the upsampling convolution layer and outputs the target contour features. Specifically, the generative network contains four upsampling convolution layers. The input dimension of the first upsampling convolution layer is M+1, the number of output channels is 256, and the upsampling scale is 2; the input channel number of the second upsampling convolution layer is 256, the number of output channels is 128, and the upsampling scale is 2; the input channel number of the third upsampling convolution layer is 128, the number of output channels is 64, and the upsampling scale is 2; the input channel number of the fourth upsampling convolution layer is 64, the number of output channels is 1, and the upsampling scale is 2. Each upsampling convolution layer is followed by a batch normalization layer and a ReLU activation function, and the last layer uses a Sigmoid activation function. The final output target contour feature dimension is 1×(H×8)×(W×8), which represents the target contour probability map.
[0151] In practical applications, thresholding can be used to convert the target contour probability map into a binary image, and the threshold can be set to 0.5. Pixels above the threshold are regarded as the inside of the target, and pixels below the threshold are regarded as the background, thus obtaining the contour of the target.
[0152] Figure 3 This is a probability distribution divergence analysis diagram of the embodiment of the present invention, which compares the probability distribution divergence performance of each method under different scene types. In the "indoor scene" type, the divergence value of this technical solution is 0.31, which is 34.0% and 27.9% lower than the 0.47 of the VAE method and the 0.43 of the Flow-based method respectively; in the "outdoor open scene" type, the divergence value of this technical solution is 0.37, which is 28.8% and 22.9% lower than the 0.52 of the VAE method and the 0.48 of the Flow-based method respectively; in the "complex interactive scene" type, the divergence value of this technical solution is 0.43, which is lower than the 0.58 of the VAE method. and 0.54 of the Flow-based method, with a reduction of 25.9% and 20.4% respectively; in the "low-light scene" type, the divergence value of this technical solution is 0.45, which is 30.8% and 26.2% lower than the 0.65 of the VAE method and 0.61 of the Flow-based method respectively; in the "dynamic scene" type, the divergence value of this technical solution is 0.48, while the VAE method and the Flow-based method are 0.69 and 0.63 respectively, with a reduction of 30.4% and 23.8%. Overall, this technical solution shows the lowest probability distribution divergence value in all kinds of scenes, with an average reduction of 29.3% in the divergence value, which proves its stability and adaptability in dealing with scenes of different complexity, especially in indoor scenes and low-light scenes.
[0153] In the prior art, target generation usually adopts a single generative adversarial network structure to directly generate target images from random noise or conditional inputs, lacking the ability to model scene spatial constraints and semantic constraints. For example, traditional GAN-based methods are difficult to ensure the consistency of the generated target with the scene environment, and are prone to produce generation results with unreasonable spatial positions and mismatched semantic attributes. The target generation method based on a probabilistic graph reasoning network proposed in the present invention, by constructing a two-stream probabilistic graph network to extract position and semantic features respectively, clearly modeling position and semantic constraints, so that the generated target can meet the spatial layout and semantic attribute requirements of the scene. Position constraints ensure that the position of the generated target conforms to the spatial distribution of the scene, and semantic constraints ensure that the semantic attributes of the generated target are coordinated with the scene semantics. In addition, the present invention combines the target generation constraints with the target description vector, so that the generation process can simultaneously meet the target description and scene constraints specified by the user.
[0154] In an optional implementation, constructing a regional attention map based on target contour features, and combining the regional attention map with the target description vector to generate a final target generation result includes:
[0155] The sliding window is used to extract the maximum response and average response features at the same time to construct spatial attention, and combined with the channel attention to form a regional attention mechanism to obtain the regional attention map. The residual connection is used to enhance the feature expression, and the regional attention map is multi-scale fused with the target description features and then decoded to generate the target result. Specifically, it includes:
[0156] Scanning the target contour features in a sliding window manner, calculating the statistical values of the features in the window based on each sliding window position, including calculating the maximum value to obtain a maximum response feature map, calculating the average value to obtain an average response feature map, performing convolution operations on the maximum response feature map and the average response feature map, adding them together and performing normalization processing to obtain a spatial attention map;
[0157] Average pooling is performed on the target contour feature in the spatial dimension and a channel attention map is obtained through a fully connected layer, and the spatial attention map, the channel attention map and the target contour feature are multiplied to obtain a regional attention map;
[0158] Performing convolution processing on the regional attention map to obtain local features, and performing residual connection between the local features and the regional attention map to obtain enhanced features;
[0159] Upsampling and expanding the target description vector, concatenating the expanded features with the enhanced features in the feature dimension, and performing multi-scale processing on the concatenated features through convolution layers with different convolution kernel sizes;
[0160] The multi-scale processed features are input into the decoder, which decodes the features through deconvolution layers and skip connections to generate target results.
[0161] Exemplarily, a sliding window method is used to scan the target contour features. A sliding window is a way to continuously move and extract local features within a specific area. In this process, the window covers different areas of the target contour, and for each window position, the statistical values of all the features therein are calculated. Specifically, the maximum eigenvalue within the window constitutes the maximum response feature map, which reflects the most significant feature of the area. In addition, the average eigenvalue within the window constitutes the average response feature map, which represents the overall information of the area. The maximum response feature map and the average response feature map are processed by convolution operations respectively to extract more detailed local information, and then they are added and adjusted by a normalization step to obtain a spatial attention map. The spatial attention map reflects the importance of different areas in the image, helps to highlight key areas, and improves the model's perception of the target.
[0162] The target contour features are average pooled in the spatial dimension. Average pooling is a downsampling operation used to reduce the spatial dimension of the feature map while retaining the most important information. The pooled features are converted into channel attention maps through a fully connected layer. The channel attention map represents the degree of significance of each feature channel. After the channel attention map is combined with the spatial attention map, a regional attention map is generated by element-by-element multiplication. The regional attention map combines the information of the spatial and channel dimensions and reflects the importance distribution of the target contour features in space and channels.
[0163] The regional attention map is further processed by convolution operations to extract local features. These local features represent fine-grained information of the target features and can help identify smaller, critical details. Subsequently, the local features are residually connected to the regional attention map. Residual connection is a skip connection technology that adds the input features to the convolutional features to make the flow of feature information smoother, effectively alleviating the gradient vanishing problem in deep neural networks, thereby enhancing the expressiveness of features.
[0164] The object description vector is a high-level description of the object, which is upsampled to a higher resolution. The upsampling operation is usually performed by interpolating low-resolution features to high resolution to recover more spatial details. The expanded object description features are concatenated with the enhanced regional features. The concatenation operation merges the two feature vectors in the feature dimension to form a richer feature representation. Next, the concatenated features are processed through multiple convolutional layers with different convolution kernel sizes to extract multi-scale features. This multi-scale processing can effectively capture the details of objects of different sizes, allowing the model to better adapt to the diversity and complexity of objects.
[0165] The multi-scale processed features are input into the decoder for decoding. The decoder upsamples the features through the deconvolution layer to restore the original image or generate the target result. The deconvolution layer is similar to the convolution layer, but its purpose is to restore the low-resolution feature map to a high-resolution image. Skip connections are used to combine early-level features in the decoding process with later-level features in order to maintain the high-frequency details of the image and ensure the fineness of the target generation result. The output of the decoder is the final target generation result, which can be a bounding box of the target, a segmentation result, or other forms of images.
[0166] Figure 4 Schematic diagram of the decoder structure of the embodiment of the present invention. Figure 4As shown in the figure, the decoder structure and its decoding process are shown in detail. The decoder receives the multi-scale processed features (64×64×256) as input, and gradually restores the spatial details of the target through a four-level deconvolution and jump connection structure. The first-level deconvolution upsamples the features to 128×128×128 and performs a jump connection with the encoder's third-layer features (128×128×128), with a detail fidelity of 95.2%; the second-level deconvolution further upsamples to 256×256×64, and combines with the encoder's second-layer features (256×256×64), and the edge positioning accuracy is improved to 0.86 pixels; the third-level deconvolution outputs a 512×512×32 feature map, which is fused with the encoder's first-layer features (512×512×32), and the texture restoration reaches 93.8%; the final first-level deconvolution generates a 1024×1024×16 high-resolution feature, which is mapped to the number of channels of the target category through a 1×1 convolution to complete the target generation. Compared with the ordinary deconvolution decoder, this technical solution improves the detail fidelity of the target with complexity level 7 by 19.1%; compared with the SegNet decoder, it improves by 8.2%; compared with the DeepLab decoder, it improves by 6.6%. Especially when dealing with difficult targets with complexity levels 9-10, this solution still maintains 87-89% detail fidelity, which is mainly due to the combination of low-level feature information provided by the jump connection and the efficient upsampling capability of the deconvolution layer.
[0167] In this embodiment, by extracting the maximum response and average response features through a sliding window and combining the spatial attention mechanism, it is possible to accurately focus on important areas in the image and enhance the model's perception of key features. The combination of spatial and channel attention provides a more detailed attention mechanism, so that the features of different regions and different channels can be effectively weighted, thereby optimizing the representation of the target. Residual connections and multi-scale processing further enhance the expressiveness of features, ensure that the details of local features are not lost, and adapt to the diversity and complexity of targets. Through deconvolution and jump connections, the generated results are clearer and rich in details, adapting to the generation requirements of different resolutions. Therefore, this technical solution not only improves the accuracy of target recognition, but also generates high-quality target results in complex environments, with strong robustness and wide application potential.
[0168] Figure 5 Schematic diagram of the structure of the adaptive scene analysis and target generation system based on deep learning according to an embodiment of the present invention. Figure 5 As shown, the system comprises:
[0169] The first unit is used to extract the depth information of each frame image in the scene image sequence, construct three-dimensional point cloud data according to the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the three-dimensional point cloud data, and combine the distance matrix and the direction matrix to obtain the scene space feature vector;
[0170] The second unit is used to perform semantic segmentation on the scene image sequence to obtain a semantic label map, construct a semantic association matrix based on the labels in the semantic label map according to the label co-occurrence frequency, and calculate the scene semantic vector based on the semantic association matrix;
[0171] The third unit is used to input the scene space feature vector and the scene semantic vector into the variational autoencoder, obtain the scene feature distribution through the encoding process of the variational autoencoder, sample the target feature vector from the scene feature distribution, calculate the attention score based on the target feature vector, and update the weight of the target feature vector according to the attention score to generate a weighted target description vector;
[0172] Unit 4: Construct a probabilistic graph reasoning network, map the scene spatial feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour feature according to the target generation constraint;
[0173] In the fifth unit, a regional attention map is constructed based on the target contour features, and the regional attention map is combined with the target description vector to generate the final target generation result.
[0174] According to a third aspect of the embodiments of the present invention,
[0175] An electronic device is provided, comprising:
[0176] processor;
[0177] a memory for storing processor-executable instructions;
[0178] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0179] According to a fourth aspect of the embodiments of the present invention,
[0180] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0181] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive scene analysis and target generation method based on deep learning, characterized in that: include: Extract the depth information of each frame in the scene image sequence, construct 3D point cloud data based on the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the 3D point cloud data, and combine the distance matrix and direction matrix to obtain the scene space feature vector; The scene image sequence is semantically segmented to obtain a semantic label map, the labels in the semantic label map are used to construct a semantic association matrix according to the label co-occurrence frequency, and the scene semantic vector is calculated based on the semantic association matrix; The scene spatial feature vector and the scene semantic vector are input into the variational autoencoder, the scene feature distribution is obtained through the encoding process of the variational autoencoder, the target feature vector is sampled from the scene feature distribution, the attention score is calculated based on the target feature vector, the weight of the target feature vector is updated according to the attention score, and a weighted target description vector is generated; Construct a probabilistic graph reasoning network, map the scene spatial feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour feature according to the target generation constraint; A regional attention map is constructed based on the target contour features, and the regional attention map is combined with the target description vector to generate the final target generation result.
2. The method according to claim 1, characterized in that Extract the depth information of each frame in the scene image sequence, construct 3D point cloud data based on the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the 3D point cloud data, and combine the distance matrix and direction matrix to obtain the scene space feature vector including: Calculate the pixel mapping relationship of adjacent frames to obtain the displacement field, perform depth information restoration based on the displacement field to obtain a 3D point cloud, perform dual feature extraction of normal vector and curvature field on the 3D point cloud to achieve object segmentation, calculate the direction matrix between objects and the distance matrix considering the overlap of convex hulls, and fuse the direction matrix and distance matrix through the attention mechanism and residual connection to obtain the scene space features, including: Acquire a scene image sequence, perform time-series completion on the depth information of each frame in the scene image sequence, calculate the corresponding relationship between the pixel positions of the current frame and the previous frame to obtain a forward mapping matrix, calculate the corresponding relationship between the pixel positions of the current frame and the next frame to obtain a backward mapping matrix, and calculate the displacement field of the pixel position based on the forward mapping matrix and the backward mapping matrix; Calculating a motion consistency score according to the displacement field, using the motion consistency score as a weight coefficient to perform weighted combination on depth values of adjacent frames to obtain a depth completion function, using the depth completion function to repair the missing depth value of the current frame to obtain a repaired depth map, and multiplying the pixel coordinates and depth values in the repaired depth map with the camera intrinsic parameter matrix to obtain a three-dimensional point cloud; A search radius is constructed for each point in the three-dimensional point cloud, a neighborhood point set is extracted within the search radius, a mean vector of the neighborhood point set is calculated, a covariance matrix is constructed using the neighborhood point set and the mean vector, an eigenvalue decomposition is performed on the covariance matrix to obtain a normal vector, an angle between the normal vector and the normal vector of the neighborhood point is calculated to obtain a normal deviation value, and points whose normal deviation values are greater than a preset deviation threshold are removed to obtain a filtered point cloud; Calculate the normal change rate of each point in the filtered point cloud to obtain a normal vector field, use the normal vector field to calculate the local curvature to obtain a curvature field, input the normal vector field and the curvature field into a region growing algorithm to segment and obtain a plane point set, calculate the spatial distance of points outside the plane point set to obtain cluster labels, and segment the point cloud into multiple object point sets according to the cluster labels; The centroid position and the point cloud covariance matrix of the object point set are respectively calculated, the point cloud covariance matrix is decomposed by eigenvalue to obtain a main direction vector, the main direction vector is projected on three coordinate planes to obtain a projection vector, the angle between objects is calculated based on the projection vector to obtain a direction matrix, the distance value between objects is calculated using the centroid position, the convex hull of the object point set is constructed and the convex hull overlap is calculated, the distance value is combined with the convex hull overlap to obtain a distance matrix, the direction matrix and the distance matrix are input into a multi-layer perceptron, the attention score is calculated on each layer of the feature map of the multi-layer perceptron, the features are selectively fused according to the attention score, and the fused features are transferred through residual connections to obtain a scene space feature vector.
3. The method according to claim 1, characterized in that The scene image sequence is semantically segmented to obtain a semantic label map, and the labels in the semantic label map are used to construct a semantic association matrix according to the label co-occurrence frequency. The scene semantic vector is calculated based on the semantic association matrix, including: Construct a temporal consistency loss optimized segmentation network to obtain a semantic label sequence. Calculate the conditional entropy of the semantic label sequence to construct an information entropy weight matrix. Multiply the information entropy weight matrix with the co-occurrence frequency matrix to obtain a semantic association matrix. Construct a graph structure with semantic categories as nodes and semantic association matrices as edge weights. Use a multi-head attention mechanism and jump connections to fuse node features to obtain a scene semantic vector. Specifically, it includes: A deep convolutional neural network is used to perform semantic segmentation on a scene image sequence to obtain a semantic label sequence, a forward mapping matrix and a backward mapping matrix are calculated according to the corresponding relationship between pixel positions between adjacent frames, the semantic prediction result of the current frame is aligned with the adjacent frames using the forward mapping matrix and the backward mapping matrix, the difference in the aligned prediction results is calculated to construct a temporal consistency loss function, and the semantic segmentation network parameters are optimized by minimizing the temporal consistency loss function to obtain a semantic label sequence with temporal consistency; Extracting semantic categories from the semantic labels of each frame in the semantic label sequence, counting the co-occurrence frequencies of each pair of semantic categories in the same scene to construct a co-occurrence frequency matrix, calculating the joint probability and edge probability of the semantic category pairs in the co-occurrence frequency matrix to obtain conditional entropy, using the conditional entropy to construct an information entropy weight matrix, multiplying the information entropy weight matrix with the co-occurrence frequency matrix and normalizing them to obtain a semantic association matrix; The semantic categories and semantic association matrix are respectively constructed as node sets and edge weights of a graph structure, and neighborhood node features are extracted for each node in the graph structure. The attention coefficient between nodes is calculated using a multi-head attention mechanism, and the neighborhood node features are weightedly aggregated according to the attention coefficient to obtain updated features of the nodes. The updated features of multiple layers of nodes are fused through jump connections, and the fused node features are concatenated to obtain a scene semantic vector containing object category information and semantic association patterns.
4. The method according to claim 1, characterized in that: The scene space feature vector and the scene semantic vector are input into the variational autoencoder, the scene feature distribution is obtained through the encoding process of the variational autoencoder, the target feature vector is sampled from the scene feature distribution, the attention score is calculated based on the target feature vector, the target feature vector is weighted and updated according to the attention score, and the target description vector with weight is generated, including: Input the scene space feature vector and the scene semantic vector into the feature decomposition unit of the variational autoencoder, obtain the prior probability distribution parameters through Bayesian inference, construct the conditional prior distribution based on the prior probability distribution parameters, extract the shared information and independent information in the scene space feature vector and the scene semantic vector according to the conditional prior distribution, respectively, obtain the feature shared subspace and the feature independent subspace, calculate the weights of the feature shared subspace and the feature independent subspace based on the mutual information maximization criterion, and perform weighted combination of the feature subspaces according to the weights to obtain the scene feature distribution; The scene feature distribution is respectively input into multiple probability distribution branches of the variational autoencoder, distribution parameters are calculated for each branch based on the statistical characteristics of the scene feature distribution, information entropy values of each branch are calculated using the distribution parameters, and a target feature vector is obtained by sampling from the scene feature distribution according to the distribution parameters of the branch with the smallest information entropy value; Each dimension of the target feature vector is constructed as a node in a conditional random field, the feature similarity between adjacent nodes is calculated to obtain a pairwise potential function, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, the pairwise potential function and the combined potential function are constructed as an energy function, and the energy function is minimized by iterative optimization to obtain an attention score; The attention score is multiplied by each dimension of the target feature vector to obtain a weighted feature vector, the statistical value of the weighted feature vector is calculated to obtain a gating threshold, and the feature dimensions in the weighted feature vector that are greater than the gating threshold are retained to obtain a weighted target description vector.
5. The method according to claim 4, characterized in that Each dimension of the target feature vector is constructed as a node in a conditional random field, the feature similarity between adjacent nodes is calculated to obtain a pairwise potential function, the combined similarity of all nodes in the node group is calculated to obtain a combined potential function, and the pairwise potential function and the combined potential function are constructed as an energy function, including: Constructing each dimension feature value of the target feature vector as a node in a conditional random field, calculating the correlation coefficient between the nodes, setting a connection threshold based on the correlation coefficient, establishing a connection relationship between node pairs whose correlation coefficient is greater than the connection threshold to obtain an adaptive node connection graph, wherein the node connection graph includes a node set and an edge set; Calculate the Euclidean distances of the corresponding feature values of adjacent nodes in the node connection graph, and calculate the feature similarity between the adjacent nodes based on the Euclidean distances to obtain a pairwise potential function; Calculating the mean of all node feature values in the node group in the node connection graph, calculating the distance between each node feature value and the mean based on the mean, and calculating the combined similarity of the node group according to the distance to obtain a combined potential function; The first energy term is obtained by calculating the logarithmic value of the paired potential function, the second energy term is obtained by calculating the logarithmic value of the combined potential function, and the energy function is obtained by summing up the first energy term and the second energy term after respectively assigning weight coefficients.
6. The method according to claim 1, characterized in that Construct a probabilistic graph reasoning network, map the scene space feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour features according to the target generation constraint, including: Construct a two-stream probability graph network to extract position and semantic features respectively, form target generation constraints by calculating the divergence between position probability distribution and uniform distribution, and the divergence between semantic probability distribution and preset distribution, concatenate the target generation constraints with the target description vector in the feature dimension, and output the target contour features through the generation network, specifically including: Constructing a probabilistic graph reasoning network, the probabilistic graph reasoning network includes a position feature stream and a semantic feature stream, the position feature stream processes the scene space feature vector through multi-layer graph convolution to obtain a position feature map, and the semantic feature stream weights the scene semantic vector by calculating the attention weight to obtain a semantic feature; Performing a normalization operation on the position feature map to obtain a position probability distribution, and performing a nonlinear transformation on the semantic feature to obtain a semantic probability distribution; Calculate the divergence value between the position probability distribution and the uniform distribution to obtain the position constraint item, calculate the divergence value between the semantic probability distribution and the preset semantic distribution to obtain the semantic constraint item, and perform weighted summation of the position constraint item and the semantic constraint item to obtain the target generation constraint; The target generation constraint and the target description vector are spliced in the feature dimension, and the spliced features are input into the generation network. The generation network amplifies the spatial scale of the input features through an upsampling convolution layer and outputs the target contour features.
7. The method according to claim 1, characterized in that Based on the target contour features, a regional attention map is constructed, and the regional attention map is combined with the target description vector to generate the final target generation results including: The sliding window is used to extract the maximum response and average response features at the same time to construct spatial attention, and combined with the channel attention to form a regional attention mechanism to obtain the regional attention map. The residual connection is used to enhance the feature expression, and the regional attention map is multi-scale fused with the target description features and then decoded to generate the target result. Specifically, it includes: Scanning the target contour features in a sliding window manner, calculating the statistical values of the features in the window based on each sliding window position, including calculating the maximum value to obtain a maximum response feature map, calculating the average value to obtain an average response feature map, performing convolution operations on the maximum response feature map and the average response feature map, adding them together and performing normalization processing to obtain a spatial attention map; Average pooling is performed on the target contour feature in the spatial dimension and a channel attention map is obtained through a fully connected layer, and the spatial attention map, the channel attention map and the target contour feature are multiplied to obtain a regional attention map; Performing convolution processing on the regional attention map to obtain local features, and performing residual connection between the local features and the regional attention map to obtain enhanced features; Upsampling and expanding the target description vector, concatenating the expanded features with the enhanced features in the feature dimension, and performing multi-scale processing on the concatenated features through convolution layers with different convolution kernel sizes; The multi-scale processed features are input into the decoder, which decodes the features through deconvolution layers and skip connections to generate target results.
8. An adaptive scene analysis and target generation system based on deep learning, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to extract the depth information of each frame image in the scene image sequence, construct three-dimensional point cloud data according to the depth information, calculate the distance matrix and direction matrix between objects in the scene based on the three-dimensional point cloud data, and combine the distance matrix and the direction matrix to obtain the scene space feature vector; The second unit is used to perform semantic segmentation on the scene image sequence to obtain a semantic label map, construct a semantic association matrix based on the labels in the semantic label map according to the label co-occurrence frequency, and calculate the scene semantic vector based on the semantic association matrix; The third unit is used to input the scene space feature vector and the scene semantic vector into the variational autoencoder, obtain the scene feature distribution through the encoding process of the variational autoencoder, sample the target feature vector from the scene feature distribution, calculate the attention score based on the target feature vector, and update the weight of the target feature vector according to the attention score to generate a weighted target description vector; Unit 4: Construct a probabilistic graph reasoning network, map the scene spatial feature vector to the position probability distribution, map the scene semantic vector to the semantic probability distribution, construct the target generation constraint based on the position probability distribution and the semantic probability distribution, input the target generation constraint and the target description vector into the generation network, and the generation network outputs the target contour feature according to the target generation constraint; In the fifth unit, a regional attention map is constructed based on the target contour features, and the regional attention map is combined with the target description vector to generate the final target generation result.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Remote sensing image content description method based on variational self-attention reinforcement learning
CN111126282A
Mechanical fault intelligent diagnosis method based on multi-sensor information fusion
CN113792602A
Scene depth completion method based on conditional variation auto-encoder and geometric guidance
CN116468768A
Instance-aware monocular semantic scene completion method, medium and equipment
CN117422629A
Training method of small sample image detection model and image detection method
CN119649136A
Cited By
Space-time context extraction method and system for real world model training
CN120451882A
A spatiotemporal context extraction method and system for real-world model training
CN120451882B
Special image detection optimization method based on artificial intelligence cloud service
CN120599334A
A special image detection optimization method based on artificial intelligence cloud service
CN120599334B
Guide rail precision evaluation method and system based on multi-sensor fusion
CN120639615A