Scene identification method for 3D and AI visual sensing visible light movement

Through the 3D scene recognition method of multimodal feature fusion and optimization, the problems of insufficient fusion of multimodal feature and lack of semantic consistency in the prior art are solved, cross-modal semantic correlation and open set scene recognition are realized, and recognition accuracy and model adaptability are improved.

CN120236187AInactive Publication Date: 2025-07-01SHENZHEN KEAN DIGITAL CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510366907.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing 3D scene recognition technology has problems such as insufficient multimodal feature fusion, lack of semantic consistency in feature coding and limited processing capabilities for objects in unknown categories, resulting in identification failure or incorrect classification.

Method used

Multimodal feature extraction, automatic encoding compression, homologous loss and double reconstruction loss calculation, hypergraph construction and memory bank alignment are adopted to achieve cross-modal semantic association and open set scene recognition through the fusion and optimization of multi-view, point cloud and voxel features.

Benefits of technology

It improves the accuracy of scene recognition and the generalization ability of the model, can handle scenes with no categories, and enhances the adaptability and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236187A_ABST
    Figure CN120236187A_ABST
Patent Text Reader

Abstract

The invention provides a scene identification method for a 3D and AI visual sensing visible light movement, and the method comprises the steps: carrying out the feature extraction of the multi-modal features of a 3D object, and obtaining a basic feature set containing a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; performing automatic coding compression on the basic feature set to obtain a potential spatial feature code; carrying out homologous loss and dual reconstruction loss calculation on the potential spatial feature codes to obtain optimized 3D object embedding representation; aggregating the plurality of modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; performing hypergraph construction and hypergraph convolution on the unified 3D object embedding matrix to obtain structure perception embedding representation; and carrying out memory bank alignment on the structure perception embedding representation to obtain an alignment embedding representation, and carrying out scene classification identification according to the alignment embedding representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a scene recognition method for a 3D and AI vision sensing visible light module. Background Art

[0002] Currently, in modern AI vision sensing technology, 3D scene recognition based on visible light modules faces many challenges. Traditional scene recognition methods mainly rely on the closed-set recognition mode, which can only recognize a limited number of pre-trained object categories, while real-world scenes often contain a large number of unseen objects and complex environments. This limitation severely restricts the performance and adaptability of AI vision systems in practical applications. The current 3D scene recognition technologies generally have the following problems: First, the multi-modal feature fusion is insufficient, and different representations such as multi-view, point cloud, and voxel cannot be effectively integrated; second, the feature encoding process lacks semantic consistency, making it difficult to capture deep semantic associations across modalities; third, the existing methods have extremely limited ability to handle unknown category objects, and often encounter recognition failures or misclassifications when encountering unfamiliar scenes.. Summary of the Invention

[0003] This application provides a scene recognition method for a 3D and AI vision sensing visible light module to improve the accuracy of scene recognition.

[0004] In a first aspect, an embodiment of this application provides a scene recognition method for a 3D and AI vision sensing visible light module, the method including: Extract features from the multi-modal features of 3D objects to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; Automatically encode and compress the basic feature set to obtain a latent space feature encoding; Calculate the homology loss and the double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation; Aggregate multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; Construct a hypergraph and perform hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation; Perform memory bank alignment on the structure-aware embedding representation to obtain an aligned embedding representation, and perform scene classification and recognition according to the aligned embedding representation.

[0005] In a second aspect, an embodiment of this application provides a scene recognition device for a 3D and AI vision sensing visible light module, the device including: A feature extraction module for extracting multi-modal features of a 3D object to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; An encoding and compression module for automatically encoding and compressing the basic feature set to obtain a latent space feature encoding; A loss calculation module for calculating a homology loss and a double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation; A feature aggregation module for aggregating multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; A hypergraph calculation module for constructing a hypergraph and performing hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation; A result output module for aligning the structure-aware embedding representation with a memory bank to obtain an aligned embedding representation, and performing scene classification and recognition according to the aligned embedding representation.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, which includes a memory and a [device name not provided in the original]; The memory is used for storing a computer program; The [device name not provided in the original], is used for executing the computer program and implementing the scene recognition method of any one of the 3D and AI vision sensing visible light machine cores in the embodiments of the present application when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a [device name not provided in the original], the [device name not provided in the original] is caused to implement the scene recognition method of any one of the 3D and AI vision sensing visible light machine cores in the embodiments of the present application.

[0008] An embodiment of the present application provides a scene recognition method for a 3D and AI vision sensing visible light module, the method comprising: extracting features from the multi-modal features of a 3D object to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; automatically encoding and compressing the basic feature set to obtain a latent space feature encoding; calculating a homology loss and a double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation; aggregating multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; constructing a hypergraph and performing hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation; performing memory bank alignment on the structure-aware embedding representation to obtain an aligned embedding representation, and performing scene classification and recognition according to the aligned embedding representation. Through the above method, by combining multi-view, point cloud, and voxel features, using automatic encoding compression and double reconstruction loss, feature representations are effectively extracted and optimized. Through hypergraph convolution and memory bank alignment, complex structural relationships between objects can be deeply mined, and accurate screening of semantic saliency anchors can be achieved. At the same time, through dynamic boundary adaptation and multi-stage confidence evaluation, the open set scene recognition ability is available, unseen categories can be processed, the generalization ability of the model is enhanced, and the recognition accuracy of the scene is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a schematic flowchart of a scene recognition method for a 3D and AI vision sensing visible light module provided by an embodiment of the present application; Figure 2 It is a schematic block diagram of a scene recognition device for a 3D and AI vision sensing visible light module provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0012] The flowcharts shown in the accompanying drawings are merely illustrative examples, and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0013] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0014] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0015] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a scene recognition method for a 3D and AI vision sensing visible light module provided by an embodiment of this application. As Figure 1 shown, the steps include: S101 - S106.

[0016] S101. Extract features from the multi-modal features of 3D objects to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix.

[0017] Exemplarily, design an AI vision sensing visible light module with multi-angle and multi-dimensional sensing capabilities, which can synchronously capture multi-view images, point cloud data, and voxel representations of 3D objects in a scene. For different modal feature extractors, semantic features are extracted from multi-view images, point clouds, and voxel data respectively. The multi-view image feature extractor is based on a deep convolutional neural network, the point cloud feature extractor uses a geometric transformation network, and the voxel feature extractor uses a three-dimensional convolutional neural network. Design a feature normalization and alignment algorithm to eliminate the semantic gap between different modalities and lay a foundation for subsequent multi-modal fusion. Establish a multi-modal feature matrix, where each row represents the multi-modal feature representation of a 3D object, providing a unified feature space for subsequent processing.

[0018] S102. Automatically encode and compress the basic feature set to obtain a latent space feature encoding.

[0019] Exemplarily, by designing a dedicated multi-modal autoencoder, high-dimensional heterogeneous features are mapped to a unified low-dimensional latent space. The encoder adopts a hierarchical feature compression strategy to gradually extract the abstract representation of features through multiple layers of non-linear transformations; the decoder is responsible for reconstructing the original features from the compressed latent space. Regularization techniques are introduced, such as the KL divergence constraint of the variational autoencoder (VAE), to enhance the structural and semantic continuity of feature compression. By minimizing the reconstruction error and the regularization loss, a lossy but information-preserving compression of the multi-modal features of 3D objects is achieved.

[0020] S103. Calculate the homology loss and the double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation.

[0021] Exemplarily, the homology loss constructs a compact feature representation space by constraining the distances between different modal embeddings of the same object; the double reconstruction loss includes single-modal and cross-modal reconstructions, balancing the information reconstruction ability and generalization performance of the model. The design of the loss function introduces an adaptive weight mechanism to dynamically adjust the importance of different loss terms during the training process. Through contrastive learning and reconstruction constraints, the semantic differences between modalities are effectively reduced, generating a more discriminative and robust 3D object embedding representation.

[0022] S104. Aggregate multiple modal features based on the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix.

[0023] Exemplarily, the aggregation strategy is based on the attention mechanism to adaptively learn the importance weights of each modal feature. A cross-modal information interaction module is introduced, and through self-attention and cross-attention mechanisms, the potential semantic associations between modalities are mined. The aggregation process combines weighted average and non-linear transformation to generate a unified embedding matrix that can comprehensively represent the semantic features of 3D objects. Through this multi-modal feature aggregation method, the feature expression ability of subsequent scene understanding tasks is significantly improved.

[0024] S105. Perform hypergraph construction and hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation.

[0025] Exemplarily, the hypergraph construction is based on the K-nearest neighbor algorithm, treating each 3D object as a vertex and constructing undirected hyperedges through semantic similarity. The hypergraph convolutional network designs an adaptive information propagation mechanism that can capture the complex structural dependencies between vertices. The convolution process introduces attention weights to dynamically adjust the information aggregation intensity according to the semantic relevance between objects. Through this structure-aware learning paradigm, the latent semantic information across objects and categories is effectively extracted.

[0026] S106. Align the structure-aware embedding representation with the memory bank to obtain an aligned embedding representation, and perform scene classification and recognition based on the aligned embedding representation.

[0027] Exemplarily, the memory bank consists of multiple semantic anchors and stores typical object representations. The alignment process generates normalized activation weights by calculating the similarity between the input embedding and the memory anchors. An alignment embedding reconstruction strategy is innovatively proposed to map the original embedding to the typical representation space of the memory bank. Scene recognition is based on a similarity metric and a confidence adaptive mechanism and can handle known and unknown categories. Through the dynamic update of the memory bank, the generalization ability and recognition performance of the model are continuously improved.

[0028] The embodiment of the present application provides a scene recognition method for a 3D and AI vision sensing visible light module. The method includes: extracting features from the multi-modal features of a 3D object to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; automatically encoding and compressing the basic feature set to obtain a latent space feature encoding; calculating a homology loss and a double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation; aggregating multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; constructing a hypergraph and performing hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation; aligning the structure-aware embedding representation with the memory bank to obtain an aligned embedding representation, and performing scene classification and recognition based on the aligned embedding representation. Through the above method, by combining multi-view, point cloud, and voxel features, using automatic encoding compression and double reconstruction loss, feature representations are effectively extracted and optimized. Through hypergraph convolution and memory bank alignment, complex structural relationships between objects can be deeply mined, and accurate screening of semantic significance anchors can be achieved. At the same time, through dynamic boundary adaptation and multi-stage confidence evaluation, the open-set scene recognition ability is available, unseen categories can be handled, the generalization ability of the model is enhanced, and the recognition accuracy of the scene is improved.

[0029] To more clearly introduce the technical solution of the present application, the technical solution of the present application will also be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the description of the technical solution of the present application and do not limit the present application.

[0030] In some embodiments, automatically encoding and compressing the basic feature set to obtain a latent space feature encoding includes: S1021 - S1026.

[0031] S1021. Perform deep convolutional encoding on the multi-view feature matrix to map the multi-view features to the first latent space features.

[0032] Exemplarily, separable convolution kernels are used for the input multi-view feature matrix to design convolutional layers with multi-scale receptive fields; a channel attention mechanism and a spatial attention mechanism are introduced into each convolutional layer to adaptively enhance the features of key perspectives; through residual connections and feature normalization, the feature extraction ability of the deep convolutional network is stabilized.

[0033] Specifically, the dimension of the input multi-view feature matrix is [B, V, C, H, W] (batch, number of views, number of channels, height, width). During the separable convolution process, first perform 3×3 depth convolution (channel-wise), then perform 1×1 point convolution to reduce the computational amount, and then use three convolution kernels of [3×3, 5×5, 7×7] for parallel processing to obtain a multi-channel encoding matrix. Then, perform attention mechanism processing on the multi-channel encoding matrix to obtain the first latent space features.

[0034] The specific formula for the channel attention mechanism is: M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) = σ(W2(δ(W1(F avg ))) + W2(δ(W1(F max )))); Where, σ is the sigmoid function, and δ is the ReLU function.

[0035] The specific formula for the spatial attention mechanism is: M s (F) = σ(f 7×7 ([AvgPool c (F); MaxPool c (F)])) where f 7×7 represents a 7×7 convolution operation, [;] represents concatenation in the channel dimension, and the output is: F' = M c (F) ⊗ F F'' = M s (F') ⊗ F' Output = F + F''.

[0036] S1022. Perform geometric transformation encoding on the point cloud feature matrix to map the point cloud features to the second latent space features.

[0037] Exemplarily, use a k-nearest neighbor graph or a ball radius neighborhood to construct a feature transformation operator based on the local geometric structure, and capture the spatial topological relationship of the point cloud data by calculating the relative positions and geometric features of the points within the local region of the point cloud feature matrix. Design a dynamic point weight modulation mechanism to adjust the feature weights according to geometric attributes such as the local density and curvature of the local structure of the point cloud. Finally, through a symmetric function and an invariance transformation, map the point cloud features to the second latent space features to enhance the geometric transformation insensitivity of the point cloud features.

[0038] S1023. Perform three-dimensional spatial transformation encoding on the voxel feature matrix to map the voxel features to the third latent space features.

[0039] Exemplarily, use a spatially dilated convolutional kernel to increase the receptive field and capture multi-scale voxel features. Design a depth decomposition and reconstruction module for voxel features to decouple the local and global semantic information in the voxel space, decompose the multi-scale voxel features into local detail features and global semantic features, and process them separately through independent encoders. Establish residual connections between feature maps at different scales to fuse feature information at multiple levels and enhance the multi-scale representation ability of voxel features.

[0040] S1024. According to the first latent space features, the second latent space features, and the third latent space features, construct a multi-modal feature interaction mapping network, and achieve feature cross-correlation through a cross-modal attention mechanism and an information fusion module.

[0041] Exemplarily, for the three input modal features, {the first latent space feature (F1), the second latent space feature (F2), and the third latent space feature (F3)}, F1 ∈ R (d×n) , F2 ∈ R (d×n) , F3 ∈ R (d×n) (where d is the feature dimension and n is the number of samples), the specific calculation steps are as follows: 1) Calculate the mutual information maximization of the first latent space feature (F1), the second latent space feature (F2), and the third latent space feature (F3). The specific process includes: Calculate the mutual information matrix S between the feature representations. The specific calculation formula is: S(i, j) = exp(F1(:, i) T F2(:, j) / τ), where τ is the temperature parameter, F1(:, i) is the first modal feature vector of the i-th sample, and F2(:, j) is the second modal feature vector of the j-th sample.

[0042] Calculate the mutual information I pos of the positive sample pairs. The specific calculation formula is: I pos = -log(diag(S) / Σ j S(i, j)); Calculate the mutual information I neg of the negative sample pairs. The specific calculation formula is: I neg = -log(Σ i≠j S(i, j) / (n 2 - n)); Obtain the total mutual information loss. The specific calculation formula is: L MI= mean(I pos ) + λ * mean(I neg ); Among them, λ is a trade - off parameter.

[0043] 2) Calculate the dynamic weight gating of the first latent space feature (F1), the second latent space feature (F2), and the third latent space feature (F3). The specific calculation process is as follows: Concatenate the three - modality features: F cat = [F1; F2; F3] ∈ R (3d×n) Calculate the attention score: A = MLP(F cat ) ∈ R (3×n) , where MLP contains two fully - connected layers, specifically expressed as: Z1 = ReLU(W1F cat + b1); A = Softmax(W2Z1 + b2); Among them, W1 ∈ R (h×3d) , is the weight matrix of the first fully - connected layer; b1 ∈ R h , is the bias vector of the first layer; h is the hidden - layer dimension of W1, generally taking 256 or 512; W2 ∈ R (3×h) , is the weight matrix of the second fully - connected layer; b2 ∈ R 3 , is the bias vector of the second layer, A ∈ R (3×n) , is the attention - score matrix of the three modalities.

[0044] 3) Conduct adversarial learning optimization. The specific calculation process is as follows: First, construct the generator network G. The specific calculation formula is: H = ReLU(W g1 F fused + b g1 )F gen = tanh(W g2 H + b g2 ); Among them, W g1 ∈ R (h×d) , is the weight matrix of the first layer of the generator; b g1 ∈ R h , is the bias vector of the first layer of the generator; W g2 ∈ R (d×h) , is the weight matrix of the second layer of the generator; b g2 ∈ R d , is the bias vector of the second layer of the generator.

[0045] Construct the discriminator network D. The specific calculation formula is: H d= ReLU(W d1 F + b d1 )D_ out = sigmoid(W d2 H d + b d2 ); Among them, W d1 ∈ R (h×d) , is the weight matrix of the first layer of the discriminator; b d1 ∈ R h , is the bias vector of the first layer of the discriminator; W d2 ∈ R (1×h) , is the weight matrix of the second layer of the discriminator; b d2 ∈ R, is the bias scalar of the second layer of the discriminator.

[0046] For the real sample F_real and the generated sample F_gen, calculate the discriminator loss. The specific calculation formula is: L D = -E[log D(F real )] - E[log(1 - D(F gen ))]; Calculate the generator loss. The specific calculation formula is: L G = -E[log D(F gen )]; 4) Perform feature fusion optimization. The specific calculation process: Combine the above three loss terms: L total = α1L MI + α2L D + α3 * LG, where α1, α2, α3 are the weight coefficients of each loss term. The value range of α1 is: 0.1 ~ 0.5, the value range of α2 is: 0.3 ~ 0.7, the value range of α3 is: 0.3 ~ 0.7, and η is 0.001.

[0047] Optimize the network parameters through backpropagation: θ new = θ old - η∇ θ L total ; Among them, η is the learning rate, θ new is the optimized network parameter, and is the network parameter before optimization.

[0048] The output optimized feature table of the multi-modal feature interaction mapping network is: F_ out = G(F_ fused ) Through the above process, semantic consistency between different modality features is ensured by maximizing mutual information, the importance of each modality is adaptively adjusted through a dynamic weight gating mechanism, the discriminability of feature representations is further enhanced through adversarial learning, and various optimization objectives are comprehensively considered through the feature fusion optimization process.

[0049] S1025. Perform cross-modal semantic compression on the output of the multi-modal feature interaction mapping network to generate low-dimensional consistent latent space features.

[0050] Exemplarily, design an orthogonal constraint non-linear projection transformation to minimize redundant information between modalities. Introduce the information bottleneck principle to control the information loss degree of feature compression; enhance the stability of feature compression through logarithmic mapping and geometric distance regularization.

[0051] S1026. Align the probability distributions of the low-dimensional consistent latent space features to obtain latent space feature encodings.

[0052] Exemplarily, use the Sinkhorn algorithm to solve the optimal transport problem to achieve soft alignment of probability distributions , control the smoothness of the probability distribution through information entropy constraint, and minimize the Wasserstein distance or KL divergence between the probability distributions of different modality features.

[0053] In some embodiments, calculate the homologous loss and double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation, including: S3011 - S3015.

[0054] S3011. Calculate the homologous loss for the latent space feature encoding to obtain the cross-modal semantic association loss.

[0055] Based on the Euclidean distance metric, calculate the distance between the latent space feature encodings of the same object in different modalities, and construct a similarity constraint matrix for the feature encodings between modalities. Achieve cross-modal semantic alignment by minimizing the Euclidean distance between the encodings of the same object in different modalities. Perform non-linear normalization processing on the similarity constraint matrix to generate the cross-modal semantic association loss.

[0056] S3012. According to the cross-modal semantic association loss, perform single-modal reconstruction on the latent space feature encoding to obtain the intra-modal semantic fidelity loss.

[0057] Perform independent reconstruction processing on the latent space feature encoding of each modality respectively, and construct an intra-modal feature reconstruction error evaluation model. By minimizing the reconstruction loss between the reconstructed feature and the original feature, introduce a regularization constraint term to control the information compression rate and feature fidelity during the reconstruction process.

[0058] S3013. Perform cross-modal reconstruction on the semantic fidelity loss within the modality to obtain the multi-modal collaborative reconstruction loss.

[0059] Based on the cross-modal reconstruction algorithm, map the latent space features of one modality to the feature space of another modality; design a cross-modal reconstruction information transfer network to achieve semantic bridging and feature transformation between modalities; optimize the semantic relevance between modalities by minimizing the cross-modal reconstruction error.

[0060] S3014. According to the multi-modal collaborative reconstruction loss, jointly optimize the latent space feature encoding to obtain a weight-balanced comprehensive loss function.

[0061] Construct a weighted combination model of the homologous loss, single-modal reconstruction loss, and cross-modal reconstruction loss; dynamically adjust the contribution ratio of different loss terms through an adaptive weight allocation mechanism; balance the reconstruction ability and generalization performance of the model based on information theory constraints.

[0062] S3015. Perform gradient backpropagation and parameter update on the weight-balanced comprehensive loss function to obtain the 3D object embedding representation.

[0063] Design a parameter update strategy based on gradient accumulation to alleviate the gradient instability in mini-batch training; introduce an adaptive learning rate decay mechanism to optimize the parameter convergence process; generate a semantically consistent and information-compressed 3D object embedding representation through multiple rounds of iterative optimization.

[0064] In some embodiments, aggregate the multi-modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix, including: S1041 - S1044.

[0065] S1041. Perform heterogeneous weight allocation on the optimized 3D object embedding representation to obtain an adaptive weight matrix.

[0066] S1042. Perform weighted combination on the 3D object embedding representation according to the adaptive weight matrix to obtain a multi-level semantic feature vector.

[0067] S1043. Perform inter-modal correlation analysis on the multi-level semantic feature vector to obtain an inter-modal correlation degree matrix, and hierarchically aggregate the multi-level semantic feature vector according to the inter-modal correlation degree matrix to obtain an intermediate embedding representation.

[0068] S1044. Calculate the inter-modal collaboration coefficient according to the intermediate embedding representation to obtain a unified 3D object embedding matrix.

[0069] By introducing an adaptive weight allocation and an inter-modal cooperation mechanism, refined processing of different modal features is achieved. Through heterogeneous weight allocation, the importance of each modal feature can be automatically learned and adjusted, avoiding the limitations of traditional static weight methods. The generation of multi-level semantic feature vectors and the inter-modal correlation analysis enable the system to capture complex cross-modal semantic associations, thereby obtaining a more robust and information-rich intermediate embedding representation. By calculating the inter-modal cooperation coefficient and constructing a unified 3D object embedding matrix, not only the comprehensiveness of the feature representation is improved, but also the fusion depth and accuracy of multi-modal information are enhanced, providing stronger technical support for subsequent 3D object recognition, classification, and understanding.

[0070] In some embodiments, hypergraph construction and hypergraph convolution are performed on the unified 3D object embedding matrix to obtain a structure-aware embedding representation, including: S1051 - S1055.

[0071] S1051. Evaluate the vertex correlation of the unified 3D object embedding matrix to obtain a hypergraph correlation matrix.

[0072] Exemplarily, the input 3D object embedding matrix contains feature vectors of multiple objects, such as object A [0.2, 0.5, 0.3], object B [0.3, 0.4, 0.3], etc. Calculate the Euclidean distance or cosine similarity for each object feature vector. For example, the cosine similarity between A and B is 0.95. Select K = 5 nearest neighbor points to construct a hyperedge. For example, object A forms a hyperedge with B, C, D, E, F, generating a hypergraph correlation matrix H, where H[i, j] = 1 indicates that vertex i belongs to hyperedge j.

[0073] S1052. Calculate the vertex degree and hyperedge degree of the hypergraph correlation matrix to obtain a normalized weight mapping matrix.

[0074] Exemplarily, calculate the degree of each vertex. For example, the degree of vertex A is the number of hyperedges connected to it, which is 5. Calculate the number of vertices included in each hyperedge. For example, hyperedge 1 contains 6 vertices. Construct a vertex degree diagonal matrix Dv and a hyperedge degree diagonal matrix De to obtain a normalized weight mapping matrix. The specific calculation formula is: W = Dv (-1 / 2) H De (-1) H T Dv (-1 / 2) ; S1053. According to the normalized weight mapping matrix, perform non-linear feature propagation on the unified 3D object embedding matrix to obtain a preliminary structure-aware feature.

[0075] Perform weighted propagation on each vertex feature vector. For example, A' = ∑(w i *v i ), w iThe propagation result is processed by nonlinear activation functions such as ReLU, and residual connection is introduced to retain the original feature information to obtain a feature vector containing structural information, such as A'[0.25, 0.45, 0.3].

[0076] S1054. Perform multi-scale information extraction on the preliminary structural perception features to obtain a multi-level structural feature representation.

[0077] A three-layer cascaded hypergraph convolutional network is designed. The first layer uses 1-hop neighborhood information to obtain local structural features, the second layer uses 2-hop neighborhood information to obtain medium-range structural information, and the third layer uses 3-hop neighborhood information to capture global structural information. The features of each layer are fused through jump connections.

[0078] S1055. Based on the multi-level structural feature representation, perform lateral semantic association constraints on the feature vector to obtain a structure-aware embedding representation.

[0079] Construct positive sample pairs (same type of objects) and negative sample pairs (different type of objects), calculate the contrast loss between feature vectors, shorten the distance between positive sample pairs, and expand the distance between negative sample pairs. Use dynamic queues to maintain historical feature information and obtain structure-aware feature representation.

[0080] Through the above method, the topological relationship between objects is captured by the association matrix, and the degree matrix is ​​used for reasonable normalization to achieve effective feature propagation and optimization, and obtain high-quality feature representation.

[0081] In some embodiments, the structure-aware embedding representation is aligned to a memory library to obtain an aligned embedding representation, and scene classification and recognition are performed based on the aligned embedding representation, including: S1061-S1065.

[0082] S1061. Calculate the memory bank activation score for the structure-aware embedding representation to obtain a semantic association activation matrix.

[0083] Based on the cosine similarity metric, the semantic similarity between the structure-aware embedding representation and the typical anchor points of the memory library is calculated; the memory library anchor point activation weight vector is constructed, and a semantic association activation matrix of probability distribution is generated through softmax normalization processing; the semantic association activation matrix is ​​adaptively thresholded and memory anchor points with high semantic significance are screened.

[0084] S1062. According to the semantic association activation matrix, the structure-aware embedding representation is reconstructed in memory to obtain a preliminary aligned embedding representation.

[0085] Through a weighted fusion strategy, semantic overlap is performed between the structure-aware embedding representation and the filtered memory anchors; a non-linear mapping function is designed to achieve spatial transformation of the structure-aware embedding and the memory anchor representation; a reconstruction loss function based on an attention mechanism is constructed to constrain the semantic consistency during the reconstruction process.

[0086] S1063. Perform semantic debiasing processing on the preliminary aligned embedding representation to obtain a debiased aligned embedding representation.

[0087] Construct a cross-category semantic difference evaluation model to quantify the semantic bias between different categories; design an entropy-based semantic debiasing constraint term to suppress the semantic shift between categories; through a contrastive learning strategy, enhance the category discrimination ability of the aligned embedding representation.

[0088] S1064. Based on the debiased aligned embedding representation, construct an open-set scene classification decision boundary to obtain a category discrimination matrix.

[0089] Based on the prototype network theory, calculate the distance between the debiased aligned embedding representation and the prototypes of known categories; design a dynamic boundary adaptive algorithm to adjust the classification threshold according to the distribution characteristics of the embedding representation; introduce Bayesian probability inference to generate the posterior probability distribution of multiple categories.

[0090] S1065. Perform confidence evaluation and classification processing on the category discrimination matrix to obtain the scene recognition result.

[0091] Construct a multi-stage confidence evaluation model to stratify the credibility of the prediction results of the category discrimination matrix; design a probability threshold classification strategy to directly output the high-confidence results, trigger secondary semantic reasoning for the medium-confidence results; mark the low-confidence regions with the "unknown" category to achieve open-set scene recognition.

[0092] Please refer to Figure 2 , Figure 2 FIG. is a schematic block diagram of a scene recognition device for a 3D and AI vision sensing visible light module according to an embodiment of the present application. The scene recognition device 200 for the 3D and AI vision sensing visible light module is used to execute the foregoing scene recognition method for the 3D and AI vision sensing visible light module. Among them, the scene recognition device 200 for the 3D and AI vision sensing visible light module can be configured in a server.

[0093] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0094] As shown Figure 2 in the figure, the scene recognition device 200 of the 3D and AI vision sensing visible light core includes: a feature extraction module 201, an encoding and compression module 202, a loss calculation module 203, a feature aggregation module 204, a hypergraph calculation module 205, and a result output module 206.

[0095] The feature extraction module 201 is configured to extract features from the multi-modal features of the 3D object to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix.

[0096] The encoding and compression module 202 is configured to automatically encode and compress the basic feature set to obtain a latent space feature encoding.

[0097] The loss calculation module 203 is configured to calculate the homology loss and the double reconstruction loss for the latent space feature encoding to obtain an optimized 3D object embedding representation.

[0098] The feature aggregation module 204 is configured to aggregate multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix.

[0099] The hypergraph calculation module 205 is configured to perform hypergraph construction and hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation.

[0100] The result output module 206 is configured to perform memory bank alignment on the structure-aware embedding representation to obtain an aligned embedding representation, and perform scene classification and recognition according to the aligned embedding representation.

[0101] An embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the 3D and AI vision sensing visible light core scene recognition method according to any one of the embodiments of the present application when executing the computer program.

[0102] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the 3D and AI vision sensing visible light core scene recognition method according to any one of the embodiments of the present application.

[0103] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present application, and these modifications or substitutions should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A scene recognition method for 3D and AI visual sensing visible light movement, characterized in that: The method comprises: Extracting multimodal features of 3D objects to obtain a basic feature set including a multi-view feature matrix, a point cloud feature matrix, and a voxel feature matrix; Automatically encoding and compressing the basic feature set to obtain a latent space feature code; Calculating homology loss and dual reconstruction loss on the latent space feature encoding to obtain an optimized 3D object embedding representation; aggregating multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix; Performing hypergraph construction and hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation; The structure-aware embedding representation is memory-bank aligned to obtain an aligned embedding representation, and scene classification and recognition is performed based on the aligned embedding representation.

2. The scene recognition method of 3D and AI visual sensing visible light core according to claim 1, characterized in that: The step of automatically encoding and compressing the basic feature set to obtain a latent space feature encoding includes: Performing deep convolution encoding on the multi-view feature matrix to map the multi-view features to first latent space features; Performing geometric transformation encoding on the point cloud feature matrix to map the point cloud features to second latent space features; Performing three-dimensional spatial transformation encoding on the voxel feature matrix to map the voxel features to third latent space features; According to the first latent space feature, the second latent space feature, and the third latent space feature, a multimodal feature interactive mapping network is constructed, and feature cross-correlation is achieved through a cross-modal attention mechanism and an information fusion module; Performing cross-modal semantic compression on the output of the multimodal feature interaction mapping network to generate low-dimensional consistent latent space features; Probability distribution alignment is performed on the low-dimensional consistent latent space features to obtain the latent space feature encoding.

3. The scene recognition method of 3D and AI visual sensing visible light core according to claim 1, characterized in that: The step of performing homology loss and dual reconstruction loss calculation on the latent space feature encoding to obtain an optimized 3D object embedding representation includes: Performing homology loss calculation on the latent space feature encoding to obtain cross-modal semantic association loss; According to the cross-modal semantic association loss, the latent space feature encoding is reconstructed in a unimodal manner to obtain an intra-modal semantic fidelity loss; Performing cross-modal reconstruction on the intra-modal semantic fidelity loss to obtain a multi-modal collaborative reconstruction loss; According to the multimodal collaborative reconstruction loss, the latent space feature encoding is jointly optimized to obtain a weight-balanced comprehensive loss function; The weight-balanced comprehensive loss function is subjected to gradient back-propagation and parameter updating to obtain the 3D object embedding representation.

4. The scene recognition method of 3D and AI visual sensing visible light core according to claim 1, characterized in that: The step of aggregating multiple modal features according to the optimized 3D object embedding representation to obtain a unified 3D object embedding matrix includes: Performing heterogeneous weight assignment on the optimized 3D object embedding representation to obtain an adaptive weight matrix; Performing weighted combination on the 3D object embedding representation according to the adaptive weight matrix to obtain a multi-level semantic feature vector; Performing inter-modal correlation analysis on the multi-level semantic feature vectors to obtain a modal correlation matrix, and hierarchically aggregating the multi-level semantic feature vectors according to the modal correlation matrix to obtain an intermediate embedding representation; The inter-modal synergy coefficient is calculated according to the intermediate embedding representation to obtain the unified 3D object embedding matrix.

5. The scene recognition method of 3D and AI visual sensing visible light core as claimed in claim 4, characterized in that: The performing hypergraph construction and hypergraph convolution on the unified 3D object embedding matrix to obtain a structure-aware embedding representation includes: Performing vertex association evaluation on the unified 3D object embedding matrix to obtain a hypergraph association matrix; Calculating the vertex degree and the hyperedge degree of the hypergraph association matrix to obtain a normalized weight mapping matrix; According to the normalized weight mapping matrix, nonlinear feature propagation is performed on the unified 3D object embedding matrix to obtain preliminary structure perception features; Performing multi-scale information extraction on the preliminary structure perception features to obtain a multi-level structure feature representation; Based on the multi-level structural feature representation, lateral semantic association constraints are performed on feature vectors to obtain the structure-aware embedding representation.

6. The scene recognition method of 3D and AI visual sensing visible light core as claimed in claim 5, characterized in that: The step of performing memory bank alignment on the structure-aware embedding representation to obtain an aligned embedding representation, and performing scene classification and recognition according to the aligned embedding representation includes: Calculating a memory bank activation score for the structure-aware embedding representation to obtain a semantic association activation matrix; According to the semantic association activation matrix, the structure-aware embedding representation is subjected to memory bank reconstruction processing to obtain a preliminary aligned embedding representation; Performing semantic debiasing processing on the preliminary aligned embedding representation to obtain a debiased aligned embedding representation; According to the debiased aligned embedding representation, an open set scene classification decision boundary is constructed to obtain a category discrimination matrix; Confidence evaluation and classification processing are performed on the category discrimination matrix to obtain a scene recognition result.