Article identification system based on computer vision

Through multimodal data acquisition and deep learning algorithms, combined with adaptive perception modules and filtering prediction technology, the problem of recognition accuracy in fruit occlusion scenarios was solved, and high-precision three-dimensional spatial coordinate recognition of fruits was achieved.

CN120747459AInactive Publication Date: 2025-10-03HENAN LANOU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510814862.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In scenes where fruits are densely packed and obscured by multiple layers of leaves, existing technologies have complex spatial relationships between fruits, insufficient Repulsion Loss recall rate, and difficulty in modeling the hierarchical relationship of obscured objects, resulting in insufficient fruit recognition accuracy.

Method used

By adopting multimodal data acquisition, feature extraction, multimodal fusion and adaptive perception modules, combined with deep learning algorithms and filtering prediction technology, an environmental perception parameter set is generated to realize the three-dimensional spatial coordinate recognition of fruits.

Benefits of technology

It improves the recognizability of occluded target features in fruit recognition, reduces light noise interference, and achieves millimeter-level recognition accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747459A_ABST
    Figure CN120747459A_ABST
Patent Text Reader

Abstract

The invention discloses an article recognition system based on computer vision. The article recognition system comprises a multi-modal data acquisition module, a multi-modal data processing module and a computer vision processing module, wherein the multi-modal data acquisition module is used for acquiring multi-modal data through a multi-modal sensor array; the data preprocessing module is used for standardizing a multi-modal data format and generating a time-space aligned multi-modal tensor; the feature extraction module is used for respectively extracting modal specific features from texture, spectrum and geometric dimensions by adopting ResNet50, 3D-CNN and PointNet + +; the multi-modal fusion module is used for constructing cross-modal joint representation; the adaptive sensing module is used for modeling illumination invariance and scene dynamics based on self-supervised comparative learning and a 3D-STMN space-time memory network, predicting a shielded target trajectory by using Kalman filtering in combination with the shielding sensing propagation module, and generating an environment sensing parameter set; and the recognition engine module is used for integrating YOLOv8 detection, Mask R-CNN segmentation and multi-modal decision tree classification, outputting a target bounding box, a category and confidence in combination with the depth data, and generating three-dimensional space coordinates combined with the depth data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object recognition, and in particular to an object recognition system based on computer vision. Background Art

[0002] In the process of agricultural modernization, harvesting robots, a core development direction for intelligent agricultural machinery, are facing technical bottlenecks in achieving efficient and precise operation. Computer vision-based object recognition systems offer an innovative approach. This system uses deep learning algorithms to extract and classify features from fruit images, employing convolutional neural network (CNN) models to identify the morphology, color, and maturity of different fruit varieties. It also incorporates multispectral imaging technology to enhance environmental adaptability. The system can locate the target fruit in real time, plan the optimal harvesting path, and achieve millimeter-level recognition accuracy in complex natural environments, providing key technical support for automated agricultural robotic harvesting.

[0003] In existing technologies, in scenarios where grape bunches, cherries, and other fruits are densely packed and obscured by multiple layers of leaves, the surface features of the fruits are severely obscured, complicating the spatial relationships between the fruits. Repulsion Loss has insufficient recall in scenarios with multiple layers of leaf occlusion, making it difficult to model the hierarchical relationships of the occluded objects. Therefore, a computer vision-based object recognition system is proposed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an object recognition system based on computer vision.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A computer vision-based object recognition system, comprising:

[0007] Multimodal data acquisition module: acquires multimodal data through a multimodal sensor array, including visible light, spectrum, depth, and motion data;

[0008] Data preprocessing module: standardizes the multimodal data format, aligns the timestamps of the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor;

[0009] Feature extraction module: ResNet50, 3D-CNN, and PointNet++ are used to extract modality-specific features from texture, spectrum, and geometry dimensions, respectively, to form a set of multimodal feature vectors;

[0010] Multimodal fusion module: HyperTransformer is used for early feature fusion, combined with GBDT weighted integration for late decision fusion, dynamic reinforcement learning is used to optimize weight distribution, and a cross-modal joint representation is constructed;

[0011] Adaptive Perception Module: Based on self-supervised contrastive learning (SimCLR) and 3D-STMN spatiotemporal memory network, it models illumination invariance and scene dynamics. Combined with the Occlusion Aware Propagation (OAP) module, it uses Kalman filtering to predict the trajectory of occluded objects and generate a set of environmental perception parameters.

[0012] Recognition engine module: Integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, combines depth data to output target bounding boxes, categories, and confidence levels, and generates three-dimensional spatial coordinates combined with depth data.

[0013] The above technical solution further includes:

[0014] Furthermore, the multimodal data acquisition module collects visible light through a visible light camera, collects spectrum through a hyperspectral camera, collects depth data through a structured light / ToF depth sensor, and collects motion data through an IMU.

[0015] Furthermore, the data preprocessing module standardizes the multimodal data format, performs timestamp alignment on the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor.

[0016] Data normalization: Normalize the visible light image to [0,1]. The calculation formula is expressed as The whiteboard reference method is used to calibrate the spectral data. The calculation formula is expressed as Perform median filtering to denoise the depth data, expressed as D filtered (x,y)=median(D(x±k,y±l)) k,l∈[—w,w];

[0017] Timestamp alignment: Use a synchronization signal generator (such as NI-DAQ) to generate PPS (pulse per second) to trigger all sensors, downsample high-frequency sensors (IMU), and perform linear interpolation on low-frequency sensors (spectrometer);

[0018] Spatial coordinate system alignment: Use Zhang Zhengyou calibration method to obtain the intrinsic parameter matrix K and distortion coefficient D, convert the depth map into a point cloud, and then project it into the RGB camera coordinate system;

[0019] Spatiotemporal alignment tensor generation: Create a sliding time window and use 3D convolution to process spatiotemporal features, denoted as F temporal =Conv3D(K 3×3×3 , stride=1), project the spectral vector (1×N) onto the image grid, expressed as Concatenate multimodal features to obtain

[0020] Furthermore, the feature extraction module uses ResNet50, 3D-CNN, and PointNet++ to extract modality-specific features from texture, spectrum, and geometric dimensions, respectively, to form a multimodal feature vector set, including the following steps:

[0021] Extracting modality-specific features from the texture dimension:

[0022] Initial convolution: extract the underlying features of the image (edges / textures), expand the receptive field through large convolution kernel + downsampling, input RGB image, perform convolution, and output

[0023] Bottleneck residual block: alleviates gradient vanishing through residual learning, performs three-stage downsampling to gradually capture high-level semantic features (such as object parts → whole), and the first stage outputs Second stage output The third stage output

[0024] Feature Pyramid Network (FPN): Fusing multi-scale features and performing 1×1 convolution on F4 Upsample P4 and add it to F3, and generate it through 3×3 convolution Upsample P3 and add it to F2, and generate it through 5×5 convolution

[0025] Extract modality-specific features from the spectral dimension:

[0026] 3D convolutional network: Simultaneously modeling the spectral dimension (material reflectance characteristics) and spatiotemporal dimension (dynamic lighting / target motion), extracting spectral-spatiotemporal joint features, and inputting the spectral cube I spec ∈R H×W×C×T , where C is the number of spectral channels and T is the time frame, and a 3D convolution operation is performed. The convolution kernel is represented as The output feature map is represented as

[0027] Where σ is the ReLU activation function, and the multi-layer 3D convolution is flattened into a feature vector sequence {f1,f2,...,f T};

[0028] Transformer encoder: It establishes long-distance dependencies through the self-attention mechanism, captures non-local correlations in the spectral sequence, and generates global spectral features through multi-layer Transformers. The generated global spectral features are expressed as where c t The context vector is represented as α tj The attention weight is expressed as

[0029] Q, K, V represent query / key / value matrices respectively;

[0030] Extract modality-specific features from geometric dimensions:

[0031] Layered sampling and grouping: gradually focus on the local structure of the point cloud and use the farthest point sampling to select the center point Ball query generates local area B r (p i );

[0032] Multi-scale feature aggregation: Feature aggregation for each local area is expressed as Multi-layer iteration generates multi-scale features

[0033] Global pooling: Splicing the features of each layer and performing global maximum pooling is expressed as in

[0034] Multimodal feature set: Integrating three types of heterogeneous features, texture, spectrum, and geometry, the generated submodal feature matrix is

[0035] Furthermore, the multimodal fusion module implements early feature fusion through HyperTransformer, combines GBDT weighted integration for late decision fusion, optimizes weight distribution through dynamic reinforcement learning, and constructs cross-modal joint representation, including the following steps:

[0036] HyperTransformer processing: Modeling long-range dependencies between modalities through self-attention mechanisms, enhancing key features (such as the association between spectral materials and geometric shapes), suppressing redundant information, and generating more robust joint feature representations;

[0037] GBDT weighted integration: It integrates the prediction results of each modality classifier and uses the decision tree to balance the prediction confidence of different modalities. The integration formula is expressed as:

[0038]

[0039] where η i is the weight of the i-th tree, T i is the i-th decision tree;

[0040] Dynamic weight optimization: adjust modal weights according to environmental conditions;

[0041] Joint representation generation: The deep features output by HyperTransformer are combined with the GBDT integration results to form a final representation that contains low-dimensional details and high-dimensional decision information.

[0042] Furthermore, the adaptive perception module models illumination invariance and scene dynamics based on self-supervised contrastive learning (SimCLR) and 3D-STMN spatiotemporal memory network, including the following steps:

[0043] Data enhancement: Randomly enhance the input frame x_i (lighting adjustment, color jitter, etc.) to generate positive sample pairs and negative sample set

[0044] Encoder and Projection Head: Use ResNet encoder f(·) and MLP projection head g(·) to extract features, denoted as z i =g(f(x i )),

[0045] Contrastive loss calculation: InfoNCE loss is used to optimize the feature space, expressed as:

[0046]

[0047] Where τ is the temperature hyperparameter;

[0048] Spatiotemporal feature encoding: Input SimCLR features into the 3D convolutional network to extract the spatiotemporal feature cube C∈R T×H×W×D ;

[0049] Memory reading and updating: From the memory matrix M∈R through the attention mechanism N×D Retrieve relevant memories from the source and then fuse the current features with the memories;

[0050] Dynamic context prediction: Generating spatiotemporal enhanced features.

[0051] Furthermore, the Occlusion Aware Propagation (OAP) module uses Kalman filtering to predict the trajectory of the occluded target and generate an environment perception parameter set, including the following steps:

[0052] Occlusion detection: The occlusion state is determined by the confidence threshold of the target detector;

[0053] Kalman filter prediction: define the target state vector s t =[x,y,v x ,v y ] T , make a prediction, the prediction expression is Where F is the state transfer matrix, Q is the process noise covariance;

[0054] Trajectory compensation: Map the predicted trajectory to the image plane to generate a virtual bounding box b virtual ;

[0055] Feature fusion: Combine SimCLR illumination invariant features, 3D-STMN spatiotemporal features, and OAP trajectory prediction results;

[0056] Parameter decoding: Generate environmental parameters through a fully connected network, ultimately achieving robust multimodal environmental perception.

[0057] Furthermore, the recognition engine module integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, combines depth data to output target bounding boxes, categories, and confidence levels, and generates three-dimensional spatial coordinates combined with depth data.

[0058] YOLOv8 object detection: Uses the YOLOv8 network to quickly detect objects in RGB images, outputs coarse-grained bounding boxes and preliminary category confidences, and uses non-maximum suppression (NMS) to filter overlapping boxes.

[0059] Mask R-CNN instance segmentation: Mask R-CNN performs ROIAlign operation on the target area detected by YOLO, extracts local features, generates binary masks, and segments the target contour;

[0060] Depth data fusion and 3D coordinate calculation: Apply a binary mask to the depth map to extract the depth value of the target area, and use the camera intrinsic parameter matrix to convert the pixel coordinates into 3D space coordinates;

[0061] Multimodal decision tree classification: A multimodal decision tree is used to fuse visual features (category vectors output by YOLO), shape features (binary mask geometric parameters), and spatial features (3D spatial coordinate statistics) to construct a decision tree. Hierarchical classification is performed based on feature importance, and hierarchical classification rules are used to eliminate unimodal ambiguity.

[0062] Output structured results: The final output includes structured recognition results including 3D coordinates, confidence, and instance masks.

[0063] The present invention has the following beneficial effects:

[0064] In the present invention, self-supervised contrastive learning (SimCLR) generates positive sample pairs of the same fruit and negative sample pairs of different fruits through data enhancement such as random cropping and color jittering. The feature extractor is trained using contrastive loss (InfoNCE) to make the feature distribution of occluded fruits consistent with that of unoccluded fruits, improve the recognizability of occluded target features, and reduce illumination noise interference. The 3D-STMN spatiotemporal memory network captures the spatial relationship between leaves and fruits through 3D convolution, uses LSTM to model the dynamic trajectory of the occluded target, predicts the spatiotemporal pattern of leaf occlusion, and corrects the detection frame position in advance. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a system block diagram of a computer vision-based object recognition system proposed in the present invention. DETAILED DESCRIPTION

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0067] See also Figure 1 As shown, the present invention is an object recognition system based on computer vision, comprising:

[0068] Multimodal data acquisition module: acquires multimodal data through a multimodal sensor array, including visible light, spectrum, depth, and motion data;

[0069] Data preprocessing module: standardizes the multimodal data format, aligns the timestamps of the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor;

[0070] Feature extraction module: ResNet50, 3D-CNN, and PointNet++ are used to extract modality-specific features from texture, spectrum, and geometry dimensions, respectively, to form a set of multimodal feature vectors;

[0071] Multimodal fusion module: HyperTransformer is used for early feature fusion, combined with GBDT weighted integration for late decision fusion, dynamic reinforcement learning is used to optimize weight distribution, and a cross-modal joint representation is constructed;

[0072] Adaptive Perception Module: Based on self-supervised contrastive learning (SimCLR) and 3D-STMN spatiotemporal memory network, it models illumination invariance and scene dynamics. Combined with the Occlusion Aware Propagation (OAP) module, it uses Kalman filtering to predict the trajectory of occluded objects and generate a set of environmental perception parameters.

[0073] Recognition engine module: Integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, combines depth data to output target bounding boxes, categories, and confidence levels, and generates three-dimensional spatial coordinates combined with depth data.

[0074] In one embodiment, the multimodal data acquisition module collects visible light through a visible light camera, collects spectrum through a hyperspectral camera, collects depth data through a structured light / ToF depth sensor, and collects motion data through an IMU.

[0075] In one embodiment, the data preprocessing module standardizes the multimodal data format, performs timestamp alignment on the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor. Specific steps include:

[0076] Data normalization: Normalize the visible light image to [0,1]. The calculation formula is expressed as The whiteboard reference method is used to calibrate the spectral data. The calculation formula is expressed as Perform median filtering to denoise the depth data, expressed as D filtered (x,y)=median(D(x±k,y±l)) k,l∈[-w,w];

[0077] Timestamp alignment: Use a synchronization signal generator (such as NI-DAQ) to generate PPS (pulse per second) to trigger all sensors, downsample high-frequency sensors (IMU), and perform linear interpolation on low-frequency sensors (spectrometer);

[0078] Spatial coordinate system alignment: Use Zhang Zhengyou calibration method to obtain the intrinsic parameter matrix K and distortion coefficient D, convert the depth map into a point cloud, and then project it into the RGB camera coordinate system;

[0079] Spatiotemporal alignment tensor generation: Create a sliding time window and use 3D convolution to process spatiotemporal features, denoted as F temporal =Conv3D(K 3×3×3 , stride=1) projects the spectral vector (1×N) onto the image grid, expressed as Concatenate multimodal features to obtain

[0080] In one embodiment, the feature extraction module uses ResNet50, 3D-CNN, and PointNet++ to extract modality-specific features from texture, spectrum, and geometric dimensions respectively to form a multimodal feature vector set, including the following steps:

[0081] Extracting modality-specific features from the texture dimension:

[0082] Initial convolution: extract the underlying features of the image (edges / textures), expand the receptive field through large convolution kernel + downsampling, input RGB image, perform convolution, and output

[0083] Bottleneck residual block: alleviates gradient vanishing through residual learning, performs three-stage downsampling to gradually capture high-level semantic features (such as object parts → whole), and the first stage outputs Second stage output The third stage output

[0084] Feature Pyramid Network (FPN): Fusing multi-scale features and performing 1×1 convolution on F4 Upsample P4 and add it to F3, and generate it through 3×3 convolution Upsample P3 and add it to F2, and generate it through 5×5 convolution

[0085] Extract modality-specific features from the spectral dimension:

[0086] 3D convolutional network: Simultaneously modeling the spectral dimension (material reflectance characteristics) and spatiotemporal dimension (dynamic lighting / target motion), extracting spectral-spatiotemporal joint features, and inputting the spectral cube I Spec ∈R H×W×C×T , where C is the number of spectral channels and T is the time frame, and a 3D convolution operation is performed. The convolution kernel is represented as The output feature map is represented as Where σ is the ReLU activation function, and the multi-layer 3D convolution is flattened into a feature vector sequence {f1,f2,...,f T};

[0087] Transformer encoder: It establishes long-distance dependencies through the self-attention mechanism, captures non-local correlations in the spectral sequence, and generates global spectral features through multi-layer Transformers. The generated global spectral features are expressed as where c t The context vector is represented as α tj The attention weight is expressed as Q, K, V represent query / key / value matrices respectively;

[0088] Extract modality-specific features from geometric dimensions:

[0089] Layered sampling and grouping: gradually focus on the local structure of the point cloud and use the farthest point sampling to select the center point Ball query generates local area B r (p i );

[0090] Multi-scale feature aggregation: Feature aggregation for each local area is expressed as Multi-layer iteration generates multi-scale features

[0091] Global pooling: Splicing the features of each layer and performing global maximum pooling is expressed as in

[0092] Multimodal feature set: Integrating three types of heterogeneous features, texture, spectrum, and geometry, the generated submodal feature matrix is

[0093] In one embodiment, the multimodal fusion module implements early feature fusion through HyperTransformer, combines GBDT weighted integration for late decision fusion, optimizes weight distribution through dynamic reinforcement learning, and constructs a cross-modal joint representation, including the following steps:

[0094] HyperTransformer processing: Modeling long-range dependencies between modalities through self-attention mechanisms, enhancing key features (such as the association between spectral materials and geometric shapes), suppressing redundant information, and generating more robust joint feature representations;

[0095] GBDT weighted integration: It integrates the prediction results of each modality classifier and uses the decision tree to balance the prediction confidence of different modalities. The integration formula is expressed as:

[0096]

[0097] where η i is the weight of the i-th tree, T i is the i-th decision tree;

[0098] Dynamic weight optimization: adjust modal weights according to environmental conditions;

[0099] Joint representation generation: The deep features output by HyperTransformer are combined with the GBDT integration results to form a final representation that contains low-dimensional details and high-dimensional decision information.

[0100] In one embodiment, the adaptive perception module models illumination invariance and scene dynamics based on self-supervised contrastive learning (SimCLR) and 3D-STMN spatiotemporal memory network, including the following steps:

[0101] Data enhancement: Randomly enhance the input frame x_i (lighting adjustment, color jitter, etc.) to generate positive sample pairs and negative sample set

[0102] Encoder and Projection Head: Use ResNet encoder f(·) and MLP projection head g(·) to extract features, denoted as z i =g(f(x i )),

[0103] Contrastive loss calculation: InfoNCE loss is used to optimize the feature space, expressed as:

[0104]

[0105] Where τ is the temperature hyperparameter;

[0106] Spatiotemporal feature encoding: Input SimCLR features into the 3D convolutional network to extract the spatiotemporal feature cube C∈R T×H×W×D ;

[0107] Memory reading and updating: From the memory matrix M∈R through the attention mechanism N×D Retrieve relevant memories from the source and then fuse the current features with the memories;

[0108] Dynamic context prediction: Generating spatiotemporal enhanced features.

[0109] In one embodiment, the Occlusion Aware Propagation (OAP) module uses Kalman filtering to predict the trajectory of the occluded target and generate an environment perception parameter set, including the following steps:

[0110] Occlusion detection: The occlusion state is determined by the confidence threshold of the target detector;

[0111] Kalman filter prediction: define the target state vector s t =[x,y,v x ,v y ], make a prediction, the prediction expression is Where F is the state transfer matrix, Q is the process noise covariance;

[0112] Trajectory compensation: Map the predicted trajectory to the image plane to generate a virtual bounding box b virtual ;

[0113] Feature fusion: Combine SimCLR illumination invariant features, 3D-STMN spatiotemporal features, and OAP trajectory prediction results;

[0114] Parameter decoding: Generate environmental parameters through a fully connected network, ultimately achieving robust multimodal environmental perception.

[0115] In one embodiment, the recognition engine module integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, combines depth data to output target bounding boxes, categories, and confidence levels, and generates three-dimensional spatial coordinates combined with depth data.

[0116] YOLOv8 object detection: Uses the YOLOv8 network to quickly detect objects in RGB images, outputs coarse-grained bounding boxes and preliminary category confidences, and uses non-maximum suppression (NMS) to filter overlapping boxes.

[0117] Mask R-CNN instance segmentation: Mask R-CNN performs ROIAlign operation on the target area detected by YOLO, extracts local features, generates binary masks, and segments the target contour;

[0118] Depth data fusion and 3D coordinate calculation: Apply a binary mask to the depth map to extract the depth value of the target area, and use the camera intrinsic parameter matrix to convert the pixel coordinates into 3D space coordinates;

[0119] Multimodal decision tree classification: A multimodal decision tree is used to fuse visual features (category vectors output by YOLO), shape features (binary mask geometric parameters), and spatial features (3D spatial coordinate statistics) to construct a decision tree. Hierarchical classification is performed based on feature importance, and hierarchical classification rules are used to eliminate unimodal ambiguity.

[0120] Output structured results: The final output includes structured recognition results including 3D coordinates, confidence, and instance masks.

[0121] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An object recognition system based on computer vision, characterized in that: include: Multimodal data acquisition module: acquires multimodal data through a multimodal sensor array, including visible light, spectrum, depth, and motion data; Data preprocessing module: standardizes the multimodal data format, aligns the timestamps of the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor; Feature extraction module: ResNet50, 3D-CNN, and PointNet++ are used to extract modality-specific features from texture, spectrum, and geometry dimensions, respectively, to form a set of multimodal feature vectors; Multimodal fusion module: HyperTransformer is used for early feature fusion, combined with GBDT weighted integration for late decision fusion, dynamic reinforcement learning is used to optimize weight distribution, and a cross-modal joint representation is constructed; Adaptive Perception Module: Based on self-supervised contrastive learning and 3D-STMN spatiotemporal memory network modeling of illumination invariance and scene dynamics, combined with the occlusion perception propagation module using Kalman filtering to predict the trajectory of occluded targets and generate an environmental perception parameter set; Recognition engine module: Integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, combines depth data to output target bounding boxes, categories, and confidence levels, and generates three-dimensional spatial coordinates combined with depth data.

2. The object recognition system based on computer vision according to claim 1, characterized in that: The multimodal data acquisition module collects visible light through a visible light camera, collects spectrum through a hyperspectral camera, collects depth data through a structured light / ToF depth sensor, and collects motion data through an IMU.

3. The object recognition system based on computer vision according to claim 1, characterized in that: The data preprocessing module standardizes the multimodal data format, performs timestamp alignment on the standardized multimodal data, and generates a spatiotemporally aligned multimodal tensor. Data normalization: Normalize the visible light image to [0,1]. The calculation formula is expressed as The whiteboard reference method is used to calibrate the spectral data. The calculation formula is expressed as Perform median filtering to denoise the depth data, expressed as D filtered (x,y)=median(D(x±k,y±l))k,l∈[-w,w]; Timestamp alignment: Use a synchronization signal generator to generate PPS to trigger all sensors, downsample high-frequency sensors, and perform linear interpolation on low-frequency sensors; Spatial coordinate system alignment: Use Zhang Zhengyou calibration method to obtain the intrinsic parameter matrix K and distortion coefficient D, convert the depth map into a point cloud, and then project it into the RGB camera coordinate system; Spatiotemporal alignment tensor generation: Create a sliding time window and use 3D convolution to process spatiotemporal features, denoted as F temporal =Conv3D(K 3×3×3 , stride = 1), projecting the spectral vector onto the image grid, expressed as Concatenate multimodal features to obtain 4. The object recognition system based on computer vision according to claim 1, characterized in that: The feature extraction module uses ResNet50, 3D-CNN, and PointNet++ to extract modality-specific features from texture, spectrum, and geometric dimensions respectively to form a multimodal feature vector set, including the following steps: Extracting modality-specific features from the texture dimension: Initial convolution: extract the underlying features of the image, expand the receptive field through large convolution kernel + downsampling, input RGB image, perform convolution, and output Bottleneck residual block: alleviates gradient vanishing through residual learning, performs three-stage downsampling to gradually capture high-level semantic features, and the first stage outputs Second stage output The third stage output Feature Pyramid Network: Fusion of multi-scale features, 1×1 convolution of F4 Upsample P4 and add it to F3, and generate it through 3×3 convolution Upsample P3 and add it to F2, and generate it through 5×5 convolution Extract modality-specific features from the spectral dimension: 3D Convolutional Network: Simultaneously model the spectral dimension and spatiotemporal dimension, extract the spectral-spatiotemporal joint features, and input the spectral cube I Spec ∈R H×W×C×T , where C is the number of spectral channels and T is the time frame, a 3D convolution operation is performed and the convolution kernel is expressed as The output feature map is represented as Where σ is the ReLU activation function, and the multi-layer 3D convolution is flattened into a feature vector sequence {f1, f2, ..., f T }; Transformer encoder: It establishes long-distance dependencies through the self-attention mechanism, captures non-local correlations in the spectral sequence, and generates global spectral features through multi-layer Transformers. The generated global spectral features are expressed as where c t The context vector is represented as α tj The attention weight is expressed as Q, K, V represent query / key / value matrices respectively; Extract modality-specific features from geometric dimensions: Layered sampling and grouping: gradually focus on the local structure of the point cloud and use the farthest point sampling to select the center point Ball query generates local area B r (p i ); Multi-scale feature aggregation: Feature aggregation for each local area is expressed as Multi-layer iteration generates multi-scale features Global pooling: Splicing the features of each layer and performing global maximum pooling is expressed as in Multimodal feature set: Integrating three types of heterogeneous features, texture, spectrum, and geometry, the generated submodal feature matrix is 5. The object recognition system based on computer vision according to claim 1, characterized in that: The multimodal fusion module uses HyperTransformer to achieve early feature fusion, combines GBDT weighted integration for late decision fusion, optimizes weight distribution through dynamic reinforcement learning, and constructs a cross-modal joint representation. It includes the following steps: HyperTransformer processing: Modeling long-distance dependencies between modalities through self-attention mechanism; GBDT weighted integration: It integrates the prediction results of each modality classifier and uses the decision tree to balance the prediction confidence of different modalities. The integration formula is expressed as: where η i is the weight of the i-th tree, T i is the i-th decision tree; Dynamic weight optimization: adjust modal weights according to environmental conditions; Joint representation generation: The deep features output by HyperTransformer are combined with the GBDT integration results to form a final representation that contains low-dimensional details and high-dimensional decision information.

6. The object recognition system based on computer vision according to claim 1, characterized in that: The adaptive perception module models illumination invariance and scene dynamics based on self-supervised contrastive learning and 3D-STMN spatiotemporal memory network, including the following steps: Data enhancement: Randomly enhance the input frame x_i (lighting adjustment, color jitter, etc.) to generate positive sample pairs and negative sample set Encoder and Projection Head: Use ResNet encoder f(·) and MLP projection head g(·) to extract features, denoted as z i =g(f(x i )) Contrastive loss calculation: InfoNCE loss is used to optimize the feature space, expressed as: Where τ is the temperature hyperparameter; Spatiotemporal feature encoding: Input SimCLR features into the 3D convolutional network to extract the spatiotemporal feature cube C∈R T×H×W×D ; Memory reading and updating: From the memory matrix M∈R through the attention mechanism N×D Retrieve relevant memories from the source and then fuse the current features with the memories; Dynamic context prediction: Generating spatiotemporal enhanced features.

7. The object recognition system based on computer vision according to claim 6, characterized in that: The method of combining the occlusion perception propagation module with the Kalman filter to predict the trajectory of the occluded target and generate the environment perception parameter set includes the following steps: Occlusion detection: The occlusion state is determined by the confidence threshold of the target detector; Kalman filter prediction: define the target state vector Make a prediction, the prediction expression is Where F is the state transfer matrix, Q is the process noise covariance; Trajectory compensation: Map the predicted trajectory to the image plane to generate a virtual bounding box b virtual ; Feature fusion: Combine SimCLR illumination invariant features, 3D-STMN spatiotemporal features, and OAP trajectory prediction results; Parameter decoding: Generate environmental parameters through a fully connected network, ultimately achieving robust multimodal environmental perception.

8. The object recognition system based on computer vision according to claim 1, characterized in that: The recognition engine module integrates YOLOv8 detection, Mask R-CNN segmentation, and multimodal decision tree classification, and outputs the target bounding box, category, and confidence level in combination with depth data. The specific steps for generating three-dimensional spatial coordinates combined with depth data are as follows: YOLOv8 object detection: Uses the YOLOv8 network to quickly detect objects in RGB images, outputs coarse-grained bounding boxes and preliminary category confidences, and uses non-maximum suppression to filter overlapping boxes. Mask R-CNN instance segmentation: Mask R-CNN performs ROI Align on the target area detected by YOLO, extracts local features, generates a binary mask, and segments the target outline; Depth data fusion and 3D coordinate calculation: Apply a binary mask to the depth map to extract the depth value of the target area, and use the camera intrinsic parameter matrix to convert the pixel coordinates into 3D space coordinates; Multimodal decision tree classification: A multimodal decision tree is used to fuse visual features (category vectors output by YOLO), shape features (binary mask geometric parameters), and spatial features (3D spatial coordinate statistics) to construct a decision tree and perform hierarchical classification based on feature importance. Output structured results: The final output includes structured recognition results including 3D coordinates, confidence, and instance masks.

Citation Information

Cited By

  • Spraying robot crack recognition and self-adaptive spraying method and equipment based on machine vision and medium

    CN120900911A

  • A machine vision-based spraying robot crack identification and adaptive spraying method, device and medium

    CN120900911B

  • Artificial intelligence classification method for object appearance detection

    CN121010828A

  • Method and system for generating surrounding rock structural plane model based on three-dimensional machine vision

    CN121053329A

  • Method and system for generating surrounding rock structure surface model based on three-dimensional machine vision

    CN121053329B