Material tracking multi-target identification method and system based on 3D +2D high-dimensional feature fusion

By integrating 3D and 2D features in the material tracking system and combining advanced target recognition and tracking algorithms, the problems of distance information loss and environmental interference in traditional methods are solved, and accurate material tracking and high-quality product production are achieved.

CN120147791APending Publication Date: 2025-06-13UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510064468.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-13

Smart Images

  • Figure CN120147791A_ABST
    Figure CN120147791A_ABST
Patent Text Reader

Abstract

The invention provides a 3D + 2D high-dimensional feature fusion-based material tracking multi-target identification method and system. The method comprises the steps of performing target identification on a to-be-tracked material by using a YOLO algorithm; performing multi-target real-time tracking on the material based on a Deep SORT algorithm; generating three-dimensional point cloud data of a target by adopting a multi-view data calibration and alignment algorithm, and carrying out noise reduction, sampling and data enhancement preprocessing operation; performing three-dimensional reconstruction on the preprocessed three-dimensional point cloud data by using a Gaussian Splitting method, and introducing time sequence information at the same time; segmenting a dynamic three-dimensional target and a static background through a Mask R-CNN technology; aligning the position of the target in the two-dimensional image with the coordinates in the three-dimensional space by using a feature fusion method; a PNP algorithm is adopted to estimate attitude and depth information of a target in a three-dimensional space, accurate positioning and state restoration of materials are realized, and efficient and reliable technical support is provided for material management and tracking in an industrial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent manufacturing, and particularly to a method and system for multi-target recognition of material tracking based on 3D+2D high-dimensional feature fusion. Background Art

[0002] Traditional material tracking methods usually rely on planar 2D images. Although they can quickly identify targets, they have significant limitations: they cannot obtain the far and near distance information of the targets. The lack of distance data makes it more difficult to achieve precise control in the production manufacturing process, unable to accurately judge the spatial relationship between objects and equipment, which in turn leads to inaccurate control of equipment actions, affecting production efficiency and product quality. At the same time, the field of view of a single camera limits the comprehensive capture of objects, and cross-camera tracking faces the challenge that the appearances of materials are almost the same, making it difficult to accurately match the ID of the same material between different cameras through appearance features.

[0003] In addition, the production environment in industrial sites is usually complex and messy. Ordinary target recognition methods are easily interfered by the surrounding environment and often face problems such as object occlusion, which easily leads to the disruption of the unique identification ID of the tracking object. For example, production equipment and objects may be similar to the background environment in color or texture, and traditional recognition algorithms are difficult to effectively distinguish objects from the environmental background, thus posing great challenges to target recognition and material tracking. Summary of the Invention

[0004] Aiming at the above problems, the purpose of the present invention is to provide a method and system for multi-target recognition of material tracking based on 3D+2D high-dimensional feature fusion, so as to achieve precise tracking in the industrial material production process, effectively improve the robustness and anti-interference ability of the system, and ensure the continuous tracking and precise positioning of materials in a complex production environment.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] On the one hand, a method for multi-target recognition of material tracking based on 3D+2D high-dimensional feature fusion is provided. The method includes the following steps:

[0007] S1. Use the YOLO algorithm to perform target recognition on the materials to be tracked in the scene;

[0008] S2. Based on the Deep SORT algorithm, perform multi-target real-time tracking on the materials. Through the processing of the Kalman filter and the Hungarian algorithm, ensure the continuous tracking and identity consistency of the materials at different time points;

[0009] S3. Adopt a multi-view data calibration and alignment algorithm to generate three-dimensional point cloud data of the target, and perform preprocessing operations such as noise reduction, sampling, and data enhancement;

[0010] S4. Use the Gaussian Splatting method to perform 3D reconstruction on the preprocessed 3D point cloud data to generate a 3D model;

[0011] S5. During the 3D reconstruction process, introduce time series information to capture the dynamic changes of the target during movement;

[0012] S6. Segment the dynamic 3D target and the static background through the Mask R-CNN technology to obtain a 3D target image;

[0013] S7. Use the feature fusion method to align the position of the target in the 2D image with the coordinates in the 3D space, and adopt the PNP algorithm to estimate the pose and depth information of the target in the 3D space to achieve accurate positioning and state restoration of the material.

[0014] Optionally, step S1 specifically includes:

[0015] S11. Collect images of a specific scene through an industrial camera;

[0016] S12. Input the real-time collected images into a pre-trained YOLO network for convolutional feature extraction;

[0017] The YOLO network divides the input image into an S×S grid and generates a prediction box and its corresponding class probability based on each grid to detect the position and class of the target material;

[0018] P(B i ) = σ(x i )·σ(y i )·σ(w i )·σ(h i )·P(object)

[0019] where B i represents the prediction box of target i, (x i , y i ) are the center coordinates of the prediction box, (w i , h i ) are the width and height of the prediction box; σ is the Sigmoid function, representing the normalization of the prediction value; P(object) is the existence probability of each target, used to describe whether there is material;

[0020] Select the most likely target according to the predicted confidence and class probability. The confidence Conf(B i ) represents the degree of trust in the recognition and judgment of this target i, and is calculated through the following formula:

[0021] Conf(B i) = P(object)·IoU(B i , GT)

[0022] where IoU(B i , GT) is the intersection over union between the predicted bounding box of target i and the ground truth box; GT represents the ground truth box;

[0023] S13. Use non-maximum suppression (NMS) to remove duplicate predicted bounding boxes; NMS compares the confidence scores of each predicted bounding box and only keeps the one with the highest confidence score, while removing overlapping boxes based on the IoU threshold;

[0024]

[0025] where IoU(B i , B j ) is the intersection over union between the predicted bounding box of target i and the predicted bounding box of target j.

[0026] Optionally, step S2 specifically includes:

[0027] S21. Use Kalman filtering to estimate the state of the target and predict the state of the target in the next frame. The state of the target includes the position and velocity of the target;

[0028] The state transition formula is:

[0029] x k = Ax k-1 + Bu k + w k

[0030] where x k represents the state of the target at time step k, x k-1 represents the state of the target at time step k - 1, A is the state transition matrix, B is the control input matrix, u k is the control input, and w k is the process noise;

[0031] The update formula for Kalman filtering:

[0032] x k = x k|k-1 + K k (z k - H k x k|k-1 )

[0033] where z k is the observation, H k is the observation matrix, K k is the Kalman gain; x k is the estimated value at time step k, and x k|k-1 is the estimated value at time step k - 1;

[0034] S22. Extract the depth features of the target using the ReID model for matching to ensure correct association of each target at different time points;

[0035] The matching score formula is:

[0036]

[0037] where, f i and f j are the depth features of targets i and j respectively, ||f i || is the L2 norm of the depth feature of target i, ||f j || is the L2 norm of the depth feature of target j, and · represents the dot product operation;

[0038] S23. In each frame, predict the state of the target through Kalman filtering and associate it with the current detection result, specifically including the following steps:

[0039] According to the predicted position of Kalman filtering and the position of the target prediction box, calculate the cost matrix between the two; the cost is jointly determined by the target feature matching score and the spatial position difference;

[0040] Use the Hungarian algorithm to solve the optimal assignment of the cost matrix and determine the matching relationship between the target prediction box and the tracking trajectory;

[0041] The un-matched prediction boxes are added to the tracking list as new targets, and the un-matched tracking trajectories are predicted and updated according to Kalman filtering. If not matched for multiple consecutive frames, the tracking trajectories are removed;

[0042] S24. According to the assignment result of the Hungarian algorithm, update the state and depth features of the matched targets to ensure the consistency of target identities; initialize the Kalman filtering state for the un-matched new targets and start a new tracking path. At the same time, perform trajectory termination judgment on the un-matched tracking trajectories to ensure the efficiency and stability of the tracking system.

[0043] Optionally, the step S3 specifically includes:

[0044] S31. First, perform camera calibration to calculate the internal and external parameters of the camera. By collecting multiple images containing the calibration board, extract the corner coordinates of the checkerboard in each image, and then optimize the camera parameters by the least squares method to make the two-dimensional points in the image match the known calibration points in the three-dimensional space, obtaining the internal and external parameters of the camera;

[0045] Camera projection formula:

[0046] P = K[R|t]X

[0047] Among them, K is the camera internal parameter matrix, expressed as f x , f y is the focal length in the pixel coordinate system, c x , c y is the pixel coordinate of the optical center;

[0048] [R|t] is the camera external parameter matrix; R is the rotation matrix, t is the translation vector, X is the three-dimensional coordinate of the object; P is the projection of the object in the image;

[0049] S32. Use multiple cameras to capture the target from different angles to obtain multi-view images, extract feature points and descriptors using ORB, perform feature matching using the Hamming distance, and select the initial matching point pairs; eliminate the mismatched point pairs through the Random Sample Consensus algorithm, calculate the fundamental matrix and the essential matrix, decompose the camera projection matrix from the essential matrix, and align the multi-view images according to the camera projection matrix to generate dense three-dimensional point cloud data;

[0050] S33. For each pair of matching points, solve the three-dimensional coordinates through the camera projection matrix, optimize the density of the three-dimensional point cloud according to the disparity map, and generate a dense depth map using the SGBM stereo matching algorithm in OpenCV; based on the triangulation principle, calculate the depth information from the disparity through the internal parameters of the camera and the baseline length, and the formula is:

[0051]

[0052] where f is the focal length of the camera, B is the baseline length between the two cameras, and d is the disparity;

[0053] S34. Use the depth information of each matching point to convert the two-dimensional pixel coordinates (x, y) into three-dimensional space coordinates (X, Y, Z), and the conversion formula is:

[0054]

[0055] S35. Use the ICP point cloud registration technology to optimize the point cloud alignment and fuse the three-dimensional point cloud data from different perspectives into a unified coordinate system;

[0056] S36. After the point cloud is generated, perform preprocessing operations such as denoising, sampling, and data enhancement on the point cloud;

[0057] First, apply three-dimensional Gaussian filtering to remove isolated noise points for denoising:

[0058] P filtered ={p i ∈P|the number of points in the neighborhood>T}

[0059] where T is the threshold for noise filtering; pi is the point cloud of each frame, P is the set of all point clouds, P filtered is the filtered point cloud;

[0060] Secondly, uniform sampling is used to reduce the scale of the point cloud data while retaining the main structural features;

[0061] Finally, random rotation, scaling, and translation operations are performed on the point cloud to enhance data diversity.

[0062] Optionally, step S4 specifically includes:

[0063] S41. In three-dimensional space, each point is represented by a Gaussian kernel as a Gaussian distribution with position and uncertainty; assume the position of a certain point is x i =(x i , y i , z i ), and its Gaussian distribution representation is:

[0064]

[0065] where x i is the three-dimensional coordinate of the point, ∑ i is the covariance matrix of the point, representing the uncertainty of the point in space and determining the extent and shape of the point cloud; x is the position of any point in three-dimensional space, used to calculate the probability density of the point under the Gaussian distribution;

[0066] S42. Weighted superposition of the Gaussian distributions of all points in the point cloud is performed to generate a three-dimensional reconstruction result;

[0067] The calculation formula for the total Gaussian weight value is:

[0068]

[0069] where w i is the weight of each point, assigned according to the confidence or observation intensity of the point;

[0070] In the target three-dimensional space, each voxel is sampled, and the Gaussian weight value G total is calculated to generate a dense three-dimensional reconstruction image.

[0071] Optionally, step S5 specifically includes:

[0072] S51. Timestamp association is performed on the point cloud frames, and a timestamp τ i is attached to each frame of point cloud data, representing the acquisition time; a time-series point cloud data set {(p i , τ i )} is constructed, where p i is the point cloud of each frame;

[0073] S52. Detect the object motion in the dynamic scene by comparing the change amount Δp = p i+1 - p i between adjacent time - frame point clouds:

[0074] Use the ICP algorithm to register adjacent point - cloud frames, solve the rigid - transformation matrix, and minimize the registration error:

[0075]

[0076] where, R is the rotation matrix; t is the translation vector;

[0077] Judge the overall dynamic - change characteristics according to the registration error, and further extract the changing area of the moving target;

[0078] S53. Generate a dynamic model according to the time - series point cloud to capture the motion trajectory and morphological changes of the target;

[0079]

[0080] where, is the current - state prediction, A is the state - transition matrix, B is the control - input matrix, and w t is the process noise;

[0081] Generate the target motion trajectory according to the time - series information of the point - cloud position, capture its dynamic - behavior characteristics, and provide support for subsequent dynamic - behavior analysis.

[0082] Optionally, the step S6 specifically includes:

[0083] S61. First, use the ResNet convolutional neural network to extract features from the input image to generate a feature map representing the spatial information of the input image; then use the target - detection module to identify the target object and determine the candidate - box region;

[0084] S62. For each target region, use the fully - convolutional network FCN to generate a binary segmentation mask to separate the target from the background; specifically including:

[0085] Use the RoIAlign operation to align the region corresponding to the candidate box in the feature map F to a fixed size;

[0086] Input the aligned feature map into the instance - segmentation branch to generate the segmentation mask M(x,y) of each candidate box, where M(x,y) ∈ [0,1] represents the probability that the pixel belongs to the target;

[0087] S63. Fuse the generated segmentation mask with the three - dimensional reconstruction image, and only retain the dynamic - target part; and perform morphological post - processing on the segmentation mask to eliminate noise and artifacts and further improve the segmentation accuracy.

[0088] Optionally, step S7 specifically includes:

[0089] S71. Use a feature fusion method to fuse the position of the target in the two-dimensional image with the coordinates in the three-dimensional space to obtain comprehensive target information;

[0090] S72. Based on the known three-dimensional point cloud coordinates {X i =(X i , Y i , Z i )} and the corresponding two-dimensional image points {P i =(u i , v i )}, use the PNP algorithm for three-dimensional positioning and pose estimation of the object; the specific steps are as follows:

[0091] Input the three-dimensional point cloud coordinates {X i =(X i , Y i , Z i )} and the two-dimensional image points {P i =(u i , v i )} into the PNP algorithm. The PNP algorithm estimates the rotation and translation of the target based on the following projection model:

[0092] p i =K·(R·X i +t)

[0093] where P i is the coordinate of the i-th two-dimensional image point, K is the camera internal parameter matrix; R is the rotation matrix, representing the rotation relationship of the target in the three-dimensional space; t is the translation vector, representing the displacement of the target in the three-dimensional space; X i is the coordinate of the i-th three-dimensional point cloud point;

[0094] Use the objective function to minimize to solve the rotation and translation of the target:

[0095]

[0096] where n is the number of matching pairs of two-dimensional points and three-dimensional points, |||| 2 is the square of the Euclidean distance, used to measure the projection error.

[0097] On the other hand, a multi-target recognition system for material tracking based on 3D+2D high-dimensional feature fusion is provided for implementing the method described in any one of the above, and the system includes:

[0098] A target recognition module for using the YOLO algorithm to perform target recognition on the materials to be tracked in the scene;

[0099] A multi-object real-time tracking module, which is used to perform multi-object real-time tracking on materials based on the Deep SORT algorithm. Through the processing of the Kalman filter and the Hungarian algorithm, it ensures the continuous tracking and identity consistency of materials at different time points;

[0100] A three-dimensional point cloud data generation and preprocessing module, which is used to generate three-dimensional point cloud data of the target by using a multi-view data calibration and alignment algorithm, and perform preprocessing operations such as noise reduction, sampling, and data augmentation;

[0101] A three-dimensional reconstruction module, which is used to perform three-dimensional reconstruction on the preprocessed three-dimensional point cloud data by using the Gaussian Splatting method to generate a three-dimensional model;

[0102] A dynamic change capture module, which is used to introduce time series information during the three-dimensional reconstruction process to capture the dynamic changes of the target during movement;

[0103] A dynamic target segmentation module, which is used to segment the dynamic three-dimensional target and the static background through the Mask R-CNN technology to obtain a three-dimensional target image;

[0104] A three-dimensional positioning and pose estimation module, which is used to use a feature fusion method to align the position of the target in the two-dimensional image with the coordinates in the three-dimensional space, and use the PNP algorithm to estimate the pose and depth information of the target in the three-dimensional space to achieve precise positioning and state restoration of the materials.

[0105] On the other hand, an electronic device is provided, and the electronic device includes:

[0106] A processor;

[0107] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the multi-object recognition method for material tracking as described above are implemented.

[0108] On the other hand, a computer-readable storage medium is provided, in which program code is stored. The program code can be called by the processor to execute the steps of the multi-object recognition method for material tracking as described above.

[0109] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0110] 1. Solve the problem of missing spatial information. By introducing multi-dimensional data collection means, the present invention can not only quickly identify material targets, but also obtain the spatial distance information of the targets, so as to accurately judge the spatial relationship between the object and the equipment. This improvement significantly improves the accuracy of equipment motion control in the production manufacturing process, reduces errors caused by inaccurate distance judgment, and improves production efficiency and product quality.

[0111] 2. Enhance the recognition ability in complex environments. The present invention adopts advanced target recognition technology, which can effectively cope with the complex and messy production environment in industrial sites. By optimizing algorithms and multi-modal fusion processing, it solves the problem that traditional recognition methods are easily affected by environmental interference and occlusion, effectively ensures that the unique identification ID of the tracked material is not disrupted, and thus improves the stability and accuracy of the material tracking system in complex scenarios.

[0112] 3. Improve the accuracy of cross-shot tracking. Aiming at the problem that cross-shot tracking in traditional methods cannot accurately match the ID due to the appearance consistency of materials, the present invention combines feature analysis and multi-shot collaborative algorithms to achieve continuous identity matching of materials under different shots, significantly improves the effect of cross-shot tracking, and ensures the information coherence during the material flow process.

[0113] 4. Enhance the system robustness and fault tolerance. The present invention solves the impact of single-point failure on system stability through the combination of multiple tracking methods and redundant design. Using multi-source data fusion technology, information missing can be compensated through other channels when data is abnormal or nodes fail, thus significantly improving the reliability and anti-interference ability of the material tracking system.

[0114] 5. Support accurate and continuous material tracking. The present invention combines multi-scene and multi-level tracking strategies to ensure continuous tracking of materials throughout the production process. At the same time, through precise positioning and real-time information association, it provides an efficient and reliable solution for the whole-process material management of intelligent manufacturing. Description of the Drawings

[0115] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0116] Figure 1 It is a flowchart of a multi-target recognition method for material tracking based on 3D+2D high-dimensional feature fusion provided by the embodiments of the present invention;

[0117] Figure 2It is a schematic diagram of the principle of the multi-object recognition method for material tracking provided by an embodiment of the present invention. Detailed implementation manners

[0118] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0119] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner.

[0120] The embodiments of the present invention provide a multi-object recognition method for material tracking based on 3D+2D high-dimensional feature fusion. This method can be implemented by an electronic device, which can be a terminal or a server. Refer to Figure 1 and Figure 2 As shown, the processing flow of this method may include the following steps:

[0121] S1. Use the YOLO algorithm to perform target recognition on the materials to be tracked in the scene. In a complex industrial environment, by quickly capturing the material targets, ensure the accuracy and reliability of the basic data for subsequent processing.

[0122] The specific steps of step S1 include:

[0123] S11. Collect images of a specific scene through an industrial camera.

[0124] S12. Real-time image processing and target recognition.

[0125] Input the real-time collected images into a pre-trained YOLO network for convolutional feature extraction; the YOLO network divides the input image into S×S grids and generates prediction boxes and their corresponding class probabilities based on each grid to detect the position and class of the target materials;

[0126] P(B i ) = σ(x i )·σ(y i )·σ(w i )·σ(h i )·P(object)

[0127] where Bi Indicates the predicted bounding box of target i, where (x i , y i ) are the center coordinates of the predicted bounding box, and (w i , h i ) are the width and height of the predicted bounding box; σ is the Sigmoid function, representing the normalization of the predicted value; P(object) is the existence probability of each target, used to describe whether there is material present.

[0128] Filter out the most likely targets based on the predicted confidence and class probability. The confidence Conf(B i ) represents the degree of trust in the recognition and judgment of target i, and is calculated by the following formula:

[0129] Conf(B i ) = P(object)·IoU(B i , GT)

[0130] where IoU(B i , GT) is the intersection over union between the predicted bounding box of target i and the ground truth box; GT represents the ground truth box.

[0131] S13. Use Non-Maximum Suppression (NMS) to remove duplicate predicted bounding boxes. NMS compares the confidence of each predicted bounding box and only retains the predicted bounding box with the highest confidence, and removes overlapping boxes according to the IoU threshold;

[0132]

[0133] where IoU(B i , B j ) is the intersection over union between the predicted bounding box of target i and the predicted bounding box of target j.

[0134] S2. Based on the Deep SORT algorithm, perform multi-object real-time tracking on the materials. Through the processing of Kalman filtering and the Hungarian algorithm, ensure the continuous tracking and identity consistency of the materials at different time points.

[0135] The specific steps of step S2 include:

[0136] S21. Use Kalman filtering to estimate the state of the target and predict the state of the target in the next frame. The state of the target includes the position and speed of the target, and the state transition formula is:

[0137] x k = Ax k-1 + Bu k + w k

[0138] where xk Represents the predicted state of the target at time k, x k-1 Represents the predicted state of the target at time k-1, A is the state transition matrix, B is the control input matrix, u k is the control input, w k is the process noise.

[0139] Update formula of Kalman filter:

[0140] x k = x k|k-1 + K k (z k - H k x k|k-1 )

[0141] where, z k is the observed value, H k is the observation matrix, K k is the Kalman gain; x k is the estimated value at time k, x k|k-1 is the estimated value at time k-1;

[0142] S22. Use the ReID model to extract the depth features of the target for matching to ensure correct association of each target at different time points. The matching score formula is:

[0143]

[0144] where, f i and f j are the depth features of targets i and j respectively, ||f i || is the L2 norm of the depth feature of target i, ||f j || is the L2 norm of the depth feature of target j, · represents the dot product operation;.

[0145] S23. Associate the detection result with the tracking trajectory.

[0146] In each frame, predict the state of the target through Kalman filter and associate it with the current detection result, which specifically includes the following steps:

[0147] According to the predicted position of the Kalman filter and the position of the target prediction box, calculate the cost matrix between the two; the cost is jointly determined by the target feature matching score and the spatial position difference;

[0148] Use the Hungarian algorithm to solve the optimal assignment of the cost matrix to determine the matching relationship between the target prediction box and the tracking trajectory;

[0149] The unmatched predicted bounding boxes are added as new targets to the tracking list. The unmatched tracking trajectories are continuously predicted and updated according to the Kalman filter. If they are not matched for multiple consecutive frames, the tracking trajectories are removed.

[0150] S24. Update the multi-object tracking path.

[0151] According to the assignment result of the Hungarian algorithm, update the states and depth features of the matched targets to ensure the consistency of target identities; initialize the Kalman filter states for the unmatched new targets and start new tracking paths. At the same time, perform trajectory termination judgment on the unmatched tracking trajectories to ensure the efficiency and stability of the tracking system.

[0152] In the embodiments of the present invention, the state of the target in the next frame is predicted based on the Kalman filter, and combined with the depth feature matching result, the identity of each target is determined. In each frame of image, by comparing the predicted position with the detected position, the tracking path of the target is updated to ensure the temporal continuity and accuracy of the target, thereby realizing the stable tracking and precise association of multiple targets.

[0153] S3. Use the multi-view data calibration and alignment algorithm to generate the three-dimensional point cloud data of the target, and perform preprocessing operations such as noise reduction, sampling, and data enhancement to optimize the three-dimensional point cloud data, improve the data quality, and reduce the interference of environmental noise on the reconstruction effect.

[0154] The step S3 specifically includes:

[0155] S31. Camera calibration.

[0156] First, perform camera calibration to calculate the internal and external parameters of the camera. By collecting multiple images containing the calibration board, extract the corner coordinates of the checkerboard in each image, and then optimize the camera parameters by the least squares method so that the two-dimensional points in the image are matched to the known calibration points in the three-dimensional space to obtain the internal and external parameters of the camera.

[0157] Camera projection formula:

[0158] P = K[R|t]X

[0159] Among them, K is the camera internal parameter matrix; expressed as f x ,f y is the focal length in the pixel coordinate system, c x ,c y is the pixel coordinate of the optical center;

[0160] [R|t] is the camera external parameter matrix; R is the rotation matrix, t is the translation vector, X is the three-dimensional coordinate of the object; P is the projection of the object in the image.

[0161] S32. Use multiple cameras to capture the target from different angles to obtain multi-view images. Extract feature points and descriptors using ORB, perform feature matching using the Hamming distance, and select the initial matching point pairs. Eliminate the mismatched point pairs through the Random Sample Consensus algorithm (RANSAC), calculate the fundamental matrix and the essential matrix, decompose the camera projection matrix from the essential matrix, and align the multi-view images according to the camera projection matrix to generate dense three-dimensional point cloud data.

[0162] Specifically, for feature point detection: Use FAST to detect potential feature points, calculate the corner strength through the Harris response function, and select the local maximum as the final feature point.

[0163] Harris response function formula:

[0164] R = det(M) - k·(trace(M)) 2

[0165] M is the image gradient matrix:

[0166]

[0167] where I x , I y are the horizontal and vertical gradients of the image, and k is an empirical constant.

[0168] Feature direction calculation: Calculate the direction of each feature point through the centroid of the pixel intensities within the image patch:

[0169]

[0170] where (x, y) are the coordinates of the pixels within the neighborhood of the feature point, I(x, y) is the pixel intensity, and W is the local window of the feature point.

[0171] Use BRIEF as the feature descriptor: Randomly select a pair of pixel points (p 1 , p 2 ) within the neighborhood of the feature point and compare their gray values:

[0172]

[0173] Repeat this process n times to generate an n-bit binary descriptor.

[0174] Measure the similarity between two feature descriptors through the Hamming distance:

[0175]

[0176] where d 1 , d 2They are two binary descriptors, and ⊕ represents the bitwise XOR operation. The result is the matching Hamming distance, and the smaller the distance, the better the match.

[0177] RANSAC to eliminate false matches: Use the Random Sample Consensus algorithm (RANSAC) to eliminate false match point pairs that do not conform to geometric constraints, and calculate the fundamental matrix F:

[0178]

[0179] where P 1 and P 2 are the coordinates of the matching points on the images 1 and 2 to be matched.

[0180] S33. For each pair of matching points, solve the three-dimensional coordinates through the camera projection matrix, optimize the density of the three-dimensional point cloud according to the disparity map, and generate a dense depth map using the SGBM stereo matching algorithm of OpenCV; Based on the principle of triangulation, calculate the depth information from the disparity through the internal parameters and baseline length of the camera. The formula is:

[0181]

[0182] where f is the focal length of the camera, B is the baseline length between the two cameras, and d is the disparity.

[0183] S34. Use the depth information of each matching point to convert the two-dimensional pixel coordinates (x, y) into three-dimensional space coordinates (X, Y, Z). The conversion formula is:

[0184]

[0185] S35. Use the ICP point cloud registration technology to optimize the point cloud alignment and fuse the three-dimensional point cloud data from different perspectives into a unified coordinate system.

[0186] S36. After the point cloud is generated, perform preprocessing operations such as noise reduction, sampling, and data enhancement on the point cloud.

[0187] First, apply three-dimensional Gaussian filtering to remove isolated noise points for noise reduction:

[0188] P filtered ={p i ∈P|the number of points in the neighborhood>T}

[0189] where T is the threshold for noise filtering; p i is each frame of point cloud, P is the set of all point clouds, and P filtered is the filtered point cloud;

[0190] Secondly, use uniform sampling to reduce the scale of the point cloud data and retain the main structural features;

[0191] Finally, perform random rotation, scaling, and translation operations on the point cloud to enhance data diversity.

[0192] S4. Use the Gaussian Splatting method to perform 3D reconstruction on the preprocessed 3D point cloud data to generate a 3D model with rich details and clear structure.

[0193] The specific steps of step S4 include:

[0194] S41. Gaussian kernel point representation.

[0195] In 3D space, each point is represented by a Gaussian kernel as a Gaussian distribution with position and uncertainty; assume the position of a certain point is x i =(x i , y i , z i ), and its Gaussian distribution representation form is:

[0196]

[0197] Among them, x i is the 3D coordinate of the point, ∑ i is the covariance matrix of the point, representing the uncertainty of the point in space and determining the extent of expansion and shape of the point cloud; x is the position of any point in 3D space, used to calculate the probability density of the point under the Gaussian distribution.

[0198] S42. Gaussian point cloud fusion.

[0199] Perform weighted superposition of the Gaussian distributions of all points in the point cloud to generate the 3D reconstruction result; the calculation formula for the total Gaussian weight value is:

[0200]

[0201] Among them, w i is the weight of each point, assigned according to the confidence or observation intensity of the point.

[0202] In the target 3D space, sample each voxel, calculate the Gaussian weight value G total , and generate a dense 3D reconstruction image.

[0203] In the Gaussian weighted superposition, calculate the color value:

[0204]

[0205] Among them, c i represents the color of the point cloud point.

[0206] S5. During the 3D reconstruction process, introduce time series information to capture the dynamic changes of the target during movement.

[0207] To capture the motion state of materials in a dynamic scene, it is necessary to integrate the time dimension into the point cloud representation and capture the changing features through time series analysis. This invention accurately restores the motion trajectory and state of materials through data analysis of consecutive time points, providing comprehensive support for the analysis of dynamic behaviors.

[0208] The specific steps of step S5 include:

[0209] S51. Timestamp association.

[0210] Perform timestamp association on the point cloud frames, and attach a timestamp τ to each frame of point cloud data i , representing the acquisition time; construct a time series point cloud data set {(p i , τ i )}, where p i is each frame of point cloud, providing a basis for capturing dynamic changes through association on the time axis.

[0211] S52. Dynamic change detection.

[0212] Detect the movement of objects in a dynamic scene by comparing the change amount Δp = p i+1 - p i between adjacent time frame point clouds: Use the ICP algorithm to register adjacent point cloud frames and solve the rigid transformation matrix to minimize the registration error:

[0213]

[0214] where R is the rotation matrix; t is the translation vector.

[0215] Based on the registration error, judge the overall dynamic change characteristics and further extract the changing area of the moving target.

[0216] S53. Generate a dynamic model according to the time series point cloud to capture the motion trajectory and morphological changes of the target.

[0217]

[0218] where is the current state prediction, A is the state transition matrix, B is the control input matrix, and w t is the process noise.

[0219] Generate the target motion trajectory based on the time series information of the point cloud position, capture its dynamic behavior characteristics, and provide support for subsequent dynamic behavior analysis.

[0220] S6. Segment the dynamic 3D objects and the static background through the Mask R-CNN technology, effectively removing background interference and retaining the 3D contours and feature information of the objects, so as to obtain the 3D object images.

[0221] The specific steps of step S6 include:

[0222] S61. First, use the ResNet convolutional neural network to extract features from the input image and generate a feature map representing the spatial information of the input image; then use the object detection module to identify the target object and determine the candidate box area.

[0223] S62. Generate a binary mask and segment the target.

[0224] For each target area, use the fully convolutional network FCN to generate a binary segmentation mask to separate the target from the background; specifically including:

[0225] Use the RoIAlign operation to align the area corresponding to the candidate box in the feature map F to a fixed size;

[0226] Input the aligned feature map into the instance segmentation branch to generate the segmentation mask M(x,y) for each candidate box, where M(x,y) ∈ [0,1] represents the probability that the pixel belongs to the target.

[0227] S63. Fuse the generated segmentation mask with the 3D reconstruction image, only retaining the dynamic target part; and perform morphological post-processing on the segmentation mask to eliminate noise and artifacts, further improving the segmentation accuracy and achieving high-precision segmentation of the dynamic 3D objects and the static background.

[0228] S7. Use the feature fusion method to align the position of the target in the 2D image with the coordinates in the 3D space, and use the PNP (Perspective-n-Point) algorithm to estimate the pose and depth information of the target in the 3D space, realizing the accurate positioning and state restoration of the materials.

[0229] The specific steps of step S7 include:

[0230] S71. Feature fusion.

[0231] Use the feature fusion method to fuse the position of the target in the 2D image with the coordinates in the 3D space to obtain comprehensive target information. This fusion can combine the target position information in the 2D image with the geometric coordinates in the 3D space, providing support for the 3D positioning and pose estimation of the object.

[0232] S72. Use the PNP algorithm for 3D positioning and pose estimation.

[0233] Based on the known 3D point cloud coordinates {X i=(X i , Y i , Z i )} and the corresponding two - dimensional image points {P i =(u i , v i )}, use the PNP algorithm for three - dimensional positioning and pose estimation of the object; the specific steps are as follows:

[0234] Input the three - dimensional point cloud coordinates {X i =(X i , Y i , Z i )} and the two - dimensional image points {P i =(u i , v i )} into the PNP algorithm. The PNP algorithm estimates the rotation and translation of the target based on the following projection model:

[0235] p i =K·(R·X i +t)

[0236] where, P i is the coordinate of the i - th two - dimensional image point, K is the camera intrinsic parameter matrix; R is the rotation matrix, representing the rotation relationship of the target in three - dimensional space; t is the translation vector, representing the displacement of the target in three - dimensional space; X i is the coordinate of the i - th three - dimensional point cloud point.

[0237] Then, use the objective function to minimize to solve for the rotation and translation of the target:

[0238]

[0239] where, n is the number of matching pairs of two - dimensional and three - dimensional points, || || 2 is the square of the Euclidean distance, used to measure the projection error.

[0240] In the embodiments of the present invention, aiming at the material characteristics in a specific industrial site environment, first, the YOLO algorithm is used for image recognition to quickly locate the target material. Subsequently, the Deep SORT algorithm is used to achieve multi-target real-time tracking of production materials, ensuring the continuity and accuracy of tracking. On this basis, a multi-view data calibration and alignment algorithm is adopted to obtain the three-dimensional point cloud data of the target, and preprocessing operations such as noise reduction, sampling, and data augmentation are performed on it to optimize the data quality. Then, the Gaussian Splatting method is used to perform three-dimensional reconstruction on the three-dimensional point cloud data, and at the same time, time series information is introduced to capture the details of the dynamic changes of the target. Further, the Mask R-CNN technology is used to segment the dynamic three-dimensional image and the static background, eliminating background interference and obtaining an accurate three-dimensional target image. Subsequently, a feature fusion method is used to spatially align the target position information in the two-dimensional image with the depth information in the three-dimensional point cloud. Finally, the PNP algorithm is used to estimate the pose and depth of the target in the three-dimensional space, realizing the accurate positioning and state restoration of the material, and providing efficient and reliable technical support for material management and tracking in the industrial environment.

[0241] Correspondingly, an embodiment of the present invention further provides a multi-target recognition system for material tracking based on 3D+2D high-dimensional feature fusion, and the system includes:

[0242] A target recognition module, which is used to perform target recognition on the materials to be tracked in the scene by using the YOLO algorithm;

[0243] A multi-target real-time tracking module, which is used to perform multi-target real-time tracking on the materials based on the Deep SORT algorithm. Through the processing of the Kalman filter and the Hungarian algorithm, it ensures the continuous tracking and identity consistency of the materials at different time points;

[0244] A three-dimensional point cloud data generation and preprocessing module, which is used to generate the three-dimensional point cloud data of the target by using a multi-view data calibration and alignment algorithm, and perform preprocessing operations such as noise reduction, sampling, and data augmentation;

[0245] A three-dimensional reconstruction module, which is used to perform three-dimensional reconstruction on the preprocessed three-dimensional point cloud data by using the Gaussian Splatting method to generate a three-dimensional model;

[0246] A dynamic change capture module, which is used to introduce time series information during the three-dimensional reconstruction process to capture the dynamic changes of the target during the movement process;

[0247] A dynamic target segmentation module, which is used to segment the dynamic three-dimensional target and the static background by using the Mask R-CNN technology to obtain a three-dimensional target image;

[0248] A three-dimensional positioning and attitude estimation module, which is used to align the position of the target in the two-dimensional image with the coordinates in the three-dimensional space by using the feature fusion method, and adopts the PNP algorithm to estimate the attitude and depth information of the target in the three-dimensional space, so as to realize the accurate positioning and state restoration of the material.

[0249] The system of this embodiment can be used to execute Figure 1 the technical solutions of the method embodiments shown. Their implementation principles and technical effects are similar and will not be elaborated here.

[0250] In an exemplary embodiment, the present invention further provides an electronic device, which includes:

[0251] A processor;

[0252] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the above-mentioned material tracking multi-target recognition method are realized.

[0253] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by the processor to realize the steps of the above-mentioned material tracking multi-target recognition method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0254] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the element.

[0255] When referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. in the specification, it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific feature, structure or characteristic. In addition, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0256] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, and specific understanding can be made by referring to the context before and after.

[0257] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0258] It should be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0259] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0260] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0261] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0262] When the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0263] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without these detailed descriptions. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0264] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A material tracking multi-target recognition method based on 3D+2D high-dimensional feature fusion, characterized in that: The following steps are involved: S1. Use the YOLO algorithm to identify the target material to be tracked in the scene; S2. Real-time multi-target tracking of materials based on Deep SORT algorithm, and through Kalman filtering and Hungarian algorithm processing, ensure continuous tracking and identity consistency of materials at different time points; S3, using multi-view data calibration and alignment algorithm to generate the three-dimensional point cloud data of the target, and perform noise reduction, sampling and data enhancement preprocessing operations; S4, using the Gaussian Splatting method to perform 3D reconstruction on the preprocessed 3D point cloud data to generate a 3D model; S5. In the 3D reconstruction process, time series information is introduced to capture the dynamic changes of the target during movement; S6. Segment the dynamic three-dimensional target and the static background through the Mask R-CNN technology to obtain a three-dimensional target image; S7. Use the feature fusion method to align the position of the target in the two-dimensional image with the coordinates in the three-dimensional space, and use the PNP algorithm to estimate the posture and depth information of the target in the three-dimensional space to achieve accurate positioning and state restoration of the material.

2. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S1 specifically includes: S11. Collect images in specific scenes through industrial cameras; S12, input the real-time collected image into the pre-trained YOLO network to perform convolution feature extraction; The YOLO network divides the input image into S×S grids and generates a prediction box and its corresponding category probability based on each grid to detect the location and category of the target material; P(B i )=σ(x i )·σ(y i )·σ(w i )·σ(h i )·P(object) Among them, B i represents the predicted box of target i, (x i ,y i ) is the center coordinate of the prediction box, (w i ,h i ) is the width and height of the prediction box; σ is the Sigmoid function, which represents the normalization of the prediction value; P(object) is the existence probability of each target, which is used to describe whether there is a material; The most likely target is selected based on the predicted confidence and category probability. i ) represents the degree of trust in the identification judgment of the target i, which is calculated by the following formula: Conf(B i )=P(object)·IoU(B i ,GT) Among them, IoU(B i ,GT) is the intersection-over-union ratio between the predicted box and the real box of target i; GT represents the real box; S13, use non-maximum suppression NMS to remove duplicate prediction boxes; NMS compares the confidence of each prediction box, retains only the prediction box with the highest confidence, and removes overlapping boxes according to the IoU threshold; Among them, IoU(B i ,B j ) is the intersection-over-union ratio between the prediction box of target i and the prediction box of target j.

3. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S2 specifically includes: S21, using Kalman filtering to estimate the state of the target and predict the state of the target in the next frame, wherein the state of the target includes the position and speed of the target; The state transition formula is: x k =Ax k-1 +Bu k +w k Among them, x k represents the state of the target at time k, x k-1 represents the state of the target at time k-1, A is the state transfer matrix, B is the control input matrix, u k is the control input, w k is the process noise; The update formula of Kalman filter is: x k =x k|k-1 +K k (z k -H k x k|k-1 ) Among them, z k is the observed value, H k is the observation matrix, K k is the Kalman gain; x k is the estimated value at time k, x k|k-1 is the estimated value at time k-1; S22, use the ReID model to extract the deep features of the target for matching to ensure that each target is correctly associated at different time points; The matching score formula is: Among them, f i and f j are the deep features of targets i and j respectively, ||f i || is the L2 norm of the deep feature of target i, ||f j || is the L2 norm of the deep feature of target j, · represents the dot product operation; S23, in each frame, predicting the state of the target through Kalman filtering and associating it with the current detection result, specifically including the following steps: Based on the predicted position of the Kalman filter and the position of the target prediction box, the cost matrix between the two is calculated; the cost is determined by the target feature matching score and the spatial position difference; The Hungarian algorithm is used to solve the optimal allocation of the cost matrix and determine the matching relationship between the target prediction box and the tracking trajectory; The unmatched prediction box is added to the tracking list as a new target. The unmatched tracking trajectory continues to be predicted and updated according to the Kalman filter. If there is no match for multiple consecutive frames, the tracking trajectory is removed. S24. According to the allocation results of the Hungarian algorithm, the state and depth features of the matched targets are updated to ensure the consistency of the target identity; the Kalman filter state is initialized for the unmatched new targets, and a new tracking path is started. At the same time, the trajectory termination judgment is performed on the unmatched tracking trajectory to ensure the efficiency and stability of the tracking system.

4. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S3 specifically includes: S31, firstly, perform camera calibration to calculate the intrinsic and extrinsic parameters of the camera, by collecting multiple images containing the calibration plate, extracting the coordinates of the corner points of the chessboard in each image, and then optimizing the camera parameters by the least square method, so that the two-dimensional points in the image are matched to the known calibration points in the three-dimensional space, and the intrinsic and extrinsic parameters of the camera are obtained; Camera projection formula: P=K[R|t]X Among them, K is the camera internal parameter matrix; it is expressed as f x , f y is the focal length in pixel coordinate system, c x , c y are the pixel coordinates of the optical center; [R|t] is the camera extrinsic matrix; R is the rotation matrix, t is the translation vector, X is the 3D coordinate of the object; P is the projection of the object in the image; S32. Use multiple cameras to shoot the target from different angles to obtain multi-view images, use ORB to extract feature points and descriptors, use Hamming distance to perform feature matching, and select initial matching point pairs; use random sampling consistency algorithm to eliminate mismatched point pairs, calculate the basic matrix and the essential matrix, decompose the camera projection matrix from the essential matrix, align the multi-view images according to the camera projection matrix, and generate dense three-dimensional point cloud data; S33. For each pair of matching points, the three-dimensional coordinates are solved by the camera projection matrix, the density of the three-dimensional point cloud is optimized according to the disparity map, and the dense depth map is generated by the SGBM stereo matching algorithm of OpenCV; based on the principle of triangulation, the depth information is calculated from the disparity through the intrinsic parameters and baseline length of the camera, and the formula is: Where f is the focal length of the camera, B is the baseline length between the two cameras, and d is the parallax; S34, using the depth information of each matching point, convert the two-dimensional pixel coordinates (x, y) into three-dimensional space coordinates (X, Y, Z), and the conversion formula is: S35. Use ICP point cloud registration technology to optimize point cloud alignment and fuse 3D point cloud data from different perspectives into a unified coordinate system; S36, after the point cloud is generated, performing denoising, sampling and data enhancement preprocessing operations on the point cloud; First, apply three-dimensional Gaussian filtering to remove isolated noise points for noise reduction: P filtered ={p i ∈P|Number of points in the neighborhood>T} Where T is the threshold of noise filtering; p i is the point cloud of each frame, P is the set of all point clouds, P filtered is the filtered point cloud; Secondly, uniform sampling is used to reduce the size of point cloud data and retain the main structural features; Finally, the point cloud is randomly rotated, scaled, and translated to enhance data diversity.

5. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S4 specifically includes: S41. In three-dimensional space, each point is represented by a Gaussian kernel as a Gaussian distribution with position and uncertainty. Suppose the position of a point is x i =(x i ,y i , z i ), whose Gaussian distribution is expressed as: Among them, x i is the three-dimensional coordinate of the point, ∑ i is the covariance matrix of the point, which indicates the uncertainty of the point in space and determines the expansion and shape of the point cloud; x is the position of any point in three-dimensional space, which is used to calculate the probability density of the point under Gaussian distribution; S42, performing weighted superposition on the Gaussian distribution of all points in the point cloud to generate a three-dimensional reconstruction result; The total Gaussian weight value is calculated as: Among them, w i A weight is assigned to each point based on the confidence or observation strength of the point; In the target three-dimensional space, sample each voxel and calculate the Gaussian weight value G total , generating dense 3D reconstructed images.

6. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S5 specifically includes: S51, associate the point cloud frames with timestamps, and add a timestamp τ to each frame of point cloud data i , represents the acquisition time; construct a time series point cloud dataset {(p i , τ i )}, where p i Point cloud for each frame; S52, by comparing the change amount of point clouds in adjacent time frames Δp=p i+1 -p i , detect object motion in dynamic scenes: Use the ICP algorithm to register adjacent point cloud frames and solve the rigid transformation matrix to minimize the registration error: Where R is the rotation matrix; t is the translation vector; According to the registration error, the overall dynamic change characteristics are judged, and the change area of ​​the moving target is further extracted; S53, generating a dynamic model based on the time series point cloud to capture the motion trajectory and morphological changes of the target; in, is the current state prediction, A is the state transfer matrix, B is the control input matrix, w t is the process noise; According to the time series information of the point cloud position, the target motion trajectory is generated to capture its dynamic behavior characteristics and provide support for subsequent dynamic behavior analysis.

7. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S6 specifically includes: S61. First, a ResNet convolutional neural network is used to extract features from the input image to generate a feature map representing the spatial information of the input image; then, a target detection module is used to identify the target object and determine the candidate frame area; S62. For each target area, a fully convolutional network (FCN) is used to generate a binary segmentation mask to separate the target from the background; specifically, the following steps are included: The RoIAlign operation is used to align the area corresponding to the candidate box in the feature map F to a fixed size; The aligned feature map is input into the instance segmentation branch to generate a segmentation mask M(x,y) for each candidate box, where M(x,y)∈[0,1] represents the probability that the pixel belongs to the target; S63, fusing the generated segmentation mask with the three-dimensional reconstructed image to retain only the dynamic target part; and performing morphological post-processing on the segmentation mask to eliminate noise and artifacts to further improve the segmentation accuracy.

8. The material tracking multi-target identification method according to claim 1 is characterized in that: The step S7 specifically includes: S71, using a feature fusion method to fuse the position of the target in the two-dimensional image with the coordinates in the three-dimensional space to obtain comprehensive target information; S72, based on the known three-dimensional point cloud coordinates {X i =(X i , Y i , Z i )} and the corresponding two-dimensional image point {P i =(u i , v i )}, use the PNP algorithm to perform 3D positioning and posture estimation of the object; the specific steps are as follows: The three-dimensional point cloud coordinates {X i =(X i , Y i , Z i )} and the two-dimensional image point {P i =(u i , v i )} is input into the PNP algorithm, which estimates the rotation and translation of the target based on the following projection model: p i =K·(R·X i +t) Among them, P i is the coordinate of the i-th two-dimensional image point, K is the camera internal parameter matrix; R is the rotation matrix, which represents the rotation relationship of the target in three-dimensional space; t is the translation vector, which represents the displacement of the target in three-dimensional space; X i is the coordinate of the i-th 3D point cloud point; Use the objective function to minimize the rotation and translation of the solution: Where n is the number of matching pairs of 2D points and 3D points, |||| 2 is the square of the Euclidean distance, which is used to measure the projection error.

9. A material tracking multi-target recognition system based on 3D+2D high-dimensional feature fusion, the system is used to implement the method according to any one of claims 1 to 8, characterized in that: The system comprises: The target recognition module is used to use the YOLO algorithm to identify the target of the material to be tracked in the scene; Multi-target real-time tracking module, used to track materials in real time based on Deep SORT algorithm, and ensure continuous tracking and identity consistency of materials at different time points through Kalman filtering and Hungarian algorithm processing; The 3D point cloud data generation and preprocessing module is used to generate the 3D point cloud data of the target using a multi-view data calibration and alignment algorithm, and perform denoising, sampling and data enhancement preprocessing operations; The 3D reconstruction module is used to perform 3D reconstruction on the pre-processed 3D point cloud data using the Gaussian Splatting method to generate a 3D model; The dynamic change capture module is used to introduce time series information in the 3D reconstruction process to capture the dynamic changes of the target during movement; The dynamic target segmentation module is used to segment the dynamic 3D target from the static background using the Mask R-CNN technology to obtain a 3D target image. The three-dimensional positioning and posture estimation module is used to align the position of the target in the two-dimensional image with the coordinates in the three-dimensional space using the feature fusion method, and uses the PNP algorithm to estimate the posture and depth information of the target in the three-dimensional space to achieve accurate positioning and state restoration of the material.

10. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when loaded and executed by the processor, implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-modal target tracking method based on decision-making level fusion

    CN120823238A

  • Pitching side type material taking machine material positioning method based on image recognition

    CN120823267A

  • Image recognition based material positioning method for luffing side-type material handler

    CN120823267B

  • Wharf cross-camera multi-target tracking method and system based on three-dimensional map

    CN120876543A

  • Three-dimensional model construction method and device based on image sequence processing

    CN121170166A