Digital twinning-based stacker visual data auxiliary control system and method
By building a digital twin model and combining lidar point cloud data and image acquisition equipment, the stacker visual data assisted control system is solved, the problem of misjudgment of stacker obstacle identification is achieved, more accurate obstacle identification is achieved and unnecessary downtime is reduced, and the operating efficiency and equipment life of the warehousing system are improved.
Patent Information
- Application Number
- CN202510424591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
The existing vision-based stacker obstacle identification technology is prone to misjudgment, which causes the stacker to stop running unnecessary, affecting the efficiency of warehousing and logistics and losing equipment.
Combining lidar point cloud data, ultra-wideband positioning labels and image acquisition equipment, a digital twin model is built, through the fusion of image features and spatial features, and a classification model is built using a support vector mechanism to identify obstacles.
It effectively reduces the misidentification of non-obstructions, ensures the normal operation of the stacker, and improves the operating efficiency and equipment life of the warehousing system.
Smart Images

Figure HDA0005346335430000011 
Figure HDA0005346335430000021
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a visual data assisted control system and method for a stacker based on digital twin. Background Art
[0002] In modern warehousing and logistics systems, the stacker is a core device, and the efficiency and safety of its operation are crucial. With the rapid development of digital technology, digital twin technology has gradually been applied to the warehousing field. By constructing a digital twin model highly consistent with the real warehousing environment, the state and behavior of physical entities can be reflected in real time. At the same time, vision technology also plays an important role in the operation monitoring of stackers. Image acquisition devices can obtain image information on the operation path of stackers, providing a data basis for subsequent analysis and decision-making.
[0003] However, the existing vision-based obstacle recognition technology for stackers has significant defects. Traditional methods often rely only on picture features for judgment, and this single-dimensional analysis method is extremely prone to misjudgment. In the actual operation environment of stackers, objects such as projections generated by light changes, newspapers fluttering in the wind, or plastic bags accidentally falling, the picture features they present are easily confused with real obstacles. When the system misidentifies these items that do not actually affect the normal operation of the stacker as obstacles, the stacker will stop running unnecessarily, which not only reduces the operation efficiency of warehousing and logistics, but also the frequent start-stop operations may cause additional wear and tear on the stacker equipment, seriously affecting the normal operation of the entire warehousing system. Summary of the Invention
[0004] The purpose of the present invention is to provide a visual data assisted control system and method for a stacker based on digital twin to solve the problems raised in the prior art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A visual data assisted control method for a stacker based on digital twin, the method comprising the following steps:
[0006] Step 1, arranging sensors to construct a data twin scenario of the warehousing environment;
[0007] Step 2, arranging an image acquisition device on the stacker to acquire images at a fixed angle;
[0008] Step 3, manually annotating the acquired images and extracting image features;
[0009] Step 4, mapping the acquired images into the data twin scenario and extracting spatial features;
[0010] Step 5, fusing the image features and spatial features, and constructing a classification model in combination with the manual annotation situation.
[0011] In step 1, deploy lidar and ultra-wideband positioning tags in the warehousing environment to build a digital twin model;
[0012] The lidar scans the environment and generates point cloud data P=(x j ,y j ,z j ), where N represents the number of point cloud data, j∈{1,2,…,N}, representing the point cloud data sequence; (x j ,y j ,z j ) represents the three-dimensional space coordinates;
[0013] The ultra-wideband positioning tags are arranged on the shelves and stacker tracks to collect the position of the stacker P(t)=(x(t),y(t),z(t)) in real time; where, (t) represents the timestamp, and (x(t),y(t),z(t)) represents the position of the stacker at time t;
[0014] Perform noise reduction on the point cloud data: use statistical outlier filtering to remove noise points;
[0015] The condition for retaining points is: 1 / k·Σ i=1 k ||p j -p i ||<d th ; where k represents the number of neighborhood points, i∈{1,2,…,k}, representing the neighborhood point sequence; d th represents the distance threshold;
[0016] Convert the point cloud to a mesh model M through the Poisson surface reconstruction algorithm; update the mesh model M in real time.
[0017] In step 2, install industrial-grade cameras at the front end and both sides of the stacker as image acquisition devices; set the fixed acquisition angle of the cameras to fully cover the running path of the stacker;
[0018] Set the image acquisition frequency to f. At the normal running speed v of the stacker, the overlapping area between two adjacent frames of images is not less than 50%, that is, the formula is satisfied: v / f≤L / 2; where L represents the effective coverage length of a single frame of the camera image in the running direction of the stacker;
[0019] The acquired image sequence is represented as I a , a = 1,2,…; Each frame of image I a is a matrix of m×n pixels, where m and n represent the height and width of the image respectively;
[0020] In step 3, manual annotation is divided into two types: obstacles and non-obstacles;
[0021] Precisely annotate the collected images by manual annotation, and mark the obstacle and non-obstacle areas;
[0022] Use the improved Scale-Invariant Feature Transform (SIFT) algorithm to extract image features: For the neighborhood of 16×16 pixels around the key points, divide it into 4×4 small regions, calculate the gradient histograms in 8 directions within each small region, and obtain a 4×4×8 = 128-dimensional descriptor vector;
[0023] Enhance the robustness to illumination changes, and normalize the gradient magnitude of each small region to satisfy the formula: Σ c=1 128 e c 2 = 1;
[0024] where e c is the c-th element of the descriptor vector;
[0025] The extracted image feature vector is denoted as F img,a , which corresponds one-to-one with the image I a .
[0026] In step 4, map the collected image I a to the digital twin scenario through coordinate transformation; use the position information P(t) of the stacker and the installation parameters of the camera: installation height h, pitch angle α, and horizontal angle β to establish the mapping relationship between the pixel coordinates (u, v) of the image obstacle or non-obstacle marking area and the three-dimensional coordinates (X, Y, Z) of the digital twin scenario:
[0027] X = x(t) + (u - m / 2)·cosβ·cosα·Δx + (v - n / 2)·sinβ·cosα·Δy;
[0028] Y = y(t) + (u - m / 2)·(-sinβ)·Δx + (v - n / 2)·cosβ·Δy;
[0029] Z = z(t) + h + (u - m / 2)·sinα·Δx + (v - n / 2)·sinα·Δy;
[0030] where Δx and Δy respectively represent the actual physical sizes corresponding to the image pixels in the X and Y directions, which are determined by the focal length and the sensor size;
[0031] In the digital twin scenario, based on the mapped coordinates, extract the spatial features of the mapped area;
[0032] Using the grid model M constructed from lidar point cloud data, calculate the point cloud density ρ of the mapped area, where ρ = S / V; here, S represents the number of points in the mapped area, and V represents the volume of the mapped area.
[0033] Calculate the distance D and relative angle θ between the mapped area and the current position of the stacker crane.
[0034] The extracted spatial feature vector is denoted as F space,a = [ρ, D, θ], corresponding to the image I a after mapping.
[0035] In step 5, fuse the image features and spatial features; the fused feature vector is denoted as: F a = [F img,a , F space,a ; normalize F a and, combined with the manual annotation situation, use the support vector machine SVM to construct a classification model; during the SVM training process, adopt the cross-validation method to select the optimal kernel function parameters and penalty factors to improve the classification accuracy of the model.
[0036] The trained classification model is used to determine in real time whether the object on the running path of the stacker crane is an obstacle; when it is determined to be a non-obstacle, the stacker crane continues to run; when it is determined to be an obstacle, the stacker crane immediately stops running and issues a warning.
[0037] A stacker crane vision data-assisted control system based on digital twin, which includes an environment modeling module, a vision acquisition module, a spatial mapping module, and an intelligent control module;
[0038] The environment modeling module is used to arrange sensors to construct a digital twin scenario of the warehousing environment; the vision acquisition module is used to arrange image acquisition devices on the stacker crane to collect images at a fixed angle, manually annotate the collected images, and extract image features; the spatial mapping module is used to map the collected images into the digital twin scenario and extract spatial features; the intelligent control module is used to fuse image features and spatial features, and construct a classification model combined with the manual annotation situation;
[0039] The output end of the environment modeling module is connected to the input end of the vision acquisition module; the output end of the vision acquisition module is connected to the input end of the spatial mapping module; the output end of the spatial mapping module is connected to the input end of the intelligent control module.
[0040] The environment modeling module includes a sensor deployment unit, a point cloud noise reduction unit, and a model generation unit;
[0041] The sensor deployment unit is used to deploy lidar and ultra-wideband tags, and collect point cloud data and position information in real time; the point cloud denoising unit is used to remove noise points by applying the statistical outlier filtering algorithm and retain the effective point cloud data; the model generation unit is used to construct a dynamically updated three-dimensional grid model through the Poisson surface reconstruction algorithm;
[0042] The output end of the sensor deployment unit is connected to the input end of the point cloud denoising unit; the output end of the point cloud denoising unit is connected to the input end of the model generation unit; the output end of the model generation unit is connected to the input end of the visual acquisition module.
[0043] The visual acquisition module includes a device installation unit, an image acquisition unit, and an image feature extraction unit;
[0044] The device installation unit is used to install an industrial camera at a specified position of the stacker and set a fixed acquisition angle; the image acquisition unit is used to acquire an image sequence at a preset frequency; the image feature extraction unit is used to extract image feature vectors by using an improved SIFT algorithm to enhance the illumination robustness;
[0045] The output end of the device installation unit is connected to the input end of the image acquisition unit; the output end of the image acquisition unit is connected to the input end of the image feature extraction unit; the output end of the image feature extraction unit is connected to the input end of the space mapping module.
[0046] The space mapping module includes a coordinate mapping unit, a spatial feature calculation unit, and a spatial feature integration unit;
[0047] The coordinate mapping unit is used to establish a conversion model from pixel coordinates to three-dimensional space based on camera parameters; the spatial feature calculation unit is used to calculate three-dimensional spatial features of the point cloud density, relative distance, and angle in the mapped area; the spatial feature integration unit is used to encode the spatial features into a standardized vector;
[0048] The output end of the coordinate mapping unit is connected to the input end of the spatial feature calculation unit; the output end of the spatial feature calculation unit is connected to the input end of the spatial feature integration unit; the output end of the spatial feature integration unit is connected to the input end of the intelligent control module.
[0049] The intelligent control module includes a feature fusion unit, a model training unit, and a real-time control unit;
[0050] The feature fusion unit is used to fuse spatial features and image features; the model training unit is used to optimize SVM parameters by using cross-validation and construct an obstacle classification decision model; the real-time control unit is used to trigger stacker commands according to the classification results and complete closed-loop control;
[0051] The output end of the feature fusion unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the real-time control unit.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention comprehensively utilizes the point cloud data of the lidar, the position data of the ultra-wideband positioning tags, and the image data of the image acquisition device, organically combines these different types of data, describes and analyzes the operating environment of the stacker from multiple dimensions, fully excavates the value of the data, and improves the utilization efficiency of the data; The present invention not only extracts image features, but also combines spatial features (such as point cloud density, distance from the current position of the stacker, and relative angle, etc.), constructs a classification model by fusing image features and spatial features, makes judgments by integrating information from multiple aspects, effectively improves the misidentification of non-obstacle items, reduces the unnecessary stops of the stacker caused by misidentification, and ensures the normal operation of the stacker. Description of the Drawings
[0053] Figure 1 It is a schematic diagram of the steps of the stacker vision data assisted control method based on digital twin of the present invention;
[0054] Figure 2 It is a schematic diagram of the process of the stacker vision data assisted control system based on digital twin of the present invention. Detailed Embodiments
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a stacker vision data assisted control method based on digital twin, and the method includes the following steps:
[0057] Step 1, arrange sensors to construct a data twin scenario of the warehousing environment;
[0058] Step 2, arrange an image acquisition device on the stacker to acquire images at a fixed angle;
[0059] Step 3, perform manual annotation on the acquired images and extract image features;
[0060] Step 4, map the acquired images to the data twin scenario and extract spatial features;
[0061] Step 5: Integrate image features and spatial features, and construct a classification model in combination with manual annotation.
[0062] In Step 1, deploy a lidar and ultra-wideband positioning tags in the warehousing environment to construct a digital twin model;
[0063] The lidar scans the environment and generates point cloud data P=(x j ,y j ,z j ), where N represents the number of point cloud data, j∈{1,2,…,N}, representing the point cloud data sequence; (x j ,y j ,z j ) represents the three-dimensional spatial coordinates;
[0064] The ultra-wideband positioning tags are arranged on the shelves and stacker tracks to collect the position of the stacker P(t)=(x(t),y(t),z(t)) in real time; where, (t) represents the timestamp, and (x(t),y(t),z(t)) represents the position of the stacker at time t;
[0065] Denoise the point cloud data: Use statistical outlier filtering to remove noise points;
[0066] The condition for retaining points is: 1 / k·Σ i=1 k ||p j -p i ||<d th ; where k represents the number of neighborhood points, i∈{1,2,…,k}, representing the neighborhood point sequence; d th represents the distance threshold;
[0067] Convert the point cloud to a mesh model M through the Poisson surface reconstruction algorithm; Update the mesh model M in real time.
[0068] In Step 2, install industrial-grade cameras at the front end and both sides of the stacker as image acquisition devices; Set the fixed acquisition angle of the cameras to fully cover the operation path of the stacker;
[0069] Set the image acquisition frequency to f. At the normal running speed v of the stacker, the overlapping area between two adjacent frames of images is not less than 50%, that is, it satisfies the formula: v / f≤L / 2; where L represents the effective coverage length of a single-frame image of the camera in the running direction of the stacker;
[0070] The acquired image sequence is represented as I a , a = 1,2,…; Each frame of image I a is a matrix of m×n pixels, where m and n represent the height and width of the image respectively;
[0071] In step 3, manual annotation is divided into two types: obstacles and non-obstacles;
[0072] The collected images are accurately annotated in a manual annotation manner, and the obstacle and non-obstacle areas are marked;
[0073] The improved Scale-Invariant Feature Transform (SIFT) algorithm is used to extract image features: for the 16×16 pixel neighborhood around the key points, it is divided into 4×4 small regions, and the gradient histograms in 8 directions are calculated in each small region to obtain a 4×4×8 = 128-dimensional descriptor vector;
[0074] To enhance the robustness to illumination changes, the gradient amplitude of each small region is normalized to satisfy the formula: Σ c=1 128 e c 2 = 1;
[0075] where e c is the c-th element of the descriptor vector;
[0076] The extracted image feature vector is denoted as F img,a and corresponds one-to-one with the image I a ;
[0077] In step 4, the collected image I a is mapped into the data twin scenario through coordinate transformation; using the position information P(t) of the stacker and the installation parameters of the camera: installation height h, pitch angle α, and horizontal angle β, the mapping relationship between the pixel coordinates (u, v) of the image obstacle or non-obstacle marked area and the three-dimensional coordinates (X, Y, Z) of the data twin scenario is established:
[0078] X = x(t) + (u - m / 2)·cosβ·cosα·Δx + (v - n / 2)·sinβ·cosα·Δy;
[0079] Y = y(t) + (u - m / 2)·(-sinβ)·Δx + (v - n / 2)·cosβ·Δy;
[0080] Z = z(t) + h + (u - m / 2)·sinα·Δx + (v - n / 2)·sinα·Δy;
[0081] where Δx and Δy respectively represent the actual physical sizes corresponding to the image pixels in the X and Y directions, which are determined by the focal length and the sensor size;
[0082] In the data twin scenario, based on the mapped coordinates, the spatial features of the mapped area are extracted;
[0083] Using the grid model M constructed from lidar point cloud data, calculate the point cloud density ρ of the mapped area, ρ = S / V; where S represents the number of point clouds in the mapped area, and V represents the volume of the mapped area;
[0084] Calculate the distance D and relative angle θ between the mapped area and the current position of the stacker;
[0085] The extracted spatial feature vector is denoted as F space,a = [ρ, D, θ], corresponding to the image I a The mapped area.
[0086] In step 5, fuse the image features and spatial features; the fused feature vector is denoted as: F a = [F img,a , F space,a ; Standardize F a and, combined with the manual annotation situation, use the support vector machine SVM to construct a classification model; during the SVM training process, adopt the method of cross-validation to select the optimal kernel function parameters and penalty factors to improve the classification accuracy of the model;
[0087] The trained classification model is used to determine in real time whether the objects on the running path of the stacker are obstacles; when it is determined that they are not obstacles, the stacker continues to run; when it is determined that they are obstacles, the stacker immediately stops running and issues a warning.
[0088] A stacker vision data assisted control system based on digital twin, which includes an environment modeling module, a vision acquisition module, a spatial mapping module and an intelligent control module;
[0089] The environment modeling module is used to arrange sensors to construct a data twin scenario of the warehouse environment; the vision acquisition module is used to arrange image acquisition devices on the stacker to acquire images at a fixed angle, manually annotate the acquired images and extract image features; the spatial mapping module is used to map the acquired images into the data twin scenario and extract spatial features; the intelligent control module is used to fuse image features and spatial features and construct a classification model in combination with the manual annotation situation;
[0090] The output end of the environment modeling module is connected to the input end of the vision acquisition module; the output end of the vision acquisition module is connected to the input end of the spatial mapping module; the output end of the spatial mapping module is connected to the input end of the intelligent control module.
[0091] The environment modeling module includes a sensor deployment unit, a point cloud noise reduction unit and a model generation unit;
[0092] The sensor deployment unit is used to deploy lidar and ultra-wideband tags, and collect point cloud data and position information in real time; the point cloud denoising unit is used to remove noise points by applying the statistical outlier filtering algorithm and retain the effective point cloud data; the model generation unit is used to construct a dynamically updated three-dimensional grid model through the Poisson surface reconstruction algorithm;
[0093] The output end of the sensor deployment unit is connected to the input end of the point cloud denoising unit; the output end of the point cloud denoising unit is connected to the input end of the model generation unit; the output end of the model generation unit is connected to the input end of the visual acquisition module.
[0094] The visual acquisition module includes a device installation unit, an image acquisition unit, and an image feature extraction unit;
[0095] The device installation unit is used to install an industrial camera at a specified position on the stacker and set a fixed acquisition angle; the image acquisition unit is used to acquire an image sequence at a preset frequency; the image feature extraction unit is used to extract image feature vectors by using an improved SIFT algorithm to enhance the illumination robustness;
[0096] The output end of the device installation unit is connected to the input end of the image acquisition unit; the output end of the image acquisition unit is connected to the input end of the image feature extraction unit; the output end of the image feature extraction unit is connected to the input end of the space mapping module.
[0097] The space mapping module includes a coordinate mapping unit, a space feature calculation unit, and a space feature integration unit;
[0098] The coordinate mapping unit is used to establish a conversion model from pixel coordinates to three-dimensional space based on camera parameters; the space feature calculation unit is used to calculate three-dimensional space features of point cloud density, relative distance, and angle in the mapped area; the space feature integration unit is used to encode the space features into a standardized vector;
[0099] The output end of the coordinate mapping unit is connected to the input end of the space feature calculation unit; the output end of the space feature calculation unit is connected to the input end of the space feature integration unit; the output end of the space feature integration unit is connected to the input end of the intelligent control module.
[0100] The intelligent control module includes a feature fusion unit, a model training unit, and a real-time control unit;
[0101] The feature fusion unit is used to fuse space features and image features; the model training unit is used to optimize SVM parameters by using cross-validation and construct an obstacle classification decision model; the real-time control unit is used to trigger stacker commands according to the classification results to complete closed-loop control;
[0102] The output end of the feature fusion unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the real-time control unit.
[0103] In this embodiment, in an automated warehousing center, the stacker crane runs along the track and needs to identify obstacles (such as scattered goods) on the path in real time. A data twin system is constructed through lidar, ultra-wideband positioning tags, and cameras to achieve obstacle classification and early warning.
[0104] Step 1: Construct a data twin environment;
[0105] Sensor deployment: The lidar (such as Velodyne HDL-64E) is installed on the top of the warehouse, and the scanning range covers the entire storage area, generating point cloud data P=(xj, yj, zj), N = 100,000 points / second; Ultra-wideband positioning tags (UWB) are arranged on the shelves (spacing 5m) and stacker crane tracks (spacing 2m) to collect the position of the stacker crane P(t)=(x(t), y(t), z(t)) in real time, with a sampling frequency of 10Hz;
[0106] Point cloud noise reduction and reconstruction: Statistical outlier filtering parameters: k = 50 neighborhood points, dth = 0.5m.
[0107] Poisson surface reconstruction generates a mesh model M, which is updated every 2 seconds.
[0108] Step 2: Image acquisition and parameter setting;
[0109] Camera configuration: Industrial cameras (Basler acA2040-90um) are installed at the front end of the stacker crane (height 2m), one on each side, with a pitch angle α = 30° and a horizontal angle β = ±45°.
[0110] The acquisition frequency f = 10Hz, the running speed of the stacker crane v = 1m / s, and the single-frame coverage length L = 0.2m, satisfying v / f = 0.1m ≤ L / 2 = 0.1m.
[0111] Image sequence: Each frame of image Ia is 2048×2048 pixels, with a resolution of 0.1mm / pixel.
[0112] Step 3: Image annotation and feature extraction;
[0113] Manual annotation: The annotator performs pixel-level annotation on the types of image obstacles (obstacles and non-obstacles) and the obstacles (such as cardboard boxes) and non-obstacles (such as newspapers) in the image.
[0114] Improved SIFT feature extraction: The neighborhood of the key point is 16×16 pixels, divided into 4×4 sub-regions, calculate the 8-direction gradient histogram, and obtain the 128-dimensional descriptor vector Fimg,a after normalization.
[0115] Step 4: Mapping from image to digital twin;
[0116] Coordinate transformation parameters: The installation height of the camera h = 2m, the focal length f = 12mm, the sensor size is 1 / 1.8 inch, calculate Δx and Δy. The position of the stacker P(t) = (x = 10m, y = 5m, z = 0.5m) (t = 10:00:00).
[0117] Mapping formula: The coordinates of a certain pixel of the obstacle in the image (u = 1000, v = 1500), and convert it to three-dimensional coordinates according to the formula;
[0118] Spatial feature calculation: Point cloud density ρ: The number of point clouds S = 200 in the mapping area, the volume V = 0.1m 3 →ρ = 2000 points / m 3 .
[0119] Distance D: The Euclidean distance from the position of the stacker to the obstacle ≈ 1.2m.
[0120] Relative angle θ: The obstacle is directly in front of the stacker (θ = 0°).
[0121] Spatial feature vector: Fspace,a = [2000, 1.2, 0].
[0122] Step 5: Feature fusion and model training;
[0123] Feature fusion: Combine image features and spatial features: Fa = [Fimg,a, Fspace,a].
[0124] Standardization processing: The mean of each dimension is 0, and the standard deviation is 1.
[0125] SVM model training: Training set: 100,000 labeled images (obstacle: non-obstacle = 1:3); Parameter selection: RBF kernel, cross-validation selects C = 10, γ = 0.1.
[0126] Real-time application: When the stacker is running, each frame of the image is processed and input into the model. If it is determined to be an obstacle, it will immediately stop running and give an early warning.
[0127] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A visual data assisted control method for a stacker based on digital twin, characterized in that: The method includes the following steps: Step 1: Arrange sensors to construct a data twin scenario for the warehousing environment; Step 2: Arrange image acquisition devices on the stacker to acquire images at a fixed angle; Step 3: Manually annotate the acquired images and extract image features; Step 4: Map the acquired images into the data twin scenario and extract spatial features; Step 5: Integrate the image features and spatial features, and construct a classification model in combination with the manual annotation situation.
2. The vision data assisted control method of the stacker based on digital twin according to claim 1, wherein: In Step 1, deploy lidar and ultra-wideband positioning tags in the warehousing environment to construct a digital twin model; The lidar scans the environment and generates point cloud data P=(x j , y j , z j ), where N represents the number of point cloud data, j∈{1, 2, …, N}, representing the point cloud data sequence; (x j , y j , z j ) represents the three-dimensional space coordinates; The ultra-wideband positioning tags are arranged on the shelves and stacker tracks to collect the position of the stacker in real time P(t)=(x(t), y(t), z(t)); where, (t) represents the timestamp, and (x(t), y(t), z(t)) represents the position of the stacker at time t; Denoise the point cloud data: Use statistical outlier filtering to remove noise points; Convert the point cloud into a mesh model M through the Poisson surface reconstruction algorithm; Update the mesh model M in real time; 3. The vision data assisted control method of the stacker based on digital twin according to claim 2, wherein: In Step 2, install industrial cameras at the front end and both sides of the stacker as image acquisition devices; Set the fixed acquisition angle of the camera to fully cover the running path of the stacker; Set the image acquisition frequency; the acquired image sequence is denoted as I a , where a = 1, 2, …; each frame of image I a is a matrix of m×n pixels, where m and n represent the height and width of the image, respectively; In Step 3, the manual annotation is divided into two types: obstacles and non-obstacles; Use the method of manual annotation to annotate the acquired images, and mark the obstacle and non-obstacle areas; Use the improved Scale-Invariant Feature Transform (SIFT) algorithm to extract image features and obtain image feature vectors; Enhance the robustness to illumination changes and normalize the gradient amplitude of each small area; The extracted image feature vector is denoted as F img,a , corresponding to the image I a .
4. The vision data assisted control method for the stacker based on digital twin according to claim 3, wherein: In step 4, the acquired image I a is mapped to the digital twin scenario through coordinate transformation; Utilize the position information P(t) of the stacker and the installation parameters of the camera: installation height h, pitch angle α, and horizontal angle β to establish the mapping relationship between the pixel coordinates (u, v) of the marked area of the image obstacle or non-obstacle and the three-dimensional coordinates (X, Y, Z) of the data twin scenario; In the data twin scenario, based on the mapped coordinates, extract the spatial features of the mapped area; Use the mesh model M constructed by the lidar point cloud data to calculate the point cloud density ρ of the mapped area, ρ = S / V; where, S represents the number of point clouds in the mapped area, and V represents the volume of the mapped area; Calculate the distance D and relative angle θ between the mapped area and the current position of the stacker; The extracted spatial feature vector is denoted as F space,a = [ρ, D, θ], corresponding to the region after mapping of the image I a 5. The vision data assisted control method of the stacker based on digital twin according to claim 4, wherein: In Step 5, integrate the image features and spatial features; The fused feature vector is represented as: F a = [F img,a , F space,a ; Normalize F a , and combine with the manual annotation situation to construct a classification model using the support vector machine SVM; During the SVM training process, use the cross-validation method to select the optimal kernel function parameters and penalty factors to improve the classification accuracy of the model; The trained classification model is used to judge in real time whether the objects on the running path of the stacker are obstacles; when judged as non-obstacles, the stacker continues to run; when judged as obstacles, the stacker immediately stops running and issues a warning.
6. A visual data assisted control system for a stacker based on digital twin, which is applied to the visual data assisted control method for a stacker based on digital twin according to any one of claims 1-5, characterized in that: The system includes an environment modeling module, a visual acquisition module, a spatial mapping module, and an intelligent control module; The environment modeling module is used to arrange sensors to construct a data twin scenario for the warehousing environment; the visual acquisition module is used to arrange image acquisition devices on the stacker to acquire images at a fixed angle, manually annotate the acquired images and extract image features; the spatial mapping module is used to map the acquired images into the data twin scenario and extract spatial features; The intelligent control module is used to fuse image features and spatial features, and construct a classification model in combination with the manual annotation situation; The output end of the environment modeling module is connected to the input end of the visual acquisition module; the output end of the visual acquisition module is connected to the input end of the spatial mapping module; the output end of the spatial mapping module is connected to the input end of the intelligent control module.
7. The visual data assisted control system of the stacker based on digital twin according to claim 6, characterized in that: The environment modeling module includes a sensor deployment unit, a point cloud denoising unit, and a model generation unit; The sensor deployment unit is used to deploy lidar and ultra-wideband tags to collect point cloud data and position information in real time; the point cloud denoising unit is used to apply a statistical outlier filtering algorithm to remove noise points and retain effective point cloud data; the model generation unit is used to construct a dynamically updated three-dimensional mesh model through the Poisson surface reconstruction algorithm; The output end of the sensor deployment unit is connected to the input end of the point cloud denoising unit; the output end of the point cloud denoising unit is connected to the input end of the model generation unit; the output end of the model generation unit is connected to the input end of the visual acquisition module.
8. The visual data assisted control system of the stacker based on digital twin according to claim 7, wherein: The visual acquisition module includes a device installation unit, an image acquisition unit, and an image feature extraction unit; The device installation unit is used to install an industrial camera at a specified position of the stacker and set a fixed acquisition angle; the image acquisition unit is used to collect an image sequence at a preset frequency; the image feature extraction unit is used to extract an image feature vector by using an improved SIFT algorithm to enhance the illumination robustness; The output end of the device installation unit is connected to the input end of the image acquisition unit; the output end of the image acquisition unit is connected to the input end of the image feature extraction unit; the output end of the image feature extraction unit is connected to the input end of the spatial mapping module.
9. The vision data assisted control system of the stacker based on digital twin according to claim 8, wherein: The spatial mapping module includes a coordinate mapping unit, a spatial feature calculation unit, and a spatial feature integration unit; The coordinate mapping unit is used to establish a conversion model from pixel coordinates to three-dimensional space based on camera parameters; the spatial feature calculation unit is used to calculate three-dimensional spatial features of the point cloud density, relative distance, and angle in the mapped area; The spatial feature integration unit is used to encode the spatial features into a standardized vector; The output end of the coordinate mapping unit is connected to the input end of the spatial feature calculation unit; the output end of the spatial feature calculation unit is connected to the input end of the spatial feature integration unit; the output end of the spatial feature integration unit is connected to the input end of the intelligent control module.
10. The vision data assisted control system of the stacker based on digital twin according to claim 9, characterized in that: The intelligent control module includes a feature fusion unit, a model training unit, and a real-time control unit; The feature fusion unit is used to fuse spatial features and image features; the model training unit is used to optimize the SVM parameters by using cross-validation and construct an obstacle classification decision model; the real-time control unit is used to trigger a stacker instruction according to the classification result to complete closed-loop control; The output end of the feature fusion unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the real-time control unit.