Method and system for monitoring material motion behavior based on monocular vision

By using monocular vision technology, instance segmentation and temporal deep learning to identify collision events, and combining camera parameters to calculate the three-dimensional motion trajectory, the problem of automated monitoring of material movement behavior in solid waste sorting equipment has been solved, achieving low-cost and high-precision monitoring of material movement behavior.

CN121661101BActive Publication Date: 2026-04-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-02-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve automatic identification of material movement behavior and 3D trajectory reconstruction based on monocular vision in solid waste sorting equipment, resulting in high hardware costs, difficulties in on-site calibration and maintenance, and an inability to meet the needs of automation and real-time monitoring.

Method used

A monocular high-speed camera is used to capture video of material movement. Combined with instance segmentation and time-series tracking technology, a time-series deep learning model is used to automatically identify collision events. The three-dimensional motion trajectory is then calculated using camera parameters and equipment reference surfaces to solve for dynamic parameters.

Benefits of technology

It achieves low-cost, high-precision automated monitoring of material movement behavior, overcoming the technical bottlenecks of high hardware cost of traditional binocular vision solutions and lack of depth information in monocular vision, and meeting the automation and real-time requirements of industrial sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661101B_ABST
    Figure CN121661101B_ABST
Patent Text Reader

Abstract

The application discloses a monocular vision-based material motion behavior monitoring method and system, relates to the field of material motion monitoring, and first utilizes a monocular high-speed camera to collect a material motion video, adopts instance segmentation and time sequence tracking technology to extract a continuous state sequence of the material on a two-dimensional image plane. The core lies in introducing a time sequence deep learning model to analyze the state sequence, automatically identifying a key event of collision between the material and a working surface of equipment, so that a collision moment and a contact point coordinate are accurately acquired. Thus, the collision contact point is used as a geometric anchor point for three-dimensional space inverse calculation, the two-dimensional image trajectory is reversely projected and reconstructed into a three-dimensional space motion trajectory in combination with camera internal parameters and physical parameters of a device reference surface, and then key dynamic parameters such as speed and a restitution coefficient are solved, so that low-cost, high-precision automatic online monitoring is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of material motion monitoring, and more specifically, to a method and system for monitoring material motion behavior based on monocular vision. Background Technology

[0002] In solid waste disposal processes, core equipment such as air separators and tension screens primarily rely on the differences in the trajectories of materials under the influence of external forces such as gravity and airflow drag to achieve sorting. It is worth noting that the collisions and friction between materials and the equipment's working surfaces (such as guide plates and screen surfaces) during their movement significantly alter their theoretical trajectories; this interaction is a key factor determining the final sorting accuracy. However, solid waste is characterized by diverse forms and complex compositions, resulting in highly random movement behavior. To overcome the limitations of traditional equipment debugging and design relying on manual experience, it is crucial to introduce digital methods for real-time monitoring of the material movement within the equipment. This allows for the acquisition of precise motion data to guide the adaptive adjustment and optimization of equipment operating parameters.

[0003] While machine vision technology is considered an ideal path for achieving the aforementioned monitoring due to its non-contact nature and rich information content, existing technologies still have significant shortcomings in monitoring the 3D trajectory of high-speed moving materials. Traditional monitoring solutions often use binocular stereo vision or RGB+depth cameras to acquire 3D trajectories. Although these solutions are theoretically mature, they are expensive in hardware and have extremely stringent requirements for the synchronous triggering, installation accuracy, and calibration process of the two cameras. This makes them difficult to deploy and maintain stably for a long time in solid waste treatment sites with high vibration and dust. In contrast, while monocular vision solutions have the inherent advantages of low hardware cost and convenient installation, there are few mature applications for monitoring material movement behavior in existing technologies. This is mainly because monocular vision inherently lacks depth information. To reconstruct a 3D trajectory, it is usually necessary to capture the collision point between the material and a known equipment surface as a geometric inverse calculation reference. However, in high-speed, continuous video streams, existing image processing technologies struggle to handle the complex movements of irregular solid waste, cannot automatically and stably identify the material's centroid, and find it difficult to accurately pinpoint the instantaneous frame of collision and contact position from thousands of images. Due to the lack of effective means to automatically detect collision events and analyze them in conjunction with temporal features, monocular systems often cannot automatically establish the mapping relationship between two-dimensional pixels and three-dimensional space, which forces them to rely on cumbersome manual annotation or complex laboratory presets, and cannot meet the needs of industrial sites for automated and real-time monitoring.

[0004] Therefore, developing a monitoring method that can automatically identify collision behavior and reconstruct three-dimensional trajectories based on monocular video has become a technical bottleneck that the industry urgently needs to overcome. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this application provides a method for monitoring material movement behavior based on monocular vision, comprising:

[0006] Acquire video footage of material movement captured by a monocular high-speed camera;

[0007] Material motion videos are segmented into material instances and time-series tracked to obtain a set of material state sequences;

[0008] Collision event identification based on temporal deep learning is performed on the material state sequence set to obtain collision event information;

[0009] Based on the camera parameter set, the collision event information and material state sequence set are used to reconstruct the three-dimensional motion trajectory to obtain the three-dimensional motion trajectory.

[0010] The kinematic parameters of the three-dimensional motion trajectory are calculated to obtain the set of dynamic parameters.

[0011] This application also provides a material movement behavior monitoring system based on monocular vision, which includes:

[0012] The material motion video acquisition module is used to acquire material motion videos captured by a monocular high-speed camera.

[0013] The material status analysis module is used to perform material instance segmentation and time-series tracking on material motion videos to obtain a set of material status sequences.

[0014] The collision event recognition module is used to perform collision event recognition on the material state sequence set based on time-series deep learning to obtain collision event information;

[0015] The 3D motion trajectory reconstruction module is used to reconstruct the 3D motion trajectory based on the camera parameter set, collision event information and material state sequence set to obtain the 3D motion trajectory;

[0016] The kinematic parameter calculation module is used to calculate the kinematic parameters of a three-dimensional motion trajectory to obtain a set of dynamic parameters.

[0017] Compared with existing technologies, this application provides a material motion behavior monitoring method and system based on monocular vision, solving the problem of difficult monitoring of material motion behavior inside solid waste sorting equipment. The scheme first uses a monocular high-speed camera to acquire material motion video, and employs instance segmentation and temporal tracking techniques to extract the continuous state sequence of the material on a two-dimensional image plane. The core lies in introducing a temporal deep learning model to analyze the state sequence, automatically identifying key events of collision between the material and the equipment's working surface, thereby accurately obtaining the collision time and contact point coordinates. Using the collision contact point as a geometric anchor point for three-dimensional spatial inversion, combined with camera intrinsic parameters and the physical parameters of the equipment's reference surface, the two-dimensional image trajectory is back-projected and reconstructed into a three-dimensional spatial motion trajectory, thereby calculating key dynamic parameters such as velocity and coefficient of restitution. This effectively overcomes the shortcomings of traditional binocular vision solutions, such as high hardware costs and difficult on-site calibration and maintenance, while also solving the technical bottleneck of existing monocular technology, which lacks depth information and cannot automatically reconstruct three-dimensional trajectories, and is unable to cope with the random collision behavior of irregular materials, achieving low-cost, high-precision automated online monitoring. Attached Figure Description

[0018] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.

[0019] Figure 1 This is a flowchart of a material movement behavior monitoring method based on monocular vision according to an embodiment of this application.

[0020] Figure 2 This is a schematic diagram of data flow in a material movement behavior monitoring method based on monocular vision according to an embodiment of this application.

[0021] Figure 3 This is a flowchart of step 2 in the material movement behavior monitoring method based on monocular vision according to an embodiment of this application.

[0022] Figure 4 This is a schematic diagram of the data flow in step 4 of the material movement behavior monitoring method based on monocular vision according to an embodiment of this application.

[0023] Figure 5 This is a block diagram of a material movement behavior monitoring system based on monocular vision according to an embodiment of this application. Detailed Implementation

[0024] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0025] In view of the shortcomings in the above-mentioned technical fields, this application proposes a method for monitoring material movement behavior based on monocular vision. Figure 1 This is a flowchart of a material movement behavior monitoring method based on monocular vision according to an embodiment of this application. Figure 2 This is a data flow diagram illustrating a material movement behavior monitoring method based on monocular vision according to an embodiment of this application. Figure 1 and Figure 2 As shown, the material motion behavior monitoring method based on monocular vision according to an embodiment of this application includes: Step 1, acquiring material motion video captured by a monocular high-speed camera; Step 2, performing material instance segmentation and temporal tracking on the material motion video to obtain a material state sequence set; Step 3, performing collision event recognition based on temporal deep learning on the material state sequence set to obtain collision event information; Step 4, reconstructing a three-dimensional motion trajectory based on the collision event information and the material state sequence set using a camera parameter set to obtain a three-dimensional motion trajectory; Step 5, calculating kinematic parameters on the three-dimensional motion trajectory to obtain a dynamic parameter set.

[0026] In step 1, material motion video is acquired using a monocular high-speed camera. It should be understood that in industrial scenarios involving solid waste sorting and processing, the movement of materials in wind-driven or vibrating equipment is highly transient and random. Especially due to the diverse shapes and complex compositions of solid waste, its trajectory under the combined influence of wind drag and gravity often changes drastically within millisecond timescales. To achieve quantitative analysis of this complex physical process, this application first establishes a high temporal resolution observation method to capture high-speed dynamic details that are imperceptible to the human eye. By deploying a monocular high-speed camera, the fleeting material motion process in three-dimensional space can be transformed into a continuous two-dimensional digital image sequence without contact. This not only freezes the instantaneous posture of the material during high-speed motion but also completely records its entire evolution over time.

[0027] In one embodiment of step 1, the specific processing is as follows: First, it involves the precise deployment of hardware equipment and the setup of the environment. A monocular industrial high-speed camera is installed at the key observation window of the solid waste sorting equipment (such as an air classifier or a tension screen). This camera needs to be equipped with a large-aperture fixed-focus lens to cover the preset material movement area, and supplemented with a high-intensity industrial LED supplementary light source to ensure that the image still has sufficient brightness and contrast during short exposure time, thereby avoiding motion blur caused by high-speed movement. Before formally acquiring video of material movement, a camera calibration process needs to be performed to obtain the set of camera parameters required for subsequent steps. A calibration board with a checkerboard pattern is placed at different positions and angles within the camera's field of view to capture multiple images. Image processing algorithms are used to extract corner points and calculate the camera's intrinsic parameter matrix (including focal length and principal point coordinates) and distortion coefficient vector (radial and tangential distortion). Subsequently, a reference mesh plate with known physical dimensions is placed on the actual working surface of the device (i.e., the expected collision plane). By capturing an image of this mesh plate and applying the PnP algorithm, the rotation matrix and translation vector of the camera relative to the device coordinate system are calculated, thus establishing a mathematical mapping relationship between the two-dimensional pixel coordinate system and the three-dimensional physical world coordinate system. Based on this, the parameters of the device reference surface are further calibrated, that is, the equation of the projection line of the device working surface on the two-dimensional image plane or the definition of the effective area is determined. Specifically, this is done by extracting the edge pixel features at the point where the reference mesh plate and the working surface are attached, or by directly identifying the inherent structural boundaries of the device, and applying edge detection and line fitting algorithms to calculate the linear equation coefficients of the working surface in the image coordinate system, thereby providing a clear two-dimensional boundary benchmark for the geometric constraints in subsequent collision event recognition.

[0028] After calibration, the high-speed camera is activated for data acquisition. Considering the high speed of material movement in equipment such as the air separator, the camera's recording frame rate is set to 240 or 480 frames per second, and the image resolution is set to a high-definition specification, such as 1920×1080 pixels, to capture subtle trajectory changes and collision moments. When the sorting equipment starts operating and the material flow enters the field of view, the camera synchronously triggers recording, converting continuous optical signals into digital signals. During the acquisition process, the camera's internal photosensitive element (such as a CMOS sensor) exposes the scene according to the preset frame rate, generating a series of digital images strictly arranged in chronological order.

[0029] After acquisition, these image data are transmitted to the computing unit via a high-speed data interface (such as GigE or USB 3.0) and encoded and packaged into a common video file format (such as AVI or MP4), thus forming a material motion video. This video is essentially a four-dimensional tensor data (time, height, width, number of channels), in which each frame of the image completely records the spatial distribution of all materials within the field of view at a certain moment.

[0030] In step 2, material instance segmentation and temporal tracking are performed on the material motion video to obtain a material state sequence set. Correspondingly, in the industrial operating environment of solid waste sorting equipment, although the raw video captured by a monocular high-speed camera digitally records the entire process of the material on the wind field or vibrating screen, it is essentially still a continuous data stream composed of massive unstructured pixels. Computers cannot directly understand the physical concepts of the material from these RGB pixel matrices, let alone directly obtain trajectory coordinates or collision features for dynamic analysis. To transform this intuitive visual record into structured data that can be processed by a physics engine or algorithm model, this application performs deep semantic parsing on the video stream, effectively separating the irregular solid waste targets in the foreground from the background environment (such as screens and conveyor belts), and establishing an identity correspondence between the same material in different frames in the temporal dimension. Therefore, the material instance segmentation and temporal tracking steps aim to transform the video pixel stream into an ordered material state sequence set with a unique identifier (ID) containing precise contour and location information, providing an indispensable atomic-level data foundation for subsequent identification of microsecond-level collision events and inverse calculation of three-dimensional spatial dynamic parameters.

[0031] Figure 3 This is a flowchart of step 2 in the material movement behavior monitoring method based on monocular vision according to an embodiment of this application. Figure 3 As shown, in one embodiment of step 2, material instance segmentation and temporal tracking are performed on the material motion video to obtain a material state sequence set, including: step 21, performing frame stream decoding on the material motion video to obtain an image frame sequence; step 22, performing deep learning-based instance segmentation on each image frame in the image frame sequence to obtain a single-frame detection set, wherein the single-frame detection includes corrected bounding boxes, confidence scores, binary pixel masks, and two-dimensional centroid coordinates; step 23, performing temporal association and ID allocation based on the ByteTrack algorithm on the single-frame detection set and the historical trajectory set to obtain an updated trajectory set; step 24, aggregating and serializing the updated trajectory set to encapsulate the material state data to obtain a material state sequence set.

[0032] In the above implementation, step 2 is specifically processed as follows: Step 21: At the start of processing, the computing unit first reads the video file through the video decoder and performs frame-by-frame decoding on the compressed data stream encapsulated in the video container according to the linear order of timestamp t. Since videos collected in industrial sites typically have extremely high temporal resolution, such as containing 240 images per second, the decoding process needs to accurately extract the original optical image corresponding to each moment to ensure that there are no dropped frames or temporal errors. If the duration of the collected video is 5 seconds, after this decoding process, a set of original frames consisting of 1200 discrete two-dimensional digital images will be obtained. Each image retains the original resolution at the time of acquisition, such as 1920×1080 pixels and color depth such as 8-bit RGB, completely preserving the spatial distribution details of the material at the corresponding moment. Subsequently, in order to adapt to the fixed requirements of the input tensor dimension of subsequent deep learning models (such as the YOLO-seg model used in this solution), it is necessary to standardize the size scaling of each parsed original image. Considering that the aspect ratio of the original image (16:9) is inconsistent with the input size during model training (usually a square, such as 640×640 pixels), direct stretching would cause distortion of the material shape, thus affecting the accuracy of subsequent physical parameter inversion. Therefore, an adaptive fill scaling strategy is usually adopted in practice: First, the scaling ratio between the original image and the target size is calculated, and the long side (1920 pixels) of the original image is proportionally reduced to the target size (640 pixels), and the short side (1080 pixels) is also proportionally reduced (approximately 360 pixels). Then, the scaled image is placed in the center of a 640×640 gray canvas, and the empty areas above and below are filled with a specific value, such as RGB value 114 gray, thus perfectly maintaining the geometric proportions of the material while changing the physical size of the image. Finally, the scaled image is normalized to eliminate the influence of the absolute value of the illumination intensity on the stability of the model's numerical calculation. The pixel values ​​of the original image are usually distributed in the integer range [0, 255], and are obtained through a linear transformation formula. This maps the value of each pixel channel to a floating-point range of [0,1]. These are the original pixel values ​​in the scaled image. These are the normalized pixel values. This process transforms the discrete integer pixel matrix into a continuous floating-point tensor. After all the above processing, the original video stream is transformed into a normalized sequence of image frames, which is a four-dimensional tensor with dimensions (N, 3, 640, 640), where N is the number of frames (1200).

[0033] In one embodiment of step 22, deep learning-based instance segmentation is performed on each image frame in the image frame sequence to obtain a single-frame detection set, including: step 221, performing size transformation, color space conversion and normalization on the image frames to obtain a normalized image tensor; step 222, inputting the normalized image tensor into a pre-trained YOLO-seg model to obtain target instances, the target instances including filtered bounding boxes, confidence scores and their corresponding binary pixel masks; step 223, centroid calculation is performed on the target instances to obtain two-dimensional centroid coordinates.

[0034] Step 221: As mentioned above, the original decoded image frame has a resolution of 1920×1080 pixels and is stored in memory in BGR (blue-green-red) channel order, commonly used in computer vision libraries. The data type is 8-bit unsigned integers, and the pixel values ​​range from [0, 255]. To adapt to the strict requirements of the input data for the subsequent pre-trained YOLO-seg model, tensor-level preprocessing of these original image data is required. In practice, color space conversion is performed first. Since most deep learning models (including YOLO-seg used in this solution) use RGB (red-green-blue) format image data during training, the image data in memory needs to be rearranged from the BGR color space to the RGB color space. This step is completed through matrix channel swapping to ensure that the color feature distribution of the input model is consistent with the feature distribution learned by the model weights. Then, pixel value normalization is performed. The pixel values ​​of the original image are large and discretely distributed; directly inputting them into the neural network can easily lead to gradient vanishing or exploding, affecting the inference convergence speed and accuracy. This is achieved through the linear mapping formula... The value of each pixel is linearly compressed from the integer range [0, 255] to the floating-point range [0.0, 1.0]. These are the original pixel values ​​in the original image after color space conversion. This is the normalized target value. For example, the original pixel value 128 will be converted to approximately 0.50196. This operation transforms a discrete integer matrix into a continuous floating-point tensor, significantly improving the stability of numerical computation. Next, the channel dimension is adjusted. The standard image storage format is height × width × number of channels (H × W × C), for example, 640 × 640 × 3. However, to utilize the parallel computing power of GPUs (Graphics Processing Units), deep learning frameworks require the input tensor to use a memory layout of number of channels × height × width (C × H × W). Therefore, through a dimension permutation operation, the image tensor is reshaped into a 3 × 640 × 640 format. Finally, to support batch inference, a dimension is added to the 0th dimension of the tensor, forming a four-dimensional normalized image tensor of batch size × number of channels × height × width (B × C × H × W). If processing one frame at a time, the tensor shape is 1 × 3 × 640 × 640. At this point, the original image frame has been transformed into a mathematical expression that fully conforms to the model input standard.

[0035] Step 222: The YOLO-seg model used in this embodiment, such as YOLOv8-seg, is a single-stage instance segmentation network with a carefully designed architecture to balance inference speed and segmentation accuracy. This model is pre-trained on a large general dataset (such as COCO) and fine-tuned on a specific solid waste particle dataset, thereby gaining the ability to extract and discriminate features from irregular materials. The weights and bias parameters in the model are the product of this training process; they store knowledge of material morphology and texture in the form of millions of floating-point numbers. The specific processing flow begins with the model's backbone network. The normalized image tensor is first input into the backbone network based on the CSPDarknet architecture. This network consists of a series of convolutional layers, batch normalization layers, and SiLU activation functions stacked together. During data flow, the network continuously reduces the spatial resolution of the feature map (downsampling) through convolutional operations with a stride of 2, while increasing the channel depth, thereby extracting image features at different scales. For example, the network sequentially generates feature maps at three scales relative to the original image size: 1 / 8 (P3), 1 / 16 (P4), and 1 / 32 (P5). The shallow P3 feature map retains rich geometric texture details, which helps identify small-sized debris; while the deep P5 feature map contains highly abstract semantic information, which helps determine the category attributes of the material. The extracted multi-scale feature maps are then fed into the neck network. This part adopts the PANet structure to solve the problem of feature fusion. PANet contains two information flow paths: one is a top-down path, which passes deep high-level semantic features to the shallow layer through upsampling operations to enhance the model's ability to locate objects; the other is a bottom-up path, which passes shallow fine-grained features to the deep layer to enhance the model's ability to segment object edges. Through this bidirectional fusion, the model generates aggregated feature maps containing rich semantic and spatial information, which can effectively cope with the challenge of huge differences in material size (from fine particles of a few millimeters to blocks of tens of centimeters) in wind sorting equipment. The fused feature maps are finally fed into the decoupling head for prediction. Unlike traditional coupling heads, decoupling heads separate the object detection task from the instance segmentation task. For each feature point, the detection branch predicts a bounding box vector, including center coordinate offset, width and height corrections, and object confidence. The confidence indicates the probability of material presence at that location and the degree of overlap between the predicted and ground truth bounding boxes. Meanwhile, the segmentation branch does not directly output a large full-image mask, but instead predicts a set of mask coefficients. If the model sets 32 prototype masks, this branch will output 32 floating-point coefficients to indicate how the current target is linearly composed of these prototypes. Furthermore, the model includes a separate prototype branch that directly generates a set of global prototype masks from deep features, with a resolution of 1 / 4 of the input image, such as 160×160 pixels.After obtaining the raw prediction results, post-processing is required to filter out the final target instances. First, confidence filtering is performed, iterating through all predicted boxes and eliminating low-quality predictions with confidence scores below a preset threshold (e.g., 0.5), retaining only high-confidence candidate targets. Then, Non-Maximum Suppression (NMS) is applied to address the problem of the same object being detected multiple times. This algorithm calculates the Intersection over Union (IoU) between candidate boxes; when the IoU between two boxes is higher than a set threshold (e.g., 0.45), it is considered a duplicate detection, and only the one with the highest confidence score is retained, thus removing redundancy. Finally, the crucial mask assembly stage occurs. For each target retained after NMS filtering, its 32 predicted mask coefficients are treated as a weight vector and combined linearly with the 32 prototype masks generated by the prototype branch, i.e., matrix multiplication. The result is a single-channel floating-point image, called the soft mask image. To obtain a definite binary segmentation result, the Sigmoid activation function is applied to the soft mask image to map it to the (0,1) interval, and a binarization threshold of 0.5 is set. Pixels with a value greater than 0.5 are marked as foreground (material area) with a value of 1, while those less than 0.5 are marked as background with a value of 0. To further improve accuracy, the generated full-image mask is cropped using the predicted bounding boxes, forcing the area outside the bounding boxes to zero to eliminate background noise interference. Finally, step 222 outputs a set containing several target instances. Each target instance is bound to fine-grained data: the bounding box (…). The scope of the material is precisely defined; the confidence level value, such as 0.85, quantifies the reliability of the recognition; and the binary pixel mask depicts the irregular outline of the material with pixel-level precision, such as a 160×160 0 / 1 matrix, which can be restored to the original image size after upsampling.

[0036] Step 223: The binary pixel mask is a two-dimensional matrix corresponding to the model input size, such as 640×640 pixels, where areas with a pixel value of 1 represent the foreground material and areas with a value of 0 represent the background. To obtain the precise geometric center of the material, the image geometric moments of this binary pixel mask are first calculated. This is done by traversing each pixel in the mask matrix. Using the discretized integral formula Solve for the geometric space moments. Indicates the mask at coordinates The value at that location is either 0 or 1. Represents horizontal coordinates The order of Represents the corresponding vertical coordinate The order of the zeroth moment. Specifically, first calculate the zeroth moment. This is the sum of all foreground pixel values, and its physical meaning is equivalent to the projected area of ​​the material on the image plane. Next, the first moment is calculated. and , It is the weighted sum of the x-coordinates of all foreground pixels, reflecting the cumulative distribution of pixels in the horizontal direction; It is the weighted sum of the y-coordinates of all foreground pixels, reflecting the cumulative distribution of pixels in the vertical direction. After calculating the geometric moments, the centroid coordinates are derived using the centroid calculation formula. (The formula is used to...) and Dividing the first moment by the zero moment allows us to calculate the two-dimensional centroid coordinates of the material at the model input scale. This coordinate, in a physical sense, represents the geometric equilibrium point of the binary mask region. Compared to methods that only take the geometric center of the bounding box, this calculation result can more accurately reflect the actual location of L-shaped, strip-shaped, or irregular sheet-like solid waste. For example, if an irregular plastic sheet is detected at a certain moment, its mask's zero-order moment... 500, first moment It is 160,000. If the value is 120000, then the calculated centroid coordinates at the model scale are (320, 240). Coordinate mapping and encapsulation are then performed. Because the current centroid coordinates... The bounding box coordinates are obtained from a preprocessed (e.g., scaled and padded) normalized image tensor, such as 640×640. To recover their true physical position in the original video frame, an inverse transformation is performed using the original size parameters recorded in step 21 (including the original resolution 1920×1080 and the scaling factor). If padded was performed during preprocessing, the padded offset must be subtracted first, and then divided by the scaling factor. For example, if the scaling factor is 0.333, the centroid coordinates (320, 240) will be mapped back to (960, 300) at the original resolution. Finally, the corrected bounding boxes obtained by inverse transformation and mapping the selected bounding boxes from step 222 back to the original image resolution, the confidence scores and original binary pixel masks directly output by the model inference in step 222, and the aforementioned calculated and mapped two-dimensional centroid coordinates are aggregated and encapsulated into a standard single-frame detection object. This object fully represents the geometric and physical properties of the material in the current frame. By performing this operation frame by frame on the image frame sequence, a single-frame detection set is finally generated.

[0037] In particular, in specific applications of solid waste sorting, the materials to be processed often exhibit highly complex physical characteristics. Examples include engineering plastics with metal inserts, locally absorbent and damp waste cardboard, or composite materials with adhered oil stains. The actual mass distribution of these materials typically exhibits significant non-uniformity, causing a substantial deviation between their physical rotation center and geometric contour center. Simultaneously, in the high-speed transmission environment of industrial sites, the material images captured by cameras are inevitably affected by motion blur. The inherent pixel-level prediction jitter of the superimposed instance segmentation algorithm when handling irregular edges poses a severe challenge to traditional homogenization centroid calculation methods based on binary masks. These methods treat all pixels within an object as equally weighted 1s, failing to perceive density gradients caused by material differences or remove observation noise from edges. This results in high-frequency random jitter in the calculated centroid coordinates, which does not match the actual physical motion of the object. This jitter is drastically amplified when performing second-order differentiation to calculate dynamic parameters such as acceleration, severely interfering with the judgment of key sorting indicators such as the collision recovery coefficient. Therefore, in order to overcome the technical defects caused by the above-mentioned homogeneous assumption and edge equal weighting, and to provide high signal-to-noise ratio basic data for subsequent kinematic parameter calculation, this embodiment adopts a robust physical centroid estimation method based on visual pseudo-density field and edge confidence. The original binary mask is replaced by constructing a non-uniform weight field that can characterize the internal mass distribution and edge reliability.

[0038] In a preferred embodiment of step 223, centroid calculation is performed on the target instance to obtain two-dimensional centroid coordinates, including: step 2231, extracting feature images and soft segmentation masks from the target instance; step 2232, constructing a fusion weight matrix based on the feature images and soft segmentation masks; step 2233, performing adaptive higher-order moment calculation on the pixel coordinate network based on the fusion weight matrix to obtain robust higher-order moments; and step 2234, deriving and correcting robust physical centroid coordinates based on the robust higher-order moments to obtain two-dimensional centroid coordinates.

[0039] Step 2231: This step aims to acquire raw data containing rich probabilistic information and texture details, rather than truncated binarized information. Specifically, an unthresholded soft segmentation mask is extracted directly from the inference results of the deep learning model. This mask is a floating-point matrix with values ​​ranging from zero to one, where the value of each element quantifies the confidence probability that the corresponding pixel belongs to the material foreground. Simultaneously, the region image corresponding to the target instance is cropped from the original video frame, and its grayscale brightness map or texture energy map is extracted as a feature image. The differences in pixel values ​​in the feature image indirectly reflect the material variations on the material surface. For example, dark or complex textured areas usually correspond to denser physical structures, providing a visual basis for subsequent simulation of the non-uniform mass distribution within the material.

[0040] Step 2232: To simultaneously address the issues of edge noise interference and uneven mass distribution within the material, a fusion weight matrix is ​​needed to characterize the mass contribution of each pixel within the material. Visual information is transformed into estimated physical mass weights through mathematical modeling, thereby generating a pixel-wise weight field. In one embodiment of step 2232, the fusion weight matrix is ​​constructed based on the feature image and soft segmentation mask, including: constructing the fusion weight matrix using the following formula:

[0041] in, At the pixel The fusion weight value at the location, It is a soft segmentation mask at the point The probability value at that point is the soft segmentation mask extracted in step 2231. This is the confidence sharpening index, which is set to a value greater than 1. It is empirically set based on pre-statistical analysis of the edge prediction jitter variance of the instance segmentation model in high-speed motion scenes, for example, set to 2.0. It is the feature image at the point Pixel value at that location, The vision-quality coupling coefficient is an adjustable hyperparameter used to control the contribution of visual features to quality estimation. It is determined by analyzing the correlation between the photometric features of images and the actual physical density distribution of a specific type of solid waste through offline calibration experiments. and These are the pixel mean and standard deviation of the feature image within the material region, used to standardize pixel values ​​to eliminate the absolute influence of illumination intensity. It is a very small constant, such as 1e-6, to prevent the denominator from being zero. In this formula, the first term... This is the edge reliability term, a noise reduction filter that tells the algorithm to only trust the definite parts at the center of the object and ignore the blurry parts at the edges. Through exponential operations, the confidence level can be non-linearly stretched; for blurry pixels with low edge processing confidence, such as around 0.5, their weights are significantly suppressed. ), while the central high-confidence region, such as the region with a weight of 0.95, is retained. This effectively reduces edge noise interference. The second term in the formula is a pseudo-density simulation term, used to adjust the weights based on visual features. In other words, it's a quality adder that tells the algorithm that although this is an image, areas with darker colors (or heavier textures) are more important, thus shifting the centroid in that direction. For example, when sorting scrap circuit board fragments, the density of dark-colored chip areas is much greater than that of light-colored substrate areas. In this case, the density can be adjusted accordingly. Setting it to a positive value (such as 0.5) gives dark areas (assuming that low grayscale values ​​are inverted or mapped to high values) a higher computational weight.

[0042] In step 2233: Considering that traditional moment calculation based on binary masks cannot reflect weight differences, after obtaining the fusion weight matrix that can characterize the quality distribution and confidence level, it is necessary to use this weight matrix for weighted integration to solve for higher-order moments that reflect physical characteristics. This step utilizes the constructed non-uniform weight field to perform weighted integration on the spatial distribution of the material, thereby calculating the geometric moments that reflect physical characteristics. The calculation process traverses every pixel within the material region, multiplying and summing its spatial coordinates with the corresponding fusion weight value, as shown in the following formula:

[0043] in, To calculate the robust (p+q) order geometric moments, , These are the column and row coordinate indices of the pixel, respectively. , Let be the order of the moment. When =0, When =0, the calculated This represents the weighted virtual total mass of the material; when =1, =0 or =0, When =1, the calculated and These represent the weighted static moments of the material in the horizontal and vertical directions, respectively. Furthermore, Represents image coordinates The trainable fusion weights are calculated at the image space. This calculation process essentially simulates the mechanical integration process of a non-uniform density thin plate, so that the final moment value no longer depends solely on the contour area of ​​the material, but is mainly determined by the main body region with high confidence and significant texture, fundamentally suppressing the numerical fluctuations caused by edge jitter.

[0044] Step 2234: Transform the abstract moment features into specific spatial coordinates. Using the ratio of the first moment to the zeroth moment, calculate the position of the weighted physical centroid in the image coordinate system. The calculation formula is as follows:

[0045]

[0046] The calculation here This is the robust physical centroid after edge denoising and density-weighted correction. Finally, combining the original image scaling and padding parameters recorded in step 21, this coordinate is inversely mapped back to the coordinate system of the original video frame. For example, for a discarded plastic bottle tumbling in the air, the density of the cap is significantly greater than that of the body. The geometric center calculated by traditional methods will exhibit violent spiral swings as the bottle rotates. However, the centroid calculated using this method will be more tightly anchored near the physical center of gravity of the cap, and its generated trajectory will be closer to a smooth parabola. This processing significantly improves the smoothness and realism of the centroid data over time, ensuring that subsequent parameters such as collision velocity vectors and energy loss rates calculated based on this coordinate have extremely high physical reliability.

[0047] Step 23 primarily addresses the issue of cross-frame target identity consistency. The input to this step includes two key sets: one is the single-frame detection set of the current frame t generated in step 22; the other is the historical trajectory set inherited from the previous frame t-1, which records the motion state and identity ID of all previously identified materials. To achieve robust tracking in a high-speed, mutually occluded solid waste sorting environment, this embodiment employs the ByteTrack algorithm, an advanced multi-target tracking strategy that fully utilizes low-confidence detection boxes for association. First, a position prediction stage is performed. For each existing trajectory object in the historical trajectory set, a Kalman filter is used for state updates. The Kalman filter is a recursive linear minimum variance estimator that internally maintains a state vector containing the target's center coordinates, aspect ratio, height, and their corresponding rates of change. Based on the posterior estimation of the previous frame, the Kalman filter uses the physical motion equations to deduce the prior state of the current frame t, i.e., predicting the theoretical position and bounding box where the trajectory might appear in the current frame. This prediction process effectively utilizes the inertial characteristics of the material moving in the air separator, providing prior constraints on the search range for subsequent matching. Next, the classification and detection box operation is performed. To distinguish high-quality targets from potentially occluded targets, the algorithm introduces a key parameter: the confidence threshold. This threshold is determined during the YOLO-seg model training phase by analyzing the precision-recall (PR) curve on the validation set, aiming to balance the false negative and false positive detection rates. In this embodiment, the high-score threshold is set to 0.5, and the low-score threshold is set to 0.1. The single-frame detection set of the current frame is traversed, and detection boxes with a confidence score higher than 0.5 are classified as high-score detection boxes. This represents clearly visible material; detection boxes with confidence levels between 0.1 and 0.5 are classified as low-scoring detection boxes. These bounding boxes often correspond to material fragments with severe motion blur or partial occlusion. The process then proceeds to the first-level cascaded matching stage. This stage aims to match the high-resolution detection bounding boxes... Trajectory predicted by Kalman filter The algorithm performs the association. It calculates the Intersection over Union (IoU) distance matrix between the two bounding boxes. IoU is a geometric metric that measures the degree of overlap between two bounding boxes, and its calculation formula is... .in, This represents the detection bounding box output by the model in the current frame. This represents the bounding box predicted by the Kalman blobs. (Numerator of the formula) Calculate the area of ​​the geometric intersection of two rectangles, i.e., the overlapping region; the denominator is... The area of ​​the geometric union of the two bounding boxes is calculated, which represents the total coverage area. The closer this ratio is to 1, the higher the overlap between the predicted and detected locations, and the greater the likelihood that they belong to the same target. Based on this distance matrix, the Hungarian algorithm is used to solve the maximum matching problem in the bipartite graph, thereby determining the optimal pairing relationship. For successfully matched trajectories, their Kalman filter states are updated using the current high-resolution detection boxes; unmatched high-resolution detection boxes are considered potential new targets; and unmatched predicted trajectories proceed to the next round. The second-level cascaded matching, the core step of byte association, is then executed. In this step, the algorithm attempts to match the remaining unmatched trajectories from the previous round. With low-scoring detection box set Association is performed, also based on IoU distance and the Hungarian algorithm. In traditional tracking algorithms, low-resolution bounding boxes are usually discarded directly. However, in solid waste sorting scenarios, materials often experience a temporary drop in confidence due to rolling or mutual occlusion. This round of spot-checking effectively recovers material trajectories that are blurry but definitely exist, significantly reducing trajectory breaks caused by occlusion. After completing two levels of matching, ID management is performed. For trajectories that are successfully matched in the first or second level, their original unique identifier (ID) remains unchanged, and their position parameters are updated. For high-resolution detection boxes that do not find a corresponding trajectory in either of the two matches, they are determined to be newly entered materials. The system initializes a new Kalman filter state for them and assigns an auto-incrementing globally unique ID, such as ID:1024. For trajectories that have not matched any detection boxes for a long time (e.g., more than 30 frames), it is determined that the material has left the field of view or has completely disappeared, and it is removed from the tracking list. At this point, an updated set of trajectories containing the states of all active materials in the current frame is generated.

[0048] Step 24: First, data extraction is performed. Each trajectory object in the updated trajectory set is traversed, and key fields are extracted. These fields include not only the frame timestamp (accurate to milliseconds) and the unique material ID, but also, specifically, the precise two-dimensional centroid coordinates calculated in Step 223 and the corresponding binary pixel mask. For example, for a plastic bottle with ID 1024, its centroid coordinates and contour mask in frame 50 are extracted. Next, sequence construction is performed. The system maintains a hash table structure in memory with ID as the primary key. For each material data extracted in the current frame, it is encapsulated as a state node and appended to the corresponding ID's data bucket (List or Vector) in chronological order. This process is similar to writing a motion log for each material in real time. Finally, integrity verification and output are performed. When video processing ends, or a trajectory with a certain ID is marked as deleted, the system extracts the entire linked list corresponding to that ID and performs an integrity check (e.g., removing noisy trajectories shorter than 5 frames). After successful verification, the object containing the material's entire lifecycle data is encapsulated into the final material state sequence set. This collection is a highly structured dataset that fully records all spatiotemporal information of each material from its entry into the field of view to its exit from the field of view.

[0049] In step 3, collision event recognition based on temporal deep learning is performed on the material state sequence set to obtain collision event information. It is understandable that in industrial applications of solid waste sorting and processing, while monocular vision solutions offer significant advantages in terms of low cost and easy installation, their core challenge lies in the lack of depth information. To calculate an accurate three-dimensional spatial trajectory from a two-dimensional image plane, a geometric constraint with known physical coordinates is needed as a reference. In equipment such as air separators or tension screens, the moment of physical collision between the material and the equipment's working surface (such as screens or guide plates) constitutes the ideal geometric anchor point. However, when faced with solid waste of varying shapes and high-speed tumbling, relying solely on simple kinematic features (such as sudden velocity changes) or traditional image differencing techniques is insufficient to distinguish between real collision behavior and non-contact tumbling motion in complex backgrounds with significant visual noise and occlusion. If the microsecond-level moment of collision cannot be accurately pinpointed, subsequent three-dimensional reconstruction will completely fail due to the lack of accurate anchor points. Therefore, a collision event recognition step based on temporal deep learning is introduced to utilize artificial intelligence models to perform deep semantic analysis on the spatiotemporal evolution characteristics of material movement, thereby automatically and robustly identifying extremely subtle collision features from continuous video streams, accurately locating collision key frames, and providing indispensable temporal and spatial boundary conditions for establishing the mapping relationship between two-dimensional pixels and three-dimensional space.

[0050] In one embodiment of step 3, collision event identification based on temporal deep learning is performed on the material state sequence set to obtain collision event information, including: step 31, extracting and serializing trajectory segments from the material state sequence set to obtain a normalized temporal tensor set; step 32, performing collision keyframe inference based on CNN-LSTM on the normalized temporal tensor set to obtain a keyframe index and collision confidence; step 33, based on the keyframe index and device reference surface parameters, performing precise localization of contact points under geometric constraints on the material state sequence set to obtain collision event information, wherein the collision event information includes material ID, keyframe index, and two-dimensional coordinates of the collision contact point.

[0051] In the above implementation, step 3 is specifically processed as follows: Step 31: The input to this step is the material state sequence set output in step 2. This set stores the full lifecycle data of each tracked material (e.g., a waste plastic bottle with a unique ID of 1024) within the field of view, including the timestamp of each frame from entering to leaving the field of view, the corrected bounding box, the two-dimensional centroid coordinates, and the binary pixel mask. In order for the deep learning model to understand the dynamic behavior of the material, the discrete frame data first needs to be reconstructed into a tensor format containing spatiotemporal context information. In practice, the processing unit traverses each material ID in the material state sequence set. For a specific ID, such as ID:1024, a fixed sliding time window is set, denoted as W. The size of this window is set according to the camera's frame rate and the typical duration of the material collision process. For example, at a frame rate of 240fps, the collision process (including approach, contact, and bounce) usually occurs within 50 milliseconds, so W is set to 16 frames (approximately 66 milliseconds) to ensure that the window can completely cover the temporal features before and after the collision. The sliding window moves across the entire trajectory sequence with a step size of 1, generating a time segment to be analyzed with each slide. Within each time window, image segments of the region of interest (ROI) are extracted. Although the full image resolution is as high as 1920×1080, full image analysis is not only computationally intensive but also introduces background noise. Therefore, the algorithm uses the two-dimensional centroid coordinates of the material in the current frame. Centered on, based on the size of its bounding box Outward cropping is performed. To accommodate the deformation that may occur at the moment of collision and to preserve information about the surrounding local environment (such as shadow changes and relative distance to the screen surface), the expansion ratio is set to 1.5 times. That is, the width of the cropping area is 1.5w and the height is 1.5h. These W consecutive local images are extracted from the original high-definition video frames to form a short film clip focusing on the details of the material's movement. Then, serialization and tensor normalization are performed. Since neural networks require input data to have a strictly uniform dimension, the extracted W ROI image clips (the original size may vary depending on the material size, such as 50×50 or 80×80 pixels) need to be uniformly scaled to a fixed resolution, such as 128×128 pixels, using a bilinear interpolation algorithm. Next, the image pixel values ​​are channel normalized, mapping RGB values ​​from [0,255] to the [0,1] interval, and adjusted according to a standardized distribution (such as subtracting the mean and dividing by the standard deviation) to accelerate model convergence. Finally, these W processed images are arranged in chronological order. Stacking them along the channel or time dimension, encapsulating them into a five-dimensional tensor, whose dimensional structure is as follows: That is, (1,16,3,128,128). This highly structured normalized temporal tensor set completely preserves the appearance and motion changes of the material within a specific time period, serving as the standard input for subsequent CNN-LSTM model inference.

[0052] Step 32: This step leverages the powerful feature extraction and temporal reasoning capabilities of deep neural networks to solve the core challenge of collision recognition. The model architecture used is a typical CNN-LSTM hybrid architecture. The model's weights and biases were obtained offline through supervised learning using a large number of manually annotated collision video datasets (such as the data annotated via an interactive interface mentioned in the background). First, spatial feature encoding is performed. The normalized temporal tensor set generated in Step 31 is input into the model's spatial feature extraction module. This module consists of a convolutional neural network (CNN) that shares weights across time steps, using lightweight backbone variants such as ResNet-18 or EfficientNet-B0. Each frame of the input tensor, with a size of 3×128×128, is independently and in parallel fed into the CNN. The CNN performs filtering operations on the image through multiple convolutional kernels to extract the edges, textures, and morphological features of the material; pooling layers reduce the spatial resolution of the feature map to suppress noise. Through layers of abstraction, each original frame of the image is mapped to a high-dimensional one-dimensional feature vector, such as a vector of length 512. This vector encapsulates the static visual features of the material at that moment, such as whether the material has undergone compression deformation or whether the material edges have become blurred. Next, temporal evolution analysis is performed. The continuous feature vector sequence output by the CNN is input into the Long Short-Term Memory (LSTM) network module in strict chronological order. LSTM is a special type of recurrent neural network (RNN) specifically designed to handle dependencies in long sequences of data. The LSTM unit contains three key mechanisms: a forget gate, an input gate, and an output gate, used to dynamically adjust the storage and discarding of information. Its core computation process is expressed by the formula...

[0053] Description. In this formula, This refers to the feature vector extracted above. . It is the input weight matrix, used to map spatial features to the hidden layer dimension of the LSTM; It is a cyclic weight matrix used to store the hidden state from the previous time step. Transmitted to the current moment. This is the bias vector. This represents the hidden state of the LSTM at time t. It is not merely the response to the current input, but also a memory container containing historical information about the material's movement up to the current time. For example, if... It records information that the material is moving downwards at high speed, and the current input... When the material suddenly stops or reverses, the gating mechanism inside the LSTM will capture this drastic state conflict, thereby... The LSTM encodes high-response signals that undergo sudden changes. This process simulates the cognitive process humans experience when observing an object suddenly bouncing. Finally, classification / regression prediction is performed. The hidden state output by the LSTM at each time step... The data is fed into a fully connected layer, mapping the high-dimensional feature space to a one-dimensional scalar space. Then, it undergoes Sigmoid activation, outputting a value between [0,1], which is the collision probability score. This score represents the model's confidence that a collision has occurred at the current time t. The processing unit collects the score sequence throughout the time window and uses a peak detection algorithm to find local maxima. If the predicted score at a certain time t exceeds a preset collision threshold, such as 0.85 (determined based on the F1-score of the validation set), and this time is at the peak of the score curve, then a collision event is determined to have occurred at that time. The frame number corresponding to this time is marked as a keyframe index, and the corresponding score is the collision confidence. For example, for the sequence with ID 1024, the model analysis indicates that the confidence of frame 56 is 0.92, indicating a collision has occurred. Finally, the output keyframe index and collision confidence are bound to the material ID, serving as the direct basis for precise contact point localization under geometric constraints in step 33.

[0054] Step 33: This step, as the final stage of the collision event recognition process, relies on the inference results of previous steps: namely, the keyframe index determined by the CNN-LSTM model (e.g., the material with ID 1024 collides in frame 56) and the device reference surface parameters obtained in the calibration stage of Step 1. First, a mask contour extraction operation is performed. Based on the input keyframe index, the binary pixel mask corresponding to the material at that moment is extracted from the material state sequence set residing in memory. This mask is a 0 / 1 matrix corresponding to the size of the original image, where regions with a value of 1 represent material entities. To obtain edge information for geometric calculations, an edge detection algorithm (such as the Canny operator or the findContours algorithm) is applied to scan the mask, extracting the edge contour point set around the material. This point set consists of a series of discrete pixel coordinates. The composition accurately depicts the geometric shape of the material on the image plane at the moment of collision. Next, nearest neighbor optimization and contact point locking are performed. Since monocular images lack depth information, the physical fact of the material contacting the equipment's working surface is used as a geometric constraint. The equipment reference surface parameters exist in the image coordinate system in the form of linear equations, denoted as... The coefficients of this equation This is pre-determined during the system initialization phase by photographing a calibration plate placed on the device surface or identifying the inherent edge features of the device. The algorithm iterates through each pixel in the edge contour point set, calculating its Euclidean distance to the line. To find the point that best matches the physical contact reality, an optimization objective formula for contact point localization is constructed and solved: In this formula, For the edge contour point set, Represents the coordinates of any pixel point on the material outline; This represents the absolute value of the algebraic distance after substituting the coordinates of the point into the equation of the line. This is the magnitude of the straight-line normal vector, used to standardize the algebraic distance to a geometric distance. The physical meaning of this formula is to find the one or a group of pixels on the contour that are closest to the equipment reference surface. For example, if the equipment screen surface appears as a horizontal straight line y-900=0 at the bottom of the image, i.e., a=0, b=1, c=-900, and the coordinates of the lowest point on the material contour are (320, 898), substituting these values ​​into the calculation yields a distance of 2 pixels. If multiple points with extremely small distances exist (e.g., materials in planar contact), the geometric centroid of these points is selected as the final two-dimensional coordinates of the collision contact point. The coordinates Mathematically, this represents the optimal solution; physically, it precisely characterizes the contact point where energy exchange occurs between the material and the equipment. Finally, information encapsulation is performed. The unique material identifier confirmed in previous steps, such as ID: 1024, the keyframe index obtained through temporal inference, such as frame 56, and the precisely calculated contact point coordinates, such as (320, 898), are aggregated and packaged to generate structured collision event information. This information package contains not only temporal positioning but also precise spatial positioning, serving as the unique geometric anchor point data output to the subsequent 3D trajectory reconstruction module, thus achieving a crucial leap from 2D image features to 3D physical parameters.

[0055] In step 4, based on the camera parameter set, a three-dimensional motion trajectory is reconstructed from the collision event information and the material state sequence set to obtain the three-dimensional motion trajectory. Specifically, the camera parameter set mentioned here includes the camera intrinsic matrix, distortion coefficient vector, rotation matrix, and translation vector. It should be understood that although monocular vision technology has advantages in solid waste sorting and monitoring such as low hardware cost and flexible deployment, its imaging principle essentially projects the three-dimensional physical world onto a two-dimensional image plane, inevitably leading to the loss of depth information. Directly analyzing material motion based on two-dimensional image pixel coordinates cannot obtain key dynamic parameters such as actual physical displacement, velocity (m / s), and collision energy, thus failing to accurately assess the actual operating efficiency of the sorting equipment. Although the time point (keyframe) of the collision between the material and the equipment and its pixel position on the image have been accurately identified in the previous steps using a deep learning model, this is only a two-dimensional observation result. To give these data physical meaning, a mathematical mapping between the two-dimensional pixel space and the three-dimensional physical space needs to be established. Therefore, the three-dimensional motion trajectory reconstruction step based on the camera parameter set aims to take advantage of the strong geometric constraint that the material must be located on the surface of the equipment when the collision occurs, use the collision point as the anchor point for three-dimensional spatial inverse calculation, and combine the physical plane assumption of the material motion to restore the time-varying two-dimensional pixel sequence into a real-scale three-dimensional spatial motion trajectory.

[0056] Figure 4 This is a schematic diagram of the data flow in step 4 of the material movement behavior monitoring method based on monocular vision according to an embodiment of this application. Figure 4As shown, in one embodiment of step 4, based on the camera parameter set, a three-dimensional motion trajectory is reconstructed from the collision event information and the material state sequence set to obtain the three-dimensional motion trajectory, including: step 41, based on the equipment working surface equation and the camera parameter set, a three-dimensional back-calculation of the collision anchor point based on geometric constraints is performed on the collision event information to obtain the coordinates of the three-dimensional physical contact point; step 42, based on the coordinates of the three-dimensional physical contact point and the prior assumptions of physical motion, anchoring and fixing, normal vector estimation and plane equation determination are performed to obtain the fitted motion plane parameters; step 43, based on the fitted motion plane parameters and the camera parameter set, a back-projection reconstruction and smoothing of the material state sequence set is performed on the entire link trajectory to obtain the three-dimensional motion trajectory.

[0057] In the above implementation, step 4 is specifically processed as follows: Step 41: The core of this step is to use the equipment working surface with a known physical location to recover the lost depth information in monocular vision. For example, in the calibration stage, a world coordinate system with a certain point on the equipment working surface as the origin has been established using the PnP algorithm. At this time, the equation of the equipment working surface (such as the screen plane of an air separator) is mathematically described as an infinitely extending plane, for example... This indicates that the plane is located in the world coordinate system. On a plane, or more generally represented as a point-normal equation, distortion correction is performed first. Because cameras used in industrial settings are equipped with large-aperture wide-angle lenses to cover a wide field of view, the captured images often exhibit radial distortion (barrel or pincushion distortion) and tangential distortion. Directly using the raw pixel coordinates introduces significant spatial positioning errors. Therefore, the camera intrinsic matrix is ​​used... (Including focal length) , and principal point coordinates , and distortion coefficient vector (Include , , , (etc.), for the input two-dimensional pixel coordinates of the collision contact point Distortion correction is performed. This process uses an iterative algorithm to map the geometrically distorted observation coordinates back to normalized image plane coordinates that conform to the ideal pinhole camera model. For example, the original observation coordinates (320, 898) may be adjusted to (318.5, 895.2) after correction, thus eliminating nonlinear errors caused by physical defects in the lens. Next, ray generation is performed. Based on the ideal pinhole imaging model, each pixel in the image actually corresponds to a ray in three-dimensional space originating from the optical center of the camera. The rotation matrix in the camera's extrinsic parameters is used... Translation vector The camera's optical center can be determined. The formula for calculating the absolute position in the world coordinate system is: Simultaneously, a unit direction vector is constructed from the optical center to the distorted pixel. In mathematical terms, this vector can be represented in the world coordinate system as: ,in It is in homogeneous coordinate form of the pixel. This ray represents the set of all possible 3D spatial points that can be projected onto this pixel. Finally, constraint intersection (anchor point calculation) is performed. At this point, there exists a definite line-of-sight ray and a known equation of the equipment working surface in the system. According to physical facts, at the microsecond instant of the collision, the material must have physical contact with the equipment working surface. Therefore, the geometric intersection of this ray and the equipment plane is the unique and precise position of the material at the instant of the collision. By solving the simultaneous ray equations... With plane equation , It is the starting point of the ray, that is, the optical center of the camera. It is the unit direction vector of the ray, i.e., the line-of-sight direction. It is the specific three-dimensional physical contact point where the material collides with the working surface of the equipment. It is any known point on the plane, such as the origin. It is any point on the plane. It is the normal vector of the plane, because at the intersection point, the point on the ray... It must be equal to a point on the plane. By solving the simultaneous equations, we can obtain This allows us to obtain the coordinates of the three-dimensional physical contact point. The specific calculations follow the formula: In this formula, The world coordinates representing the optical center of the camera; It is the normal vector of the working plane of the equipment. For example, for a horizontal screen surface, the normal vector might be... ; It is any known point on the plane (such as the origin). ). Denominator The dot product of the line of sight direction and the plane normal vector was calculated to determine the angle between the line of sight and the plane, ensuring the stability of the numerical solution (the denominator is not zero). The numerator represents the projection of the perpendicular distance from the optical center to the plane. For example, the calculated coordinates (150.5mm, 400.2mm, 0.0mm) not only satisfy the geometric constraints mathematically, but also serve as the absolute anchor point for the reconstruction of the material's three-dimensional trajectory physically, solving the fundamental problem of depth uncertainty in monocular vision.

[0058] Step 42: After obtaining the unique physical contact point, a priori assumption about the physical motion needs to be introduced to reconstruct the complete trajectory. This is because the motion of solid waste particles in the air separator or tension screen is not random, but rather approximately located within a specific two-dimensional plane under the combined influence of gravity, airflow drag, and initial velocity. This assumption reduces the complex three-dimensional curve reconstruction problem to a constrained planar projection problem. First, anchor points are fixed. The calculated coordinates of the three-dimensional physical contact point are then... This is set as a fixed point that the fitted motion plane must pass through. This means that regardless of the final calculated planar pose, it must include this absolute position determined by visual observation and geometric constraints. Normal vector estimation is then performed. This step requires constructing the normal vector based on the specific operating principles of the equipment. Taking an air classifier as an example, the material is subjected to a vertically downward force of gravity. and the drag force of the main airflow in the horizontal direction Function. According to the principles of mechanics, the trajectory of material movement is mainly confined within a plane spanned by the gravity vector and the airflow direction vector, or within a rebound plane perpendicular to the screen surface. If the material is subjected to the rebound effect of the screen surface and mainly moves along the longitudinal axis of the equipment, it can be deduced that the normal vector of the motion plane is perpendicular to the direction of motion. For example, if the material is in... If a object undergoes parabolic motion within a plane, then the normal vector of the plane of motion should be parallel to... Axis, i.e. In practice, the direction of the two-dimensional pixel connections in the frames before and after the collision can be analyzed, and combined with camera extrinsic parameters, to roughly estimate the projection direction of the motion trajectory on the horizontal plane, thus constructing the normal vector more accurately. Finally, the plane equation is established. Using the point normal equation principle, through known anchor points... and the estimated normal vector This uniquely determines the plane equation describing the material's motion space, i.e., the fitted motion plane parameters. Mathematically, it is described as follows: After unfolding, you get This derivation process is based on the standard mathematical steps of converting the point normal form of a plane equation to its general form in analytic geometry. Among these steps... That is, the components of the unit normal vector. Let G be the coordinates of any point on the plane. Through calculation The distance parameters obtained. This equation defines a virtual slice plane that cuts through the camera's view frustum in three-dimensional space, and the entire trajectory of the material is considered as a curve drawn on this slice plane.

[0059] Step 43: This step performs an inverse projection transformation, mapping each two-dimensional observation point in the time series back to three-dimensional space. The processing unit traverses the material state sequence set, and for each frame t in the sequence, extracts the corresponding two-dimensional centroid coordinates, i.e., the pixel coordinates calculated in step 223, denoted as . Since it has been determined that the material is always located within the fitted plane of motion. Therefore, the three-dimensional coordinates of each pixel in each frame are... This refers to the intersection of the line-of-sight ray passing through the optical center at that pixel and the fitted motion plane. To improve computational efficiency, an inverse transformation method based on the homography matrix is ​​used. The formula is described as follows: .in, Homogeneous coordinates of pixels . It is the planar homography matrix induced by the equations of motion plane and the camera projection matrix. Specifically, due to Satisfy the plane equation And satisfy camera projection This allows us to derive a direct mapping matrix from the pixel plane to the motion plane. In this embodiment, a more intuitive calculation method is to again utilize the ray-plane intersection method: for each frame t, the logic in step 41 is repeated, but this time the plane is no longer the device working plane, but the fitted motion plane determined in step 42. That is, to calculate... By performing this frame-by-frame calculation, what was originally just pixel coordinates can be transformed... The sequence is converted into a three-dimensional coordinate sequence with physical units of measurement (such as millimeters). After geometric inverse calculation, trajectory smoothing is performed. Since the original 2D detection centroid may be affected by image noise, illumination changes, or segmentation edge jitter, the directly inversely calculated 3D point set... The motion may exhibit sawtooth-like fluctuations in space, which will severely affect the accuracy of subsequent velocity and acceleration differential calculations. Therefore, Kalman filtering or least-squares polynomial fitting is applied to the discrete three-dimensional coordinate set. For example, for parabolic motion segments, quadratic polynomials can be used to fit the data. Regression smoothing is performed to filter out high-frequency noise and restore a smooth curve that conforms to physical laws. Finally, temporal concatenation is performed. The smoothed 3D coordinate points are reconnected according to the timestamp order of the original video to generate a continuous 3D motion trajectory. This trajectory dataset contains complete 3D path information of the material from entering the field of view, colliding with it, to leaving the field of view. For example, the data may show that a certain ID material is located at (100, 300, 500) at t=0ms, falls at a speed of 2.5m / s, collides with the screen surface located at Z=0 at t=200ms, and then bounces back at a speed of 1.8m / s.

[0060] In step 5, kinematic parameters are calculated on the three-dimensional trajectory to obtain a set of dynamic parameters. That is, although the three-dimensional trajectory reconstruction in the previous steps has successfully obtained the geometric coordinate sequence of the material changing over time in three-dimensional space, these discrete coordinate points only describe the kinematic phenomena of where the material is and how it moves, and do not yet touch upon the core physical mechanism of the sorting process. To further explore the mechanical properties of the material when it interacts with the working surface of equipment such as air separators or tension screens, this application extracts deeper dynamic information from the geometric data. The energy loss, rebound angle, and momentum change of the material at the moment of collision are key factors determining whether it can be successfully sorted. For example, there is a significant difference in the velocity vectors of highly elastic rubber and low-elasticity stones after collision, which is the physical basis of tension screen sorting. Therefore, the kinematic parameter calculation step aims to transform macroscopic trajectory data into microscopic dynamic indicators, accurately calculate the incident and reflection velocity vectors at the moment of collision using mathematical methods, and then derive dimensionless parameters such as the coefficient of restitution. This provides data support with clear physical meaning for quantitatively evaluating sorting efficiency, optimizing equipment operating parameters, and constructing high-precision digital twin models.

[0061] In one embodiment of step 5, kinematic parameter calculation is performed on the three-dimensional motion trajectory to obtain a set of dynamic parameters, including: step 51, based on collision event information, performing time-domain slicing and data extraction of the trajectory before and after the collision to obtain an incident trajectory point set and a reflection trajectory point set; step 52, performing instantaneous velocity vector calculation based on the least squares method on the incident trajectory point set and the reflection trajectory point set to obtain an incident velocity vector and a reflection velocity vector; step 53, performing derived dynamic parameter calculation based on the incident velocity vector and the reflection velocity vector to obtain a scalar restitution coefficient, a normal restitution coefficient, and an energy loss rate; step 54, aggregating the scalar restitution coefficient, the normal restitution coefficient, the energy loss rate, the incident velocity vector, and the reflection velocity vector to obtain a set of dynamic parameters.

[0062] In the above implementation, step 5 is specifically processed as follows: Step 51: This step is the data preparation stage for dynamic calculation. Its core is to accurately extract local data segments related only to the collision event from the long trajectory over the entire time period. In practice, the time-domain window is first defined. In order to capture the instantaneous velocity at the moment of collision, a time period immediately adjacent to the collision point needs to be selected. If the window is too short, too few data points will make the calculation susceptible to noise; if the window is too long, the trajectory curvature caused by the influence of gravity and airflow on the material will introduce nonlinear errors, violating the linear assumption in the instantaneous velocity calculation. Therefore, a suitable time window length is set based on experience. In this embodiment, for a high-speed operating condition with a sampling frame rate of 240fps, the following settings are configured: The sequence consists of 8 frames, approximately 33 milliseconds. This length is sufficient to cover the short linear motion phase before and after the collision, while masking the complex force-induced motion at a distance. Trajectory segmentation is then performed. The system indexes keyframes. Using the origin of the time axis, bidirectional slicing is performed on the full trajectory data. For the incident process, the time index is extracted at... The trajectory points within the range. Note that keyframes are not included here. Because the material is in an unsteady state under stress and deformation at the moment of collision, its coordinates may have significant calculation errors. If the keyframe is 56, then the 3D coordinate data from frames 48 to 55 are extracted to form the incident trajectory point set. For the reflection process, the time index is extracted at... The trajectory points within the range, specifically the data from frames 57 to 64, constitute the reflection trajectory point set. These point sets are organized in memory as a list, with each element containing... The four-tuple is then used for validity verification. Due to potential transient jitter or false detections in machine vision detection, the extracted point set may contain outliers deviating from the normal trajectory. By calculating the rate of change of Euclidean distance between adjacent points, if the displacement of a point abruptly exceeds a preset threshold (e.g., the displacement of adjacent frames exceeds three times the average value), it is identified as a noise point and removed. This ensures that the retained point set is spatially continuous and smooth. The two cleaned point sets are then encapsulated separately as pure input sources for subsequent velocity vector calculations.

[0063] Step 52: Calculate the instantaneous velocity at the collision interface from discrete position points using mathematical fitting. Since directly subtracting coordinates from adjacent frames (finite difference method) to calculate velocity would greatly amplify visual measurement noise, leading to drastic fluctuations in the results, this scheme employs a linear regression analysis method based on the least squares principle, utilizing the common constraints of multiple frames of data to solve for the optimal velocity estimate. First, the incident trajectory point set is processed. This point set contains N data points; in this example, N=8, denoted as... Materials within an extremely short time window Within this system, its motion can be approximated as uniform linear motion, meaning that position and time have a linear relationship. This is discussed for each of the three coordinate components. Regarding time frame count Perform independent linear regression analysis. Taking the Z-axis component as an example, establish the linear equation. ,in, The rate of change (slope) of displacement in the Z-axis direction. The intercept is used to solve for the optimal slope. Applying the principle of least squares, we find a straight line that minimizes the sum of the squared distances from all data points to that line. The calculation is performed using the following formula:

[0064]

[0065] In this formula: This is the camera's sampling frame rate, such as 240 frames / second. This is a key factor in converting image domain parameters into physical domain parameters. It is the relative frame number of the i-th trajectory point. It is the arithmetic mean of these N frame numbers. The variance (unnormalized) of the time variable is calculated and used as the denominator. It is the three-dimensional position vector of the i-th trajectory point. , It is the average value of the position vector. The calculation uses the unnormalized covariance of time and location as the numerator. The fractional part on the right-hand side of the formula is essentially the calculated slope of the best-fit line. The physical meaning of this is meters per frame. Specifically, for material ID 1024, during the incident phase (frames 48-55), its Z-axis coordinate decreases uniformly from 0.5m to 0.1m. The calculated time average... The mean value is 51.5. The Z-component is 0.3m. Substituting this into the least squares formula, the slope along the Z-axis is calculated. The value is -0.05 meters per frame. At this point, a physical unit transformation is performed. The calculated slope vector is then... Multiply by the original video frame rate .like =240, then the physical velocity component in the Z-axis direction =-0.05×240=-12m / s. Similarly, calculate the components along the X and Y axes, as shown below. =2.4m / s, =0.5 m / s. The final composite vector yields the incident velocity vector in standard physical units. =(2.4,0.5,-12)m / s. Then, the exact same calculation process is repeated for the set of reflection trajectory points. Because the material bounces after the collision, its Z-axis coordinate will increase from small to large, therefore the calculated... The value will be positive. For example, the slope calculated during the reflection phase is... =0.03 m / frame, then the Z component of the reflection velocity is 0.03 × 240 = 7.2 m / s. The final reflection velocity vector is obtained. =(2.2,0.6,7.2)m / s.

[0066] Step 53: The input data for this step comes from two key vectors: the incident velocity vector and the incident velocity vector. and reflection velocity vector These two vectors have been uniformly converted to standard physical units (m / s) to accurately describe the velocity state of the material when it contacts the screen surface moment before the collision and when it leaves the screen surface moment after the collision. First, a scalar velocity ratio calculation is performed. This calculation aims to assess the material's ability to maintain overall momentum during the collision, without considering directional deflection. First, the magnitudes (i.e., velocities) of the two velocity vectors are calculated using the Euclidean norm formula. In the embodiment of step 52, the calculated incident velocity vectors... =(2.4,0.5,-12)m / s, reflection velocity vector =(2.2,0.6,7.2)m / s. Then the incident velocity is... ≈12.25m / s, reflection rate ≈7.55 m / s. Next, using the formula... Calculate the overall scalar recovery coefficient In this formula, Represents the scalar restitution coefficient, a dimensionless value between 0 and 1 that measures the overall rate of velocity retention of materials before and after a collision. and Let L and L be the magnitudes (L2 norms) of the three-dimensional instantaneous velocity vectors before and after the collision, respectively. Substituting these values, we can calculate... =7.55 / 12.25≈0.616. This value indicates that the material retained approximately 61.6% of its translational velocity after colliding with the equipment, indirectly reflecting that the material may possess moderate inelastic characteristics. Subsequently, normal projection decomposition and normal recovery characteristic calculations were performed. In the sorting principle of air classifiers or bouncing screens, the material's rebound ability in the direction perpendicular to the screen surface is often more critical than its horizontal movement, as this directly determines whether the material will be thrown up and thus cross obstacles or be carried away by the airflow. To extract this component, the unit normal vector of the equipment's working plane needs to be obtained. This parameter is stored in the camera parameter set or device calibration configuration file established in step 1. For example, if the calibration data shows the screen surface is horizontally placed and the normal vector points to the collision side (i.e., vertically upwards), then... =(0,0,1). Using the formula Calculate the normal coefficient of restitution In this formula, The coefficient of restitution represents the normal coefficient of restitution, which specifically measures the elastic rebound ability of a material in the direction perpendicular to the surface of the equipment. This represents the vector dot product operation, used to extract the projection component of the velocity vector in the direction of the normal vector. This indicates taking the absolute value, ensuring the coefficient is positive. Based on the previous example, the component of the incident velocity in the normal direction is... =-12m / s, the negative sign indicates downward motion, and the component of the reflection velocity in the normal direction is... =7.2 m / s (the positive sign indicates an upward rebound). Substituting into the formula, we get... =|7.2| / |-12|=0.6. This value of 0.6 is less than the scalar restitution coefficient (0.616), revealing that the energy dissipation of the material is more severe in the direction of vertical impact. This is helpful in distinguishing hard aggregates ( Higher) and soft plastics ( (Lower energy levels) are crucial. Finally, kinetic energy fingerprint calculation and loss rate quantification are performed. To more intuitively reflect the material's hardness and degree of plastic deformation, the energy loss rate needs to be calculated. If the mass m of the material remains constant, the kinetic energy is directly proportional to the square of the velocity. Using the formula... Perform the calculation. In this formula, This represents the percentage of kinetic energy lost during the collision, ranging from 0 to 1. Substitute the value... =1-(0.616)^2≈1-0.379=0.621. This means that in the extremely short instant of the collision, approximately 62.1% of the mechanical energy is converted into heat, sound, or plastic deformation energy within the material. Such a high energy loss rate typically corresponds to high-damping materials such as waste textiles or soft plastics, while for glass or stone, this value is usually below 0.3. This parameter constitutes an important characteristic for material identification.

[0067] Step 54: First, create a standardized data template instance in memory. This template contains several key fields: Event ID, directly inherited from the material ID in the collision event information, such as 1024; Timestamp, recording the precise absolute time of the collision; Velocity Vector, a nested structure containing the calculated incident vector (2.4, 0.5, -12) and reflection vector (2.2, 0.6, 7.2); Comprehensive Regeneration Coefficient Set, containing the calculated scalar regeneration coefficient 0.616, normal regeneration coefficient 0.6, and energy loss rate 0.621; and a Location field, recording the three-dimensional coordinates of the collision point. After filling in the fields, the processing unit serializes the data object into a standard text format, using JSON or CSV. For example, the generated JSON data packet might look like this: {"id":1024,"Collision Keyframe Index":56,"Incident Velocity Vector":[2.4,0.5,-12.0],"Reflection Velocity Vector":[2.2,0.6,7.2],"Kinetic Indicators":{"Quantity Recovery Coefficient":0.616,"Normal Recovery Coefficient":0.600,"Energy Loss Rate":0.621}}. This final generated structured file is the kinetic parameter set. After this parameter set is generated, it will be pushed in real time to two different data links. One link leads to the device's lower-level PLC (Programmable Logic Controller), which reads the data of multiple consecutive materials. A higher average value (indicating softer feed material) will automatically increase the fan frequency of the air separator or adjust the vibration force of the tension screen according to preset logic to prevent screen clogging and improve separation purity. Another link connects to a cloud database for offline big data statistical analysis. By accumulating tens of thousands of such collision data points, engineers can analyze the distribution of physical characteristics of different batches of solid waste, thereby optimizing the screen inclination design or guide vane material selection for next-generation sorting equipment.

[0068] In summary, a material motion behavior monitoring method based on monocular vision, as described in this application, is presented, solving the problem of difficult monitoring of material motion behavior inside solid waste sorting equipment. This scheme first uses a monocular high-speed camera to acquire material motion video, and employs instance segmentation and temporal tracking techniques to extract the continuous state sequence of the material on a two-dimensional image plane. The core lies in introducing a temporal deep learning model to analyze the state sequence, automatically identifying key events of collision between the material and the equipment's working surface, thereby accurately obtaining the collision time and contact point coordinates. Using the collision contact point as a geometric anchor point for three-dimensional spatial inversion, combined with camera intrinsic parameters and the physical parameters of the equipment's reference surface, the two-dimensional image trajectory is back-projected and reconstructed into a three-dimensional spatial motion trajectory, thereby calculating key dynamic parameters such as velocity and coefficient of restitution. This effectively overcomes the shortcomings of traditional binocular vision solutions, such as high hardware costs and difficulties in on-site calibration and maintenance. It also solves the technical bottleneck of existing monocular technology, which lacks depth information and cannot automatically reconstruct three-dimensional trajectories, making it difficult to cope with the random collision behavior of irregular materials, thus achieving low-cost, high-precision automated online monitoring.

[0069] Figure 5 This is a block diagram of a material movement behavior monitoring system based on monocular vision according to an embodiment of this application. Figure 5 As shown, the material motion behavior monitoring system 100 based on monocular vision according to an embodiment of this application includes: a material motion video acquisition module 110, used to acquire material motion videos collected by a monocular high-speed camera; a material state analysis module 120, used to perform material instance segmentation and time-series tracking on the material motion videos to obtain a material state sequence set; a collision event recognition module 130, used to perform collision event recognition based on time-series deep learning on the material state sequence set to obtain collision event information; a three-dimensional motion trajectory reconstruction module 140, used to reconstruct a three-dimensional motion trajectory based on a camera parameter set, using the collision event information and the material state sequence set to obtain a three-dimensional motion trajectory; and a kinematic parameter calculation module 150, used to calculate kinematic parameters on the three-dimensional motion trajectory to obtain a dynamic parameter set.

[0070] Here, those skilled in the art will understand that the specific operations of each step in the above-described monocular vision-based material motion behavior monitoring system have been referenced above. Figures 1 to 4The method for monitoring material movement behavior based on monocular vision has been described in detail, and therefore, its repeated description will be omitted.

Claims

1. A method for monitoring material movement behavior based on monocular vision, characterized in that, include: Acquire video footage of material movement captured by a monocular high-speed camera; Material motion videos are segmented into material instances and time-series tracked to obtain a set of material state sequences; Collision event identification based on temporal deep learning is performed on the material state sequence set to obtain collision event information, where the collision event is the collision between the material and the working surface of the equipment. Based on the camera parameter set, a three-dimensional motion trajectory is reconstructed from the collision event information and the material state sequence set to obtain the three-dimensional motion trajectory. This includes: based on the equipment working surface equation and the camera parameter set, performing three-dimensional back-calculation of the collision anchor points based on geometric constraints to obtain the coordinates of the three-dimensional physical contact points; based on the coordinates of the three-dimensional physical contact points and prior assumptions about physical motion, performing anchoring fixation, normal vector estimation, and plane equation determination to obtain the fitted motion plane parameters; and based on the fitted motion plane parameters and the camera parameter set, performing back-projection reconstruction and smoothing of the entire trajectory of the material state sequence set to obtain the three-dimensional motion trajectory. The kinematic parameters of the three-dimensional motion trajectory are calculated to obtain the set of dynamic parameters.

2. The material movement behavior monitoring method based on monocular vision according to claim 1, characterized in that, Material motion videos are segmented into material instances and time-series tracked to obtain a set of material state sequences, including: Frame stream decoding is performed on the material motion video to obtain an image frame sequence; Deep learning-based instance segmentation is performed on each image frame in the image frame sequence to obtain a single-frame detection set, wherein the single-frame detection includes corrected bounding boxes, confidence scores, binary pixel masks, and two-dimensional centroid coordinates; The single-frame detection set and historical trajectory set are correlated temporally and assigned IDs based on the ByteTrack algorithm to obtain the updated trajectory set; The updated trajectory set is aggregated and serialized to encapsulate the material state data to obtain a material state sequence set.

3. The material movement behavior monitoring method based on monocular vision according to claim 2, characterized in that, Deep learning-based instance segmentation is performed on each image frame in the image frame sequence to obtain a single-frame detection set, including: Image frames are resized and their color space converted and normalized to obtain normalized image tensors; The normalized image tensor is input into the pre-trained YOLO-seg model to obtain target instances, which include filtered bounding boxes, confidence scores and their corresponding binary pixel masks. The centroid of the target instance is calculated to obtain the two-dimensional centroid coordinates.

4. The material movement behavior monitoring method based on monocular vision according to claim 3, characterized in that, The centroid of the target instance is calculated to obtain its two-dimensional centroid coordinates, including: Extract feature images and soft segmentation masks from target instances; A fusion weight matrix is ​​constructed based on feature images and soft segmentation masks; Based on the fusion weight matrix, an adaptive higher-order moment solution is performed on the pixel coordinate network to obtain robust higher-order moments; The robust physical centroid coordinates are derived and corrected based on robust higher-order moments to obtain two-dimensional centroid coordinates.

5. The material movement behavior monitoring method based on monocular vision according to claim 4, characterized in that, Based on the feature image and soft segmentation mask, a fusion weight matrix is ​​constructed, including: constructing the fusion weight matrix using the following formula, where the formula is: in, At the pixel The fusion weight value at the location, It is a soft segmentation mask at the point The probability value at that location. It is the confidence sharpening index. It is the feature image at the point Pixel value at that location, It is the vision-quality coupling coefficient. and These are the pixel mean and standard deviation of the feature image within the material region. It is a very small constant to prevent the denominator from being zero.

6. The material movement behavior monitoring method based on monocular vision according to claim 1, characterized in that, Collision event identification based on temporal deep learning is performed on the material state sequence set to obtain collision event information, including: Trajectory segments are extracted from the material state sequence set and serialized and stacked to obtain a normalized temporal tensor set; Collision keyframe inference based on CNN-LSTM is performed on the normalized temporal tensor set to obtain the keyframe index and collision confidence. Based on the keyframe index and device reference surface parameters, the contact points of the material state sequence set are precisely located under geometric constraints to obtain collision event information. The collision event information includes the material ID, keyframe index, and two-dimensional coordinates of the collision contact points.

7. The material movement behavior monitoring method based on monocular vision according to claim 1, characterized in that, The camera parameter set includes the camera intrinsic parameter matrix, distortion coefficient vector, rotation matrix, and translation vector.

8. The material movement behavior monitoring method based on monocular vision according to claim 1, characterized in that, The kinematic parameters of the three-dimensional motion trajectory are solved to obtain the set of dynamic parameters, including: Based on collision event information, the trajectory time domain slices and data extraction of the three-dimensional motion trajectory before and after the collision are performed to obtain the incident trajectory point set and the reflection trajectory point set; The incident velocity vector and the reflection velocity vector are obtained by solving the instantaneous velocity vector of the incident trajectory point set and the reflection trajectory point set based on the least squares method; Derived dynamic parameters are calculated based on the incident velocity vector and the reflection velocity vector to obtain the scalar recurve coefficient, the normal recurve coefficient, and the energy loss rate; The scalar restitution coefficient, normal restitution coefficient, energy loss rate, incident velocity vector, and reflection velocity vector are aggregated to obtain a set of dynamic parameters.

9. A material movement behavior monitoring system based on monocular vision, characterized in that, include: The material motion video acquisition module is used to acquire material motion videos captured by a monocular high-speed camera. The material status analysis module is used to perform material instance segmentation and time-series tracking on material motion videos to obtain a set of material status sequences. The collision event recognition module is used to perform collision event recognition on the material state sequence set based on time-series deep learning to obtain collision event information. The collision event is the collision between the material and the working surface of the equipment. The 3D motion trajectory reconstruction module is used to reconstruct the 3D motion trajectory based on the collision event information and the material state sequence set using a set of camera parameters. This includes: performing 3D back-calculation of collision anchor points based on geometric constraints using the equipment working surface equation and the set of camera parameters to obtain the coordinates of the 3D physical contact points; performing anchoring, normal vector estimation, and plane equation determination based on the 3D physical contact point coordinates and prior physical motion assumptions to obtain the fitted motion plane parameters; and performing back-projection reconstruction and smoothing of the entire trajectory of the material state sequence set based on the fitted motion plane parameters and the set of camera parameters to obtain the 3D motion trajectory. The kinematic parameter calculation module is used to calculate the kinematic parameters of a three-dimensional motion trajectory to obtain a set of dynamic parameters.

Citation Information

Patent Citations

  • Method, system and equipment for controlling ball mill manufactured by resistor industrial internet of things

    CN118142647A

  • Target material positioning and tracking method based on AI technology

    CN120931693A