Logistics sorting system monitoring method based on visual large model
By using a visual large model-based monitoring method for logistics sorting systems, image data is collected and processed in real time to identify the status of the sorting system and generate early warning signals. This solves the efficiency and accuracy problems of logistics sorting systems in dynamic environments and achieves efficient and accurate sorting status monitoring and logistics item positioning.
Patent Information
- Application Number
- CN202510937150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
The existing logistics sorting system is easily affected by factors such as occlusion and barcode damage, resulting in low efficiency or recognition errors. It also lacks real-time monitoring of the dynamically changing warehousing environment and precise positioning mechanism for logistics parts, making it difficult to detect and handle abnormal situations in a timely manner.
A method based on a large visual model is adopted to collect image data in real time through a visual detection device, adaptively adjust the image size and preprocess it, use a visual transformation model for feature extraction and time series positioning, combine classification networks and time series analysis technology to generate high-dimensional feature vectors, identify the sorting system status and push early warning signals, and the management terminal generates accurate sorting strategies to assist the conveyor line in completing efficient sorting.
It achieves low-latency, high-precision sorting status monitoring and logistics part positioning, significantly improves sorting efficiency, adapts to dynamic environmental changes, reduces system downtime and misjudgment rate, and improves the intelligence level of the logistics sorting system.
Smart Images

Figure CN120808270A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of logistics automation and computer vision, specifically relates to a logistics sorting system monitoring method based on a visual large model. BACKGROUND
[0002] With the rapid development of the logistics industry, the efficiency and accuracy of the logistics sorting system are crucial to warehouse operations. Traditional sorting systems rely on barcode scanning or manual operation, which is easily affected by factors such as obstruction and barcode damage, resulting in low efficiency or identification errors. Although existing sorting technologies based on computer vision have improved, they rely on pre-trained models for specific scenarios and lack generalization ability, making it difficult to adapt to dynamic warehouse environments. The existing technology lacks real-time monitoring of the running state of the sorting system and precise positioning mechanism of the logistics pieces, and it is difficult to discover and handle abnormal situations (such as goods accumulation and equipment failure) in a timely manner. The present application proposes a logistics sorting system monitoring method based on a visual large model, which realizes low-latency and high-precision sorting state monitoring and logistics piece positioning through efficient image acquisition, feature extraction, state analysis, and timing positioning technology. SUMMARY
[0003] To solve the above problems existing in the prior art, the present application provides a logistics sorting system monitoring method based on a visual large model, The purpose of the present application can be achieved by the following technical solutions: S1: Real-time image data of the logistics sorting system is collected by a visual detection device, and the image size is adaptively adjusted according to the complexity of the logistics scene. The image is preprocessed to generate an image sequence in a unified coordinate system; S2: The image sequence is input into a visual conversion model. The visual conversion model divides the frame image into fixed-size non-overlapping image blocks, arranges them in time sequence to form an input sequence, and introduces a closed-loop feedback mechanism. By comparing the real-time sorting result with the historical data, the parameters of the model are dynamically adjusted to realize online self-optimization of the model. The input sequence aggregates the output features of all image blocks through global averaging to generate a high-dimensional feature vector; S3: The high-dimensional feature vector is analyzed by a classification network to analyze the running state of the sorting system, and to identify normal sorting, abnormal accumulation, equipment failure and other states. The sorting system combines timing analysis technology to predict the state of the logistics piece, generates an early warning signal and pushes the logistics piece state to the system management terminal; S4: The management terminal generates a precise sorting strategy according to the logistics piece position information, state analysis result and priority scheduling instruction, assists the conveying line to complete efficient sorting, and reduces system downtime and misjudgment rate according to the closed-loop feedback mechanism.
[0004] Specifically, the visual inspection device includes a focusable camera for real-time image acquisition, identification of the size of the logistics object, and scanning of barcode information; the barcode is always positioned above the logistics object by controlling the shooting angle.
[0005] Specifically, the shooting angle is achieved by two small servos that respectively control the horizontal and vertical rotation of the focus-adjustable camera to adjust the shooting angle and lock the logistics part.
[0006] Specifically, the preprocessing includes Gaussian filtering denoising, adaptive histogram equalization, and multi-perspective image registration based on SIFT feature points. The Gaussian filtering denoising smoothes the image by applying a Gaussian filter to each frame of the image, thereby effectively removing the vibration sound of the conveyor belt; the adaptive histogram equalization enhances the image contrast by dividing the image into small areas and setting a contrast limit threshold, thereby highlighting the details of the logistics parts; the multi-perspective image registration based on SIFT feature points matches and aligns the key points of images from different perspectives into single-perspective equivalent images, thereby achieving spatial consistency of multi-perspective images and eliminating perspective differences.
[0007] Specifically, the visual model is a visual transformation model, which divides the image into small blocks and processes them as sequences, aggregates the output features of the image blocks, and generates high-dimensional feature vectors, thereby capturing the shape, texture, location, and equipment operating status data of the goods in logistics sorting.
[0008] Specifically, the classification network adopts a fully connected neural network, takes the high-dimensional feature vector extracted by the large visual model as input, analyzes the operating status of the sorting system through the Softmax function, identifies normal sorting, cargo accumulation, equipment failure and other states, and realizes the status monitoring of logistics parts. The calculation formula of the Softmax function is: , where p i is the output probability of the i-th class, z i is the logit value of the i-th category, exp(z i ) is the pair z i Apply the exponential function, ∑ C J=1 exp(z i ) is the sum of the exponentials of all categories, exp(z max ) is the maximum value of the applied exponential function.
[0009] Specifically, the time series analysis technology generates a continuous state sequence of logistics parts through long short-term memory network decoding, thereby realizing dynamic modeling of the operating state of the sorting system.
[0010] Specifically, the visual detection device further comprises an adaptive light compensator, which dynamically adjusts the light compensation brightness according to the real-time collected warehouse environment light intensity, enhances the image acquisition quality through infrared or visible light compensation, and improves the feature extraction accuracy of the visual large model in low light or high contrast scenes.
[0011] Specifically, the logistics piece information includes spatial coordinates and state types of the logistics piece, and the information is pushed to the management terminal through a wireless communication mode, realizing real-time monitoring of the positioning of the logistics piece and the running state of the sorting system.
[0012] Specifically, the management terminal adjusts the sorting grabbing path or the speed of the conveying belt according to the position information, generates an optimized sorting strategy, and realizes the improvement of the classification accuracy of the logistics piece.
[0013] The beneficial effects of the present application are: Through multi-view image registration and time sequence analysis, the sorting state is monitored in real time, the abnormal delay is low, the positioning accuracy of the logistics piece is accurate, the sorting efficiency is significantly improved, the adaptive light compensation and environmental noise processing module enhances the image quality, adapts to dynamic light and noise scenes, the online learning mechanism fine-tunes the model, ensures the robustness of the system to the diversity of goods and environmental changes, optimizes the delay of distributed processing and transmission, the overall response time is short, the state report is pushed fast, and the invention supports fast strategy adjustment. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to facilitate the understanding of those skilled in the art, the present application will be further described below with reference to the accompanying drawings.
[0015] Figure 1 A flowchart of a logistics sorting system monitoring method based on a visual large model. DETAILED DESCRIPTION
[0016] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purposes, the specific embodiments, structures, features and effects thereof according to the present application are described in detail below with reference to the accompanying drawings and preferred embodiments.
[0017] Please refer to Figure 1 A logistics sorting system monitoring method based on a visual large model, comprising the following steps: S1: real-time acquisition of image data of the logistics sorting system by a visual detection device, and adaptive adjustment of image size according to the complexity of the logistics scene, pre-processing of the image to generate an image sequence in a unified coordinate system; S2: input the image sequence into a visual conversion model, the visual conversion model divides the frame image into fixed-size non-overlapping image blocks, arranges them in time sequence to form an input sequence, and introduces a closed-loop feedback mechanism to dynamically adjust the parameters of the model by comparing the real-time sorting results with the historical data, realize online self-optimization of the model, the input sequence aggregates the output features of all image blocks through global averaging to generate a high-dimensional feature vector; S3: the high-dimensional feature vector analyzes the running state of the sorting system through a classification network, identifies multiple states such as normal sorting, abnormal accumulation, equipment failure, etc.; the sorting system combines time series analysis technology to predict the state of the logistics piece, generates an early warning signal and pushes the logistics piece state to the system management terminal; S4: the management terminal generates accurate sorting strategies according to the logistics piece position information, state analysis results and priority scheduling instructions, assists the conveying line to complete efficient sorting, and reduces system downtime and misjudgment rate according to the closed-loop feedback mechanism.
[0018] Specifically, the visual detection device includes an adjustable focus camera for real-time image acquisition, identification of logistics piece size and scanning of barcode information; by controlling the shooting angle, the barcode is always located above the logistics piece.
[0019] This embodiment is deployed in a cross-border e-commerce warehouse center, which processes 50,000 multilingual label packages per day. It needs to cope with high-speed assembly line, irregular package, multi-device collaboration and multi-language label recognition requirements. The system needs to realize real-time state monitoring, accurate package positioning, multi-language label analysis and exception handling, while optimizing cross-device collaboration efficiency and adapting to new language labels. Four high-resolution cameras and multi-modal sensors are deployed on the high-speed sorting assembly line to cover the full viewing angle and collect package label, conveyor belt and robot data. It is equipped with an adaptive light compensation module and an environmental noise adaptive module to cope with high-speed motion blur and label reflection. Image preprocessing includes: applying Gaussian filter to remove motion blur and spot noise, enhancing label text contrast, detecting about 1000 key points based on scale invariant feature transform, aligning multi-view images through single matrix to generate uniform coordinate system image sequence, generating 10-dimensional feature vector through adaptive filtering of infrared ranging and vibration signals, storing image sequence and sensor data to data warehouse, triggering analysis task, total time consumption is 18ms, supporting high-speed assembly line requirements, using lightweight visual conversion model, migrating knowledge from cloud ViT-Large through dynamic knowledge distillation, reducing parameter quantity by 50%, adapting to edge devices, generating 256-dimensional embedding vector, adding learnable position encoding, fusing sensor data, adjusting through dynamic cross-modal attention mechanism according to label reflection degree, integrating OCR module to analyze multi-language labels, generating text features and embedding feature vectors. Global average pooling generates a 256-dimensional feature vector, which takes 20ms, adapts to multi-language label scenarios, and the feature vector is input into a one-dimensional convolutional neural network followed by a fully connected layer, which outputs state probability through Softmax, identifies "normal sorting", "label misreading", "goods accumulation" and "device failure", and has an accuracy of 99.6%. Time series analysis uses long short-term memory network to process 5 frames of probability sequence to generate state sequence and locate package spatial coordinates, which takes 4ms, and continuous 5 frames of abnormality triggers alarm to generate report and push to management terminal through 5G. Label text and coordinates are combined to verify sorting accuracy, PPO reinforcement learning is used to optimize sorting strategy, and based on state sequence, positioning data, label information and energy consumption data, multiple types of robotic arms and sorting robots are coordinated, which takes 4ms, efficiency is improved by 23%, and energy consumption is reduced by 15%. Federated learning updates the model based on 3000 frames of data and label text every day, aggregates parameters with encryption to protect privacy, and generalizes accuracy to 99.8%. The optimization strategy is published to the control module through 5G to realize cross-device collaboration. Generative adversarial network is used to generate simulated multi-language labels and irregular package images (1500 frames per day) to support zero-shot learning, which takes 8ms, enhances new language label and irregular package recognition, and the accuracy is 99.8% after fine-tuning. The overall delay is 24ms (preprocessing 18ms, feature extraction 20ms, state analysis 4ms, optimization 4ms), the classification accuracy is 99.6%, the accuracy is 99.8% after fine-tuning, the positioning accuracy is ±1.5cm, the abnormality detection sensitivity is 99.5%, and the false positive rate is 0.5%.Adaptive light compensation and OCR module deal with label reflection and multi-lingual scenarios, GAN enhances generalization, sorting efficiency improves by 23%, energy consumption reduces by 15%. Federal learning supports cross-warehouse collaboration, adapts to the high-speed demand of cross-border e-commerce.
[0020] Specifically, the shooting angle is adjusted and locked by two small steering gears by controlling the horizontal and vertical rotation of the adjustable focus camera respectively.
[0021] Specifically, the preprocessing includes Gaussian filter denoising, adaptive histogram equalization, and multi-view image registration based on SIFT feature points. The Gaussian filter denoising smoothes the image by applying a Gaussian filter to each frame of image, effectively removing the vibration sound of the conveyor belt. The adaptive histogram equalization enhances the contrast of the image by dividing it into small regions and setting a contrast limit threshold, highlighting the details of the logistics piece. The multi-view image registration based on SIFT feature points aligns the key points of different view images as single view equivalent images, achieving spatial consistency of multi-view images and eliminating view differences.
[0022] Specifically, the visual model is a visual conversion model, which processes the image by dividing it into small blocks and treating it as a sequence, and aggregates the output features of all image blocks to generate a high-dimensional feature vector, capturing the shape, texture, position, and equipment running state data of the goods in the logistics sorting.
[0023] Specifically, the classification network uses a fully connected neural network, taking the high-dimensional feature vector extracted by the visual large model as input, and analyzes the sorting system running state through the Softmax function, identifying normal sorting, goods accumulation, equipment failure, etc. to achieve state monitoring of the logistics piece. The Softmax function calculation formula is: , where p i is the output probability of the i-th class, z i is the logit value of the i-th class, exp(z i ) is the exponential function of z i , ∑ C J=1 exp(z i ) is the sum of the exponentials of all classes, and exp(z max ) is the maximum value of the exponential function.
[0024] Specifically, the time series analysis technique decodes through a long short-term memory network to generate a continuous state sequence of the logistics piece, achieving dynamic modeling of the sorting system running state.
[0025] Specifically, the visual detection device further comprises an adaptive light compensator, which dynamically adjusts the light compensation brightness according to the real-time collected warehouse environment light intensity, enhances the image acquisition quality through infrared or visible light compensation, and improves the feature extraction accuracy of the visual large model in low light or high contrast scenes.
[0026] Specifically, the logistics piece information includes the spatial coordinates and state types of the logistics piece, and the information is pushed to the management terminal through a wireless communication mode, realizing real-time monitoring of the positioning of the logistics piece and the running state of the sorting system.
[0027] Specifically, the management terminal adjusts the sorting grabbing path or the speed of the conveying belt according to the position information, generates an optimized sorting strategy, and realizes the improvement of the classification accuracy of the logistics piece.
[0028] This embodiment is deployed in a small and medium-sized logistics sorting station, which processes 10,000 packages per day and needs to cope with dynamic environment challenges, including variable light, cargo diversity and equipment heterogeneity (such as different types of mechanical arms). The system needs to realize real-time state monitoring, accurate positioning of packages and abnormal processing, while optimizing energy consumption and supporting rapid adaptation to new scenarios. The system is deployed with 3 high-resolution cameras and infrared and vibration sensors, equipped with an adaptive light compensation module, collects image and environment data, and responds to day and night light changes. Image preprocessing includes Gaussian filter to remove noise, CLAHE to enhance contrast, SIFT to register to generate uniform coordinate system image sequences, sensor data is filtered by NLMS to generate 8-dimensional features, and sequences are stored in a data warehouse to trigger analysis, with a total time consumption of 17ms. The visual conversion model extracts features through dynamic knowledge distillation, divides the image into 16x16 pixel blocks, fuses sensor data, adjusts according to light, and generates a 192-dimensional feature vector through global average pooling, with a time consumption of 20ms. Adapt to edge devices. The features are input into a 1D-CNN, followed by a fully connected layer, and Softmax outputs states such as "normal sorting" and "goods accumulation", with an accuracy of 99.5%. LSTM processes 5-frame probability sequences to generate state sequences and locate packages, and continuously triggers alarms for 5 frames of abnormalities, pushing reports to the management terminal. PPO reinforcement learning optimizes strategies, with a time consumption of 4ms, an efficiency improvement of 20%, and an energy consumption reduction of 10%. Federated learning updates the model based on 2000 frames of data daily, aggregates parameters, protects privacy, and achieves an accuracy of 99.7%. GAN generates simulated irregular package images to support zero-shot learning. The overall delay is 23ms, the positioning accuracy is ±1.5cm, the abnormality detection sensitivity is 99.4%, the false positive rate is 0.7%, the sorting efficiency is improved by 20%, and the system adapts to dynamic environments.
[0029] The embodiment is deployed in a cold chain logistics and storage center, handles 5000 refrigerated packages per day, needs to monitor the goods accumulation, equipment failure and package positioning in a low temperature environment, and should deal with the challenges of low temperature mist, insufficient light and special-shaped packaging. By deploying 4 high-resolution cameras and infrared, vibration and temperature sensors, and equipping with adaptive infrared fill light and environmental noise module, image and environmental data are collected. The image is filtered by Gaussian filter, CLAHE, SIFT registration to generate a unified coordinate system image sequence. The sensor data is filtered by NLMS filter to generate a 10-dimensional feature vector, the total time consumption is 20 ms, the output sequence is stored in the data warehouse, triggering the analysis task, using the visual converter model, migrating knowledge from the cloud ViT-Large through dynamic knowledge distillation, reducing the parameter quantity by 50%, the inference time consumption is 18 ms, the image is divided into 16x16 pixel blocks, the embedding vector is generated and the position encoding is added, the sensor data is fused, the dynamic cross-modal attention is used, and the light variance is adjusted. The global average pooling generates a 192-dimensional feature vector, the total time consumption is 21 ms, the edge device is adapted, the feature vector is input into the 1D-CNN, followed by the fully connected layer (256 neurons, Dropout 0.3), and the Softmax outputs the state probability of "normal sorting", "goods accumulation" and "equipment failure", etc. The accuracy is 99.6%, the LSTM processes 5 frames of probability sequence to generate state sequence and locate the package, and the alarm is triggered for 5 consecutive frames of anomaly to generate a report, which is pushed to the management terminal through 5G. The PPO reinforcement learning optimization strategy is used to adjust the conveyor belt speed and the mechanical arm path based on the state sequence, positioning and energy consumption data: the time consumption is 5 ms, the efficiency is improved by 23%, and the energy consumption is reduced by 12%. The model is updated based on 1000 frames of data per day through federated learning, the parameter aggregation protects privacy, the generalization accuracy is 99.8%, the GAN generates simulated special-shaped package images, supports zero-shot learning, enhances unknown goods recognition, and the classification accuracy reaches 99.8% after fine-tuning. The overall delay is 24 ms, the classification accuracy is 99.6%, and the positioning accuracy is ±1.5 cm after fine-tuning. The anomaly detection sensitivity is 99.5%, the false alarm rate is 0.6%, the infrared fill light and GAN deal with low temperature mist and special-shaped packaging, the sorting efficiency is improved by 23%, the energy consumption is reduced by 12%, and the cold chain scene is adapted.
[0030] The above is only a preferred embodiment of the present application, not any form of limitation on the present application. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the technical solution of the present application, and any brief modification, equivalent change and modification of the above embodiment according to the technical essence of the present application are still within the scope of the technical solution of the present application.
Claims
1. A logistics sorting system monitoring method based on a large visual model, characterized in that: include: S1: Using a visual inspection device to collect image data of the logistics sorting system in real time, and adaptively adjust the image size according to the complexity of the logistics scene, pre-process the image, and generate an image sequence in a unified coordinate system; S2: Inputting the image sequence into a visual conversion model, the visual conversion model divides the frame image into non-overlapping image blocks of fixed size, arranges them in time sequence to form an input sequence, and introduces a closed-loop feedback mechanism to dynamically adjust the model parameters by comparing the real-time sorting results with historical data to achieve online self-optimization of the model. The input sequence aggregates the output features of all image blocks through global averaging to generate a high-dimensional feature vector; S3: The high-dimensional feature vector is used to analyze the operating status of the sorting system through a classification network to identify various states such as normal sorting, abnormal accumulation, and equipment failure. The sorting system combines time series analysis technology to predict the status of logistics parts, generate early warning signals, and push the status of logistics parts to the system management terminal; S4: The management terminal generates an accurate sorting strategy based on the logistics part location information, status analysis results and priority scheduling instructions, assists the conveyor line to complete efficient sorting, and reduces system downtime and misjudgment rate based on the closed-loop feedback mechanism.
2. The method according to claim 1, characterized in that The visual inspection device includes a focusable camera for real-time image acquisition, identification of the size of the logistics object and scanning of barcode information; the barcode is always located above the logistics object by controlling the shooting angle.
3. The method according to claim 1, characterized in that The shooting angle is achieved by two small servos that respectively control the horizontal and vertical rotation of the focus-adjustable camera to adjust the shooting angle and lock the logistics parts.
4. The method according to claim 1, wherein The preprocessing includes Gaussian filtering denoising, adaptive histogram equalization, and multi-perspective image registration based on SIFT feature points. The Gaussian filtering denoising smoothes the image by applying a Gaussian filter to each frame of the image, thereby effectively removing the vibration sound of the conveyor belt; the adaptive histogram equalization enhances the image contrast by dividing the image into small areas and setting a contrast limit threshold, thereby highlighting the details of the logistics parts; the multi-perspective image registration based on SIFT feature points matches and aligns the key points of images from different perspectives into single-perspective equivalent images, thereby achieving spatial consistency of multi-perspective images and eliminating perspective differences.
5. The method according to claim 4, characterized in that The visual model is a visual transformation model that divides the image into small blocks and processes them as sequences. It aggregates the output features of all image blocks to generate high-dimensional feature vectors, thereby capturing the shape, texture, location, and equipment operating status data of goods in logistics sorting.
6. The method according to claim 1, characterized in that The classification network adopts a fully connected neural network, takes the high-dimensional feature vector extracted by the large visual model as input, analyzes the operating status of the sorting system through the Softmax function, identifies normal sorting, cargo accumulation, equipment failure and other states, and realizes the status monitoring of logistics parts.
7. The method according to claim 1, characterized in that The time series analysis technology generates a continuous state sequence of logistics parts through long short-term memory network decoding, and realizes dynamic modeling of the operating state of the sorting system.
8. The method according to claim 2, characterized in that The visual inspection device also includes an adaptive fill light, which dynamically adjusts the fill light brightness according to the real-time collected warehouse environment light intensity, enhances the image acquisition quality through infrared or visible light fill light, and improves the feature extraction accuracy of the visual large model in low-light or high-contrast scenes.
9. The method according to claim 1, characterized in that The logistics piece information includes the spatial coordinates and status type of the logistics piece. The information is pushed to the management terminal via wireless communication, realizing the positioning of the logistics piece and real-time monitoring of the operating status of the sorting system.
10. The method according to claim 4, characterized in that The management terminal adjusts the sorting and grabbing path or the conveyor belt speed according to the location information, and improves the classification accuracy of logistics parts by generating an optimized sorting strategy.
Citation Information
Patent Citations
Intelligent logistics sorting system based on cloud computing
CN113759917A
Sorting center abnormal behavior identification method based on video detection technology
CN114581824A
Transform-based logistics package separation method
CN114708295A
Express delivery tracking method, device and equipment and storage medium
CN117172652A
Intelligent pneumonia prediction system based on cough sound and pulmonary respiration sound
CN119174600A