Intelligent warehouse goods recognition and positioning system based on image recognition and interactive interface
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIBET QIYUAN TECHNOLOGY CO LTD
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-07
AI Technical Summary
多数系统采用单视角或少数几个固定视角进行图像采集,当货物存在堆叠、被货架或其他货物遮挡时,无法获取完整的货物特征信息,导致识别准确率大幅下降
本发明通过多视角环形采集阵列结合自适应权重分配技术,能够同步获取货物多个角度的图像信息和深度信息,根据不同视角图像的质量和特征丰富度动态调整融合权重,当某个视角被完全遮挡时自动切换至备用视角,结合连续多帧时序特征进行平滑处理,有效提升了复杂光照和遮挡条件下的图像采集质量。采用多尺度特征提取网络融合浅层细节特征和深层语义特征,引入空间注意力和通道注意力机制增强关键特征的表达能力,结合跨域特征对齐技术解决不同设备和环境下的图像域偏移问题,实现了货物的细粒度识别,能够准确区分同型号不同批次的产品。
Smart Images

Figure CN122530692A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent warehousing image recognition technology, and in particular to an intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface. Background Technology
[0002] With the rapid development of the warehousing and logistics industry, the types and quantities of goods in warehouses are constantly increasing, and the frequency of goods entering and leaving the warehouse is constantly rising. Traditional manual methods of goods identification and positioning can no longer meet the needs of modern warehouse management. Manual methods rely on the memory and experience of staff, resulting in low efficiency, high labor intensity, and a high risk of errors, making it impossible to achieve rapid inventory and accurate management of large-scale goods. While barcode identification and RFID technologies have improved the efficiency of goods management to some extent, barcodes are easily worn and contaminated, and RFID tags are expensive. Furthermore, both require close-range contact reading, making it impossible to achieve rapid identification of batches of goods, and it is also difficult to obtain spatial location and status information of goods.
[0003] Image recognition technology, with its advantages of non-contact, speed, and batch processing, is gradually being applied to warehouse cargo management. However, existing image recognition-based cargo identification and positioning systems still have many shortcomings. Most systems use a single viewpoint or a few fixed viewpoints for image acquisition. When goods are stacked, obscured by shelves or other goods, complete cargo feature information cannot be obtained, leading to a significant drop in recognition accuracy. In terms of positioning, existing systems mostly calculate cargo positions directly based on semantic segmentation results or single depth information, without fully considering the impact of camera calibration errors and semantic segmentation confidence on positioning accuracy. For multi-layered stacked goods, it is difficult to accurately distinguish the vertical coordinates of goods on different layers, and the positioning accuracy cannot meet the needs of automated warehousing. In addition, the feature extraction and fusion methods of existing systems are relatively simple, and there are domain offset problems between images acquired from different warehouses and different camera devices, making it impossible to achieve fine-grained identification of goods of the same model but different batches or with different production dates.
[0004] The existing cargo identification and positioning systems also suffer from outdated interaction methods, primarily relying on traditional computer-based operations. This necessitates frequent back-and-forth movement of on-site personnel between the operating terminal and the cargo storage area, resulting in low work efficiency. The systems also have limited anomaly detection capabilities, only able to identify obvious anomalies such as missing cargo, failing to promptly detect subtle anomalies like tilted, damaged, or misplaced cargo, potentially leading to safety hazards and management loopholes. Furthermore, most systems employ offline training methods for their identification and positioning models, failing to automatically update and optimize based on newly added cargo data within the warehouse. This results in poor model generalization ability, and the inability to share and collaboratively optimize model parameters between different warehouses, leading to redundant training and wasted computational resources. Summary of the Invention
[0005] The present invention proposes an intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface to solve the problems mentioned in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface, comprising the following modules: The multi-view image acquisition and preprocessing module consists of a circular acquisition array composed of a high-definition industrial camera and a depth camera. It acquires images of the cargo, and simultaneously acquires the cargo's depth and texture information. It preprocesses the acquired images, extracts the foreground region of the cargo, and generates a standardized cargo image dataset. The cargo image feature extraction and enhancement module uses an improved convolutional neural network to extract cargo features and introduces a spatial attention mechanism to enhance attention to key feature regions of the cargo. The multimodal feature fusion and cargo recognition module integrates the visual features, depth features, and pre-stored standard cargo feature library of cargo. It uses a feature-level fusion algorithm to generate a unified cargo feature vector, and a classifier classifies and recognizes the cargo feature vector to output cargo information. The cargo localization module based on image semantic segmentation uses a semantic segmentation network to perform pixel-level segmentation of cargo images, identify the boundaries and contours of the cargo, and calculate the three-dimensional spatial coordinates of the cargo by combining camera calibration parameters and the warehouse three-dimensional coordinate system to determine the cargo location. The cargo status recognition and anomaly detection module identifies the cargo status by analyzing the image features of the cargo, detects abnormal cargo conditions, generates anomaly detection reports, and issues early warning information. The interactive interface and operation guidance module provides a visual interactive interface, allowing users to operate the system through touch, voice, and gestures. Augmented reality technology is used to overlay operation guidance information onto the actual warehouse scene. The identification result management and data synchronization module stores cargo identification results, location information, and status information, and establishes a cargo information database. The system performance self-optimization module continuously trains and optimizes the system's recognition and localization models based on newly added cargo image data and user annotation results.
[0007] Furthermore, it also includes a multi-view image adaptive weight allocation module, which calculates the weight coefficients of each view image based on the sharpness, lighting conditions, and feature richness of the images from different viewpoints. The calculation formula is as follows: ; in, For the first Weighting coefficients for images from each viewpoint; For the first Quality rating of images from each perspective; For the first Feature richness score for images from different perspectives; To determine the total number of perspectives involved in the fusion, the cargo features from different perspectives are weighted and fused according to the calculated weight coefficients to generate a comprehensive cargo feature vector. At the same time, a dynamic perspective switching mechanism is added, which automatically switches to a preset backup acquisition perspective when a perspective is completely blocked by cargo or shelves, and then smoothly fused by combining the temporal features of multiple consecutive frames.
[0008] Furthermore, it also includes a dynamic correction module for cargo positioning accuracy, which dynamically corrects the cargo positioning results by combining the confidence level of semantic segmentation and camera calibration error. The calculation formula is as follows: ; in, The corrected three-dimensional spatial coordinates of the cargo; The initial calculated three-dimensional spatial coordinates of the cargo; This is the camera calibration error correction coefficient; The confidence level of the semantic segmentation result; This represents the maximum positioning error value.
[0009] Furthermore, it also includes a cargo completion and recognition module for occluded scenes. When cargo is partially occluded, this module extracts local features of the unoccluded area, combines a prior shape model with contextual information to generate a complete feature representation, and completes the image features of the occluded area through a generative adversarial network. The module establishes a hierarchical occlusion processing mechanism, dividing the occlusion into three levels: mild, moderate, and severe, based on the proportion of occluded area. It adopts three different processing strategies: local feature matching, prior model completion, and contextual inference, respectively, and uses the feature distribution of cargo on the same shelf and in the same batch, as well as historical storage location data, to assist in the completion.
[0010] Furthermore, the multi-view image acquisition and preprocessing module supports automatic adjustment of the camera's exposure time, focal length, and shooting angle. It employs multi-threaded parallel processing technology to synchronously preprocess images acquired by multiple cameras, establishes an image quality assessment system, and automatically re-acquires images that do not meet the quality standards. Simultaneously, it integrates an adaptive illumination compensation system that automatically adjusts the brightness and color temperature of the fill light according to the ambient light intensity, eliminating the impact of strong light reflection and dark shadows on image quality.
[0011] Furthermore, the cargo image feature extraction and enhancement module employs a multi-scale feature extraction network to simultaneously extract cargo features at different scales, fuse shallow detail features and deep semantic features to generate a multi-scale fused feature vector, and introduces a channel attention mechanism to adaptively adjust the weights of different feature channels. This module incorporates cross-domain feature alignment technology to solve the image domain offset problem, and converts high-dimensional feature vectors into low-dimensional hash codes through feature hash encoding to achieve fine-grained cargo feature extraction and feature matching.
[0012] Furthermore, the multimodal feature fusion and cargo recognition module adopts an attention-guided multimodal feature fusion network to learn the correlation and complementarity of different modal features and assign adaptive fusion weights. It uses an ensemble learning method combined with the classifier output to generate cargo recognition results. This module establishes a dynamic feature library update mechanism, automatically adds new features and cleans up redundant and outdated data, adds anomaly feature filtering function, supports simultaneous recognition of multiple labels, and outputs multi-dimensional cargo information.
[0013] Furthermore, the image semantic segmentation-based cargo localization module uses an instance segmentation network to perform instance-level segmentation of densely stacked cargo, distinguishing different individual cargoes, extracting the centroid coordinates and bounding box information of the cargoes, and combining the depth information obtained by the depth camera to calculate the three-dimensional spatial coordinates of the cargoes, establishing a mapping relationship between the cargo positions and warehouse shelves and storage locations, and realizing storage location localization. At the same time, it integrates stacked cargo three-dimensional reconstruction and localization technology, uses multi-view depth images to reconstruct the three-dimensional point cloud model of the cargoes, calculates the spatial position and orientation of the cargoes, detects the storage location occupancy status and generates a real-time status distribution map, and outputs cargo orientation angle information.
[0014] Furthermore, the interactive interface and operation guidance module supports access from computers, mobile devices, and head-mounted augmented reality devices, providing functions such as goods inquiry, inbound and outbound registration, inventory counting, and anomaly handling. This module uses augmented reality technology to highlight the location of target goods and operation instructions in the actual scene, supports voice command control and gesture recognition operation, realizes AR virtual picking path planning and multi-person collaborative operation, and has an offline operation mode, which automatically synchronizes to the server after the network is restored.
[0015] Furthermore, the system performance self-optimization module establishes an incremental learning mechanism, using newly added cargo image data and user-annotated data to incrementally train the recognition model and the positioning model, avoiding catastrophic forgetting of the model; this module establishes a model performance evaluation system, automatically triggering retraining when the model performance is lower than a preset threshold; it supports model version management and rollback, introduces federated learning technology, combines data from multiple warehouses for model training, realizes dynamic scheduling of hardware resources, and automatically adjusts system parameters for special working conditions.
[0016] Compared with existing technologies, the beneficial effects of this invention are: This invention utilizes a multi-view circular acquisition array combined with adaptive weight allocation technology to simultaneously acquire image and depth information of goods from multiple angles. It dynamically adjusts the fusion weights based on the quality and feature richness of images from different perspectives, automatically switching to a backup perspective when one viewpoint is completely obscured. Combined with smoothing processing using sequential features from multiple consecutive frames, it effectively improves image acquisition quality under complex lighting and occlusion conditions. A multi-scale feature extraction network is employed to fuse shallow detail features and deep semantic features. Spatial attention and channel attention mechanisms are introduced to enhance the expressive power of key features. Cross-domain feature alignment technology addresses image domain offset issues under different devices and environments, enabling fine-grained identification of goods and accurately distinguishing between different batches of the same model.
[0017] This invention utilizes a dynamic cargo positioning accuracy correction module, combining semantic segmentation confidence and camera calibration errors to correct the initial positioning results. It leverages high-precision reference points on the shelves to correct the camera's global coordinate offset in real time, achieving layered positioning correction for multi-layered stacked cargo. Integrating stacked cargo 3D reconstruction and positioning technology, it reconstructs a 3D point cloud model of the cargo using multi-view depth images, accurately calculating the spatial position and orientation of each item. Simultaneously, it automatically detects the idle, occupied, and partially occupied states of shelf locations, generating a real-time location status distribution map. This provides accurate orientation information for the robotic arm's automatic grasping, improving warehouse space utilization efficiency.
[0018] This invention establishes a hierarchical occlusion handling mechanism, employing different processing strategies based on the proportion of occluded area. It combines generative adversarial networks to complete the image features of occluded areas and utilizes the feature distribution of goods from the same shelf and batch, along with historical storage location data, for contextual inference. This effectively solves the common problem of goods occlusion in warehouses and improves the system's robustness in complex warehousing scenarios. It supports multi-terminal access and augmented reality interaction, using AR technology to overlay operation instructions and goods information onto the actual scene, planning optimal picking routes, and combining voice and gesture control to enhance operational convenience. Offline operation mode and multi-user collaboration functions ensure the continuity and efficiency of large-scale warehouse operations.
[0019] This invention employs an incremental learning mechanism to continuously optimize the recognition and localization models, preventing catastrophic forgetting. It introduces federated learning collaborative optimization technology, combining data from multiple warehouses for model training, improving the model's generalization ability without leaking original business data. Dynamic scheduling of hardware resources is achieved, automatically allocating computing resources based on system load and automatically adjusting system parameters for special operating conditions such as strong light, dim light, and dust in the warehouse. Comprehensive cargo status recognition and anomaly detection functions can promptly identify various cargo anomalies. The data synchronization module achieves seamless integration with the warehouse management system, constructing a complete cargo lifecycle management system. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the overall operation of the intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface proposed in this invention. Figure 2 Flowchart for image acquisition, preprocessing, and quality control; Figure 3 Here is a flowchart of the feature extraction and occlusion completion recognition process; Figure 4 Flowchart for high-precision semantic segmentation and hierarchical localization correction; Figure 5 A flowchart for interactive guidance, data synchronization, and model self-optimization. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0023] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.
[0024] Reference Figures 1 to 5 The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface includes the following modules: The multi-view image acquisition and preprocessing module consists of a ring-shaped acquisition array composed of multiple high-definition industrial cameras and depth cameras. It simultaneously acquires images of the front, side, top, and bottom surfaces of the cargo, as well as depth and texture information. The acquired images are then processed for denoising, distortion correction, illumination normalization, and background segmentation. The foreground region of the cargo is extracted, and a standardized cargo image dataset is generated. The cargo image feature extraction and enhancement module uses an improved convolutional neural network to extract global and local features of the cargo. Global features include the overall shape, size and color distribution of the cargo, while local features include the cargo's labels, barcodes, QR codes and surface texture features. A spatial attention mechanism is introduced to enhance the attention to key feature regions of the cargo and suppress the interference of background noise and irrelevant features. The multimodal feature fusion and cargo recognition module integrates the visual features, depth features, and pre-stored standard cargo feature library of cargo. It uses a feature-level fusion algorithm to generate a unified cargo feature vector, and uses a classifier to classify and recognize the cargo feature vector, outputting the cargo category, model, and batch information. The cargo localization module based on image semantic segmentation uses a semantic segmentation network to perform pixel-level segmentation of cargo images, identify the boundaries and contours of the cargo, and calculate the three-dimensional spatial coordinates of the cargo in the warehouse by combining camera calibration parameters and the warehouse 3D coordinate system, thus determining the precise location of the cargo. The cargo status recognition and anomaly detection module identifies the stacking, tilting and damage status of cargo by analyzing the image features of the cargo, detects missing, misplaced and mixed cargo, generates anomaly detection reports and issues early warning information; The interactive interface and operation guidance module provides a visual interactive interface that displays the identification results, location information and status information of the goods. It supports users to operate the system through touch, voice and gesture. The operation guidance information is overlaid on the actual warehouse scene through augmented reality technology. The identification result management and data synchronization module stores the identification results, location information, and status information of all goods, establishes a goods information database, and realizes data synchronization and information sharing with the warehouse management system; The system performance self-optimization module continuously trains and optimizes the system's recognition and positioning models based on newly added cargo image data and user annotation results, thereby improving the system's recognition accuracy and positioning precision.
[0025] This invention also includes a multi-view image adaptive weight allocation module, which calculates the weight coefficients of each view image based on the sharpness, lighting conditions, and feature richness of the images from different viewpoints. The calculation formula is as follows: ; in, For the first Weighting coefficients for images from each viewpoint; For the first Quality rating of images from each perspective; For the first Feature richness score for images from different perspectives; To determine the total number of perspectives involved in the fusion, the cargo features from different perspectives are weighted and fused according to the calculated weight coefficients to generate a comprehensive cargo feature vector. At the same time, a dynamic perspective switching mechanism is added. When a perspective is completely blocked by cargo or shelves, it automatically switches to a preset backup acquisition perspective. Combined with the temporal features of multiple consecutive frames, smooth fusion is performed to eliminate random noise interference from single-frame images and improve the cargo recognition accuracy under complex lighting and occlusion conditions.
[0026] This invention also includes a dynamic correction module for cargo positioning accuracy, which dynamically corrects the cargo positioning results by combining the confidence level of semantic segmentation and camera calibration error. The calculation formula is as follows: ; in The corrected three-dimensional spatial coordinates of the cargo; The initial calculated three-dimensional spatial coordinates of the cargo; This is the camera calibration error correction coefficient; The confidence level of the semantic segmentation result; To minimize the positioning error, the positioning results are dynamically corrected to eliminate the impact of semantic segmentation uncertainty and camera calibration error on positioning accuracy. At the same time, shelf reference marks are introduced to assist calibration. The global coordinate offset of the camera is corrected in real time using the preset high-precision reference points on the shelf. Layered positioning correction is achieved for multi-layer stacked goods, distinguishing the vertical coordinates of goods in different layers, thereby improving the accuracy and reliability of goods positioning.
[0027] This invention also includes a cargo completion and recognition module for occluded scenes. When cargo is partially occluded, local features of the unoccluded area are extracted, and a complete feature representation of the cargo is generated by combining the cargo's prior shape model and contextual information. The image features of the occluded area are completed by a generative adversarial network, achieving accurate identification and location of partially occluded cargo. At the same time, a hierarchical occlusion processing mechanism is established, which divides the degree of occlusion into three levels: mild, moderate, and severe, based on the proportion of occluded area. Three different processing strategies are adopted: local feature matching, prior model completion, and contextual inference. The feature distribution of cargo on the same shelf and in the same batch, as well as historical storage location data, are used to assist in completion, improving the cargo recognition success rate in severe occlusion scenarios.
[0028] In this invention, the multi-view image acquisition and preprocessing module supports automatic adjustment of camera exposure time, focal length, and shooting angle to adapt to different lighting conditions and the acquisition needs of goods of different sizes. It adopts multi-threaded parallel processing technology to synchronously preprocess images acquired by multiple cameras, establishes an image quality evaluation system, and automatically re-acquires images that do not meet the quality standards. At the same time, it integrates an adaptive lighting compensation system to automatically adjust the brightness and color temperature of the supplementary light according to the ambient light intensity, eliminating the impact of strong light reflection and dark light shadows on image quality. It supports seamless stitching of multi-view images of large-sized goods to generate complete panoramic images, as well as motion blur elimination processing for moving goods, realizing real-time acquisition and preprocessing of dynamic goods, and providing stable image data for subsequent modules.
[0029] In this invention, the cargo image feature extraction and enhancement module employs a multi-scale feature extraction network to simultaneously extract cargo features at different scales. It fuses shallow detail features and deep semantic features to generate a multi-scale fused feature vector. A channel attention mechanism is introduced to adaptively adjust the weights of different feature channels, highlighting feature channels that contribute significantly to cargo recognition and suppressing the influence of useless feature channels. Simultaneously, cross-domain feature alignment technology is incorporated to address the domain offset problem between images acquired from different warehouses and camera devices, achieving fine-grained feature extraction. This allows for the differentiation of cargo of the same model, from different batches, and from different production dates. Feature hashing encoding converts high-dimensional feature vectors into low-dimensional hash codes, significantly accelerating feature matching and improving the efficiency of batch cargo recognition.
[0030] In this invention, the multimodal feature fusion and cargo recognition module employs an attention-guided multimodal feature fusion network. This network automatically learns the correlations and complementarities between different modal features, assigns adaptive fusion weights to features of different modalities, and uses an ensemble learning method to combine the outputs of multiple classifiers to generate the final cargo recognition result. This supports simultaneous recognition and classification of batches of cargo. A dynamic feature library update mechanism is also established to automatically add newly recognized cargo features to the feature library, periodically clean up redundant and outdated feature data, and incorporate an abnormal feature filtering function to eliminate interference from abnormal features caused by stains, scratches, and packaging deformation. This module supports simultaneous recognition of multiple tags and can simultaneously output information on multiple dimensions such as cargo category, model, batch, production date, and shelf life, meeting the needs of complex warehouse management.
[0031] In this invention, the image semantic segmentation-based cargo positioning module uses an instance segmentation network to perform instance-level segmentation of densely stacked cargo, distinguishing different individual cargoes, extracting the centroid coordinates and bounding box information of each cargo, and combining the depth information obtained by the depth camera to calculate the three-dimensional spatial coordinates of the cargo, establishing a mapping relationship between the cargo position and warehouse shelves and storage locations, and achieving accurate storage location positioning of the cargo. At the same time, it integrates stacked cargo three-dimensional reconstruction and positioning technology, uses multi-view depth images to reconstruct the three-dimensional point cloud model of the cargo, accurately calculates the spatial position and posture of each cargo, automatically detects the idle, occupied and half-occupied states of shelf storage locations, generates a real-time storage location status distribution map, and outputs the tilt angle and rotation angle of the cargo, providing accurate posture information for the robotic arm to automatically grasp.
[0032] In this invention, the interactive interface and operation guidance module support multi-terminal access, including computer, mobile, and head-mounted augmented reality devices. It provides functions such as goods query, inbound / outbound registration, inventory counting, and anomaly handling. Through augmented reality technology, it highlights the location of target goods and operation instructions in the actual scene, guiding staff to quickly find the target goods and complete the corresponding operations. It supports voice command control and gesture recognition operation, and also incorporates an AR virtual picking path planning function, displaying the optimal picking path from the current location to the target goods in the augmented reality interface. It supports multi-person collaborative operation, synchronizing the operation status and goods information of multiple staff members in real time. It also has an offline operation mode, caching all operation data when the network is interrupted, and automatically synchronizing to the server after the network is restored to maintain the continuity of warehouse operations.
[0033] In this invention, the system performance self-optimization module establishes an incremental learning mechanism, using newly added cargo image data and user-annotated data to incrementally train the recognition and positioning models, avoiding catastrophic forgetting of the models. A model performance evaluation system is established to periodically evaluate the model's recognition accuracy, positioning precision, and response speed. When the model performance falls below a preset threshold, a model retraining process is automatically triggered. Simultaneously, model version management and rollback are supported. Federated learning collaborative optimization technology is introduced, combining data from multiple warehouses for model training, improving the model's generalization ability without leaking original business data. Dynamic scheduling of hardware resources is achieved, automatically allocating CPU, GPU, and memory resources according to system load to optimize system response speed. System parameters are automatically adjusted for special working conditions such as strong light, dim light, and dust in the warehouse, maintaining stable system operation in complex environments.
[0034] This invention relates to an intelligent warehouse cargo identification and positioning system based on image recognition and an interactive interface. It adopts a cloud-edge-device collaborative architecture. Multi-view image acquisition devices and on-site processing units are deployed at the edge to complete image acquisition, preprocessing, and preliminary feature extraction. The core algorithm model and data storage system are deployed in the cloud to achieve feature fusion, cargo identification, positioning calculation, and model training. Users access the system functions through a multi-terminal interactive interface. The overall system flow includes multi-view image synchronous acquisition, image preprocessing and quality assessment, multi-scale feature extraction and enhancement, multi-modal feature fusion, cargo classification and identification, instance segmentation and 3D localization, status recognition and anomaly detection, result display and operation guidance, and data storage and model optimization. The invention will be further described in detail below with reference to two specific embodiments.
[0035] Example 1:
[0036] This embodiment is applied to an intelligent warehouse for small e-commerce goods. The warehouse has a total area of 5,000 square meters and is equipped with 20 sets of shelves, each set of shelves has six layers, mainly storing small items such as clothing, daily necessities, and electronic products. The single warehouse carries more than 100,000 types of goods, with an average daily inbound and outbound volume of 50,000 items. The system is deployed at the warehouse's inbound and outbound entrances, shelf aisles, and inventory areas to achieve automatic identification and positioning of the entire process of goods registration upon entry, in-stock inventory, and outbound picking.
[0037] The multi-view image acquisition and preprocessing module deploys eight high-definition industrial cameras and four depth cameras at each inlet / outlet, forming a circular acquisition array evenly distributed around the goods conveyor line to simultaneously acquire images of the front, sides, top, and bottom of the goods. A mobile acquisition robot is deployed in the shelving aisles, equipped with the same circular acquisition array, moving along the shelving track to acquire images of goods location by location. The system automatically adjusts the camera's exposure time, focal length, and shooting angle, scaling the acquisition range according to the size of the goods, and activating an adaptive lighting compensation system to adjust the brightness and color temperature of the supplementary lighting for different lighting conditions. The acquired images undergo noise reduction, distortion correction, illumination normalization, and background segmentation using multi-threaded parallel processing technology to extract the foreground region of the goods. An image quality evaluation system is established, scoring images based on three indicators: sharpness, contrast, and illumination uniformity. Images with scores below a threshold are automatically re-acquired. For large goods, the system automatically stitches multi-view images to generate a complete panoramic image. For moving goods, a motion blur removal algorithm is used to process the images to ensure stable image quality.
[0038] The cargo image feature extraction and enhancement module employs a multi-scale feature extraction network, simultaneously extracting cargo features at three different scales and fusing shallow texture and edge detail features with deep shape and semantic features. Spatial attention and channel attention mechanisms are introduced to automatically focus on key feature regions such as cargo labels, barcodes, and QR codes, adaptively adjusting the weights of different feature channels. Cross-domain feature alignment technology is incorporated, using a domain adaptive algorithm to align image features acquired from different camera devices and under different lighting conditions, eliminating the influence of domain offset. The system achieves fine-grained feature extraction, capable of distinguishing goods of the same model but different batches and production dates. Feature hashing encoding converts high-dimensional feature vectors into low-dimensional hash codes, improving feature matching speed by an order of magnitude and meeting the needs for rapid identification of batches of goods.
[0039] The multimodal feature fusion and cargo recognition module employs an attention-guided multimodal feature fusion network to automatically learn the correlation between visual and deep features, assigning adaptive fusion weights to the two modalities. An ensemble learning method is used, combining the outputs of a convolutional neural network classifier, a support vector machine classifier, and a random forest classifier to generate the final cargo recognition result. A dynamic feature library update mechanism is established, automatically adding features of newly arrived cargo to the feature library daily and cleaning up feature data of cargo exceeding its shelf life monthly. An abnormal feature filtering function is added to exclude interference from abnormal features caused by stains, scratches, and packaging deformation. The system supports simultaneous recognition of multiple tags, outputting cargo category, model, batch, production date, and shelf life information in a single recognition.
[0040] The image semantic segmentation-based cargo localization module employs an instance segmentation network to perform instance-level segmentation of densely stacked cargo, distinguishing each individual cargo and extracting its centroid coordinates and bounding box information. It then calculates the cargo's 3D spatial coordinates using depth information acquired by a depth camera, establishing a mapping relationship between cargo location and shelf / storage location. A dynamic cargo localization accuracy correction module is introduced, which corrects the initial localization results by combining semantic segmentation confidence and camera calibration errors, and uses pre-set high-precision reference points on the shelf to correct the camera's global coordinate offset in real time. For multi-layered stacked cargo, the system implements layered localization correction, distinguishing the vertical coordinates of cargo in different layers using depth information. It integrates stacked cargo 3D reconstruction and localization technology, using multi-view depth images to reconstruct the 3D point cloud model of the cargo, accurately calculating the cargo's tilt and rotation angles to provide posture information for automated robotic arm grasping. The system automatically detects the idle, occupied, and half-occupied status of each storage location, generating a real-time storage location status distribution map.
[0041] The cargo status recognition and anomaly detection module analyzes the image features of the cargo to identify the number of stacked layers, tilt angle, and degree of damage, detecting anomalies such as missing, misplaced, or mixed cargo. When an anomaly is detected, the system automatically generates an anomaly detection report, marking the location and type of the anomaly, and sends an alert to warehouse management personnel. The multi-view image adaptive weight allocation module calculates the weight coefficients for each viewpoint based on the quality score and feature richness score of the images from different viewpoints, and performs weighted fusion of features from different viewpoints. When a viewpoint is completely obscured by cargo or shelves, the system automatically switches to a preset backup acquisition viewpoint and performs smooth fusion by combining the temporal features of five consecutive frames, eliminating random noise interference from single-frame images. The occlusion scene cargo completion and recognition module establishes a hierarchical occlusion processing mechanism. Occlusion of less than 30% is considered mild occlusion, which is identified using local feature matching. Occlusion of 30% to 70% is considered moderate occlusion, which is completed by combining prior shape models and generative adversarial networks to complete the features of the occluded area. Occlusion of more than 70% is considered severe occlusion, which is inferred by using the feature distribution of goods on the same shelf and in the same batch, as well as historical storage location data, for contextual association.
[0042] The interactive interface and operation guidance module supports access from computers, mobile devices, and head-mounted augmented reality devices, providing functions such as goods query, inbound / outbound registration, inventory counting, and anomaly handling. Augmented reality technology highlights the location of target goods and provides operation guidance in real-world scenarios. An AR virtual picking route planning function is added, displaying the optimal picking route from the current location to the target goods on the interface. The system supports voice command control and gesture recognition, allowing staff to issue goods query commands via voice and confirm operations via gestures. Multi-person collaborative operation is supported, synchronizing the operation status and goods information of multiple staff members in real time. An offline operation mode is available, caching all operation data during network interruptions and automatically synchronizing it to the server after network recovery. The recognition result management and data synchronization module establishes a goods information database, storing the recognition results, location information, and status information of all goods, achieving real-time data synchronization with the warehouse management system. The system performance self-optimization module establishes an incremental learning mechanism, incrementally training the recognition and positioning models weekly using newly added goods image data and user-annotated data. A model performance evaluation system is established to assess the model's recognition accuracy, positioning precision, and response speed monthly. When model performance falls below a preset threshold, a retraining process is automatically triggered. Model version management and rollback are supported. Federated learning collaborative optimization technology is introduced, combining data from multiple similar e-commerce warehouses for model training to improve the model's generalization ability without leaking original business data. The system implements dynamic hardware resource scheduling, automatically allocating CPU, GPU, and memory resources based on system load, and automatically adjusting system parameters for special operating conditions such as strong light, low light, and dust in the warehouse.
[0043] Table 1: Performance Comparison of E-commerce Small Item Warehouse Systems Table 1 shows the data from three consecutive months of operational statistics of the warehouse in this embodiment. The comparison group consists of the traditional barcode recognition system and the manual management mode used in the warehouse before the renovation. The system of this invention significantly improves the accuracy and speed of cargo recognition through multi-view acquisition and multi-modal feature fusion technology, solving the problems of traditional barcodes being easily worn and contaminated. High-precision 3D positioning technology and anomaly detection function enable refined management of goods, effectively reducing the occurrence of misdelivery, omission, and loss. AR interaction and offline operation modes greatly improve the work efficiency of on-site staff and reduce labor intensity. The system's automatic inventory function shortens the warehouse inventory cycle from once a month to once a day, improving the real-time performance and accuracy of inventory data.
[0044] Example 2:
[0045] This embodiment is applied to an intelligent warehouse for industrial parts. The warehouse has a total area of 8,000 square meters and is equipped with 15 sets of heavy-duty racks. Each set of racks has four layers and mainly stores industrial products such as mechanical parts, hardware tools, and electrical components. The size of the goods ranges from a few centimeters to several meters. There are approximately 20,000 types of goods in a single warehouse, with an average daily inbound and outbound volume of 10,000 items. The system focuses on meeting the requirements for the identification and positioning of large-sized and irregularly shaped goods, as well as the safety status monitoring requirements for heavy goods.
[0046] The multi-view image acquisition and preprocessing module deploys twelve high-definition industrial cameras and six depth cameras at each inlet / outlet, forming a large-size circular acquisition array covering a cargo acquisition space of up to three meters by three meters by three meters. In the heavy-duty racking area, a gantry-type acquisition device is deployed. This device is equipped with a liftable acquisition platform, on which a multi-view camera array is mounted. It can move horizontally and vertically along the racking to acquire images of goods at different heights. The system automatically adjusts the camera's shooting parameters. For large-sized goods, it uses a partitioned acquisition and stitching method to generate a complete image. For reflective metal parts, polarization filtering technology is used to eliminate the effect of reflection. In the image preprocessing stage, a metal surface reflection correction algorithm is added to improve image quality in reflective areas. An image quality assessment system for industrial parts is established, focusing on evaluating edge sharpness and texture details. Images that do not meet the quality standards are automatically re-acquired after adjusting the supplementary lighting angle.
[0047] The cargo image feature extraction and enhancement module addresses the complex shapes and uniform textures of industrial parts by optimizing a multi-scale feature extraction network to enhance the extraction capability of geometric features. It introduces a shape context descriptor to extract contour and topological features of the parts, fusing them with deep features extracted by a convolutional neural network. Cross-domain feature alignment technology is incorporated to resolve appearance differences between parts from different manufacturers and batches. The system can distinguish between industrial parts of the same model but different precision levels, as well as standard parts with similar appearances but different specifications. Feature hashing encoding technology improves the speed of part feature matching, meeting the rapid identification requirements when batches of parts are put into storage.
[0048] The multimodal feature fusion and cargo recognition module optimizes the multimodal feature fusion network, increasing the weight of deep features in industrial component recognition and using depth information to distinguish components that are similar in appearance but different in size. An ensemble learning method is employed, combining the results of geometric feature matching and deep learning classification to generate the final recognition result. A dynamic feature library update mechanism is established, updating the component feature library quarterly and cleaning up feature data of obsolete models. An abnormal feature filtering function is added to eliminate interference from abnormal features caused by rust, oil stains, etc. The system supports simultaneous recognition of multiple tags, outputting the component's category, model, specifications, accuracy level, and manufacturer information.
[0049] The image semantic segmentation-based cargo localization module employs an improved instance segmentation network, optimized for large-sized, irregularly shaped components, to enhance the accuracy of segmentation boundaries. It reconstructs 3D point cloud models of components using multi-view depth information, accurately calculating their 3D spatial coordinates and orientation. A dynamic cargo localization accuracy correction module is introduced, combining semantic segmentation confidence and camera calibration errors to correct the localization results, and using high-precision reference points on the rack to correct the coordinate offset of the gantry-type acquisition device in real time. For multi-layered stacked heavy components, the system implements layered localization correction, accurately distinguishing components on different layers using depth information. It automatically detects the idle, occupied, and partially occupied states of rack locations, generating a real-time location status distribution map to provide a basis for heavy cargo storage planning. It outputs the tilt angle and center of gravity position of components, providing safety guidance for forklift and crane loading and unloading operations.
[0050] The cargo status recognition and anomaly detection module focuses on monitoring the stacking and tilting status of heavy cargo. When the tilt angle exceeds a safety threshold, the system immediately issues a warning to prevent the cargo from tipping over. It also detects quality anomalies such as rust, deformation, and damage, as well as management anomalies such as misplacement or mixed storage. The multi-view image adaptive weight allocation module automatically adjusts the weight coefficients of different viewpoints based on the shape characteristics of industrial parts. For axisymmetric parts, the weights of the front and side views are increased; for irregularly shaped parts, the weights of all viewpoints are evenly distributed. When a viewpoint is occluded, it automatically switches to a backup viewpoint and performs smooth fusion using temporal features from multiple consecutive frames. The occlusion scene cargo completion and recognition module establishes a hierarchical occlusion processing mechanism. For common local occlusion situations in industrial parts, it combines the standard 3D model of the parts and a generative adversarial network to complete the geometric features of the occluded area, achieving accurate identification and location of partially occluded parts.
[0051] The interactive interface and operation guidance module supports access from computers, mobile devices, and industrial tablets, providing functions such as cargo query, warehousing registration, outbound transfer, and equipment maintenance. Augmented reality technology highlights the location of target goods and provides loading / unloading guidance in real-world scenarios, and AR virtual path planning is added to plan optimal travel routes for forklifts and cranes. The system supports voice command control and gesture recognition, facilitating operation by on-site personnel while wearing gloves. Multi-user collaborative operation is supported, with real-time synchronization of loading / unloading progress and cargo status. An offline operation mode is available, caching all operation data during network interruptions and automatically synchronizing upon network recovery. The recognition result management and data synchronization module establishes an industrial component information database, storing the recognition results, location information, and quality status information of all components, achieving data synchronization with the enterprise resource planning system and production management system. The system performance self-optimization module establishes an incremental learning mechanism, performing incremental training every two weeks using newly added component image data and labeled data. A model performance evaluation system is established, evaluating model performance quarterly and automatically triggering a retraining process. The system supports model version management and rollback, and incorporates federated learning collaborative optimization technology. It combines warehouse data from multiple upstream and downstream enterprises for model training, improving the model's ability to identify parts from different manufacturers. The system implements dynamic scheduling of hardware resources, automatically adjusting system resource allocation based on workload. For harsh working conditions such as dust and oil stains in the warehouse, it automatically adjusts image preprocessing parameters and model inference parameters to ensure stable system operation.
[0052] Table 2: Performance Comparison of Industrial Parts Warehouse Systems Table 2 shows the data from six consecutive months of operational statistics of the warehouse in this embodiment. The comparison group consists of traditional RFID systems and manual management methods used in warehouses of similar size and in the same industry. This invention's system is optimized for the characteristics of industrial parts, accurately identifying irregularly shaped and large-sized goods. Its high-precision 3D positioning and attitude detection functions provide safety assurance for the loading and unloading of heavy goods. The system's non-contact identification method avoids the problems of RFID tags being easily damaged or susceptible to metal interference. Its flexible deployment and strong environmental adaptability enable stable operation in the harsh conditions of industrial warehouses, effectively improving the automation and intelligence level of industrial warehouse management.
[0053] Reference Figure 1 This diagram illustrates the entire closed-loop operation process from when goods enter the field of view to when the system completes self-optimization. The process begins with the simultaneous acquisition of multi-view images, followed by feature extraction and multimodal fusion using a deep learning network to identify the type of goods. Subsequently, precise spatial localization is achieved through semantic segmentation, while simultaneously detecting damage or stacking anomalies in the goods. Finally, the recognition results are presented to the operator through an AR interactive interface, and the data is synchronized to the management system. The self-optimization module at the end ensures that the system can continuously evolve itself based on newly added data.
[0054] Reference Figure 2 This diagram details the system's front-end input logic. The system acquires omnidirectional images through a circular acquisition array and incorporates an "adaptive weight allocation" mechanism. For the complex environment of the warehouse, the system automatically performs illumination compensation and distortion correction. Specifically, this process includes a quality assessment closed loop: if the acquired image fails to meet standards due to occlusion or lighting issues, the system will trigger automatic re-acquisition or viewpoint switching to ensure that subsequent recognition modules acquire high-quality, standardized image datasets.
[0055] Reference Figure 3 This image highlights the system's ability to identify cargo under complex conditions, such as partial occlusion. The system employs an improved convolutional neural network (CNN) to extract global shape and local label features in parallel. When cargo occlusion is detected, the system initiates tiered processing based on the occlusion ratio: minor occlusion uses local matching, while severe occlusion utilizes a generative adversarial network (GAN) combined with historical cargo location data and prior models to complete image features, thereby achieving complete recognition of "incomplete images."
[0056] Reference Figure 4This process describes how the system converts 2D pixels into 3D spatial coordinates. The system distinguishes densely stacked individuals using an instance segmentation network and constructs a 3D point cloud using depth information. To eliminate camera errors, the system introduces a "dynamic correction module" that calibrates coordinates in real time by combining the confidence level of semantic segmentation with high-precision reference markers on the shelf. For multi-layered goods, the system achieves layered positioning by differentiating vertical coordinates, ultimately providing precise attitude and position information for automated equipment or operators.
[0057] Reference Figure 5 This diagram illustrates the system's backend management and user interaction logic. The system supports multi-terminal access via AR, mobile devices, and computers, overlaying picking routes onto the real-world environment using augmented reality technology. A data synchronization module ensures that recognition results are shared with the WMS system. The system's core evolutionary capability lies in incremental learning: by collecting new data and user-annotated results, the system utilizes federated learning technology to collaboratively optimize the model while protecting privacy, achieving dynamic scheduling of hardware resources and automatic adaptation to special operating conditions.
[0058] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface, characterized in that, Includes the following modules: The multi-view image acquisition and preprocessing module consists of a circular acquisition array composed of a high-definition industrial camera and a depth camera. It acquires images of the cargo, and simultaneously acquires the cargo's depth and texture information. It preprocesses the acquired images, extracts the foreground region of the cargo, and generates a standardized cargo image dataset. The cargo image feature extraction and enhancement module uses an improved convolutional neural network to extract cargo features and introduces a spatial attention mechanism to enhance attention to key feature regions of the cargo. The multimodal feature fusion and cargo recognition module integrates the visual features, depth features, and pre-stored standard cargo feature library of cargo. It uses a feature-level fusion algorithm to generate a unified cargo feature vector, and a classifier classifies and recognizes the cargo feature vector to output cargo information. The cargo localization module based on image semantic segmentation uses a semantic segmentation network to perform pixel-level segmentation of cargo images, identify the boundaries and contours of the cargo, and calculate the three-dimensional spatial coordinates of the cargo by combining camera calibration parameters and the warehouse three-dimensional coordinate system to determine the cargo location. The cargo status recognition and anomaly detection module identifies the cargo status by analyzing the image features of the cargo, detects abnormal cargo conditions, generates anomaly detection reports, and issues early warning information. The interactive interface and operation guidance module provides a visual interactive interface, allowing users to operate the system through touch, voice, and gestures. Augmented reality technology is used to overlay operation guidance information onto the actual warehouse scene. The identification result management and data synchronization module stores cargo identification results, location information, and status information, and establishes a cargo information database. The system performance self-optimization module continuously trains and optimizes the system's recognition and localization models based on newly added cargo image data and user annotation results.
2. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, It also includes a multi-view image adaptive weight allocation module, which calculates the weight coefficients of each view image based on the sharpness, lighting conditions, and feature richness of the images from different viewpoints. The calculation formula is as follows: ; in, For the first Weighting coefficients for images from each viewpoint; For the first Quality rating of images from each perspective; For the first Feature richness score for images from different perspectives; To determine the total number of perspectives involved in the fusion, the cargo features from different perspectives are weighted and fused according to the calculated weight coefficients to generate a comprehensive cargo feature vector. At the same time, a dynamic perspective switching mechanism is added, which automatically switches to a preset backup acquisition perspective when a perspective is completely blocked by cargo or shelves, and then smoothly fused by combining the temporal features of multiple consecutive frames.
3. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, It also includes a dynamic correction module for cargo positioning accuracy, which dynamically corrects the cargo positioning results by combining the confidence level of semantic segmentation and camera calibration error. The calculation formula is as follows: ; in, The corrected three-dimensional spatial coordinates of the cargo; The initial calculated three-dimensional spatial coordinates of the cargo; This is the camera calibration error correction coefficient; The confidence level of the semantic segmentation result; This represents the maximum positioning error value.
4. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, It also includes a cargo completion and recognition module for occluded scenes. When cargo is partially occluded, this module extracts local features of the unoccluded area, combines a prior shape model with contextual information to generate a complete feature representation, and completes the image features of the occluded area through a generative adversarial network. The module establishes a hierarchical occlusion processing mechanism, dividing the occlusion into three levels: mild, moderate, and severe, based on the proportion of occluded area. It adopts three different processing strategies: local feature matching, prior model completion, and contextual inference, respectively, and uses the feature distribution of cargo on the same shelf and in the same batch, as well as historical storage location data, to assist in the completion.
5. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The multi-view image acquisition and preprocessing module supports automatic adjustment of camera exposure time, focal length, and shooting angle. It uses multi-threaded parallel processing technology to synchronously preprocess images acquired by multiple cameras, establishes an image quality assessment system, and automatically re-acquires images that do not meet the quality standards. At the same time, it integrates an adaptive lighting compensation system to automatically adjust the brightness and color temperature of the fill light according to the ambient light intensity, eliminating the impact of strong light reflection and dark shadows on image quality.
6. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The cargo image feature extraction and enhancement module adopts a multi-scale feature extraction network to extract cargo features at different scales simultaneously, fuse shallow detail features and deep semantic features to generate a multi-scale fused feature vector, and introduces a channel attention mechanism to adaptively adjust the weights of different feature channels. This module incorporates cross-domain feature alignment technology to address the image domain offset problem. It converts high-dimensional feature vectors into low-dimensional hash codes through feature hashing, enabling fine-grained cargo feature extraction and feature matching.
7. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The multimodal feature fusion and cargo recognition module employs an attention-guided multimodal feature fusion network to learn the correlation and complementarity of different modal features and assign adaptive fusion weights. It uses an ensemble learning method combined with the classifier output to generate cargo recognition results. This module establishes a dynamic feature library update mechanism, automatically adds new features and cleans up redundant and outdated data, adds anomaly feature filtering function, supports simultaneous recognition of multiple labels, and outputs multi-dimensional cargo information.
8. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The image semantic segmentation-based cargo localization module uses an instance segmentation network to perform instance-level segmentation of densely stacked cargo, distinguishing different individual cargoes, extracting the centroid coordinates and bounding box information of the cargoes, and combining the depth information obtained by the depth camera to calculate the three-dimensional spatial coordinates of the cargoes, establishing a mapping relationship between the cargo positions and warehouse shelves and storage locations, and realizing storage location localization. At the same time, it integrates stacked cargo three-dimensional reconstruction and localization technology, uses multi-view depth images to reconstruct the three-dimensional point cloud model of the cargoes, calculates the spatial position and orientation of the cargoes, detects the storage location occupancy status and generates a real-time status distribution map, and outputs cargo orientation angle information.
9. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The interactive interface and operation guidance module supports access from computers, mobile devices, and head-mounted augmented reality devices, providing functions such as goods inquiry, inbound and outbound registration, inventory counting, and anomaly handling. This module uses augmented reality technology to highlight the location of target goods and operation instructions in the actual scene, supports voice command control and gesture recognition operation, realizes AR virtual picking path planning and multi-person collaborative operation, and has an offline operation mode, which automatically synchronizes to the server after the network is restored.
10. The intelligent warehouse cargo identification and positioning system based on image recognition and interactive interface according to claim 1, characterized in that, The system performance self-optimization module establishes an incremental learning mechanism, which uses newly added cargo image data and user-annotated data to incrementally train the recognition model and the localization model to avoid catastrophic forgetting of the model; the module also establishes a model performance evaluation system, which automatically triggers retraining when the model performance is lower than a preset threshold. It supports model version management and rollback, introduces federated learning technology, combines data from multiple warehouses for model training, realizes dynamic scheduling of hardware resources, and automatically adjusts system parameters for special operating conditions.