Polar region unmanned storage goods identification and inventory method and system based on artificial intelligence

By leveraging the collaborative operation of unmanned vehicles, drones, and cloud systems, combined with multimodal fusion perception methods, the challenges of material identification and warehousing in harsh polar environments have been solved, enabling unmanned, end-to-end material management and improving identification accuracy and warehousing efficiency.

CN120822909AActive Publication Date: 2025-10-21POLAR RES INST OF CHINA

Patent Information

Application Number
CN202511325065.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-21
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing technologies cannot achieve automated and intelligent material identification and warehousing in harsh polar environments, and face problems such as degraded equipment performance, unreliable communication, difficulties in data collection, and low accuracy in cargo identification.

Method used

The system employs unmanned vehicles and drones working in tandem, combined with a cloud system for data collection, transmission, and analysis. It utilizes multimodal fusion perception methods for cargo identification, generates a 3D digital map, and performs dynamic warehousing planning.

Benefits of technology

It has enabled unmanned, end-to-end material identification and warehousing in polar environments, improving the accuracy of cargo identification and warehousing efficiency, ensuring the stability and reliability of the system, and meeting the material support needs of the research station.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822909A_ABST
    Figure CN120822909A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence-based polar region unmanned warehouse cargo identification and inventory method and system, and relates to the field of artificial intelligence and computer vision, and the identification and inventory method specifically comprises the steps that at least one unmanned vehicle carries out the data collection of cargos in an unloading area, obtaining video data and point cloud data of the goods; at least one unmanned aerial vehicle serves as a data relay, and video data and point cloud data collected by the unmanned aerial vehicle are transmitted to a cloud system; performing multi-modal fusion based on the pure cargo surface point cloud and the features of the video data to realize cargo instance segmentation, and identifying the content category of each cargo instance in combination with the image information in the region of interest; generating a three-dimensional digital map of the unloading area; and based on the three-dimensional digital map, performing multi-dimensional dynamic warehousing planning to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated warehousing technology, and more specifically, to unmanned intelligent logistics management technology for use in extreme environments. More specifically, the present invention relates to methods and systems that utilize artificial intelligence and computer vision to automatically identify, schedule, transport, store, and inventory large quantities of goods in harsh environments such as polar regions. Background Art

[0002] With increasing global attention to scientific exploration and resource development in the polar regions, the demand for establishing long-term or semi-permanent research stations and bases in these regions is growing. The efficient operation of these stations relies heavily on stable and reliable logistical support. However, the polar regions present an extremely harsh environment, characterized by ultra-low temperatures (reaching tens of degrees below zero Celsius), prolonged snowstorms, polar nights, and unpredictable ice conditions. This extreme natural environment poses a disruptive challenge to traditional material management and warehousing operations.

[0003] Under conventional circumstances, automated warehousing technology has made significant progress. Large logistics centers widely utilize equipment and software such as automated guided vehicles, stacking cranes, shuttles, and warehouse management systems, enabling efficient and accurate entry and exit of goods. These systems typically operate indoors under ideal conditions with controlled temperatures, flat floors, and stable network signals. Their navigation relies heavily on ground-mounted QR codes, magnetic strips, or LiDAR simultaneous localization and mapping (SLAM) in structured environments. Furthermore, the identification and counting of goods relies heavily on pre-defined identifiers such as barcodes and RFID tags, as well as stable, high-speed internal networks for data transmission and processing.

[0004] However, when these existing technologies are directly applied to polar regions, their inherent flaws become apparent. First, polar resupply typically relies on a centralized unloading model during a window period. This means that transport vessels, during a short seaworthy period, unload several months or even a year's worth of supplies onto ice or improvised sites some distance from the camp. This results in large quantities of cargo being piled up unorganized outdoors, completely lacking the structured environment of a conventional warehouse. Traditional AGV navigation and positioning technology is virtually ineffective in this dynamic, open, and sparsely characterized snowy environment. Second, the polar environment places severe demands on both equipment and personnel. Prolonged low temperatures can degrade the performance of electronic components, significantly reduce battery life, and embrittle metal materials. Blizzards not only severely impact the performance of optical sensors, resulting in poor data quality or even interruption, but also make manual outdoor operations extremely risky and inefficient. Manual inventory and handling are not only time-consuming and labor-intensive, but also pose life-threatening risks such as frostbite and getting lost. This approach fails to meet the immediate and continuous supply requirements of research stations.

[0005] More critically, data communication reliability is a challenge. Stable, long-distance, high-bandwidth wireless communication is a luxury in the vast polar regions. Severe weather, such as blizzards, can severely interfere with radio signals, making the real-time transmission of large amounts of data (such as high-definition video streams) from outdoor unloading areas to indoor control centers extremely difficult and unreliable. This information silo effect prevents the management system from timely and accurate access to comprehensive front-end information on goods, such as type, quantity, and location, making subsequent warehousing planning and scheduling impossible. Without effectively addressing the core issues of visibility, transmission, and accurate calculations in harsh outdoor environments, automated and intelligent warehouse management will become a pipe dream.

[0006] Therefore, existing technologies cannot provide a complete, automated material identification and warehousing solution that can adapt to polar conditions, from outdoor unloading areas to indoor storage. Faced with the dynamic and disordered distribution of goods outdoors, the severe constraints imposed by extreme weather on equipment and communications, and the complex decision-making required for material warehousing, a new technical solution is urgently needed to overcome the application bottlenecks of traditional warehousing technology in extreme environments. Summary of the Invention

[0007] The present invention provides an artificial intelligence-based method for identifying and counting cargo in unmanned polar warehouses, which specifically includes the following steps: At least one unmanned vehicle collects data on cargo in the unloading area to obtain video data and point cloud data of the cargo; At least one drone acts as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to the cloud system; The cloud-based system pre-processes the received point cloud data and video data to identify regions of interest corresponding to the cargo surface in the video data; performs multimodal fusion based on the features of the clean cargo surface point cloud and the video data to segment cargo instances, thereby obtaining a three-dimensional bounding box for each independent cargo instance; and identifies the content category of each cargo instance by combining image information within the region of interest; and generates a three-dimensional digital map of the unloading area. Based on the three-dimensional digital map, combined with task priority information, real-time environmental information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.

[0008] The embodiments of this specification also propose an artificial intelligence-based polar unmanned warehouse cargo identification and inventory system, which includes: Collection module: at least one unmanned vehicle collects data on the cargo in the unloading area to obtain video data and point cloud data of the cargo; at least one drone acts as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to the cloud system; 3D digital map generation module: The cloud system pre-processes the received point cloud data and video data, identifies regions of interest corresponding to the cargo surface in the video data, performs multimodal fusion based on the features of the clean cargo surface point cloud and the video data, and achieves cargo instance segmentation, thereby obtaining a 3D bounding box for each independent cargo instance. Combining the image information within the region of interest, the content category of each cargo instance is identified, and a 3D digital map of the unloading area is generated. Dynamic task queue generation module: Based on the three-dimensional digital map, combined with task priority information, real-time environmental information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.

[0009] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned artificial intelligence-based polar unmanned warehouse cargo identification and inventory method is implemented.

[0010] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned artificial intelligence-based polar unmanned warehouse cargo identification and inventory method.

[0011] Compared with the existing technology, the artificial intelligence-based polar unmanned warehouse cargo identification and inventory method and system proposed in the present invention has significant beneficial effects.

[0012] The present invention has built a complete, full-process, unmanned polar material support system, solving the fundamental problem of high risk, low efficiency, or even impossibility of manual operations in harsh polar environments. The present invention deeply couples and collaborates unmanned vehicles, drones, and cloud-based intelligent systems: at the front end, unmanned vehicles adapted to polar environments perform autonomous and comprehensive data collection, avoiding the dangers of outdoor operations for personnel; in the data transmission link, drones are used as aerial mobile relays, overcoming the bottleneck of unreliable long-distance wireless communications in polar blizzards and ensuring that massive amounts of high-dimensional data can be safely and completely transmitted back; at the back end, the cloud-based system performs intelligent analysis, planning, and scheduling. This complete closed loop of ground collection, aerial relay, and cloud-based decision-making has formed an intelligent solution capable of long-term, stable, and autonomous operation. Its systematic and innovative design is unmatched by existing single technologies or simple combinations of equipment.

[0013] At the core of cargo identification, this invention proposes a multimodal fusion perception method, significantly improving the segmentation accuracy of independent cargo instances in disordered, densely stacked scenarios. Traditional methods struggle to distinguish between closely adjacent or stacked cargo, but this invention cleverly leverages the complementary advantages of image and point cloud data. By preprocessing video data, the region of interest (ROI) on the cargo surface is precisely located. Using this as prior information, features in the image containing rich texture and boundary details are selectively and accurately elevated to three-dimensional space, where they are deeply integrated with the precise geometric structure provided by the point cloud. This strategy of using image semantic information to guide three-dimensional geometric analysis can significantly enhance the distinguishability of features between different cargo instances, especially in seams or occluded areas where point cloud data is sparse, enabling precise boundary definition and solving the industry's instance segmentation challenge.

[0014] This invention also introduces a consistency verification mechanism based on three-dimensional geometric information, providing a double guarantee for the reliability of identification results. After identifying the cargo's content category through image information, the system uses point cloud data to calculate the cargo's actual physical volume and compares it with the volume information identified from the image text. This cross-modal cross-validation method can effectively detect and flag anomalies such as identification errors or packaging discrepancies with the contents, significantly improving the accuracy and reliability of inventory counts, which is crucial for ensuring the safety of supplies during long-term polar missions.

[0015] The design of the drone data relay is itself a highly robust and practical engineering solution in extreme communication environments. Furthermore, the system dynamically generates and updates a three-dimensional digital map of the unloading area in real time. This map not only includes the location and category of each item, but also incorporates accessibility assessments, digitizing the physical stacking relationships on site. This provides unprecedented decision-making support for subsequent multi-dimensional dynamic warehousing planning, making the entire warehousing process truly intelligent and optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a flow chart of the present invention's artificial intelligence-based polar unmanned warehouse cargo identification and inventory. DETAILED DESCRIPTION

[0018] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0019] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the features in the following embodiments and embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.

[0020] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this application, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspect described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.

[0021] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples, however, one skilled in the art will appreciate that the examples can be practiced without these specific details.

[0022] In order to better understand the technical solution of the present invention, this embodiment first describes in detail the scenario and system composition to which the present invention is applied. The method and system designed by the present invention are intended to perform operations in typical polar environments such as the Arctic or Antarctic. The notable characteristics of such environments are extremely harsh climate, perennial low temperatures, and frequent long-lasting and intense snowstorms, which make it difficult for humans to perform long-term outdoor work. At the same time, it also places extremely high demands on the stability of the equipment and the reliability of communications.

[0023] In the application scenario of the present invention, material supply depends on transport ships with icebreaking capabilities to be completed within a short window period each year when weather conditions are relatively stable. In order to avoid the risk of being trapped by ice floes or encountering sudden severe weather, transport ships need to carry out rapid material unloading operations after arrival. The operation site is set as a designated area with relatively flat and open terrain far away from the main building complex of the scientific research station or base. This area is referred to as the unloading area in the present invention. The transport ship will unload large amounts of materials such as survival, scientific research and equipment spare parts for the next few months or even a whole year at one time on the ice or land in the unloading area. Due to the tight time limit for unloading operations, the goods are initially stacked in a disordered or semi-ordered state in the unloading area, with containers of different types and sizes mixed together, and no fixed, directly identifiable storage layout is formed. After completing the unloading operation, the transport ship will quickly leave, leaving these materials waiting for subsequent transshipment and warehousing.

[0024] The system deployment of this invention fully considers the specificities of the aforementioned scenarios. An unmanned storage room is built within the research station or base, serving as the final storage and management center for all supplies. This unmanned storage room is physically separated from the "unloading" site. The system's core computing and control unit, the cloud system, is deployed within the unmanned storage room to ensure its operation in a stable and secure environment. The cloud system includes high-performance servers, data storage arrays, and a software platform running artificial intelligence algorithms.

[0025] To execute subsequent operational processes, the system's hardware is deployed in a distributed manner. Several unmanned vehicles, specially designed for polar environments, are pre-positioned or deployed in the unloading area. These vehicles are equipped with low-temperature adaptability, off-road capabilities, and a variety of sensor configurations. Meanwhile, the unmanned storage base is equipped with several drones, similarly designed to withstand low temperatures and strong winds. This spatial separation of unmanned vehicles and drones is the foundation of this invention's design for the long-distance and unreliable communication environment of the polar regions, and constitutes a prerequisite for the implementation of all subsequent technical solutions of this invention.

[0026] Next, this example details the specific method for collecting ground visual information using unmanned vehicles in the unloading area. In this stage, multiple unmanned vehicles work together without human intervention to complete a comprehensive, multi-angle scan of all cargo within the unloading area, generating high-quality raw data for subsequent analysis.

[0027] After receiving activation commands from the unmanned warehouse's cloud-based system, multiple unmanned vehicles pre-deployed in the unloading area immediately take flight. To efficiently and completely explore the unknown and disordered environment, the unmanned vehicle fleet implements a collaborative coverage path planning algorithm based on multi-agent boundary exploration.

[0028] S2.1 Rasterize the entire unloading area on a two-dimensional plane to construct an environmental map model ,in is a collection of grid cells, representing discrete location points, is the set of edges connecting adjacent grids.

[0029] S2.2 Unmanned vehicle fleets autonomously plan their own routes , to maximize information gain and minimize total cost. Each unmanned vehicle The path planning can be modeled as an optimization problem, and its objective function is as follows: in, Representing unmanned vehicles A candidate path is a series of connected grid points. is the comprehensive evaluation value of the candidate path. Represents the information gain of the path, defined as the path The number of unknown grid cells that can be newly observed by all observation points on the map. This item encourages the autonomous vehicle to move towards the unknown area (i.e., the map boundary). It represents the execution cost of the path, which is the total length of the path or the estimated energy consumption. Indicates the overlap cost of the path, which is used to measure the path Set of paths planned by other unmanned vehicles The purpose of introducing this item is to prevent multiple unmanned vehicles from repeatedly exploring the same area, thereby improving the collaborative operation efficiency of the entire fleet. are the weight coefficients of information gain, execution cost, and overlap cost, respectively. They are positive real numbers and can be preset or dynamically adjusted by the cloud system according to task priority and environmental conditions to balance exploration efficiency and energy consumption.

[0030] S2.3 During operation, each unmanned vehicle independently calculates the path evaluation value of its surrounding boundary points , and selects the path with the highest evaluation value to move. Through periodic information exchange (for example, broadcasting each selected path), the unmanned vehicle fleet can dynamically and collaboratively complete a traversal scan of the entire unloading area.

[0031] S2.4 Along the planned path When driving, each unmanned vehicle uses its multi-sensor fusion module to collect data. This module includes at least one high-resolution panoramic or wide-angle camera and one solid-state laser radar. The camera is used to collect high-resolution video data streams, which are recorded as ,in The laser radar simultaneously scans the three-dimensional space and generates point cloud data describing the surrounding environment and the geometric shape of the goods, which is recorded as All sensor data undergoes strict clock synchronization to ensure that video frames and point cloud data are precisely aligned in timestamps, providing a foundation for subsequent data fusion and 3D reconstruction.

[0032] The S2.5 system calculates in real time the ratio of the number of observed grids to the total number of grids in the unloading area, i.e. the coverage rate. , when coverage Reaching the preset threshold The data collection task is completed.

[0033] in, is the number of grid cells currently covered by at least one unmanned vehicle sensor. It is the unloading area Each unmanned vehicle locally stores multi-angle, high-density video data of the area it is responsible for. and point cloud data , forming a comprehensive, redundant digital record of the cargo layout in the unloading area, awaiting the next stage of data transmission.

[0034] Next, this embodiment describes in detail the specific method of the drone in the present invention performing aerial data relay to safely and reliably transmit the large-capacity data collected by the unmanned vehicle from the unloading area to the unmanned storage room.

[0035] S3.1 When the unmanned vehicle fleet is collecting data, it periodically sends status information, including the amount of collected data, remaining power, and current coverage rate. The information is sent to the cloud system in the unmanned warehouse through a low-power, long-distance narrowband communication link.

[0036] S3.2 Based on the received status information, the cloud system starts the drone scheduling strategy decision model to determine whether to adopt a single-flight full transmission mode or a multiple-flight continuous transmission mode.

[0037] S3.2.1 The decision model evaluates the feasibility of the single flight mode and determines whether there is at least one drone in the current hangar. , its battery consumption and data storage for a single mission can meet the mission requirements. The feasibility of a single mission requires that the following two conditions are met at the same time: Storage capacity condition determination, drone Available data storage space Must be greater than or equal to the estimated total amount of data to be collected in the unloading area . The total number of grids in the unloading area With an empirical data density coefficient Multiply them together to get .

[0038] Battery life condition determination, drone Current battery level It must be sufficient to support the entire mission process, including flying to the unloading area, hovering and waiting to receive all data, and finally returning home. The calculation formula is as follows: in, is the one-way flight time of the drone, which is determined by the distance between the unmanned storage room and the unloading area. and the average cruising speed of the drone under current weather conditions Decide, . and They are the battery consumption rate per unit time when the drone is in flight and hovering state respectively. is the estimated remaining time required for the unmanned vehicle fleet to complete data collection. All data Estimated time required to transfer data from the autonomous vehicle to the drone. It is a safety redundant power set to deal with emergencies.

[0039] S3.2.2 The cloud system traverses all available drones. If any drone is found that meets the above storage and battery conditions, the first case is activated: single flight full transmission mode. In this mode, the system will assign the drone The drone will wait for the data collection task of the unmanned vehicle fleet to be completed ( After arriving above the unloading area, it receives all the video data stored by all unmanned vehicles at once through a high-bandwidth short-range communication link. and point cloud data After receiving the data, the drone immediately returns and uploads the data to the cloud system after arriving at the unmanned storage room.

[0040] If the cloud system evaluates and finds that no single drone can meet the flight time or storage requirements for a single mission, the second scenario is automatically activated: multi-flight continuous transmission mode. This mode aims to use an ant-like approach, with multiple drones working together to transfer data in batches. The scheduling algorithm for multi-flight continuous transmission mode is as follows: 1) The cloud system assigns the first available drone Immediately fly to the unloading area and upon arrival start downloading some of the data already collected from the unmanned vehicle.

[0041] 2) In During the execution of the task, the cloud system continuously monitors its two key states: the amount of data loaded and remaining power .

[0042] 3) If and only if any of the following conditions are met, The return conditions are triggered: (1) Storage space is about to run out: ,in A small storage margin; (2) The remaining power reaches the return threshold: , ensuring it has enough power to return safely.

[0043] 4) Cloud systems based on drones The status and data transmission rate of the device are used to estimate the return trigger time. To achieve seamless connection, the cloud system will calculate and assign the next drone Best time to depart : Ensured successor drones Able to Arriving precisely when preparing to leave the unloading area minimizes possible operational interruptions that may occur while the unmanned vehicle is waiting for data dumps.

[0044] 5) Steps 2)-4) are repeated, and the drone They are dispatched one after another until the unmanned vehicle fleet completes its collection mission ( ), and the last batch of collected data is The drone was successfully retrieved and uploaded to the cloud system. At this point, the entire data relay process is complete.

[0045] Next, this embodiment describes in detail how the cloud system processes the received raw point cloud data. and video data Perform preprocessing.

[0046] S4.1 Point cloud data preprocessing Raw point cloud data transmitted from the drone to the cloud system It includes multiple elements such as cargo, ground, attached snow, and snow in the air. To accurately reconstruct the 3D model of the cargo, the following purification process is required: S4.1.1 Noise and outlier filtering Snowflakes flying in the air will form a large number of sparse, isolated noise points in the laser radar scan. This paper uses a statistical outlier removal algorithm to filter out these noises. Any point in , calculate its distance to the nearest The average distance of neighboring points . Calculate the global mean and standard deviation When the average neighborhood distance of a point satisfies the following conditions, the point is considered an outlier and is removed: in: is the number of neighborhood points for the outlier removal algorithm. Is the standard deviation multiple threshold, used to control the strictness of the filter. Through this step, we can get the point cloud after preliminary filtering of air noise. .

[0047] S4.1.2 Ground segmentation In order to separate the cargo from the ground, the present invention adopts the random sampling consistency algorithm to fit the ground plane model. Iteratively randomly select the minimum set of points (preferably 3 points) to fit a plane equation , and calculate the distance from other points in the point cloud to the plane. The distance is less than the threshold The points are considered as interior points. After multiple iterations, the plane model with the most interior points is determined to be the ground. The set of all interior points belonging to the ground model is Will be separated from the point cloud, and the remaining non-ground point set It mainly includes various types of goods.

[0048] S4.1.3 Separation of snow surface and cargo surface The snow attached to the cargo surface has different geometric and physical properties from the cargo box. The present invention uses a region growing algorithm based on physical characteristics to achieve accurate separation of the two.

[0049] S4.1.3.1 For non-ground point sets Every point in , calculate its multi-dimensional physical characteristics, mainly including the laser radar reflection intensity and the local surface normal vector calculated based on the neighborhood points .

[0050] S4.1.3.2 Select seed points. Traverse all points and select the ones that meet the high reflection intensity. The local area is highly flat (the variance of the neighborhood normal vector is less than ) as seed points, which are most likely located on the surface of the artificial cargo box.

[0051] S4.1.3.3 Starting from all seed points, region growing is performed in parallel. A growing region will try to swallow up its neighboring unassigned points. A neighboring point To be incorporated into the current area, ,in: is the average normal vector of the current growing area. is the normal vector of the neighboring point to be examined. It is the normal vector similarity threshold. Its value is close to 1, which means that only points with very close normal vector directions (i.e., good coplanarity) will be merged.

[0052] S4.1.3.4 The growing process continues until no points can be added to any region. The set of all grown regions is considered the final pure cargo surface point cloud. . And the point set The remaining points that are not merged into any region are classified as irregular snow surfaces and removed.

[0053] S4.2 Video Data Preprocessing For the video data stream collected by the unmanned vehicle ,In order to effectively separate the snow-covered area and the box surface containing ,information, the present invention adopts a threshold segmentation and morphological ,processing method based on color space.

[0054] S4.2.1 Snow area segmentation based on HSV color space Visually, snow has stable color characteristics of low saturation and high brightness.

[0055] Each frame of RGB image Convert to HSV color space. Create a binary snow mask by setting a predefined threshold range. For any pixel in the image , its judgment logic is as follows: in: Pixels hue, saturation, and lightness values. This is the preset threshold parameter used to define white. 1 means the pixel belongs to snow, 0 means it does not.

[0056] S4.2.2 Mask Optimization and Region of Interest (ROI) Extraction Binarized snow mask There may be noise (such as isolated white spots or tiny holes inside). For this reason, morphological operations are performed on the binarized snow mask to optimize it.

[0057] Use the open operation to remove small noise points, and then use the closed operation to fill the holes inside the snow area to get a smoother and more complete snow area mask. .

[0058] By performing a logical inversion on this mask, the required region of interest (ROI) can be obtained. This ROI accurately identifies all non-snow-covered box surface areas in the image that may contain cargo information. The extracted ROI set is recorded as ,The set includes the ROI area image and its corresponding image frame as well as the ,position information in the image frame.

[0059] Next, this embodiment details how the cloud system analyzes preprocessed data to segment, identify content, and count individual items within the unloading area. This process is accomplished through a phased, hierarchical, multimodal cargo perception network. The network's primary task is to leverage the fusion of point cloud and image information to address the instance segmentation challenge in physical space, specifically distinguishing between closely stacked or adjacent items. A secondary task is to leverage image semantics for content recognition, which is then cross-validated using point cloud geometry.

[0060] S5.1 Construction of the overall architecture of the hierarchical multimodal cargo perception network HMCP-Net consists of two parallel processing branches (point cloud branch and image branch) and a subsequent recognition and verification module.

[0061] S5.2 Stage 1: Cargo instance segmentation based on multimodal fusion. The goal of the first stage is to generate an accurate 3D bounding box for each independent cargo box in the unloading area.

[0062] S5.2.1 Parallel feature extraction: The input of the point cloud branch is the cleaned cargo surface point cloud . It is converted into a three-dimensional voxel grid through a voxel encoder. Each non-empty voxel contains the statistical features of its internal points, and the point cloud voxel features are obtained. The input of the image branch is the original video frame . Extract deep feature maps through ResNet ,in is the height and width of the feature map, is the number of channels.

[0063] The present invention still uses raw video frames as input. This is because convolutional neural networks extract image features not by analyzing individual pixels in isolation, but rather by perceiving patterns within a region using convolution kernels of a certain size. The deep features ultimately extracted from a pixel within the ROI are determined by the large number of pixels within it and its surroundings (possibly outside the ROI). Feeding only a cropped, irregular ROI region into the network creates a large number of artificial boundaries around the edges of the goods. These false boundaries severely interfere with the convolution process, resulting in low-quality and distorted extracted edge features. However, feeding the entire frame allows the network to see the natural transition between the goods and the surrounding environment (such as snow or the ground), thereby learning more discriminative, high-quality boundary features, which is crucial for subsequently distinguishing adjacent or stacked goods.

[0064] Furthermore, polar environments experience dramatic variations in lighting and weather conditions (e.g., sunny, cloudy, light snow, and ground reflectivity). Using the entire raw video frame as input allows the network to perceive global environmental information at an early stage of feature extraction. For example, the network can adapt its feature extraction strategy for the cargo area based on information such as the brightness and color temperature of the sky and distant scenery, resulting in greater robustness to changes in lighting and contrast. This global contextual awareness, unattainable by inputting isolated ROIs, makes the model more stable in the diverse polar outdoor scenes.

[0065] More importantly, modern, mature CNN backbone networks (such as ResNet) are designed to process regular rectangular images. Inputting a cropped, irregularly shaped ROI requires complex padding or deformation operations before feeding it into the network. This not only introduces noise and artifacts but also disrupts the original spatial relationships. Using a strategy that inputs complete frames completely avoids this problem.

[0066] S5.2.2 Image Feature 3D Voxel Generation Module,To achieve cross-modal fusion, it is necessary to,upgrade 2D image features to 3D space.

[0067] S5.2.2.1 Utilizing point clouds And the camera internal and external parameter matrix, through projection and sliding window maximum interpolation method, generate a dense depth map aligned with the image .

[0068] S5.2.2.2 Image feature map Upsample to the same resolution as the original image through bilinear interpolation. Then, only the resulting set of regions of interest (ROIs) Backprojection is performed on the pixels within .

[0069] Specifically, for any pixel within the ROI , its three-dimensional space coordinates Calculated by the following formula: in, is the camera internal parameter matrix. Each generated 3D point All carry it in the feature map The eigenvector of the corresponding position All these points Constructing a semantically rich virtual point cloud .

[0070] S5.2.2.3 Virtual point cloud Perform voxel processing to obtain Image feature 3D voxels on the same spatial grid .

[0071] In step S5.2.2.2, the present invention performs back-projection operations only on pixels within the region of interest (ROI) to generate pseudo-voxels for fusion. The purpose of this design is to use semantic information to guide geometric reconstruction, achieve the focus of computing resources and improve segmentation accuracy: A single high-definition image contains millions of pixels. Back-projecting all of these pixels would generate an extremely large and redundant virtual point cloud, far exceeding the original LiDAR point cloud. Subsequent operations such as voxelization, fusion, and convolution would incur enormous computational and storage overhead, making it impractical for practical applications. The ROI (the segmented, non-snow-covered cargo surface) typically only occupies a small portion of the image. Operating on only this small portion of valid pixels can reduce computational complexity by one to two orders of magnitude, enabling the efficient execution of complex multimodal fusion algorithms.

[0072] The non-ROI areas in the image are backgrounds such as snow, sky, and distant scenery that are completely irrelevant to the cargo instance segmentation task. Back-projecting these pixels into three-dimensional space will generate a large number of semantic noise points. These noise points will seriously interfere with the subsequent instance segmentation. For example, the three-dimensional voxels of a snowy area may be difficult to distinguish from the pseudo-voxels of a white cargo box in terms of features, causing the segmentation algorithm to mistakenly merge the two or fail to find clear object boundaries. Back-projecting only the ROI pixels is equivalent to performing a strong semantic filtering before the lifting operation, ensuring that the generated camera image feature three-dimensional voxels They are all related to real goods, providing clean input for subsequent accurate segmentation.

[0073] The core value of multimodal fusion lies in information complementarity. In scenarios where goods are tightly stacked or placed side by side, the LiDAR point cloud data at the seams of the goods may be very sparse, or even have holes due to occlusion. This makes it difficult to determine whether this is the edge of a single object or the boundary between two objects based solely on point cloud geometry. In these situations, image information provides crucial clues: Providing high-density boundary information: Even at the seams of two different containers, even those of similar color, shadows, varying surface textures, or subtle color variations often exist. Image sensors can capture these subtle variations at a resolution far higher than that of LiDAR. When these ROI pixels, carrying high-frequency texture and color features, are back-projected into 3D space, they create a clear, high-density feature boundary within the gaps in the sparse LiDAR point cloud.

[0074] Filling sensor holes: LiDAR beams may not receive effective echoes on some smooth surfaces due to the angle of incidence, resulting in holes in the point cloud. Cameras, however, are not subject to this limitation. Back-projected image features can effectively fill these geometric holes, making the cargo surface complete at the feature level. This prevents segmentation algorithms from mistakenly splitting a complete cargo container into two parts due to a point cloud hole in the middle.

[0075] Therefore, by back-projecting only the pixels in the ROI area, the present invention actually utilizes the semantic segmentation results of the image to accurately inject the most effective and dense image features (especially boundaries and surface textures) into the weakest areas of the LiDAR data in three-dimensional space, thereby greatly enhancing the feature distinguishability between different cargo instances, enabling subsequent segmentation models to find their boundaries more accurately.

[0076] S5.2.3 Selective modality convolution fusion module, which is used to dynamically and selectively fuse the voxel features of the two modalities. , its fusion characteristics The calculation method is: in: and are the feature vectors of LiDAR voxels and camera voxels respectively. Indicates concatenating two feature vectors. is the weight and bias of a linear layer, is the Sigmoid activation function, which together form a selection unit and output a weight between 0 and 1. Represents element-wise multiplication. This selection unit allows the network to learn the importance of camera features based on the spliced ​​features, scale them accordingly, and then add them back to the original LiDAR features to achieve effective complementarity of information. The final fused voxel feature is .

[0077] In step S5.2.3, the modality-selective convolutional fusion module does not simply add or concatenate point cloud and image features. Instead, it employs a dynamic, learnable selection mechanism. This design is based on adaptively adjusting the fusion weights of the two modal features based on the reliability of sensor information in different scenarios and spatial locations, thereby achieving robust and efficient information complementarity.

[0078] The advantage of point cloud features is that they provide accurate 3D geometric structure and depth information, but they are sparse and lack semantic information. The advantage of camera features is that they provide dense high-level semantic information such as color, texture, and symbols, but they lack accurate depth. In the fusion process, the reliability of the two is dynamically changing. Scene 1: A cargo box with a clean surface and clear QR code and text printed on it. In this case, camera features The value of is extremely high, it directly provides the key clues for identifying the content. Although the point cloud feature can outline the contour, the information it provides is relatively basic. At this time, the input of the selection unit ( and The splicing of will show a pattern with high semantic information, and the network will learn to output a larger weight value for this choice, thus In the middle, the camera features are significantly enlarged Scenario 2: A cargo box with a partially snow-covered surface and blurred labels. In this case, the camera features There is unreliable noise in the snow area and the information value provided in the label ambiguity area is also low. However, point cloud features It can still accurately capture the flat surface and precise size of the cargo box without being affected by the color of the snow. At this point, the selection unit will learn to output a smaller weight value when the input feature shows a pattern with low semantic information but a stable geometric structure. This will effectively suppress unreliable camera features. The contribution of More retention and reliance on original point cloud features , ensuring the stability of the results.

[0079] The advantages of this design in this application scenario lie in its strong environmental adaptability and robustness. In polar environments, sensor data degradation is common (e.g., labels covered by frost, point clouds missing due to reflective objects, etc.). A fixed fusion strategy (such as simple addition) cannot cope with this dynamic change and may perform well in one situation but be severely disrupted by noisy data in another. However, after training with a large amount of diverse data (including various data degradation scenarios), the selection mechanism can intelligently determine whether to trust point clouds or images at this time and location based on the quality of the input local features. This adaptive, fine-grained fusion strategy ensures that, regardless of the specific situation, the system maximizes the use of the most reliable signal sources and suppresses noisy ones, thereby achieving segmentation accuracy and robustness far exceeding those of static fusion methods in the complex and changing polar environment.

[0080] S5.2.4 Object feature enhancement module, this module is introduced to make the network pay more attention to the area containing objects.

[0081] S5.2.4.1 A parallel network based on the PointNet++ architecture is used as a point cloud classification module to classify the original point cloud. Processing is performed to predict an object score for each point, and point-by-point category information is obtained, where the category information includes a background category or an object category.

[0082] S5.2.4.2 The fused voxel features Compress to a bird's-eye view to generate a BEV feature map At the same time, the point-by-point category information output by the point cloud classification module is also converted into a foreground BEV map .

[0083] S5.2.4.3 uses the self-attention mechanism to enhance the BEV features. Convolutional layer Generate query, key, and value matrices: Attention Map and the final enhanced BEV characteristic map The calculation is as follows: in, Used to generate and , which means that attention will be focused on the object area identified by the point cloud; From the fusion , contains rich multimodal information, enabling the network to focus computing resources on the locations most likely to contain goods.

[0084] S5.2.5 Instance Segmentation Head The enhanced BEV feature map Input a detection head, such as CenterPoint's detection head, which predicts the 3D bounding box of each independent cargo instance. The final output is a set of A collection of bounding boxes , where each bounding box defines a physically separate cargo unit in which: The coordinates of the center point of the cargo unit, is the length, width, height and yaw angle of the cargo unit.

[0085] S5.3 Phase 2: Image-based content recognition and 3D verification. In this phase, detailed content analysis is performed on each segmented cargo instance.

[0086] S5.3.1 ROI-Instance Association and Content Identification S5.3.1.1 For each segmented 3D bounding box , projecting it back into the video frame On the image plane, we get a two-dimensional box .

[0087] S5.3.1.2 In this two-dimensional frame ROI area obtained by preprocessing In the intersection of , a multi-task recognition model is executed, which performs the following operations in parallel: The optical character recognition branch extracts textual information from packaging. The barcode / QR code decoding branch reads machine-readable labels. The symbol classification branch recognizes internationally recognized logistics or hazardous materials symbols (such as the red cross and flame).

[0088] S5.3.1.3 By fusing these recognition results, the system generates Determine its content category (Examples: food, fuel, medical supplies, etc.).

[0089] S5.3.2 3D geometry verification For each 3D bounding box ,from Extract all points that fall into the box to form a point cloud subset Using the Convex Hull algorithm, according to the point cloud subset Calculate the actual three-dimensional volume of the cargo If text information about size or volume is obtained during OCR recognition in S5.3.1 , then calculate a consistency score : like Below the preset threshold , the cargo instance is marked as pending review to indicate that there may be identification errors or the packaging does not match the contents.

[0090] Next, this example details the method for generating a three-dimensional digital map of the unloading area and implementing multi-dimensional dynamic warehousing planning. The goal of this stage is to integrate the discrete recognition results into a structured global map and, based on this map and multi-dimensional dynamic factors, generate an optimal, adaptive cargo warehousing and transportation plan.

[0091] S6.1 Generation of 3D digital map of unloading area The system has obtained a A detailed list of individual cargo instances, including the 3D bounding box for each instance , content category This information is integrated into a structured, queryable three-dimensional digital map , serving as the basis for all subsequent planning and scheduling.

[0092] S6.1.1 Map data structure definition 3D digital map is constructed as a A collection of objects, . Each object Corresponds to a cargo instance and contains the following key fields: ID: unique identifier of the cargo instance, Pose: The pose information of the cargo instance, that is, the obtained three-dimensional bounding box Category: Content category of the cargo instance Status: The status of the cargo instance, combined with the verification results Divided into verified, pending review, etc. Accessibility: The accessibility score of the goods , a value between 0 and 1 that quantifies the physical difficulty of obtaining the good.

[0093] S6.1.2 Build accessibility assessment model and accessibility score It is not directly measured, but calculated by analyzing the spatial stacking and occlusion relationship between goods. , its accessibility score : in, Indicates goods right The support relationship. The bottom projection area of ​​the 3D bounding box is The top projection areas of The Z coordinate center value is greater than hour, , otherwise it is 0. This item indicates that the The quantity of goods above. Indicates goods right The system is based on the path blocking relationship. The orientation angle Determine its optimal grasping surface and extend a virtual grasping channel from the surface. The bounding box of intersects the channel, then , otherwise it is 0. This item indicates that the block The amount of cargo ahead. and It is the weight coefficient of the support relationship and the blocking relationship, which is used to adjust the impact of different constraints on the accessibility score.

[0094] score A score close to 1 indicates that the cargo is exposed and easy to grab; a score close to 0 indicates that the cargo is deeply buried or severely obstructed. This score calculation allows the map to record not only what is available, but also how easy it is to grab.

[0095] S6.2 Multi-dimensional dynamic inventory planning is the decision-making core of the cloud system. It generates a three-dimensional digital map based on , and combines four categories of dynamic information to generate the optimal warehousing task sequence for the unmanned vehicle fleet.

[0096] S6.2.1 Multidimensional dynamic information of planning input 1. Cargo attribute information: directly from the map , including the category of each cargo and accessibility .

[0097] 2. Task priority information: The cloud system maintains a task priority list for each cargo category. Defining an Urgency Score For example, medical supplies The highest value is food, followed by equipment spare parts.

[0098] 3. Real-time environmental information: The system accesses external sensor data in real time to form an environmental state vector ,in is the wind speed, For visibility, is the ambient temperature.

[0099] 4. System resource information: the number of currently available autonomous vehicles and their respective battery status.

[0100] S6.2.2 Building a Dynamic Task Prioritization Model Inventory planning is not a one-time fixed plan, but a continuous and dynamic decision-making process. At each decision moment (for example, when an unmanned vehicle completes its mission and becomes idle), the cloud system will check all currently reachable (i.e. Greater than a certain threshold ) of goods Calculate a composite task priority score : in, The value of the shipment is determined by its urgency. Shipments with higher priority receive higher scores. : Opportunity gain is determined by accessibility. Easier-to-reach items receive higher scores, encouraging the system to prioritize easier items and then move on to more challenging ones, quickly clearing the periphery. is the environmental risk cost. To evaluate the current environment Moving goods For example, when the wind speed When very high, carry a tall and light load ( big, The risk function value of the Shipping cost. This function Estimate the goods The energy consumption or time required to transport it from its current location back to the unmanned storage room is mainly related to volume and category. And the transportation distance. is the weight coefficient.

[0101] S6.2.3 Task Scheduling and Execution. At each decision moment, the cloud system performs the following operations: Update real-time environment information , calculate the task priority score for all reachable goods . Press the task Sort from high to low to generate a dynamic task queue. Assign the task at the top of the queue to an idle unmanned vehicle. Once a cargo is taken away by an unmanned vehicle, the system will immediately remove it from the map. Remove the cargo instance and recalculate the accessibility scores of the surrounding cargo affected by it (supported or blocked by it) , thereby dynamically updating the state of the entire map.

[0102] The embodiments of this specification also propose an artificial intelligence-based polar unmanned warehouse cargo identification and inventory system, which includes: Collection module: at least one unmanned vehicle collects data on the cargo in the unloading area to obtain video data and point cloud data of the cargo; at least one drone acts as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to the cloud system; 3D digital map generation module: The cloud system pre-processes the received point cloud data and video data, identifies regions of interest corresponding to the cargo surface in the video data, performs multimodal fusion based on the features of the clean cargo surface point cloud and the video data, and achieves cargo instance segmentation, thereby obtaining a 3D bounding box for each independent cargo instance. Combining the image information within the region of interest, the content category of each cargo instance is identified, and a 3D digital map of the unloading area is generated. Dynamic task queue generation module: Based on the three-dimensional digital map, combined with task priority information, real-time environmental information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.

[0103] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned artificial intelligence-based polar unmanned warehouse cargo identification and inventory method is implemented.

[0104] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned artificial intelligence-based polar unmanned warehouse cargo identification and inventory method.

[0105] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0106] In this specification, the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments described later, the description is relatively simple, and the relevant parts can be referred to the partial description of the previous embodiments.

[0107] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A polar unmanned warehouse cargo identification and inventory method based on artificial intelligence, characterized in that: The method comprises: At least one unmanned vehicle collects data on cargo in the unloading area to obtain video data and point cloud data of the cargo; At least one drone acts as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to the cloud system; The cloud-based system pre-processes the received point cloud data and video data to identify regions of interest corresponding to the cargo surface in the video data; performs multimodal fusion based on the features of the clean cargo surface point cloud and the video data to segment cargo instances, thereby obtaining a three-dimensional bounding box for each independent cargo instance; and identifies the content category of each cargo instance by combining image information within the region of interest; and generates a three-dimensional digital map of the unloading area. Based on the three-dimensional digital map, combined with task priority information, real-time environmental information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.

2. The method for identifying and counting goods in unmanned storage in polar regions based on artificial intelligence according to claim 1, characterized in that: The data collection of the cargo in the unloading area by the at least one unmanned vehicle includes: rasterizing the unloading area to construct an environmental map model, and providing each unmanned vehicle with a Plan your route , the path planning is performed by optimizing the following objective function accomplish: in, is the information gain of the path, is the execution cost of the path, is the overlap cost of the path, are the weight coefficients of information gain, execution cost and overlap cost respectively.

3. The method for identifying and counting goods in unmanned polar warehouses based on artificial intelligence according to claim 2 is characterized in that: At least one drone acting as a data relay includes: Before the UAV departs from the unmanned storage room, a feasibility assessment of the single-flight full transmission mode is conducted, including the calculation of the total power required for the mission. : in, is the one-way flight time of the drone, and are the battery consumption rates per unit time of the drone in flight and hovering states, The estimated remaining time required for the unmanned vehicle fleet to complete data collection, Estimated time required to transfer all data. It is a safe redundant power supply; if and only if the current power of the UAV meets the conditions, the single-flight full transmission mode will be executed.

4. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 3, characterized in that: The step of identifying a region of interest corresponding to the surface of the cargo in the video data comprises: Convert the RGB image frame in the video data to the HSV color space and create a binary snow mask using the following decision logic: : in, are the hue, saturation, and lightness values ​​of the pixel points, is a preset threshold parameter for defining white; and the snow mask is logically inverted to obtain the region of interest.

5. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 4, characterized in that: Process the surface point cloud of the clean cargo to obtain LiDAR voxel features , and the video data is combined with the region of interest to obtain the camera voxel features ; The two are fused by selecting the modal convolution fusion module to obtain the fused voxel features : in, and are the feature vectors of the LiDAR voxel and camera voxel at the same spatial position, Represents a splicing operation, is the Sigmoid activation function, is the weight and bias of the selection unit, This is element-wise multiplication.

6. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 5, characterized in that: The step of generating a three-dimensional digital map of the unloading area further includes: for each cargo instance in the map Calculating an accessibility score : in, To determine the goods right The support relationship, To determine the goods right The path blocking relationship, and are the weight coefficients of the support relationship and the blocking relationship respectively.

7. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 6, characterized in that: The steps of multi-dimensional dynamic warehousing planning specifically include: for each currently accessible cargo instance Calculate a composite task priority score : in, Score the urgency of the cargo category, To combine real-time environmental information Calculated environmental risk costs, For transportation costs, is the weight coefficient corresponding to each item; and the tasks are sorted according to the task priority scores to generate the dynamic task queue.

8. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 7, characterized in that: The step of segmenting the cargo instances further includes an object feature enhancement step of enhancing the fused features, and the object feature enhancement step specifically includes: The multimodal fused features are converted into a bird's-eye view to generate a BEV feature map , and use the point cloud classification module to obtain the foreground BEV map ; The BEV feature map is enhanced by a self-attention mechanism, which calculates the attention map in the following way and enhanced BEV characteristic map : in, are query, key, and value matrices respectively, is a convolutional layer; and based on the enhanced BEV feature map Obtain the three-dimensional bounding box.

9. The artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to claim 8, characterized in that: After identifying the content category of each cargo instance, 3D geometry verification is also included: , according to its 3D bounding box Extract the corresponding point cloud subset from the clean cargo surface point cloud , and calculate its three-dimensional volume ; If text information about volume is identified from the image information in the region of interest , the consistency score is calculated by the following formula : And according to the consistency score Determine whether the content category identification result is consistent with the physical size of the cargo instance.

10. An artificial intelligence-based polar unmanned warehouse cargo identification and inventory system, the system being used to execute the artificial intelligence-based polar unmanned warehouse cargo identification and inventory method according to any one of claims 1 to 9, characterized in that: The system includes: Collection module: at least one unmanned vehicle collects data on the cargo in the unloading area to obtain video data and point cloud data of the cargo; at least one drone acts as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to the cloud system; 3D digital map generation module: The cloud system pre-processes the received point cloud data and video data, identifies regions of interest corresponding to the cargo surface in the video data, performs multimodal fusion based on the features of the clean cargo surface point cloud and the video data, and achieves cargo instance segmentation, thereby obtaining a 3D bounding box for each independent cargo instance. Combining the image information within the region of interest, the content category of each cargo instance is identified, and a 3D digital map of the unloading area is generated. Dynamic task queue generation module: Based on the three-dimensional digital map, combined with task priority information, real-time environmental information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned storage room.

Citation Information

Patent Citations

  • Garbage pickup robot based on visual semantic SLAM (simultaneous localization and mapping)

    CN111360780A

  • Method and system for remotely and intelligently detecting goods in storage space

    CN115147748A

  • Container identification method based on image and point cloud combination

    CN115661064A

  • Cooperative operation system and method for unmanned aerial vehicle and unmanned loader

    CN117075629A

  • Multi-machine cooperation method and system for multimodal transport cargo receiving, unloading and transferring

    CN120410358A

Cited By

  • Pallet goods identification method and device and storage medium

    CN121640040A

  • Electronic warehouse receipt cargo authenticity verification method based on multi-mode AI

    CN121883045A