Artificial intelligence-based polar unmanned warehouse goods identification and inventory method and system

CN120822909BActive Publication Date: 2025-12-23POLAR RES INST OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511325065.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-23
Estimated Expiration
2045-09-17

Smart Images

  • Figure CN120822909B_ABST
    Figure CN120822909B_ABST
Patent Text Reader

Abstract

The application provides a polar unmanned warehouse goods identification and inventory method and system based on artificial intelligence, and relates to the fields of artificial intelligence and computer vision. The identification and inventory method specifically comprises the following steps: at least one unmanned vehicle collects data of goods in an unloading area to obtain video data and point cloud data of the goods; at least one unmanned aerial vehicle serves as a data relay and transmits the video data and point cloud data collected by the unmanned vehicle to a cloud system; multi-modal fusion is performed based on the features of pure goods surface point cloud and the video data to realize goods instance segmentation, and the content category of each goods instance is identified in combination with image information in the region of interest; a three-dimensional digital map of the unloading area is generated; and based on the three-dimensional digital map, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned warehouse room.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automated warehousing, in particular to an unmanned intelligent logistics management technology applied in extreme environments. More specifically, the present application relates to a method and system for realizing the automated identification, scheduling planning, transportation into a warehouse and inventory of large quantities of goods in harsh environments such as the polar regions by using artificial intelligence and computer vision. BACKGROUND

[0002] With increasing global attention to scientific research and resource development in polar regions, there is an increasing demand for the establishment of long-term or semi-permanent scientific research stations and bases in polar regions. The efficient operation of these sites is highly dependent on stable and reliable logistics support. However, the environment in polar regions is extremely harsh, with typical characteristics including ultra-low temperature (as low as several tens of degrees below zero), long periods of blizzards, polar night phenomenon, and unpredictable ice conditions. This extreme natural environment poses a revolutionary challenge to traditional material management and warehousing operation modes.

[0003] In conventional environments, automated warehousing technology has made great progress. Large logistics centers generally use automated guided vehicles, stackers, shuttles, and warehouse management systems to achieve efficient and accurate in-out warehouse of goods. These systems usually operate in ideal conditions such as indoors, temperature control, flat ground, and stable network signals, and their navigation methods rely on two-dimensional codes, magnetic strips laid on the ground, or laser radar simultaneous localization and mapping (SLAM) in structured environments. At the same time, the identification and inventory of goods also highly depend on pre-set identifiers such as barcodes and radio frequency identification tags, as well as stable and high-speed internal networks for data transmission and processing.

[0004] However, when these existing technologies are directly migrated to polar application scenarios, their inherent defects are exposed. First, the material supply in polar regions usually adopts a window period centralized unloading mode, i.e., a transport ship unloads several months or even a year's worth of materials on the ice or a simple site at a certain distance from the camp during a short navigation period. This results in a large number of goods being stacked disorderly outdoors, completely lacking the structured environment of a conventional warehouse, and traditional AGV navigation and positioning technologies are almost completely ineffective in such dynamic, open, and sparse snow environments. Second, the polar environment poses a severe challenge to equipment and personnel. Long periods of low temperature can cause electronic components to degrade in performance, battery life to decrease sharply, and metal materials to become brittle. Snowy weather not only severely affects the performance of optical sensors, leading to low-quality or even interrupted data collection, but also makes manual outdoor work extremely risky and inefficient. Manual inventory and handling not only consume time and effort, but also pose life-threatening risks such as frostbite and getting lost, which cannot meet the instant and continuous requirements of scientific research stations for material support.

[0005] More importantly, the reliability of data communication. In the vast polar region, stable long-distance, high-bandwidth wireless communication is a luxury. Severe weather such as snowstorms can seriously interfere with radio signals, making it extremely difficult and unreliable to transmit large amounts of data (such as high-definition video streams) in real time from the outdoor unloading area to the indoor control center. This information island effect prevents the management system from obtaining timely and accurate global information about the front-end goods, such as type, quantity, and specific location, making subsequent warehouse planning and scheduling impossible. If the core problem of seeing, transmitting, and calculating accurately in outdoor harsh environments cannot be effectively solved, the automation and intelligence of the entire warehouse management will be nothing more than an empty talk.

[0006] Therefore, the prior art cannot provide a complete, automated, and intelligent material identification and warehousing solution that can adapt to the entire process from the outdoor unloading area to the indoor warehouse in the polar region. In the face of outdoor dynamic and disordered goods layout, the severe constraints of extreme weather on equipment and communication, and the complex decision-making needs of material warehousing, a new technical solution is urgently needed to break through the application bottleneck of traditional warehouse technology in extreme environments. SUMMARY

[0007] The present application provides an artificial intelligence-based polar unmanned warehouse goods identification and inventory method, which specifically comprises the following steps:

[0008] At least one unmanned vehicle collects data on the goods in the unloading area to obtain video data and point cloud data of the goods;

[0009] At least one unmanned aerial vehicle relays the video data and point cloud data collected by the unmanned vehicle to a cloud system;

[0010] The cloud system pre-processes the received point cloud data and video data, identifies the region of interest corresponding to the surface of the goods in the video data, and performs multi-modal fusion based on the features of the pure goods surface point cloud and the video data to achieve goods instance segmentation, thereby obtaining a three-dimensional bounding box for each independent goods instance. In combination with the image information in the region of interest, the content category of each goods instance is identified; and a three-dimensional digital map of the unloading area is generated;

[0011] Based on the three-dimensional digital map, in combination with task priority information, real-time environmental information, and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned warehouse room.

[0012] The present application also provides an artificial intelligence-based polar unmanned warehouse goods identification and inventory system, which comprises:

[0013] The collection module: at least one unmanned vehicle collects data of goods in the unloading area to obtain video data and point cloud data of the goods; at least one unmanned aerial vehicle serves as a data relay and transmits the video data and point cloud data collected by the unmanned vehicle to a cloud system;

[0014] The three-dimensional digital map generation module: the cloud system pre-processes the received point cloud data and video data, identifies a region of interest corresponding to a surface of the goods in the video data, performs multi-modal fusion based on a pure goods surface point cloud and features of the video data to realize goods instance segmentation, thereby obtaining a three-dimensional bounding box of each independent goods instance, identifies a content category of each goods instance in combination with image information in the region of interest, and generates a three-dimensional digital map of the unloading area;

[0015] The dynamic task queue generation module: based on the three-dimensional digital map, in combination with task priority information, real-time environment information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned warehouse room.

[0016] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned polar unmanned warehouse goods identification and inventory method based on artificial intelligence when executing the computer program.

[0017] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned polar unmanned warehouse goods identification and inventory method based on artificial intelligence.

[0018] Compared with the prior art, the polar unmanned warehouse goods identification and inventory method and system based on artificial intelligence have significant beneficial effects.

[0019] The application constructs a complete full-process unmanned polar material guarantee system, solves the fundamental problem of high risk, low efficiency and even inability of manual operation in the polar harsh environment. The application deeply couples and cooperates unmanned vehicles, unmanned aerial vehicles and cloud intelligent systems: the front end is self-operated and comprehensive data collection by unmanned vehicles adapted to the polar environment, which avoids the danger of personnel outdoor operation; the data transmission link adopts unmanned aerial vehicles as air mobile relay, overcomes the bottleneck of unreliable long-distance wireless communication in polar blizzards, and ensures that the massive high-dimensional data can be safely and completely returned; the back end is completely analyzed, planned and dispatched by the cloud system. This complete closed loop of ground collection-air relay-cloud decision forms an intelligent solution that can run long-term, stably and autonomously, and the systematic innovative design is unmatched by simple combination of existing single technologies or devices.

[0020] In the core link of cargo identification, the application proposes a multi-modal fusion perception method, which greatly improves the segmentation accuracy of independent cargo instances in disordered and densely stacked scenes. Traditional methods are difficult to distinguish closely adjacent or stacked goods, while the application ingeniously utilizes the complementary advantages of image and point cloud data. By preprocessing the video data, the region of interest (ROI) of the cargo surface is accurately located, and this prior information is used to selectively and accurately enhance the features containing rich texture and boundary details in the image to the three-dimensional space, and deeply fuse with the accurate geometric structure provided by the point cloud. This strategy of guiding three-dimensional geometric analysis with image semantic information can significantly enhance the feature distinguishability between different cargo instances, especially in the sparse joint or occluded area of point cloud data, which can realize accurate boundary definition and solve the problem of instance segmentation in the industry.

[0021] The application also introduces a consistency verification mechanism based on three-dimensional geometric information, which provides double protection for the reliability of the identification result. After identifying the cargo content category through image information, the system calculates the actual physical volume of the cargo using point cloud data and compares it with the volume information identified from the image text. This cross-modal cross-validation method can effectively find and mark out identification errors or inconsistencies between packaging and content, significantly improving the accuracy and reliability of the inventory result, which is crucial for ensuring the safety of polar long-term tasks.

[0022] The design of the unmanned aerial vehicle data relay itself is an engineering solution with great robustness and practical value in extreme communication environment. At the same time, the system can dynamically generate and update the three-dimensional digital map of the unloading area in real time. The map not only contains the location and category of each cargo, but also introduces the accessibility evaluation, digitizes the physical stacking relationship on site, and provides unprecedented decision basis for subsequent multi-dimensional dynamic warehousing planning, so that the entire warehousing process realizes true intelligence and optimization. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0024] Figure 1 It is a flowchart of the polar unmanned warehouse cargo identification and inventory based on artificial intelligence of the present application. DETAILED DESCRIPTION

[0025] The embodiments of the present application will be described in detail below with reference to the drawings.

[0026] The embodiments of the present application will be described in detail below with reference to the drawings.

[0027] It should be noted that the various aspects of the embodiments described below are within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms and that any specific structure and / or function described herein is merely illustrative. Based on the teachings herein one skilled in the art should appreciate that an aspect described herein can be implemented independently of any other aspects and that an aspect described herein can be implemented both as any number of devices and / or as any number of methods. In addition, the various embodiments described herein can be implemented as any number of software running on any number of hardware devices.

[0028] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, one skilled in the art will understand that the examples can be practiced without these specific details.

[0029] In order to better understand the technical solutions of the present application, the present embodiment first describes the scene and system composition to which the present application is applied. The method and system designed by the present application are intended to perform work in a typical polar environment such as the North Pole or the South Pole. The significant feature of such an environment is that the climate is extremely harsh, the temperature is low all year round, and the weather is often accompanied by long duration and high intensity blizzards, which makes it difficult for humans to work outdoors for a long time, and also puts high requirements on the stability of the equipment and the reliability of the communication.

[0030] In the application scenario of the present application, the supply of materials depends on the transport ship with icebreaking capability to be completed in the short window period of relatively stable weather conditions every year. In order to avoid the risk of being surrounded by floating ice or encountering sudden severe weather, the transport ship needs to unload the materials quickly after arrival. The work site is set as a designated area away from the main building group of the scientific research station or base, which is relatively flat and open, and this area is referred to as the unloading area in the present application. The transport ship will unload large-scale materials such as survival, scientific research and equipment spare parts for future months or even a whole year at one time on the ice or land in the unloading area. Due to the urgent unloading operation, the goods in the unloading area are in an initial disordered or semi-ordered stacking state, and different types and sizes of boxes are mixed and placed without forming a fixed and directly identifiable storage layout. After completing the unloading operation, the transport ship will quickly leave, leaving these materials waiting for subsequent transfer and storage.

[0031] The system deployment of the present application fully considers the particularity of the above-mentioned scenario. A unmanned storage room is built in the scientific research station or base as the final storage and management center of all materials. There is a physical distance between the unmanned storage room and the unloading area. The core computing and control unit of the system, i.e. the cloud system, is deployed inside the unmanned storage room to ensure that it runs in a stable and safe environment. The cloud system includes high-performance servers, data storage arrays and software platforms running artificial intelligence algorithms.

[0032] To execute subsequent operational processes, the system's hardware was deployed in a distributed manner. Multiple unmanned vehicles (UAVs) specifically designed for polar environments were pre-placed or deployed in the unloading area. These UAVs possess low-temperature adaptability, off-road capability, and a variety of sensor configurations. Simultaneously, at the unmanned storage facility base, multiple drones, similarly designed to withstand low temperatures and strong winds, were deployed. This spatially separated deployment of UAVs and drones is the fundamental setting of this invention for the long-distance, unreliable communication environment of the polar regions, constituting a prerequisite for the implementation of all subsequent technical solutions of this invention.

[0033] Next, this embodiment details the specific method for unmanned vehicles to perform ground visual information collection in the unloading area according to the present invention. In this stage, multiple unmanned vehicles work collaboratively without human intervention to complete a comprehensive, multi-angle scan of all goods in the entire unloading area, generating high-quality raw data required for subsequent analysis.

[0034] Upon receiving the activation command from the cloud system inside the unmanned warehouse, multiple unmanned vehicles pre-deployed in the unloading area immediately started up. To achieve efficient and complete exploration of the unknown and disordered environment, the unmanned vehicle fleet executed a cooperative coverage path planning algorithm based on multi-agent boundary exploration.

[0035] S2.1 Rasterizes the entire unloading area on a two-dimensional plane to construct an environmental map model. ,in It is a collection of grid cells, representing discrete location points. It is the set of edges that connect adjacent grid cells.

[0036] The S2.2 unmanned vehicle fleet autonomously plans its own travel paths. To maximize information gain and minimize total cost. Each autonomous vehicle Path planning can be modeled as an optimization problem, with the objective function shown below:

[0037]

[0038] in, Represents driverless cars A candidate path is a sequence of connected grid points. It is the comprehensive evaluation value of the candidate path. The information gain of the path is defined as the path's information gain. The number of newly observed unknown grid cells across all observation points. This factor incentivizes the autonomous vehicle to move toward unknown areas (i.e., map boundaries). This represents the execution cost of the path, which can be either the total length of the path or the estimated energy consumption. The cost of path overlap is used to measure the path's cost. Collection of planned routes from other autonomous vehicles The degree of overlap. This is introduced to avoid multiple autonomous vehicles repeatedly exploring the same area, thereby improving the overall collaborative efficiency of the fleet. These are the weighting coefficients for information gain, execution cost, and overlap cost, respectively. They are positive real numbers and can be preset or dynamically adjusted by the cloud system according to task priority and environmental conditions to balance exploration efficiency and energy consumption.

[0039] S2.3 During operation, each autonomous vehicle independently calculates the path evaluation value of its surrounding boundary points. The autonomous vehicle fleet can dynamically and collaboratively scan the entire unloading area by periodically exchanging information (e.g., broadcasting their chosen paths).

[0040] S2.4 along the planned path While in motion, each autonomous vehicle collects data using its onboard multi-sensor fusion module. This module includes at least one high-resolution panoramic or wide-angle camera and one solid-state LiDAR. The camera is used to acquire high-resolution video data streams, denoted as... ,in This represents a timestamp. Simultaneously, the LiDAR performs a 3D spatial scan, generating point cloud data describing the surrounding environment and the geometry of the cargo, denoted as... All sensor data undergoes rigorous clock synchronization to ensure precise alignment of video frames and point cloud data in terms of timestamps, providing a foundation for subsequent data fusion and 3D reconstruction.

[0041] The S2.5 system calculates in real time the ratio of the number of observed graticules to the total number of graticules in the unloading area, i.e., the coverage rate. When coverage Reaching the preset threshold At that time, the data collection task was completed.

[0042]

[0043] in, It is the number of grid cells that are currently effectively covered by at least one autonomous vehicle sensor. It is the unloading area. The total number of grid cells. Each autonomous vehicle locally stores multi-angle, high-density video data of its assigned area. and point cloud data This has created a comprehensive and redundant digital record of the cargo layout in the unloading area, awaiting the next stage of data transmission.

[0044] Next, the embodiment details the specific method of the unmanned aerial vehicle performing air data relay in the present application, which safely and reliably transmits the large-capacity data collected by the unmanned vehicle from the unloading area to the unmanned warehouse room.

[0045] S3.1 While the unmanned vehicle fleet is collecting data, periodically send status information, including the amount of data collected, the remaining power, the current coverage , etc. to the cloud system of the unmanned warehouse room through a low-power, long-distance narrowband communication link.

[0046] S3.2 The cloud system starts the unmanned aerial vehicle scheduling strategy decision model based on the received status information to determine whether to use a single flight complete transmission mode or a multi-flight continuous transmission mode.

[0047] S3.2.1 The decision model assesses the feasibility of the single flight mode, and determines whether there is at least one unmanned aerial vehicle in the current hangar whose battery consumption and data storage for a single task can meet the task requirements. The feasibility of a single task must meet the following two conditions:

[0048] Storage capacity condition determination, the available data storage space of the unmanned aerial vehicle must be greater than or equal to the estimated total amount of data to be collected in the unloading area . . The total number of grids in the unloading area and an empirical data density coefficient are multiplied to obtain, that is .

[0049]

[0050] Battery endurance condition determination, the current battery level of the unmanned aerial vehicle must be sufficient to support it to complete the entire task flow, including flying to the unloading area, hovering to wait and receiving all data, and finally returning. The total amount of power required is calculated as follows:

[0051]

[0052] Among them, is the one-way flight time of the unmanned aerial vehicle, determined by the distance between the unmanned warehouse room and the unloading area and the average cruising speed of the unmanned aerial vehicle under the current weather . . and are the battery consumption rates per unit time in the flight and hovering states of the unmanned aerial vehicle, respectively. ​This is the estimated remaining time required for the autonomous vehicle fleet to complete data collection. It is to include all data Estimated time required to transfer data from the driverless car to the drone. It is a safety redundancy power supply set up to deal with emergencies.

[0053] S3.2.2 The cloud system searches through all available drones. If it finds any drone that simultaneously meets the above storage and battery requirements, it initiates the first scenario: single-flight full transfer mode. In this mode, the system will assign that drone... The drone will wait for the unmanned vehicle fleet to complete all data collection tasks. Only after receiving the signal from the unmanned storage room did it take off. Upon arriving over the unloading area, it received all the video data stored by all the unmanned vehicles in one go via a high-bandwidth short-range communication link. and point cloud data After receiving the data, the drone immediately returns to base and uploads the data to the cloud system upon arrival at the unmanned storage room.

[0054] If the cloud system's evaluation finds that no single drone can meet the endurance or storage requirements for a single mission, the second scenario will be automatically initiated: multi-drone relay transmission mode. This mode aims to use a piecemeal approach, with multiple drones collaboratively transferring data in batches. The scheduling algorithm for multi-drone relay transmission mode is as follows:

[0055] 1) The cloud system assigns the first available drone. It immediately flew to the unloading area and, upon arrival, began downloading some of the data already collected from the driverless vehicle.

[0056] 2) In During the mission, the cloud system continuously monitors two key states: the amount of data loaded. and remaining battery power .

[0057] 3) If and only if any of the following conditions are met, The return-to-home condition is triggered: (1) Storage space is about to run out: ,in (1) A small storage margin; (2) The remaining power reaches the return threshold: This ensures that the aircraft has enough power to return safely.

[0058] 4) Cloud system based on drones Based on the status and data transmission rate, the return-to-home trigger time is estimated. To achieve seamless integration, the cloud system will calculate and assign the next drone. Best departure time : Ensure the succession of drones Able to be in the previous drone The vehicle arrives precisely when it is ready to leave the unloading area, minimizing potential operational interruptions caused by the unmanned vehicle waiting for data transfer.

[0059] 5) Steps 2)-4) The scheduling process repeats cyclically, and the drone... They were successively dispatched until the unmanned vehicle convoy completed its data collection mission. ), and the last batch of collected data was used by the first The drone was successfully retrieved and uploaded to the cloud system. With this, the entire data relay process was completed.

[0060] Next, this embodiment will describe in detail how the cloud system processes the received raw point cloud data. and video data Preprocessing is performed.

[0061] S4.1 Point Cloud Data Preprocessing

[0062] Raw point cloud data transmitted from drones to cloud systems It includes various elements such as cargo, ground, attached snow, and falling snow. To accurately reconstruct the 3D model of the cargo, the following cleanup process must be performed:

[0063] S4.1.1 Noise Point and Outlier Filtering

[0064] Snowflakes drifting in the air create numerous sparse, isolated noise points during lidar scanning. This invention employs a statistical outlier removal algorithm to filter out this noise. For point clouds... any point in Calculate its distance to the nearest The average distance of each neighboring point Calculate the global mean. and standard deviation A point is considered an outlier and is removed when the average neighborhood distance of that point meets the following conditions:

[0065]

[0066] in: It represents the number of neighboring points in the outlier removal algorithm. This is the standard deviation threshold, used to control the strictness of the filtering. Through this step, we obtain the point cloud after preliminary filtering out aerial noise. .

[0067] S4.1.2 Ground Segmentation

[0068] To separate the goods from the ground, the present invention uses the random sample consensus algorithm to fit a ground plane model. The random sample consensus algorithm iteratively randomly samples a minimal point set (preferably 3 points) from the point cloud to fit a plane equation and calculates the distance of other points in the point cloud to the plane. Points with distance less than a threshold are considered inliers. After multiple iterations, the plane model with the most inliers is determined as the ground. The inlier set belonging to this ground model will be separated from the point cloud, and the remaining non-ground point set mainly contains various goods.

[0069] S4.1.3 Snow surface and goods surface separation

[0070] Snow attached to the goods surface has different geometric and physical properties from the goods box. The present invention uses a region growing algorithm based on physical features to achieve accurate separation of the two.

[0071] S4.1.3.1 For each point in the non-ground point set , calculate its multi-dimensional physical features, mainly including laser radar reflection intensity and local surface normal vector calculated based on neighborhood points .

[0072] S4.1.3.2 Seed point selection. Traverse all points, and points that meet high reflection intensity local area height flat (neighborhood normal vector variance less than ) are selected as seed points, which are most likely located on the surface of the artificial goods box.

[0073] S4.1.3.3 Starting from all seed points, region growing is performed in parallel. A growing region will try to absorb its adjacent unassigned points. An adjacent point is incorporated into the current region, which must satisfy , where: is the average normal vector of the current growing region. is the normal vector of the adjacent point to be examined. is the normal vector similarity threshold, which is close to 1, indicating that only points with very similar normal vector directions (i.e. good coplanarity) can be merged.

[0074] S4.1.3.4 The growing process continues until no point can be added to any region. The set of all grown regions is considered as the final pure goods surface point cloud . And the point set ​The remaining points that are not merged by any region are classified as irregular snow surface and are removed.

[0075] S4.2 Video data preprocessing

[0076] For the video data stream collected by the unmanned vehicle , in order to effectively separate the snow-covered region and the box surface containing information, the present application adopts a threshold segmentation and morphological processing method based on color space.

[0077] S4.2.1 Snow region segmentation based on HSV color space

[0078] Snow has a stable color feature of low saturation and high brightness in vision.

[0079] Convert each frame of RGB image to HSV color space. A binary snow mask is created by setting a pre-defined threshold range. For any pixel point in the image, the decision logic is as follows:

[0080]

[0081] wherein: H, S and V are the hue, saturation and lightness values of the pixel point , respectively. is a pre-set threshold parameter for defining white. 1 represents that the pixel belongs to snow, and 0 represents that it does not belong to snow.

[0082] S4.2.2 Mask optimization and region of interest (ROI) extraction

[0083] The binary snow mask may have noise (such as isolated white points or small internal cavities). Therefore, morphological operation is performed on the binary snow mask for optimization.

[0084] The open operation is used to remove small noise points, and then the close operation is used to fill the internal cavities of the snow region, so as to obtain a smoother and more complete snow region mask .

[0085] By logically inverting the mask, the region of interest (ROI) required can be obtained. The ROI accurately identifies all non-snow box surface regions in the image that may contain cargo information. The extracted ROI set is denoted as , which includes the ROI region image and its corresponding image frame and position information in the image frame.

[0086] Next, this embodiment elaborates on the analysis of the pre-processed data by the cloud system to achieve the segmentation, content recognition and quantity statistics of the independent goods in the unloading area. This process is completed through a phased hierarchical multi-modal goods perception network. The primary task of this network is to use the fusion information of point cloud and image to solve the instance segmentation problem in physical space, that is, to distinguish closely stacked or adjacent goods; the secondary task is to use the semantic information of the image for content recognition, and use the geometric information of the point cloud for cross verification.

[0087] S5.1 Hierarchical multi-modal goods perception network overall architecture construction

[0088] HMCP-Net contains two parallel processing branches (point cloud branch and image branch) and a subsequent identification verification module.

[0089] S5.2 First stage: goods instance segmentation based on multi-modal fusion, the first stage aims to generate an accurate three-dimensional bounding box for each independent container in the unloading area.

[0090] S5.2.1 Parallel feature extraction: the input of the point cloud branch is the purified surface point cloud of the goods . It is converted into a three-dimensional voxel grid by a voxelization encoder, and each non-empty voxel contains the statistical features of the points inside it, obtaining point cloud voxel features . The input of the image branch is the original video frame . Deep feature maps are extracted by ResNet , where is the height and width of the feature map, is the number of channels.

[0091] The input of the present invention is still the original video frame, because the convolutional neural network extracts image features is not to analyze a single pixel in isolation, but to perceive the pattern in a region through a convolution kernel with a certain size. A pixel located inside the ROI, its final deep features extracted, is determined by a large number of pixels including it and its periphery (possibly outside the ROI). If only the cropped, irregular ROI region is input into the network, a large number of false boundaries will be artificially created at the edges of the goods, which will seriously interfere with the convolution process, resulting in low-quality and distorted information of the extracted goods edge features. Inputting the complete frame allows the network to see the natural transition of the goods and the surrounding environment (such as snow and ground), thereby learning more discriminative and high-quality boundary features, which is crucial for subsequent distinguishing of closely adjacent or stacked goods.

[0092] In addition, the light, weather conditions (such as sunny, cloudy, light snow, ground reflectivity) of the polar environment change dramatically. Taking the entire original video frame as input enables the network to perceive global environmental information at an early stage of feature extraction. For example, the network can adjust the feature extraction strategy for the cargo area from the brightness, color temperature, and other information of the sky and distant view, thereby having stronger robustness to changes in light and contrast. This global context awareness capability, which cannot be obtained by inputting isolated ROI regions, makes the model more stable in the variable polar outdoor scene.

[0093] More importantly, modern mature CNN backbone networks (such as ResNet) are designed based on processing regular rectangular images. Inputting a set of cropped, irregular-shaped ROIs requires complex padding or deformation operations to send them into the network, which not only introduces noise and artifacts, but also disrupts the original spatial relationship. The strategy of inputting complete frames completely avoids this problem.

[0094] S5.2.2 Image feature three-dimensional voxel generation module. To realize cross-modal fusion, two-dimensional image features need to be lifted to three-dimensional space.

[0095] S5.2.2.1 Utilize point cloud and camera intrinsic and extrinsic parameter matrices to generate a dense depth map aligned with the image through projection and sliding window maximum interpolation .

[0096] S5.2.2.2 Upsample the image feature map to the same resolution as the original image through bilinear interpolation. Then, only the pixels within the obtained ROI set are subjected to back projection.

[0097] Specifically, for any pixel point within the ROI, its three-dimensional space coordinates are calculated by the following formula: wherein, is the camera intrinsic matrix. Each generated three-dimensional point carries the feature vector at its corresponding position in the feature map . All these points constitute a virtual point cloud rich in semantics.

[0098] S5.2.2.3 Voxelize the virtual point cloud to obtain image feature three-dimensional voxels under the same spatial grid as .

[0099] In the S5.2.2.2 step, the present application only performs the back projection operation on the pixels within the region of interest (ROI) to generate pseudo voxels for fusion. The purpose of this design is to guide the geometric reconstruction with semantic information, realize the focus of computing resources and improve the segmentation accuracy:

[0100] A frame of high-definition image contains millions of pixels. If all pixels are back projected, an extremely large and redundant virtual point cloud will be generated, whose data volume is much larger than the original LiDAR point cloud. The subsequent voxelization, fusion, convolution and other operations will bring huge computing and storage overhead, which is not of practical value. The ROI region (i.e. the segmented non-snow cargo surface) usually only accounts for a small part of the image. Only the effective pixels in this part are operated, which can reduce the amount of calculation by one to two orders of magnitude, so that the complex multi-modal fusion algorithm can run efficiently.

[0101] The non-ROI region in the image is snow, sky, long view and other backgrounds completely irrelevant to the cargo instance segmentation task. Back projecting these pixels into three-dimensional space will generate a large number of semantic noise points. These noise points will seriously interfere with the subsequent instance segmentation. For example, the three-dimensional voxels of a piece of snow may be difficult to distinguish from the pseudo voxels of a white cargo box in terms of features, causing the segmentation algorithm to incorrectly merge the two or fail to find a clear object boundary. Only back projecting the ROI pixels is equivalent to performing a strong semantic filtering before the lifting operation, ensuring that the generated camera image feature three-dimensional voxels are all related to real cargo, providing a clean input for subsequent accurate segmentation.

[0102] The core value of multi-modal fusion lies in information complementation. In the scene where goods are closely stacked or placed side by side, the data of LiDAR point cloud at the joint of goods may be very sparse, or even have holes due to occlusion, which makes it difficult to determine whether it is the edge of an object or the boundary line of two objects only by the point cloud geometric information. At this time, the image information provides a decisive clue:

[0103] Provide high-density boundary information: the joint of two different cargo boxes, even if the colors are similar, often has shadows, different surface textures or subtle color differences. The image sensor can capture these subtle changes at a much higher resolution than LiDAR. When these ROI pixels carrying high-frequency texture and color features are back projected into three-dimensional space, they will draw a clear and high-density feature boundary in the gap of sparse LiDAR point cloud.

[0104] Fill sensor holes: LiDAR beams can not receive valid echoes on some smooth surfaces due to the problem of incident angle, resulting in point cloud holes. However, cameras are not limited by this. The back-projected image features can effectively fill these geometric holes, making the surface of the goods complete at the feature level, thereby preventing the segmentation algorithm from incorrectly dividing a complete container into two parts due to the middle point cloud hole.

[0105] Therefore, by only back-projecting the pixels in the ROI region, the present application actually uses the semantic segmentation results of the image to accurately inject the most effective and dense image features (especially boundaries and surface textures) into the weakest place of the LiDAR data in three-dimensional space, thereby greatly enhancing the feature distinguishability between different instances of goods, enabling the subsequent segmentation model to more accurately find their boundaries.

[0106] S5.2.3 Select a modal convolution fusion module, which is used to dynamically and selectively fuse the voxel features of the two modalities. For the voxel at the same position in space , the calculation method of its fused feature is:

[0107] Where: and are the feature vectors of the LiDAR voxel and the camera voxel, respectively. represents concatenating the two feature vectors. is the weight and bias of a linear layer, is a sigmoid activation function, which together constitute a selection unit that outputs a weight between 0 and 1. represents element-wise multiplication. This selection unit allows the network to learn the importance of camera features based on the concatenated features and scale them accordingly, then add them back to the original LiDAR features to achieve effective complementation of information. The final fused voxel feature is obtained.

[0108] In the S5.2.3 step, the modal convolution fusion module is not simply adding or concatenating the point cloud features and image features, but a dynamic and learnable selection mechanism is designed. The basis of this design is to adaptively adjust the fusion weight of the two modal features according to the reliability of the sensor information in different scenes and different spatial positions, thereby achieving robust and efficient information complementation.

[0109] Point cloud features offer the advantage of providing accurate 3D geometry and depth information, but they are sparse and lack semantic information. Camera features, on the other hand, provide dense high-level semantic information such as color, texture, and symbols, but lack precise depth. The reliability of both during fusion is dynamically changing, as illustrated in scenario one: a clean cargo box with a clear QR code and text. In this case, camera features... The value of point cloud features is extremely high, as they directly provide key clues for content recognition. While point cloud features can outline contours, they provide relatively basic information. Therefore, the input of the selection unit ( and The concatenation of these elements will present a pattern with high semantic information. The network will learn to assign a larger weight value to this choice, thereby improving the fusion of features. In the middle, the camera features are significantly magnified. The contribution of [the camera]. Scenario 2: A cargo container with its surface partially covered by snow and its labels blurred. In this case, camera features [are important]. In snow-covered areas, it is unreliable noise, and in areas with blurred labels, it provides less valuable information. However, point cloud features... It remains unaffected by the color of the snow, accurately capturing the smooth surface and precise dimensions of the cargo box. At this point, the selection unit learns to output a smaller weight value when the input features exhibit patterns with low semantic information but stable geometric structure. This effectively suppresses unreliable camera features. The contribution of [the system / mechanism] makes the final fusion feature [adjusted / improved]. More of the original point cloud features are preserved and relied upon. This ensures the stability of the results.

[0110] The advantages of this design in this application scenario are manifested in its strong environmental adaptability and robustness. In polar environments, sensor data degradation is commonplace (e.g., tags covered by frost, point cloud gaps caused by container reflections). A fixed fusion strategy (such as simple addition) cannot cope with such dynamic changes, and may perform well in one situation but be severely interfered with by noisy data in another. The selection mechanism, trained on a large amount of diverse data (including various data degradation scenarios), can intelligently determine, based on the quality of the local features of the input, whether to trust the point cloud or the image more in this particular situation. This adaptive, fine-grained fusion strategy ensures that, regardless of the specific circumstances, the system can maximize the use of the most reliable information source and suppress noise sources, thereby achieving segmentation accuracy and robustness far exceeding static fusion methods in the complex and ever-changing polar environment.

[0111] S5.2.4 Object Feature Enhancement Module: This module is introduced to make the network pay more attention to regions containing objects.

[0112] S5.2.4.1 One parallel network based on PointNet++ architecture as the point cloud classification module processes the original point cloud to predict an object score for each point, obtaining point-wise class information, which includes background class or object class.

[0113] S5.2.4.2 The fused voxel features are compressed into a bird's eye view to generate a BEV feature map . At the same time, the point-wise class information output by the point cloud classification module is also converted into a foreground BEV map .

[0114] S5.2.4.3 The BEV features are enhanced using a self-attention mechanism. Three independent convolutional layers generate query, key and value matrices:

[0115]

[0116]

[0117]

[0118] attention map and the final enhanced BEV feature map are calculated as follows:

[0119]

[0120] where, is used to generate and , meaning that attention will be focused on the object area identified by the point cloud; while comes from the fused , containing rich multi-modal information, so that the network can focus computing resources on the most likely location containing goods.

[0121] S5.2.5 Instance segmentation head

[0122] The enhanced BEV feature map is input into a detection head, such as the CenterPoint detection head, which will predict a three-dimensional bounding box for each independent goods instance. The final output is a set of bounding boxes , where each bounding box defines a physically independent goods unit, where, the center point coordinates of the goods unit, ​​Length, width, height and yaw angle of the cargo unit.

[0123] S5.3 Second stage: image-based content recognition and 3D verification, in which detailed content analysis is performed on each segmented cargo instance.

[0124] S5.3.1 ROI-instance association and content recognition

[0125] S5.3.1.1 For each 3D bounding box segmented , project it back to the image plane of the video frame , and get a 2D box .

[0126] S5.3.1.2 Within the intersection of the 2D box and the ROI region obtained in pre-processing , perform a multi-task recognition model which does the following in parallel:

[0127] An optical character recognition branch extracts textual information on the package. A barcode / QR code decoding branch reads machine-readable labels. A symbol classification branch identifies internationally recognized logistics or hazardous goods signs (e.g. Red Cross, flame sign, etc.).

[0128] S5.3.1.3 By fusing the recognition results, the system determines the content category for each cargo instance (e.g. food, fuel, medical supplies, etc.).

[0129] S5.3.2 3D geometric verification

[0130] For each 3D bounding box , extract all points falling into the box from to form a point cloud subset . Use the Convex Hull algorithm to calculate the actual 3D volume of the cargo from the point cloud subset . If textual information about dimensions or volume is obtained in the OCR recognition of S5.3.1 , calculate a consistency score :

[0131] If is below a pre-set threshold , mark the cargo instance as pending review to indicate possible recognition errors or package-content mismatch.

[0132] Next, this embodiment details the specific method for generating a three-dimensional digital map of the unloading area and a multi-dimensional dynamic warehousing planning step in this invention. The goal of this stage is to integrate the discrete identification results into a structured global map, and based on this map and multi-dimensional dynamic factors, generate an optimal, adaptive cargo warehousing and transportation plan.

[0133] S6.1 Generation of 3D Digital Map of Unloading Area

[0134] The system has obtained a containing A detailed list of individual cargo instances, including the 3D bounding box of each instance. Content Categories This information is integrated into a structured, queryable 3D digital map. This serves as the foundation for all subsequent planning and scheduling.

[0135] S6.1.1 Map Data Structure Definition

[0136] 3D digital map Constructed as a containing A collection of objects Each object Each item corresponds to a specific item instance and contains the following key fields: ID: A unique identifier for the item instance. Pose: The pose information of the cargo instance, i.e., the resulting 3D bounding box. Category: Content category of goods instances Status: The status of the cargo instance, combined with the verification results. Divided into verified, pending verification, etc. Accessibility: Goods accessibility score. A value between 0 and 1, used to quantify the physical difficulty of obtaining the goods.

[0137] S6.1.2 Construct an accessibility assessment model and assign accessibility scores. It is not obtained through direct measurement, but rather calculated by analyzing the spatial stacking and occlusion relationships between goods. For any instance of goods in the map... Its accessibility score :

[0138]

[0139] in, Indicates goods right The supporting relationship. When The bottom projection area of ​​the 3D bounding box and The top projection areas overlap, and the Z coordinate center value of is greater than , otherwise 0. This term represents the number of goods above . represents the path blocking relationship of goods to . The system determines the optimal grasping face of according to its orientation angle , and extends a virtual grasping channel from this face. If the bounding box of intersects with this channel, then , otherwise 0. This term represents the number of goods in front of . and are weight coefficients of support and blocking relationships, used to adjust the influence of different constraints on reachability scores.

[0140] The score approaches 1 indicates that the goods are exposed and easy to grasp; approaching 0 indicates that the goods are deeply buried or severely blocked. The calculation of this score makes the map not only record what is there, but also how easy it is to take.

[0141] S6.2 Multi-dimensional dynamic warehouse planning is the decision-making core of the cloud system, which generates the optimal warehouse task sequence for the unmanned vehicle fleet according to the generated three-dimensional digital map , combined with four categories of dynamic information.

[0142] S6.2.1 Planning input multi-dimensional dynamic information

[0143] 1. Goods attribute information: directly from the map , including the category and reachability of each good.

[0144] 2. Task priority information: the cloud system maintains a task priority list, defining an urgency score for each category of goods . For example, the value of medical supplies is the highest, followed by food, and then equipment spare parts.

[0145] 3. Real-time environmental information: the system accesses external sensor data in real time to form an environmental state vector , where is the wind speed, is the visibility, is the ambient temperature.

[0146] 4. System resource information: the number of unmanned vehicles currently available and their respective battery states.

[0147] S6.2.2 Building a dynamic task priority ranking model

[0148] The warehouse planning is not a one-time generation of a fixed plan, but a continuous and dynamic decision-making process. At each decision-making moment (e.g., when a car is idle after completing a task), the cloud system will calculate a comprehensive task priority score for all currently reachable (i.e. greater than a certain threshold : :

[0149]

[0150] where, value benefit determined by the urgency of the goods. High-priority goods will get higher scores. : opportunity benefit determined by accessibility. Goods that are easier to get will get higher scores, encouraging the system to take the easy ones first and clear the periphery quickly. environmental risk cost. This function is used to assess the risk of moving goods in the current environment . For example, when the wind speed is very high, the risk function value of moving a high and light goods (large, small) will increase sharply. transportation cost. This function estimates the energy consumption or time required to transport the goods from their current location back to the unmanned warehouse, which is mainly related to volume and category and transportation distance. is the weight coefficient.

[0151] S6.2.3 Task scheduling and execution, at each decision-making moment, the cloud system performs the following operations:

[0152] update real-time environmental information , calculate the task priority score for all reachable goods . Sort the tasks from high to low, generate a dynamic task queue. Assign the top task in the queue to an idle unmanned car. Once a certain goods is taken away by the unmanned car, the system will immediately remove the goods instance from the map and recalculate the accessibility score of the surrounding goods affected (supported or blocked) by it , dynamically updating the state of the entire map. ​​

[0153] The embodiment of the present specification also provides a polar unmanned warehouse goods identification and inventory system based on artificial intelligence, which comprises:

[0154] The acquisition module: at least one unmanned vehicle acquires data of goods in the unloading area to obtain video data and point cloud data of the goods; at least one unmanned aerial vehicle serves as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to a cloud system;

[0155] The three-dimensional digital map generation module: the cloud system pre-processes the received point cloud data and video data, identifies a region of interest corresponding to a surface of the goods in the video data, performs multi-modal fusion based on pure goods surface point cloud and features of the video data to realize goods instance segmentation, thereby obtaining a three-dimensional bounding box of each independent goods instance, identifies a content category of each goods instance in combination with image information in the region of interest, and generates a three-dimensional digital map of the unloading area;

[0156] The dynamic task queue generation module: based on the three-dimensional digital map, in combination with task priority information, real-time environment information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting goods from the unloading area to the unmanned warehouse room.

[0157] An electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned polar unmanned warehouse goods identification and inventory method based on artificial intelligence when executing the computer program.

[0158] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned polar unmanned warehouse goods identification and inventory method based on artificial intelligence.

[0159] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0160] In the specification, the same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the difference from other embodiments. Especially, for the embodiments described later, the description is simple, and the relevant part can be referred to the part of the description of the foregoing embodiments.

[0161] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An artificial intelligence-based polar unmanned warehousing goods identification and inventory method, characterized in that, The method aims to perform work in a polar environment typical of the North Pole or the South Pole, and the work site is set as a designated area relatively flat and open away from the main building group of a scientific research station or base, and the method comprises: At least one unmanned vehicle collects data of the goods in the unloading area to obtain video data and point cloud data of the goods; At least one unmanned aerial vehicle serves as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to a cloud system; The cloud system pre-processes the received point cloud data and video data, identifies the region of interest corresponding to the surface of the goods in the video data, performs multi-modal fusion based on the characteristics of the pure goods surface point cloud and the video data to realize goods instance segmentation, thereby obtaining a three-dimensional bounding box of each independent goods instance, identifies the content category of each goods instance in combination with the image information in the region of interest, and generates a three-dimensional digital map of the unloading area; The step of identifying the region of interest corresponding to the surface of the goods in the video data specifically comprises: Converting RGB image frames in video data to HSV color space, creating a binary snow mask and logically inverting the snow mask to get a region of interest; The pure cargo surface point cloud is processed to obtain LiDAR voxel features , and the video data is combined with the region of interest to obtain camera voxel features ; the two are fused by selecting a modal convolution fusion module to obtain fused voxel features : wherein, and are feature vectors of LiDAR voxels and camera voxels at the same spatial location, respectively, denotes a concatenation operation, is a Sigmoid activation function, are weights and biases of the selection unit, is an element-wise multiplication; the selection unit allows the network to learn the importance of the camera features from the concatenated features and scale the camera features according to their importance before adding them back to the original LiDAR features; Converting the multi-modal fused features into a bird's eye view to generate a BEV feature map , and obtaining a foreground BEV map by using a point cloud classification module ; enhancing the BEV feature map through a self-attention mechanism, and obtaining a three-dimensional bounding box based on the enhanced BEV feature map ​ Based on the three-dimensional digital map, in combination with task priority information, real-time environment information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting the goods from the unloading area to the unmanned warehouse room.

2. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 1, characterized in that, The at least one unmanned vehicle collects data on the goods in the unloading area, including: rasterizing the unloading area to construct an environment map model, and for each unmanned vehicle Planning a travel path The path planning is achieved by optimizing the following objective function : wherein, is an information gain for the path, is an execution cost for the path, is an overlap cost for the path, are weight coefficients for the information gain, the execution cost and the overlap cost, respectively.

3. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 2, characterized in that, The at least one unmanned aerial vehicle serving as a data relay comprises: Before the UAV departs from the unmanned warehouse, the feasibility of the single-pass full transmission mode is evaluated, and the evaluation includes calculating the total amount of electricity required for the task : wherein, T is the one-way flight time of the UAV, and are the battery consumption rate per unit time in flight and hover state of the UAV, respectively, T is the estimated remaining time for the UAV fleet to complete data collection, T is the estimated time for transmitting all data, T is the safety redundancy power; the one-time complete transmission mode is executed only when the current power of the UAV satisfies the condition.

4. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 3, characterized in that: The step of identifying a region of interest in the video data corresponding to a surface of the item specifically includes creating a binarized snow mask by decision logic : wherein, Hs, and is a predetermined threshold parameter for defining white color.

5. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 4, characterized in that: The step of generating a three-dimensional digital map of the unloading zone further comprises: for each instance of cargo in the map computing a reachability score : wherein, for determining a support relationship of a cargo to a support relationship, for determining a support relationship of a cargo to a path blocking relationship, and are weight coefficients of the support relationship and the blocking relationship, respectively.

6. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 5, characterized in that: Steps of a multi-dimensional dynamic in-lab planning, specifically including: for each currently reachable instance of a cargo computing an overall task priority score : wherein, is an urgency score of the cargo category, is combined with real-time environmental information calculated environmental risk cost, is a transportation cost, is a corresponding weight coefficient; and the tasks are ranked according to the task priority scores to generate a dynamic task queue.

7. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 6, characterized in that: The self-attention mechanism computes the attention map by and the enhanced BEV feature map : wherein, are query, key and value matrices, respectively, is a convolutional layer.

8. The polar region unmanned warehousing goods identification and inventory method based on artificial intelligence according to claim 7, characterized in that: After identifying the content class of each instance of goods, a 3D geometry verification is also included: for each instance of goods , a corresponding subset of points is extracted from the clean goods surface point cloud , according to its three-dimensional bounding box , and its three-dimensional volume is computed ; If textual information about the volume is identified from the image information within the region of interest then a consistency score is calculated by the following equation : and according to the consistency score determining whether the content category recognition result is consistent with the physical dimensions of the instance of the item.

9. An artificial intelligence based polar unmanned warehousing goods identification and inventory system, the system is used to execute an artificial intelligence based polar unmanned warehousing goods identification and inventory method according to any one of claims 1-8, characterized in that, The system comprises: The acquisition module: at least one unmanned vehicle collects data of the goods in the unloading area to obtain video data and point cloud data of the goods; at least one unmanned aerial vehicle serves as a data relay to transmit the video data and point cloud data collected by the unmanned vehicle to a cloud system; The three-dimensional digital map generation module: the cloud system pre-processes the received point cloud data and video data, identifies the region of interest corresponding to the surface of the goods in the video data, performs multi-modal fusion based on the characteristics of the pure goods surface point cloud and the video data to realize goods instance segmentation, thereby obtaining a three-dimensional bounding box of each independent goods instance, identifies the content category of each goods instance in combination with the image information in the region of interest, and generates a three-dimensional digital map of the unloading area; The dynamic task queue generation module: based on the three-dimensional digital map, in combination with task priority information, real-time environment information and system resource information, multi-dimensional dynamic warehousing planning is performed to generate a dynamic task queue for transporting the goods from the unloading area to the unmanned warehouse room.

Citation Information

Patent Citations

  • Cooperative operation system and method for unmanned aerial vehicle and unmanned loader

    CN117075629A

  • Method and system for realizing AR navigation based on cold chain warehouse AI remote control

    CN120445229A