Express delivery cabinet and pick-up method and system based on large language model and visual identification

By combining AR glasses or a fixed camera with large language models and visual recognition technology, the system can automatically identify express delivery slips and complete the delivery locker operation, solving the problems of cumbersome operation and low efficiency in existing technologies, and realizing an efficient and safe express delivery and pickup process.

CN121938013APending Publication Date: 2026-04-28SHANGHAI DIHUANG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI DIHUANG INTELLIGENT TECH CO LTD
Filing Date
2025-09-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing parcel locker delivery operations are cumbersome and inefficient. They rely on external equipment and are easily affected by network issues and unfamiliarity with the operation. Furthermore, they cannot quickly determine the locker location that matches the package size, resulting in slow delivery speeds for couriers and potential resource shortages.

Method used

By using AR glasses or fixed cameras combined with large language models and visual recognition technology, the system can automatically identify express delivery slips and complete the delivery locker operation, achieving volume estimation and multimodal decision-making, automatically matching locker locations, and reducing manual intervention.

Benefits of technology

It enables a hands-free, touchless, efficient, and safe last-mile delivery and pickup process, improving operational efficiency, reducing human error, and adapting to diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938013A_ABST
    Figure CN121938013A_ABST
Patent Text Reader

Abstract

The invention relates to and discloses an express delivery cabinet delivery and pickup method and system based on a large language model and visual identification. The method comprises the following steps: collecting an appearance image of an express delivery package and an express sheet image of the express delivery package; preprocessing the express package appearance image and the express sheet image to obtain the appearance size of the express package and the preprocessed express sheet image; performing character recognition on the preprocessed express sheet image so as to obtain preliminary structured text information; and performing field standardization, error correction and cabinet opening matching decision on the preliminary structured text information through a large language model, and generating a delivery operation instruction. According to the method and the system, dynamic volume estimation and lattice matching are carried out by combining computer vision with a large language model, and an iterative optimization decision is realized. The method has the basic technical effects that the recognition accuracy is improved, the cabinet opening distribution is automatic, and the method adapts to diversified scenes through privacy protection model updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of combining artificial intelligence and logistics, and in particular to a method and system for express delivery locker drop-off and pickup based on large language models and visual recognition. Background Technology

[0002] The current parcel delivery via express lockers mainly relies on the following process: courier or station staff scan the barcode / QR code on the parcel's waybill using a mobile app, or enter the waybill information into the locker's control panel, then select the locker's storage compartment to complete the delivery. This process is cumbersome, inefficient, and prone to inaccurate compartment selection. Especially during peak delivery periods, slow courier delivery speeds can lead to locker resource shortages and impact last-mile delivery efficiency.

[0003] Existing technologies have the following shortcomings: 1) Cumbersome operation: Delivery personnel need to hold a mobile phone or operate the locker's touchscreen to complete multiple steps, affecting efficiency; 2) Dependence on external equipment: Mobile apps or locker control computer interfaces are the main interaction methods, which are easily affected by factors such as network issues, screen damage, or unfamiliarity with operation; 3) Low locker recognition efficiency: Existing technologies cannot quickly determine the locker location that matches the package size, requiring manual judgment, which is time-consuming. Furthermore, existing OCR recognition technologies are mostly limited to basic label extraction, lacking deep integration with LLM (Local Management Model), and cannot handle dynamic volume estimation and uncertainty optimization.

[0004] Therefore, there is an urgent need in this field to develop a method and system for express delivery locker drop-off and pickup based on large language models and visual recognition. This method and system can automatically identify waybills and complete locker operations by combining innovative volume estimation and multimodal LLM decision-making with AR glasses or fixed cameras, overcoming the limitations of existing technologies and realizing a hand-free, touch-free, efficient and safe last-mile express delivery and pickup process. Summary of the Invention

[0005] The purpose of this application is to provide a method and system for express delivery locker drop-off and pickup based on large language model and visual recognition. This method and system can automatically identify waybills and complete locker operations by combining innovative volume estimation and multimodal LLM decision through AR glasses or fixed cameras. It overcomes the limitations of the prior art and realizes a hand-free, touch-free, efficient and safe last-mile express delivery and pickup process.

[0006] The first aspect of this application provides a method for express delivery locker drop-off and pickup based on a large language model and visual recognition, including the following steps:

[0007] Images of the package's appearance and the package's waybill are captured using AR glasses or a camera mounted on the locker.

[0008] The appearance image of the express package and the express waybill image are preprocessed to obtain the external dimensions of the express package and the preprocessed express waybill image;

[0009] The preprocessed express waybill image is subjected to text recognition to obtain preliminary structured text information, which includes, but is not limited to, the following fields: waybill number, recipient's name, contact information, and express company name;

[0010] The preliminary structured text information is standardized, error corrected, and locker matching is performed using a large language model, and a delivery operation instruction is generated, which includes the locker locker number to which the express package is to be delivered.

[0011] According to the delivery operation instruction, the corresponding express locker compartment is automatically opened, and the delivery operation instruction is provided to the delivery person through voice prompts, visual interface, or AR glasses interface, thereby completing the delivery event.

[0012] In another preferred embodiment, the preprocessed express waybill image is used to perform text recognition using an OCR model to obtain preliminary structured text information.

[0013] In another preferred embodiment, the preprocessed express waybill image is subjected to text recognition using a large language model to obtain preliminary structured text information.

[0014] In another preferred embodiment, the preprocessed express waybill image is used to perform text recognition through an OCR model and a large language model to obtain preliminary structured text information.

[0015] In another preferred embodiment, preprocessing the appearance image of the express package and the waybill image includes: using a computer vision algorithm to extract the boundary and depth information of the express package from the appearance image of the express package, and calculating the volume of the express package.

[0016] In another preferred embodiment, the system also records the delivery event through the express locker system and generates a pickup notification on the server side and sends it to the user.

[0017] The user's identity is verified through the camera and facial recognition. Once the verification is successful, the corresponding parcel locker compartment is automatically opened, thus completing the parcel retrieval event.

[0018] In another preferred embodiment, the field standardization of the preliminary structured text information through the large language model includes semantic parsing of the preliminary structured text information and outputting standardized field information, which includes, but is not limited to, the following information: waybill number, recipient information, and courier company name.

[0019] In another preferred embodiment, error correction of the preliminary structured text information by the large language model includes: when the preliminary structured text information has missing or ambiguous fields, the large language model intelligently completes the missing or ambiguous fields based on contextual reasoning, outputs the inference result, and evaluates the confidence level of the inference result.

[0020] In another preferred embodiment, assessing the confidence level of the inference result includes providing a quantitative numerical value for the correctness of the inference result;

[0021] When the accuracy of the quantification is lower than a preset threshold, the inference result is not adopted.

[0022] In another preferred embodiment, the cabinet matching decision based on the preliminary structured text information using the large language model includes: the large language model combining the external dimensions and type of the express parcel with the real-time status data of the express cabinet and delivery rules to determine the available express cabinet compartments that match the express parcel.

[0023] In another preferred embodiment, the delivery rule refers to a large language model that automatically selects a slot matching the package size based on the identified package size and the available slots in the locker. Based on the identified package type and the available slot types, it automatically selects a standard slot, a refrigerated slot, or an insulated slot. For large items, non-standard items, or items without available slots, the user is directly prompted that the package cannot be delivered.

[0024] In another preferred embodiment, the type of the express parcel is obtained by combining the express waybill data with volume estimation, wherein the volume estimation adopts depth estimation of camera images and reference object calibration method, and the large language model guides iterative optimization to handle uncertainty.

[0025] In another preferred embodiment, the types of express parcels include ordinary parcels, large parcels, non-standard parcels, and fresh cold chain parcels.

[0026] A second aspect of this application provides a system for parcel locker drop-off and pickup based on a large language model and visual recognition, the system comprising:

[0027] AR glasses or a camera installed on the express delivery locker are configured to capture images of the appearance of the express package and the express delivery label of the express package.

[0028] The preprocessing module is configured to preprocess the appearance image of the express package and the express waybill image, use computer vision algorithms to extract the boundary and depth information of the express package from the appearance image of the express package, calculate the volume of the express package, and thus obtain the external dimensions and volume of the express package, as well as the preprocessed express waybill image.

[0029] A visual recognition module is configured to perform text recognition on the preprocessed express waybill image to obtain preliminary structured text information, including but not limited to the following fields: waybill number, recipient's name, contact information, and express company name.

[0030] The large language model processing module is configured to perform field format standardization, error correction, and locker matching decisions on the preliminary structured text information through the large language model, and generate delivery operation instructions, the delivery operation instructions including the locker locker number to which the express package is to be delivered.

[0031] The cabinet control prompt module, according to the delivery operation instruction, controls the corresponding express cabinet compartment to open automatically, and provides the delivery operation instruction to the delivery person through voice prompts, visual interface or AR glasses interface.

[0032] In another preferred embodiment, the visual recognition module is configured to use the large language model to perform multimodal processing, perform end-to-end semantic understanding and size estimation by directly inputting image data, reduce OCR dependence, and achieve robust recognition of blurry or non-standard waybills.

[0033] In another preferred embodiment, the visual recognition module uses the large language model to fuse confidence assessment and iteratively adjust the estimated parameters (the external dimensions or volume of the express package) to achieve accurate matching of the compartment type, i.e., multimodal processing: LLM directly processes the image.

[0034] In another preferred embodiment, the visual recognition module is configured to perform text recognition on the preprocessed express waybill image using an OCR model.

[0035] In another preferred embodiment, text recognition is performed on the preprocessed express waybill image using an OCR model and a large language model.

[0036] In another preferred embodiment, the system further includes a volume estimation module. This module uses a depth camera or monocular vision combined with reference object calibration to extract 3D data of the package. It then iteratively optimizes the volume of the package by fusing multi-source confidence scores through the Large Language Model (LLM). Specifically, it uses a camera to capture images, extracts boundaries and depth using Computer Vision (CV), and iteratively optimizes the volume by fusing confidence scores using LLM to achieve grid matching.

[0037] A third aspect of this application provides an electronic device, comprising:

[0038] Memory, used to store computer-executable instructions; and,

[0039] A processor, coupled to the memory, is configured to implement the steps of the method described above when executing the computer-executable instructions.

[0040] A fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method described above.

[0041] The fifth aspect of this application provides a computer program product including computer-executable instructions, characterized in that the computer-executable instructions, when executed by a processor, implement the steps in the above-described method.

[0042] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. It should be understood that the accompanying drawings described below are merely some implementation examples of the present invention, and those skilled in the art can obtain other implementation examples based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram of the express delivery locker drop-off and pickup method based on a large language model and visual recognition according to the first embodiment of this application. Detailed Implementation

[0045] Through extensive and in-depth research, the inventors have developed for the first time a method and system for parcel locker delivery and retrieval based on a large language model and visual recognition. This method and system captures parcel images in real time using AR glasses or a camera, utilizes computer vision combined with a large language model for dynamic volume estimation and locker matching, achieves iterative optimization decisions, improves recognition accuracy, automates locker allocation, and adapts to diverse scenarios through privacy-preserving model updates. This method and system automatically recognizes waybills and completes locker operations using AR glasses or a fixed camera, realizing a hands-free, touch-free, efficient, and secure last-mile parcel delivery and retrieval process.

[0046] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.

[0047] Explanation of some concepts:

[0048] A "platform" refers to a backend service system or cloud infrastructure used to support the core functions of the parcel locker delivery and retrieval system, including data processing, storage, communication, and business logic execution. It acts as an intermediary layer connecting the parcel locker hardware (cameras, AR glasses, locker control modules, etc.), Large Language Modeling (LLM), visual recognition modules, and user / delivery personnel terminals. The platform interacts with the parcel lockers, servers, and terminal devices via a network (wired or wireless) to coordinate the entire delivery and retrieval process.

[0049] A server is a computer system on a network that provides services to other devices. The objects served by a server are usually called terminals or clients, and communication between the server and terminals can be wired or wireless. Servers can be implemented in various ways, ranging from a single computer device to a combination of multiple computer devices (such as a cluster server or cloud server). In some application scenarios, servers may also be referred to as server-side applications or cloud environments.

[0050] Correspondence: refers to the corresponding relationship between two or more pieces of data, which is usually stored on storage devices (such as storage servers, hard drives, memory, etc.). The specific storage format can be diverse, for example, it can be a file representing the correspondence, or a table in a database, etc.

[0051] Terminal: Also known as terminal device, it is a device located at the outermost edge of a computer or communication network, primarily used for user information input and output of processing results. Besides input / output functions, terminals can also perform certain calculations and processing, implementing some system functions. Terminals can be, for example, smartphones, tablets, laptops, desktop computers, smartwatches, smart bracelets, televisions, projectors with input capabilities, and personal digital assistants (PDAs), etc.

[0052] This application has at least one of the following advantages:

[0053] (a) The system and method of this application enable hands-free operation, that is, without the need for a mobile APP or cabinet touch screen, the operation can be completed directly through the camera and voice / visual prompts, which greatly improves efficiency;

[0054] (b) The system and method of this application can automatically parse the waybill information of different express companies and match the cabinet location based on vision + LLM intelligent recognition, reducing human judgment errors, that is, realizing automatic recognition and cabinet location matching; in particular, through the innovative method of volume estimation, accurate compartment allocation is achieved.

[0055] (c) The system and method of this application can be deployed in AR glasses scenarios and can also be applied to express delivery lockers and station cameras. It has good scalability and realizes multi-terminal adaptation.

[0056] (d) The system and method of this application can perform semantic repair on blurry / incomplete waybills, improve the recognition success rate, and support abnormal waybills and diverse formats.

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that these are merely examples provided to the reader of possible implementations of the present invention and are not intended to limit the scope of the invention.

[0058] The first embodiment of this application relates to a method for express delivery locker drop-off and pickup using a large language model and visual recognition, the process of which is as follows: Figure 1 As shown, the method includes the following steps:

[0059] (a) Collection and preprocessing of express parcels

[0060] Images of the package's appearance and the package's waybill are captured using AR glasses or a camera mounted on the locker.

[0061] The appearance image and waybill image of the express package are preprocessed. Computer vision algorithms are used to extract the boundary and depth information of the express package from the appearance image of the express package, calculate the volume of the express package, and obtain the external dimensions of the express package and the preprocessed waybill image.

[0062] Specifically, the delivery person holds the package within the camera's field of view or places it on a workbench area covered by a fixed camera. After the camera captures the image, the system automatically tracks and locates the package, completing the process.

[0063] Package dimensions measurement (length, width, height, and three-dimensional inspection);

[0064] The location and cropping of the express delivery waybill area are used to generate a standard image that can be processed by OCR.

[0065] The package's external dimensions measurement (3D detection of length, width, and height) involves using a depth camera or monocular vision combined with a reference object (such as a QR code of known size) for calibration, extracting 3D point cloud data, and calculating the volume. The innovation lies in introducing LLM-guided iterative estimation. When the initial estimate has high uncertainty, LLM generates prompts requiring adjustment of the package angle and recapture of the image, achieving a non-obvious optimization loop.

[0066] Standard image: An image generated after accurately identifying and correcting regular or irregular areas of a courier waybill, which is then suitable for OCR processing.

[0067] (b) Visual recognition and preliminary analysis

[0068] The preprocessed express waybill image is subjected to text recognition using an OCR model or a large language model to obtain preliminary structured text information, which includes, but is not limited to, the following fields: waybill number, recipient's name, contact information, and express company name;

[0069] Specifically, by using the large language model to employ multimodal processing, end-to-end semantic understanding and size estimation are performed by directly inputting image data, reducing OCR dependence and achieving robust recognition of blurry or non-standard waybills.

[0070] The OCR model is used to recognize text on the shipping label, outputting preliminary structured text information, including the tracking number, recipient's name, contact information, and courier company. However, this data often has the following problems:

[0071] The waybill formats of different courier companies are not standardized;

[0072] Blurry or missing characters caused field loss.

[0073] Some content was misrecognized (e.g., numbers and letters were confused).

[0074] (c) Multi-stage processing involving large language models

[0075] The large language model is invoked to perform field standardization, error correction, and locker matching decisions on the preliminary structured text information, and to generate delivery operation instructions, which include the locker locker number to which the express package is to be delivered.

[0076] Specifically, the large language model plays a key role in the following aspects:

[0077] Field standardization: Semantic parsing of OCR output is performed to identify fields such as waybill number, recipient information, and company name, unifying the expression methods of different formats.

[0078] Error correction and completion: When the recognition result is ambiguous or missing, LLM performs intelligent repair based on contextual reasoning. For example, it identifies "SF" as "SF Express", infers the missing middle digit of the phone number, and gives a quantitative value for the accuracy of the recognition and inference results. Developers can set a threshold based on this, and results below the accuracy threshold will not be adopted.

[0079] Business rule matching: Based on the real-time status of the express locker returned by the server, determine the package type (ordinary package, large package, cold chain package, etc.) and filter available lockers that match the package type according to the delivery rules.

[0080] Natural language generation: Generate operation instructions for delivery personnel or pick-up personnel, such as "This item is from SF Express, please put it in compartment 56 of locker 3".

[0081] Multimodal LLM is used to directly process the original image, fusing visual and text features to achieve end-to-end recognition and reduce the propagation of OCR errors.

[0082] Volume prediction matching grid decision: The LLM takes the predicted volume as input (e.g., length × width × height calculation), integrates the confidence level (e.g., uncertainty score of illumination influence), and compares it with the preset length, width, and height values ​​for grid specifications (small, medium, large). If the confidence level is low, the LLM triggers an iteration: generating a prompt "rotate the package and retake the photo," updating the prediction, and achieving an accurate match. This process is superior to traditional static measurement.

[0083] When multiple people are making deliveries, multimodal LLM posture recognition, facial recognition, and package label recognition can be used to achieve single or multiple batch deliveries. Specifically, after identifying the delivery person through facial recognition, the system predicts their delivery operation based on the number of packages they are carrying and their posture as they approach the locker, automatically opening multiple available lockers. As the user delivers a package, the system automatically identifies the package label information based on the acquired images of the package's appearance and the package label, and identifies the locker's location. The system then binds the delivery person's identity information with the label information and locker location information to generate a corresponding business order, which is submitted to the platform for subsequent processing.

[0084] LLM (Limited Language Management) is used to recognize and process gestures. Some gestures are predefined, such as tapping the locker door twice to open it. When a mail carrier wants to open a locker, they can specify the locker door to be opened by tapping it twice with their palm. The LLM recognizes this gesture, determines the location of the taps, verifies the mail carrier's identity via facial recognition, checks if a package is in the locker, and determines if the package is intended for the current mail carrier. If so, the LLM issues a command to open the locker door at the tapped location. If the locker is empty, it can be determined that the mail carrier may be about to deliver a package, and the locker door can be opened directly.

[0085] (d) Counter location selection and control execution

[0086] According to the delivery operation instruction, the corresponding express locker compartment is automatically opened, and the delivery operation instruction is provided to the delivery person through voice prompts, visual interface, or AR glasses interface, thereby completing the delivery event.

[0087] Specifically, the system determines the optimal cabinet slot location based on the LLM output and issues an instruction to open the cabinet slot.

[0088] The corresponding compartment door of the express delivery locker automatically pops open.

[0089] At the same time, prompts are given to the delivery person via voice module or AR glasses interface.

[0090] By using multimodal LLM to identify user gesture behavior, it can determine whether there are any anomalies in the delivery service and generate backup solutions.

[0091] (e) Delivery confirmation and record uploading and package pickup process

[0092] The delivery event is recorded by the express locker system, and a pickup notification is generated by the server and sent to the user.

[0093] The user's identity is verified through the camera and facial recognition. Once the verification is successful, the corresponding parcel locker compartment is automatically opened, thus completing the parcel retrieval event.

[0094] Specifically, the delivery person places the package and closes the locker door; the system automatically records the delivery event (waybill number, locker number, timestamp, etc.); the server generates a pickup notification and sends it to the user.

[0095] When a user arrives at the locker, they can complete identity verification through facial recognition. The identity data collected by the camera is compared with the recipient information of the package order currently stored in the locker. Once confirmed, the corresponding locker compartment door is opened. In other cases, such as when the user's identity is uncertain, the user is prompted to enter a retrieval code to retrieve the package.

[0096] After the package is picked up, the system updates the record synchronously and, with the help of LLM, observes the user's operation behavior through a camera and then provides the platform with a confirmation message that the package has been picked up.

[0097] In addition, similar to the delivery operation, LLM can also be used to perform corresponding operations based on gesture language recognition. Some gesture languages ​​are predefined, such as knocking twice on the locker door to open it. When a user wants to open a locker, they can specify the locker door to be opened by tapping it twice with their palm. After the LLM recognizes this gesture language, it determines the location of the tapping, verifies the sender's identity through facial recognition, checks if there is a package in the locker, and determines if the package belongs to the current user. If the information matches, the LLM issues a command to open the locker door at the tapping location; if the information does not match, or is insufficient to verify the user's information, it prompts the user to provide additional information to complete the retrieval operation.

[0098] It should be added that the visual recognition model (OCR + object detection) performs preliminary parsing of the order slip, extracting text and structured data. This type of data often contains format differences, recognition errors, or missing fields, therefore further semantic processing using a Large Language Model (LLM) is required. Specifically:

[0099] (1) Enhanced label parsing: LLM performs semantic analysis and field standardization on the OCR recognition results (such as unifying the label format of different express companies and automatically recognizing the waybill number, recipient's name and contact information) to avoid manual selection and input.

[0100] (2) Abnormal information completion: When the waybill is blurry, missing, or dirty, LLM can automatically infer the missing or blurry fields (such as contact address information) based on the context, thereby improving the recognition success rate.

[0101] (3) Intelligent cabinet matching: LLM combines the parcel size measurement results and the database of available cabinet slots returned by the platform, and directly outputs the best size matching cabinet slot location instruction according to delivery rules and business logic, replacing manual judgment of slot size and location.

[0102] (4) Interaction Optimization: LLM generates natural language prompts for display on the voice module or AR glasses, guiding delivery personnel to complete operations. For example, when insufficient lighting makes it difficult to recognize waybills, the system reminds users to improve lighting conditions; when the waybill is not properly positioned, the system reminds users to adjust the package's position to improve the accuracy of package information recognition; after recognition, the system can provide voice prompts for users to proceed with subsequent steps. Compared to traditional methods that rely on mobile apps or cabinet touchscreens, delivery personnel do not need to manually confirm.

[0103] Throughout the process, LLM simplifies and replaces traditional manual operations: there's no longer a need to manually select a courier company, enter a tracking number, or manually determine the locker location; delivery personnel can simply follow natural language or augmented reality prompts to complete the operation. The entire process transforms from "manual operation-driven" to "intelligent model-driven," achieving a natural operation method that is hands-free, touchless, automated in decision-making, and protects privacy.

[0104] After successful delivery, the system automatically uploads the record and sends a receipt reminder to the user (SMS / Mini Program). When picking up the package, the user can verify their identity via facial recognition (using a camera) and the LLM (Low-Level Management System) before opening the locker. See Table 1 below, which compares traditional process-driven parcel delivery events with large-scale model process-driven parcel delivery events.

[0105]

[0106] The second embodiment of this application relates to a system for parcel locker delivery and retrieval based on large language models and visual recognition. This system utilizes AR glasses or a fixed camera + LLM intelligent recognition for parcel locker delivery and retrieval. The main technical solution is as follows:

[0107] Hardware Components

[0108] AR glasses or fixed cameras: used to capture real-time images of package appearance and waybills. Cameras can be deployed around parcel lockers, at parcel delivery station counters, etc.

[0109] Parcel locker / server: Used to receive image data and call visual recognition models and large language models for information processing.

[0110] Voice / visual interaction terminal:

[0111] For parcel lockers, prompts are displayed directly through the locker's voice module or screen.

[0112] For delivery personnel wearing AR glasses, the AR glasses interact with the recipient through visual cues (text / icons) or voice prompts.

[0113] AR glasses scenario: Delivery personnel wearing AR glasses can scan packages and waybills by simply raising their hands while delivering parcels. The glasses display locker location information, and the locker door opens automatically. The AR glasses application connects to platform services and LLM servers.

[0114] Fixed camera scenario: Fixed cameras at the front end of the station or express locker capture real-time images. The data is connected to the industrial control computer of the locker. At the same time, the industrial control computer is connected to the platform service and LLM server. The deliveryman can place the package on the workbench / locker door to identify and complete the locker delivery.

[0115] Package pickup scenario: The user stands in front of the locker, the system confirms the user's identity through facial recognition, and the user can directly open the locker to pick up the package.

[0116] Specifically, this application provides a system for express delivery locker drop-off and pickup based on a large language model and visual recognition, the system comprising:

[0117] AR glasses or a camera installed on the express delivery locker are configured to capture images of the appearance of the express package and the express delivery label of the express package.

[0118] The preprocessing module is configured to preprocess the appearance image of the express package and the express waybill image, use computer vision algorithms to extract the boundary and depth information of the express package from the appearance image of the express package, calculate the volume of the express package, and thus obtain the external dimensions and volume of the express package, as well as the preprocessed express waybill image.

[0119] A visual recognition module is configured to perform text recognition on the preprocessed express waybill image to obtain preliminary structured text information, including but not limited to the following fields: waybill number, recipient's name, contact information, and express company name.

[0120] The large language model processing module is configured to perform field format standardization, error correction, and locker matching decisions on the preliminary structured text information through the large language model, and generate delivery operation instructions, the delivery operation instructions including the locker locker number to which the express package is to be delivered.

[0121] The cabinet control prompt module, according to the delivery operation instruction, controls the corresponding express cabinet compartment to open automatically, and provides the delivery operation instruction to the delivery person through voice prompts, visual interface or AR glasses interface.

[0122] Preferably, the visual recognition module uses the large language model to fuse confidence assessment and iteratively adjust the estimated parameters (the external dimensions or volume of the express package) to achieve accurate matching of the compartment type.

[0123] Preferably, the visual recognition module uses the large language model to perform multimodal processing, and performs end-to-end semantic understanding and size estimation by directly inputting image data, reducing OCR dependence and achieving robust recognition of blurry or non-standard waybills.

[0124] Preferably, it also includes a volume estimation module, which uses a depth camera or monocular vision combined with reference object calibration to extract 3D data; the large language model integrates multi-source confidence to achieve iterative volume optimization.

[0125] The first embodiment is a method embodiment corresponding to this embodiment. The technical details in the first embodiment can be applied to this embodiment, and the technical details in this embodiment can also be applied to the first embodiment.

[0126] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the various method embodiments of this application. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.

[0127] Furthermore, embodiments of this application also provide an electronic device, including a memory for storing computer-executable instructions, and a processor; the processor is used to implement the steps in the above-described method embodiments when executing the computer-executable instructions in the memory. The processor may be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), a Microcontroller Unit (MCU), a Neural Processing Unit (NPU), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices. The aforementioned memory may be read-only memory (ROM), random access memory (RAM), flash memory, a hard disk, or a solid-state drive, etc. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0128] Furthermore, embodiments of this application also provide a computer program product, including computer-executable instructions that, when executed by a processor, implement the steps in the above-described method embodiments.

[0129] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.

[0130] All references to this specification are considered to be incorporated integrally into the disclosure of this application so that they can serve as the basis for modifications if necessary. Furthermore, it should be understood that the above descriptions are merely preferred embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the scope of protection of one or more embodiments of this specification.

[0131] In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for express delivery locker drop-off and pickup based on a large language model and visual recognition, characterized in that, Includes the following steps: Images of the package's appearance and the package's waybill are captured using AR glasses or a camera mounted on the locker. The appearance image of the express package and the express waybill image are preprocessed to obtain the external dimensions of the express package and the preprocessed express waybill image; The preprocessed express waybill image is subjected to text recognition to obtain preliminary structured text information, which includes, but is not limited to, the following fields: waybill number, recipient's name, contact information, and express company name; The preliminary structured text information is standardized, error corrected, and locker matching is performed using a large language model, and a delivery operation instruction is generated, which includes the locker locker number to which the express package is to be delivered. According to the delivery operation instruction, the corresponding express locker compartment is automatically opened, and the delivery operation instruction is provided to the delivery person through voice prompts, visual interface, or AR glasses interface, thereby completing the delivery event.

2. The method as described in claim 1, characterized in that, Preprocessing the appearance image and waybill image of the express package includes: using computer vision algorithms to extract the boundary and depth information of the express package from the appearance image of the express package, and calculating the volume of the express package.

3. The method as described in claim 1, characterized in that, It also includes recording the delivery event through the express locker system, generating a pickup notification through the server and sending it to the user; The user's identity is verified through the camera and facial recognition. Once the verification is successful, the corresponding parcel locker compartment is automatically opened, thus completing the parcel retrieval event.

4. The method as described in claim 1, characterized in that, The large language model is used to standardize the preliminary structured text information by performing semantic parsing and outputting standardized field information, which includes, but is not limited to, the following information: waybill number, recipient information, and courier company name.

5. The method as described in claim 4, characterized in that, Error correction of the preliminary structured text information by the large language model includes: when the preliminary structured text information has missing or ambiguous fields, the large language model intelligently completes the missing or ambiguous fields based on contextual reasoning, outputs the inference result, and evaluates the confidence of the inference result.

6. The method as described in claim 5, characterized in that, The large language model is used to make locker matching decisions based on the preliminary structured text information. This includes: the large language model combining the dimensions and type of the express package with the real-time status data of the express locker and the delivery rules to determine the available locker slots that match the express package.

7. A system for express delivery locker drop-off and pickup based on a large language model and visual recognition, characterized in that, The system includes: AR glasses or a camera installed on the express delivery locker are configured to capture images of the appearance of the express package and the express delivery label of the express package. The preprocessing module is configured to preprocess the appearance image of the express package and the express waybill image, use computer vision algorithms to extract the boundary and depth information of the express package from the appearance image of the express package, calculate the volume of the express package, and thus obtain the external dimensions and volume of the express package, as well as the preprocessed express waybill image. A visual recognition module is configured to perform text recognition on the preprocessed express waybill image to obtain preliminary structured text information, including but not limited to the following fields: waybill number, recipient's name, contact information, and express company name. The large language model processing module is configured to perform field format standardization, error correction, and locker matching decisions on the preliminary structured text information through the large language model, and generate delivery operation instructions, the delivery operation instructions including the locker locker number to which the express package is to be delivered. The cabinet control prompt module, according to the delivery operation instruction, controls the corresponding express cabinet compartment to open automatically, and provides the delivery operation instruction to the delivery person through voice prompts, visual interface or AR glasses interface.

8. The system as described in claim 7, characterized in that, It also includes a volume estimation module, which uses a depth camera or monocular vision combined with reference object calibration to extract the 3D data of the package, and iterates and optimizes the volume of the express package by fusing multi-source confidence through the large language model.

9. An electronic device, characterized in that, include: Memory is used to store executable instructions for a computer; as well as, A processor, coupled to the memory, is configured to implement the steps of the method as described in any one of claims 1 to 6 when executing the computer-executable instructions.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 6.