Intelligent receiving cabinet receiving method based on image recognition technology

By setting up multiple sets of cameras on the top of the smart cabinet and combining advanced image recognition algorithms, the existing smart cabinet has solved the problem of high cost and poor recognition effect, and convenient, efficient and safe item collection management has been achieved, and the level of intelligence and automation has been improved.

CN120107896APending Publication Date: 2025-06-06JIANGSU ANFANG ELECTRIC POWER TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510250444.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing smart requisition cabinet has high cost and poor recognition effect, making it difficult to meet the actual requisition needs.

Method used

By setting up multiple sets of cameras at the top of the body of the collection cabinet, image recognition of stored items is performed, and combined with the YOLOv5 algorithm, Retinex algorithm, multi-view fusion and occlusion reasoning, high-speed motion target tracking and customer position recognition, accurate identification and management of items is achieved.

Benefits of technology

It lowers the threshold for new users, improves the convenience and efficiency of storage and collection, ensures the security of item storage, ensures data accuracy and timeliness, solves the problem of item positioning when vouchers are lost, and significantly improves the intelligence and automation level of smart collection cabinets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107896A_ABST
    Figure CN120107896A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of receiving cabinets, discloses an intelligent receiving cabinet receiving method based on an image recognition technology, and solves the problems that an existing intelligent receiving cabinet receiving method is high in cost and poor in recognition effect. According to the invention, multiple functions are provided, so that a new user is guided and helped to quickly master for the first time; storage and fetching processes are standard, operation is recorded through cameras, a system binds a voucher and updates a database, accurate recording of the operation is guaranteed, real-time monitoring of articles monitors abnormity with the help of a YOLOv5 algorithm, when a user loses the voucher and the lost articles, various algorithms such as light self-adaption and multi-view fusion are used for searching, the articles are accurately positioned in combination with multi-camera images, and the user experience is improved. While hardware and maintenance costs are reduced, complex environments are effectively dealt with, image recognition accuracy and efficiency are greatly improved, user experience and article storage security are improved, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to a method for collecting using an intelligent collection cabinet based on image recognition technology. Background Art

[0002] In the current field of item storage, personal item lockers are widely used in various places, such as shopping malls, libraries, gyms, etc. Traditional personal item lockers rely on keys or paper barcodes to control the opening and closing of the cabinet doors. Although this method is simple to operate, it has obvious disadvantages. Once the user loses the key or barcode, it will take a lot of time and manpower to solve the problem of collecting items.

[0003] When image recognition technology was first introduced, in order to accurately locate objects, cameras were installed inside each individual cabinet. On the one hand, the hardware procurement cost was extremely high because cameras were installed inside hundreds or thousands of cabinets. On the other hand, the subsequent maintenance, inspection, and data storage and processing costs were also high, which severely restricted its large-scale promotion and use. In addition, early image recognition algorithms performed poorly in complex environments. When the weather changed and the light changed, or when objects were obscured or moved quickly, the accuracy and efficiency of image recognition dropped significantly, making it difficult to meet actual use needs.

[0004] In view of the above problems, an innovative design is carried out based on the original intelligent collection cabinet collection method based on image recognition technology. Summary of the invention

[0005] The purpose of the present invention is to provide a smart dispensing cabinet dispensing method based on image recognition technology, and adopt this method to work, thereby solving the problems of high cost and poor recognition effect of existing smart dispensing cabinet dispensing methods.

[0006] To achieve the above object, the present invention provides the following technical solution: a method for collecting items from an intelligent collection cabinet based on image recognition technology, which uses multiple groups of cameras arranged at equal intervals on the top of the collection cabinet body to perform image recognition on stored collected items, including the following steps:

[0007] S1: Initial collection guidance: When a new user collects the voucher for the first time, he / she will be guided through the touch screen next to the collection cabinet to learn how to select the cabinet and obtain the voucher in the form of pictures, texts and voice;

[0008] S2: Execute the deposit operation: the customer selects an empty cabinet to deposit the item, the external camera records the entire process, the system binds the voucher to the cabinet, and records the cabinet occupancy status;

[0009] S3: Carry out real-time monitoring of items: The camera maintains a continuous monitoring state, transmits the captured images to the background in real time, and uses the YOLOv5 algorithm to analyze the monitoring images to detect whether there are any abnormal situations;

[0010] S4: Implement data storage and update: The system stores the deposit data and the collection data, and can update these data in real time to ensure the accuracy and timeliness of the data;

[0011] S5: Carry out normal identity verification for picking up items: When a customer comes to pick up items, he / she inserts the key or scans the code at the verification terminal. The system verifies the customer's operation information to determine whether he / she has the authority to pick up items.

[0012] S6: Complete normal item collection confirmation: If the system verification is successful, the corresponding cabinet door is opened, and the customer closes the cabinet door after receiving the item. At this time, the system updates the database and marks the cabinet as idle;

[0013] S7: Register lost property assistance information: When a customer loses his / her voucher and comes to ask for help, the staff will record in detail the characteristics of the item described by the customer and the storage time and other related information;

[0014] S8: Execute lost property search: The staff inputs the lost property information provided by the customer into the backend. The system combines the images taken by multiple cameras and uses the image recognition algorithm and the following targeted algorithms to determine the locker where the item is located:

[0015] S801: Perform light-adaptive image enhancement: Introduce the Retinex algorithm to separate brightness and reflectivity to avoid the influence of light interference on the image, so as to improve the image quality and facilitate subsequent analysis;

[0016] S802: Implement multi-view fusion and occlusion reasoning: Fusion of multi-camera view information, inference of the features of the occluded parts using deep learning technology, to obtain more comprehensive image information;

[0017] S803: Execute high-speed moving target tracking: Use Kalman filter algorithm to predict the trajectory of the target, so as to accurately track the target object that may be in a high-speed moving state;

[0018] S804: Complete customer position identification: Use computer vision algorithms to determine the customer's position, and associate it with the location of the locker to assist in determining the locker where the item is located;

[0019] S9: Execute item collection operation: After successfully locating the cabinet where the items are stored through the above operation, the staff enters a specific door opening command in the main drive control terminal. The command is sent to the system and triggers the corresponding cabinet door to open. The customer can then take away the items they have stored from the opened cabinet.

[0020] Furthermore, in step S3, the YOLOv5 algorithm, which is based on a convolutional neural network, converts the target detection task into a regression problem and directly predicts the category and position of the target in a network;

[0021] Feature extraction network: It consists of a series of convolutional layers and pooling layers, which are used to extract features from images. It uses cross-stage local dense blocks and convolution kernels of different scales to extract multi-scale features. The calculation process is expressed as:

[0022] x 1 =Conv 1 (x in )

[0023] x 2 =Conv 2 (x 1 )

[0024] x 3 =C oncatenate (x 1 , x 2 )

[0025] x out =Conv 3 (x 3 )

[0026] Among them, x in is the input image, conv 1 represents the i-th convolution operation, through which local and global features in the image are continuously extracted;

[0027] Detection head: After feature extraction, the detection head is responsible for predicting the category and location of the target. For each predicted bounding box, its confidence score and the probability of the category to which it belongs are calculated. The confidence score indicates the possibility of the existence of the target in the bounding box. The confidence score calculation formula is: Confidewnce = σ(p obj )

[0028] Among them, p obj is the probability of the target existence predicted by the model, σ is the sigmoid function, which maps the probability value to the interval [0, 1];

[0029] For the position prediction of the bounding box, an offset method is used. Assume that the coordinates of the predicted bounding box center point are (x, y), the width and height are (w, h), and the offset relative to the grid unit is (t x , t y ), the scale factor is (t w , t h ), then:

[0030] x=(cx +σ(t x ))×wy grid

[0031] y=(c y +σ(t y ))×h grid

[0032] w=e tw × anchor

[0033] h=e th ×h anchor

[0034] Among them, (c x , c y ) is the coordinate of the upper left corner of the grid cell, (w grid, h grid ) is the width and height of the grid cell, (w anchor ,h anchor ) are the preset width and height of the anchor box;

[0035] In the real-time monitoring scenario of items, we pre-train the YOLOv5 model so that it can recognize cabinet doors and common items. When the images collected by the external camera are transmitted to the background image recognition and data processing unit, the model will quickly process the images to detect whether there are abnormal openings of cabinet doors and abnormal situations of stolen items. If an anomaly is detected, the reliability of the anomaly will be judged based on the confidence score and category information, and an alarm will be issued in time to notify the staff.

[0036] Furthermore, in step S801, light adaptive image enhancement is specifically performed by the Retinex algorithm to perform adaptive illumination compensation, the core idea of ​​which is to separate the brightness and reflectivity of the image, and to improve the clarity of the image under different illumination conditions by enhancing the reflectivity. By formula calculation, the image gain is automatically reduced to avoid overexposure, and the brightness distribution of the image is dynamically adjusted by analyzing the image histogram to ensure that the color and shape characteristics of the object are not disturbed by light, so that the characteristics of the object can be clearly presented under different illumination conditions.

[0037] Furthermore, in step S802, multi-perspective fusion and occlusion reasoning specifically adopt a multi-perspective fusion and occlusion reasoning algorithm. When occlusion is detected, an occlusion reasoning network based on deep learning is used to infer the characteristics of the occluded part, combined with prior knowledge of the object, the shape and proportion of common objects. When the human body occludes part of the object, the shape and size of the occluded part are inferred by analyzing the unobstructed edges of the object and the known object category information, thereby completely identifying the object characteristics.

[0038] Furthermore, in step S803, the moving target tracking specifically uses the Kalman filter algorithm to achieve high-speed moving target detection and tracking. The Kalman filter is an optimal recursive data processing algorithm based on the linear system state space model. Through continuous iterative prediction and update steps, the algorithm analyzes the position and shape changes of objects in continuous frame images, and predicts the movement trajectory of objects. Even if the objects move quickly, their feature points can be accurately captured. For objects that are quickly taken out of a bag and put into a cabinet, the algorithm calculates their movement speed and direction based on the position changes of the objects in the previous and next frames of images, thereby stably tracking the objects and extracting their feature information.

[0039] Furthermore, in step S804, customer position recognition specifically introduces a position point recognition algorithm based on computer vision, and uses the image captured by the camera to determine the customer position. First, the camera image is preprocessed, and the edge detection algorithm Canny algorithm is used to extract the contour information in the image. The approximate contour of the human body is identified through contour analysis, and the customer's standing center position is determined by analyzing the three-dimensional coordinates of multiple key points. The average value of the three-dimensional coordinates of all key points is calculated as the standing center coordinate. After the standing center position is obtained, the layout information of the collection cabinet is combined to determine the locker position corresponding to the customer. This helps to more accurately identify the storage and retrieval behavior of items in subsequent image analysis, and assists in judging the nature and degree of occlusion when occlusion occurs.

[0040] Furthermore, in step S801, the Retinex algorithm is used to perform adaptive illumination compensation for the influence of weather light. The algorithm separates the brightness and reflectivity of the image and improves the clarity of the image under different illumination conditions by enhancing the reflectivity. The calculation formula is:

[0041]

[0042] R(x,y)=log(I(x,y))-log(F(x,y)*I(x,y))

[0043] Among them, I(x, y) is the pixel value of the original image at the position (x, y), L(x, y) is the illumination component, F(x, y) is the Gaussian low-pass filter, R(x, y) is the reflectance component, and S(x, y) is the enhanced image. On sunny days with sufficient light, the image gain is automatically reduced to avoid overexposure. On rainy days or at night when the light is dim, the reflectance component is enhanced, the image gain is increased, and the color balance is adjusted to ensure that the features of the object can be clearly presented under different lighting conditions.

[0044] Furthermore, in step S802, multiple cameras photograph the same object from different perspectives, and the image features obtained by each camera are represented as F i(i=1,2,…,n), n is the number of cameras, and the fused feature F fusion Calculated by weighted average:

[0045]

[0046] Among them, w i It is a weight that is dynamically adjusted according to the camera position, angle, and image quality factors. When occlusion is detected, the occlusion reasoning network based on deep learning is used, combined with prior knowledge of the object, the shape and proportion of common objects, to infer the characteristics of the occluded part and achieve complete recognition of the object features.

[0047] Furthermore, in step S803, during the high-speed moving target detection and tracking process, it is assumed that the state equation of the system is:

[0048] X k =AK k-1 +BU k +W k

[0049] The observation equation is: Z k =HK k +V k

[0050] Among them, X k is the state vector at time k, which contains the position and speed information of the object; A is the state transfer matrix, which describes the change of state over time; B is the control matrix, U k is the control vector, which can be set to 0 in this scenario, W k is the process noise; Z k is the observation vector at time k, that is, the position information of the object in the image collected by the camera, H is the observation matrix, which maps the state vector to the observation space; V k It is observation noise. Through continuous iterative prediction and update steps, the movement trajectory of objects is analyzed and predicted according to the position and shape changes of objects in continuous frame images, and the feature points of fast-moving objects are accurately captured. For objects that are quickly taken out of the bag and put into the cabinet, the algorithm calculates their movement speed and direction according to the position changes of the objects in the previous and next frames of images, thereby stably tracking the objects and extracting their feature information.

[0051] Furthermore, in step S804, the camera image is preprocessed using a position point recognition algorithm of computer vision, and the contour information in the image is extracted using an edge detection Canny algorithm. The rough contour of the human body is identified through contour analysis. It is assumed that the key point on the human body contour is P j (j=1, 2, ..., m), m is the number of key points, for each key point, according to its pixel coordinates (xj ,y j ), combined with the camera calibration parameters, the intrinsic matrix and the extrinsic matrix, the three-dimensional coordinates (X j , Y j , Z j ):

[0052]

[0053] Among them, (u j, v j ) is the pixel coordinate of the key point on the image plane. Then, the customer's standing center position is determined by analyzing the three-dimensional coordinates of multiple key points, and the average value of the three-dimensional coordinates of all key points is calculated as the standing center coordinate. After obtaining the standing center position, the locker position corresponding to the customer is determined in combination with the layout information of the collection cabinet. This helps to more accurately identify the storage and retrieval behavior of items in subsequent image analysis, and assists in determining the nature and degree of occlusion when occlusion occurs.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] The present invention proposes a method for collecting items from an intelligent collection cabinet based on image recognition technology. The existing intelligent collection cabinet collection method has high cost and poor recognition effect. The present invention greatly reduces the collection threshold for new users by showing the operation method in pictures, texts and voices in the first collection guide in terms of user experience, making the deposit and retrieval process more convenient and efficient, and improving user satisfaction.

[0056] From the perspective of security, real-time monitoring of items uses the YOLOv5 algorithm to continuously monitor anomalies, effectively preventing abnormal opening of cabinet doors and theft of items, and ensuring the safety of item storage. In terms of data management, data storage updates the storage and collection data in real time to ensure data accuracy and timeliness, providing a reliable basis for subsequent query and statistical analysis;

[0057] In the case of lost keys or scanned code credentials, the lost property search system combines images from multiple cameras and uses targeted algorithms such as light adaptive image enhancement, multi-perspective fusion and occlusion reasoning, high-speed motion target tracking, and customer position recognition to accurately locate the locker where the items are located, solving the user's urgent needs. Overall, this method of collecting items has greatly improved the intelligence and automation level of smart collection cabinets, reduced labor costs, and has broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a flow chart of the intelligent collection cabinet collection method based on image recognition technology of the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] The present invention will be further described below in conjunction with the embodiments.

[0061] See also Figure 1 A method for collecting items from an intelligent collection cabinet based on image recognition technology is provided. The method uses a plurality of cameras arranged at equal intervals on the top of the collection cabinet body to perform image recognition on stored collected items. The method comprises the following steps:

[0062] Step 1: First-time collection guidance: When a new user collects the voucher for the first time, he / she will be guided through the touch screen next to the collection cabinet to learn how to select the cabinet and obtain the voucher in the form of pictures, texts and voice;

[0063] Step 2: Execute the deposit operation: The customer selects an empty cabinet to deposit the item, and the external camera records the entire process. The system binds the voucher to the cabinet and records the cabinet occupancy status;

[0064] Step 3: Conduct real-time monitoring of objects: The camera maintains a continuous monitoring state, transmits the captured images to the background in real time, and uses the YOLOv5 algorithm to analyze the monitoring images to monitor whether there are any abnormalities. In step S3, the YOLOv5 algorithm, which is based on a convolutional neural network, converts the target detection task into a regression problem, and directly predicts the category and location of the target in a network;

[0065] Feature extraction network: It consists of a series of convolutional layers and pooling layers, which are used to extract features from images. It uses cross-stage local dense blocks and convolution kernels of different scales to extract multi-scale features. The calculation process is expressed as:

[0066] x 1 =Conv 1 (x in )

[0067] x 2 =Conv 2 (x 1 )

[0068] x 3 =C oncatenate (x 1 , x 2 )

[0069] x out =Conv 3 (x3 )

[0070] Among them, x in is the input image, conv 1 represents the i-th convolution operation, through which local and global features in the image are continuously extracted;

[0071] Detection head: After feature extraction, the detection head is responsible for predicting the category and location of the target. For each predicted bounding box, its confidence score and the probability of the category to which it belongs are calculated. The confidence score indicates the possibility of the existence of the target in the bounding box. The confidence score calculation formula is: Confidewnce = σ(p obj )

[0072] Among them, p obj is the probability of the target existence predicted by the model, σ is the sigmoid function, which maps the probability value to the interval [0, 1];

[0073] For the position prediction of the bounding box, an offset method is used. Assume that the coordinates of the predicted bounding box center point are (x, y), the width and height are (w, h), and the offset relative to the grid unit is (t x , t y ), the scale factor is (t w , t h ), then:

[0074] x=(c x +σ(t x ))×wy grid

[0075] y=(c y +σ(t y ))×h grid

[0076] w=e tw × anchor

[0077] h=e th ×h anchor

[0078] Among them, (c x , c y ) is the coordinate of the upper left corner of the grid cell, (w grid ,h grid ) is the width and height of the grid cell, (w anchor ,h anchor ) are the preset width and height of the anchor box;

[0079] In the real-time monitoring scenario of items, we pre-train the YOLOv5 model so that it can recognize cabinet doors and common items. When the images collected by the external camera are transmitted to the background image recognition and data processing unit, the model will quickly process the images to detect whether there are abnormal openings of cabinet doors and abnormal situations of stolen items. If an anomaly is detected, the reliability of the anomaly will be judged based on the confidence score and category information, and an alarm will be issued in time to notify the staff.

[0080] Step 4: Implement data storage and update: The system stores the deposit data and the collection data, and can update these data in real time to ensure the accuracy and timeliness of the data;

[0081] Step 5: Carry out normal identity verification for picking up items: When a customer comes to pick up items, he / she inserts the key or scans the code at the verification terminal. The system verifies the customer's operation information to determine whether he / she has the authority to pick up items.

[0082] Step 6: Complete normal item collection confirmation: If the system verification is passed, the corresponding cabinet door will be opened. After the customer receives the item, the cabinet door will be closed. At this time, the system will update the database and mark the cabinet as idle;

[0083] Step 7: Register lost property assistance information: When a customer loses his / her voucher and comes to ask for help, the staff will record in detail the characteristics of the item described by the customer, the storage time and other related information;

[0084] Step 8: Execute lost property search: The staff inputs the lost property information provided by the customer into the backend. The system combines the images taken by multiple cameras and uses the image recognition algorithm and the following targeted algorithms to determine the locker where the item is located:

[0085] Perform light-adaptive image enhancement: Introduce the Retinex algorithm to separate brightness and reflectivity to avoid the influence of light interference on the image, so as to improve the image quality and facilitate subsequent analysis. In step S801, light-adaptive image enhancement is specifically performed by the Retinex algorithm to perform adaptive illumination compensation. The core idea is to separate the brightness and reflectivity of the image, and improve the clarity of the image under different illumination conditions by enhancing the reflectivity. Through formula calculation, the image gain is automatically reduced to avoid overexposure. Through the analysis of the image histogram, the brightness distribution of the image is dynamically adjusted to ensure that the color and shape characteristics of the object are not interfered by light, so that the characteristics of the object can be clearly presented under different illumination conditions. In step S801, the Retinex algorithm is used to perform adaptive illumination compensation for the influence of weather light. The algorithm separates the brightness and reflectivity of the image, and improves the clarity of the image under different illumination conditions by enhancing the reflectivity. The calculation formula is:

[0086]

[0087] R(x,y)=log(I(x,y))-log(F(x,y)*I(x,y))

[0088] Where I(x, y) is the pixel value of the original image at the position (x, y), L(x, y) is the illumination component, F(x, y) is the Gaussian low-pass filter, R(x, y) is the reflectance component, and S(x, y) is the enhanced image. On sunny days with sufficient light, the image gain is automatically reduced to avoid overexposure. On rainy days or at night when the light is dim, the reflectance component is enhanced, the image gain is increased, and the color balance is adjusted to ensure that the features of the object can be clearly presented under different lighting conditions.

[0089] Realize multi-view fusion and occlusion reasoning: fuse the view information of multiple cameras, use deep learning technology to infer the features of the occluded part, so as to obtain more comprehensive image information. In step S802, multi-view fusion and occlusion reasoning specifically adopt multi-view fusion and occlusion reasoning algorithms. When occlusion is detected, the occlusion reasoning network based on deep learning is used to combine the prior knowledge of the object, the shape and proportion of common objects, and infer the features of the occluded part. When the human body occludes part of the object, the shape and size of the occluded part are inferred by analyzing the unobstructed edges of the object and the known object category information, so as to fully identify the features of the object. In step S802, multiple cameras shoot the same object from different viewpoints, and the image features obtained by each camera are represented as F i (i=1,2,…,n), n is the number of cameras, and the fused feature F fusion Calculated by weighted average:

[0090]

[0091] Among them, w i The weights are dynamically adjusted according to the camera position, angle, and image quality factors. When occlusion is detected, the occlusion reasoning network based on deep learning is used to combine the prior knowledge of objects, the shapes and proportions of common objects, and infer the features of the occluded parts to achieve complete recognition of the object features.

[0092] Perform high-speed moving target tracking: Use the Kalman filter algorithm to predict the trajectory of the target in order to accurately track the target object that may be in a high-speed moving state. In step S803, the moving target tracking specifically uses the Kalman filter algorithm to achieve high-speed moving target detection and tracking. The Kalman filter is an optimal recursive data processing algorithm based on the linear system state space model. Through continuous iteration of prediction and update steps, the algorithm analyzes the position and shape changes of the object in the continuous frame image and predicts the movement trajectory of the object. Even if the object moves quickly, its feature points can be accurately captured. For items that are quickly taken out of the bag and put into the cabinet, the algorithm calculates its movement speed and direction based on the position changes of the object in the previous and next frames of images, thereby stably tracking the object and extracting its feature information. In step S803, during the high-speed moving target detection and tracking process, it is assumed that the state equation of the system is:

[0093] X k =AK k-1 +BU k +W k

[0094] The observation equation is: Z k =HK k +V k

[0095] Among them, X k is the state vector at time k, which contains the position and speed information of the object; A is the state transfer matrix, which describes the change of state over time; B is the control matrix, U k is the control vector, which can be set to 0 in this scenario, W k is the process noise; Z k is the observation vector at time k, that is, the position information of the object in the image collected by the camera, H is the observation matrix, which maps the state vector to the observation space; V k It is observation noise. Through continuous iteration of prediction and update steps, the movement trajectory of objects is analyzed and predicted according to the position and shape changes of objects in consecutive frame images, and the feature points of fast-moving objects are accurately captured. For objects that are quickly taken out of a bag and put into a cabinet, the algorithm calculates the speed and direction of movement according to the position changes of the objects in the previous and next frames of images, thereby stably tracking the objects and extracting their feature information.

[0096] Complete customer position identification: Use computer vision algorithm to determine the customer's position, and associate it with the position of the locker to assist in determining the locker where the item is located. In step S804, customer position identification specifically introduces a position point recognition algorithm based on computer vision, and uses the image collected by the camera to determine the customer's position. First, the camera image is preprocessed, and the edge detection algorithm Canny algorithm is used to extract the contour information in the image. Through contour analysis, the approximate contour of the human body is identified. The three-dimensional coordinates of multiple key points are analyzed to determine the center position of the customer's position. The average value of the three-dimensional coordinates of all key points is calculated as the center coordinate of the position. After obtaining the center position of the position, combined with the layout information of the collection cabinet, the position of the locker corresponding to the customer is determined. This helps to more accurately identify the storage and retrieval behavior of items in subsequent image analysis, and assist in determining the nature and degree of occlusion when occlusion occurs. In step S804, the position point recognition algorithm of computer vision is used to preprocess the camera image first, and the edge detection Canny algorithm is used to extract the contour information in the image. Through contour analysis, the approximate contour of the human body is identified. Assuming that the key point on the human body contour is P j (j=1, 2, ..., m), m is the number of key points, for each key point, according to its pixel coordinates (x j ,y j ), combined with the camera calibration parameters, the intrinsic matrix and the extrinsic matrix, the three-dimensional coordinates (X j , Y j , Z j ):

[0097]

[0098] Among them, (u j, v j ) is the pixel coordinate of the key point on the image plane. Then, the customer's standing center position is determined by analyzing the three-dimensional coordinates of multiple key points, and the average value of the three-dimensional coordinates of all key points is calculated as the standing center coordinate. After obtaining the standing center position, the locker position corresponding to the customer is determined in combination with the layout information of the collection cabinet. This helps to more accurately identify the storage and retrieval behavior of items in subsequent image analysis, and assists in determining the nature and degree of occlusion when occlusion occurs.

[0099] Step 9: Execute the item collection operation: After successfully locating the cabinet where the items are stored through the above operation, the staff enters a specific door opening command in the main drive control terminal. The command is sent to the system and triggers the corresponding cabinet door to open. The customer can then take away the items they have stored from the opened cabinet.

[0100] Specifically, first of all, in the initial use guidance phase, when a new user first contacts the smart locker, the touch screen guide terminal next to the locker will automatically start, showing the user the specific operation methods of selecting an empty locker and obtaining a key or paper barcode in the form of pictures and texts and supplemented by clear voice prompts. This fully considers the principle of human-computer interaction, greatly reduces the threshold for new users, and allows first-time users to quickly become familiar with the basic operation process of the locker, laying a good foundation for subsequent storage operations;

[0101] Secondly, the customer enters the storage operation and selects an empty locker according to the instructions. After opening the locker door, the customer places the personal belongings in it. At this time, the high-definition wide-angle cameras installed at the four corners and the middle of the overall cabinet frame are immediately activated, using the optical imaging principle and advanced image sensing technology to accurately record the entire storage process. When the customer closes the locker door, the system will automatically bind the key or barcode to the locker through coding technology, and use database storage technology to accurately record the occupied status of the locker in the database.

[0102] Then, during the storage period, real-time monitoring of items plays a key role. External cameras continuously monitor the collection cabinet area in all directions. The collected image data is transmitted to the background in real time through efficient transmission technology. The background uses the YOLOv5 algorithm based on deep learning, the powerful feature extraction ability of the convolutional neural network and the precise prediction function of the detection head to conduct in-depth analysis of the image, so as to determine whether there are abnormal situations such as abnormal opening of the cabinet door and the theft of items, effectively ensuring the storage safety of items. At the same time, data storage updates will store the relevant data of each deposit and collection, such as deposit and retrieval time, item feature description, etc. in the database in a timely manner, and update it in real time during the entire operation process, providing reliable data support for the entire collection process;

[0103] When a customer comes to collect items, normal collection identity verification begins. The customer inserts a key or scans a paper barcode in the verification terminal, and the verification terminal quickly establishes a connection with the backend system. The system quickly queries the corresponding locker number and verifies its validity. If the verification is successful, it enters S6 normal collection confirmation. The system sends a command to open the corresponding cabinet door. After the customer takes the items and closes the cabinet door, the camera records the collection process. The system then updates the database and marks the locker as idle to ensure that the collection process is accurately recorded.

[0104] If a customer accidentally loses his / her voucher, the lost property assistance information registration will come into play. The customer needs to go to the staff for help. The staff will record the key information such as the characteristics of the item and the storage time in detail at the special information registration terminal. Then, the lost property search will start. The staff will input this information into the background system. The system will combine the images collected by multiple cameras and use a series of advanced algorithms such as light adaptive image enhancement, multi-view fusion and occlusion reasoning, high-speed motion target tracking and customer position recognition to accurately determine the locker where the item is located. These algorithms are based on the Retinex algorithm, deep learning network, Kalman filter algorithm and computer vision algorithm, which effectively overcome the influence of multiple complex factors such as light changes, object occlusion, fast movement and customer position;

[0105] Finally, when picking up items, the staff will determine the cabinet where the items are located, enter the command at the main drive, open the corresponding cabinet, and allow the customer to take out their items smoothly, thus completing the entire collection process;

[0106] Through this interlocking step-by-step method, based on the coordinated operation of multiple advanced technical principles, the intelligent dispensing cabinet has achieved efficient, safe and convenient item dispensing management, and the beneficial effects it brings are remarkable. In terms of user experience, it has greatly improved convenience and efficiency. The initial dispensing guidance and smooth storage and dispensing process have greatly improved user satisfaction. In terms of security, real-time monitoring of items has effectively prevented abnormal situations and ensured the safety of items. At the data management level, data storage and update ensure the accuracy and timeliness of data, and provide strong support for subsequent query and analysis. In the case of lost credentials, the precise positioning function of lost property search solves the user's worries. On the whole, the dispensing method of the intelligent dispensing cabinet has significantly improved the level of intelligence and automation, reduced labor costs, and has broad application prospects and promotion value.

[0107] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0108] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and alterations may be made to the embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for collecting using an intelligent collection cabinet based on image recognition technology, characterized in that: The following steps are involved: S1: Provide guidance for the first time: When a new user first uses the voucher, he can use the touch screen next to the voucher to guide the terminal and learn how to select the voucher and get the voucher in the form of pictures, texts and voice; S2: Execute the deposit operation: the customer selects an empty cabinet to deposit the item, and the external camera records the entire process. The system binds the voucher to the cabinet and records the cabinet occupancy status; S3: Carry out real-time monitoring of items: The camera continuously monitors and transmits the captured images to the background in real time, and uses the YOLOv5 algorithm to analyze and detect abnormal situations; S4: Implement data storage and update: The system stores the data of deposits and collections, and updates it in real time to ensure that the data is accurate and timely; S5: Carry out normal identity verification for picking up items: When picking up items, customers insert the key or scan the code at the verification terminal, and the system verifies the operation information and determines the right to pick up items; S6: Complete normal item collection confirmation: If the verification is passed, the corresponding cabinet door is opened, and the cabinet door is closed after the customer takes the item. The system updates the database and marks the cabinet as free; S7: Register lost property assistance information: When a customer loses his / her voucher and asks for help, the staff will record the characteristics of the item and the time it was stored; S8: Execute lost property search: The staff inputs the lost property information into the backend, and the system combines multiple camera images to determine the locker where the item is located; S801 performs light-adaptive image enhancement: Retinex algorithm is introduced to separate brightness and reflectivity to avoid light interference; S802 realizes multi-view fusion and occlusion reasoning: by fusing multi-camera view information, deep learning is used to infer the occluded features; S803 performs high-speed moving target tracking: uses the Kalman filter algorithm to predict the target trajectory; S804 completes customer position identification: uses computer vision algorithm to determine the customer position and associates it with the locker position; S9: Execute item collection operation: After successfully locating the cabinet, the staff inputs the door opening command in the main drive control terminal, triggering the cabinet door to open and the customer to take the item.

2. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 1 is characterized in that: In step S3, the YOLOv5 algorithm, which is based on a convolutional neural network, converts the target detection task into a regression problem and directly predicts the category and position of the target in a network; Feature extraction network: It consists of a series of convolutional layers and pooling layers, which are used to extract features from images. It uses cross-stage local dense blocks and convolution kernels of different scales to extract multi-scale features. The calculation process is expressed as: x3=C oncatenate (x1,x2) Among them, x in is the input image, conv1 represents the i-th convolution operation, through which local and global features in the image are continuously extracted; Detection head: After feature extraction, the detection head is responsible for predicting the category and location of the target. For each predicted bounding box, its confidence score and the probability of the category to which it belongs are calculated. The confidence score indicates the possibility of the existence of the target in the bounding box. The confidence score calculation formula is: Confidewnce = σ(p obj ) Among them, p obj is the probability of the target existence predicted by the model, σ is the sigmoid function, which maps the probability value to the interval [0, 1]; For the position prediction of the bounding box, an offset method is used. Assume that the coordinates of the predicted bounding box center point are (x, y), the width and height are (w, h), and the offset relative to the grid unit is (t x , t y ), the scale factor is (t w , t h ), then: x=(c x +σ(t x ))×wy grid y=(c y +σ(t y ))×h grid w=e tw ×w anchor h=e th ×h anchor Among them, (c x , c y ) is the coordinate of the upper left corner of the grid cell, (w grid, h grid ) is the width and height of the grid cell, (w anchor ,h anchor ) are the preset width and height of the anchor box; In the real-time monitoring scenario of items, we pre-train the YOLOv5 model so that it can recognize cabinet doors and common items. When the images collected by the external camera are transmitted to the background image recognition and data processing unit, the model will quickly process the images to detect whether there are abnormal openings of cabinet doors and abnormal situations of stolen items. If an anomaly is detected, the reliability of the anomaly will be judged based on the confidence score and category information, and an alarm will be issued in time to notify the staff.

3. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 1 is characterized in that: In step S801, light adaptive image enhancement is specifically performed by the Retinex algorithm to perform adaptive illumination compensation. The core idea is to separate the brightness and reflectivity of the image, and to improve the clarity of the image under different illumination conditions by enhancing the reflectivity. The image gain is automatically reduced through formula calculation to avoid overexposure. The brightness distribution of the image is dynamically adjusted through analysis of the image histogram to ensure that the color and shape features of the object are not disturbed by light, so that the features of the object can be clearly presented under different illumination conditions.

4. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 1 is characterized in that: In step S802, multi-perspective fusion and occlusion reasoning specifically adopts a multi-perspective fusion and occlusion reasoning algorithm. When occlusion is detected, an occlusion reasoning network based on deep learning is used to infer the characteristics of the occluded part, combined with prior knowledge of the object, the shape and proportion of common objects. When the human body occludes part of the object, the shape and size of the occluded part are inferred by analyzing the unobstructed edges of the object and the known object category information, thereby completely identifying the object characteristics.

5. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 1 is characterized in that: In step S803, the moving target tracking specifically uses the Kalman filter algorithm to achieve high-speed moving target detection and tracking. The Kalman filter is an optimal recursive data processing algorithm based on the linear system state space model. Through continuous iterative prediction and update steps, the algorithm analyzes the position and shape changes of objects in continuous frame images, and predicts the movement trajectory of objects. Even if the objects move quickly, their feature points can be accurately captured. For items that are quickly taken out of a bag and put into a cabinet, the algorithm calculates their movement speed and direction based on the position changes of the objects in the previous and next frames of images, thereby stably tracking the objects and extracting their feature information.

6. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 1 is characterized in that: In step S804, customer position identification specifically introduces a position point recognition algorithm based on computer vision, and uses the image captured by the camera to determine the customer's position. First, the camera image is preprocessed, and the edge detection algorithm Canny algorithm is used to extract the contour information in the image. Through contour analysis, the approximate contour of the human body is identified, and the three-dimensional coordinates of multiple key points are analyzed to determine the center position of the customer's position. The average value of the three-dimensional coordinates of all key points is calculated as the center coordinate of the position. After the center position of the position is obtained, the position of the locker corresponding to the customer is determined in combination with the layout information of the locker.

7. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 3 is characterized in that: In step S801, the Retinex algorithm is used to perform adaptive illumination compensation for the influence of weather light. The algorithm separates the brightness and reflectivity of the image and improves the clarity of the image under different illumination conditions by enhancing the reflectivity. The calculation formula is: R(x,y)=log(I(x,y))-log(F(x,y)*I(x,y)) Among them, I(x, y) is the pixel value of the original image at the (x, y) position, L(x, y) is the illumination component, F(x, y) is a Gaussian low-pass filter, R(x, y) is the reflectance component, and S(x, y) is the enhanced image. On sunny days with sufficient light, the image gain is automatically reduced to avoid overexposure. On rainy days or at night when the light is dim, the reflectance component is enhanced, the image gain is increased, and the color balance is adjusted to ensure that the features of the object can be clearly presented under different lighting conditions.

8. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 4 is characterized in that: In step S802, multiple cameras photograph the same object from different perspectives, and the image features obtained by each camera are represented as F i (i=1,2,…,n), n is the number of cameras, and the fused feature F fusion Calculated by weighted average: Among them, w i It is a weight that is dynamically adjusted according to the camera position, angle, and image quality factors. When occlusion is detected, the occlusion reasoning network based on deep learning is used, combined with prior knowledge of the object, the shape and proportion of common objects, to infer the characteristics of the occluded part and achieve complete recognition of the object features.

9. The method for collecting using an intelligent collecting cabinet based on image recognition technology according to claim 5 is characterized in that: In step S803, during the high-speed moving target detection and tracking process, it is assumed that the state equation of the system is: X k =AK k-1 +BU k +W k The observation equation is: Z k =HK k +V k Among them, X k is the state vector at time k, which contains the position and speed information of the object; A is the state transfer matrix, which describes the change of state over time; B is the control matrix, U k is the control vector, which can be set to 0 in this scenario, W k is the process noise; Z k is the observation vector at time k, that is, the position information of the object in the image collected by the camera, H is the observation matrix, which maps the state vector to the observation space; V k It is observation noise. Through continuous iterative prediction and update steps, the movement trajectory of objects is analyzed and predicted according to the position and shape changes of objects in continuous frame images, and the feature points of fast-moving objects are accurately captured. For objects that are quickly taken out of the bag and put into the cabinet, the algorithm calculates their movement speed and direction according to the position changes of the objects in the previous and next frames of images, thereby stably tracking the objects and extracting their feature information.

10. The intelligent collection cabinet collection method based on image recognition technology according to claim 6 is characterized in that: In step S804, the camera image is preprocessed using the position point recognition algorithm of computer vision, and the contour information in the image is extracted using the edge detection Canny algorithm. The rough contour of the human body is identified through contour analysis. It is assumed that the key point on the human body contour is P j (j=1, 2, ..., m), m is the number of key points, for each key point, according to its pixel coordinates (x j ,y j ), combined with the camera calibration parameters, the intrinsic matrix and the extrinsic matrix, the three-dimensional coordinates (X j , Y j , Z j ): Among them, (u j, v j ) is the pixel coordinate of the key point on the image plane. Then, the customer's standing center position is determined by analyzing the three-dimensional coordinates of multiple key points. The average value of the three-dimensional coordinates of all key points is calculated as the standing center coordinate. After obtaining the standing center position, the locker position corresponding to the customer is determined in combination with the layout information of the locker.

Citation Information

Patent Citations

  • Foreign matter detection method and device, electronic equipment and storage medium

    CN114724025A

  • Method and system for detecting left articles in security and protection monitoring based on YOLOv4 and Depsort

    CN115984199A

  • Locker device, locker system, and method for controlling locker device

    JP2022166507A