A luggage matching method based on image retrieval

The luggage matching method using image retrieval and weight verification addresses the inefficiencies in luggage loss by accurately identifying and retrieving lost luggage through a YoloV5 and ResNet-50 model, enhancing efficiency and reducing costs.

CN115131654BActive Publication Date: 2025-07-15KNOWYOU INFORMATION TECH SHANGHAI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210889494.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-07-15
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

During the transportation of air luggage, the luggage label is damaged, unclearly printed or not tightly hung, resulting in the inaccurate identification of the luggage, resulting in loss. The existing technology is inefficient and it is difficult to efficiently track and retrieve the lost luggage.

Method used

Using image retrieval combined with weight matching method, by constructing a YoloV5 object detection model and ResNet-50 convolutional neural network, the luggage image features are extracted and binary encoding is generated, stored in a cloud server, and matching query is used to achieve efficient identification and tracking of luggage.

Benefits of technology

It improves the accuracy and efficiency of luggage loss tracking, reduces the passive work of customer service personnel, reduces the abnormal expenses of airlines, and improves passenger satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131654B_ABST
    Figure CN115131654B_ABST
Patent Text Reader

Abstract

The present invention relates to a luggage matching method based on image retrieval. The luggage target is obtained by constructing a luggage target detection model in advance, and then the target is subjected to feature extraction through a feature extraction model. After that, the output of the feature extraction is sent to a luggage picture auto-encoding model for binary encoding, and the generated binary encoding of the luggage photo is stored in a cache database. When there is a situation of luggage tag loss, only need to take a picture of the tagged luggage and upload it. The binary encoding of the tagged luggage photo is obtained through the luggage target detection model, the feature extraction model, and the luggage picture auto-encoding model, and the binary encoding is used to query in the cache database to find the corresponding passenger information of the luggage. The method of the present invention can effectively help the lost tagged luggage to find the owner, reducing the passivity of customer service staff when looking for luggage. The present invention adopts a method combining image retrieval and weight matching, greatly improving the accuracy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of civil aviation, and particularly relates to a baggage matching method based on image retrieval, mainly adopting the method of image retrieval and matching, supplemented by weighing or other means. Background Art

[0002] During the process of air baggage transportation, there are many links in the baggage handling and transportation, and there are many transferred baggages, which often leads to the problem of baggage loss. Therefore, the accurate tracking of baggages, especially the retrieval of lost baggages, is particularly important for reducing the cost of abnormal baggage handling. The main reason for baggage loss is that the baggage label is damaged, the printing is unclear, or it is not firmly attached, resulting in the inability to identify the baggage, causing the baggage to be sorted out and not loaded onto the plane or sent to the wrong flight at a certain link in the transportation. The lost baggage may appear at any airport, and finally the airport or the airline customer service department will analyze its origin. For passengers and carriers of lost baggages, they need to report at the baggage inquiry desk or actively call airports across the country to find their own baggages. At present, after a passenger reports the loss of their baggage, the airline will search one flight after another according to the flight the passenger took, with low efficiency. Summary of the Invention

[0003] Aiming at the problems existing in the prior art, the present invention provides a baggage matching method based on image retrieval, including the following steps:

[0004] (1) Construct a baggage target detection model; construct a YoloV5 target detection model, and the constructed model is stored on a cloud server;

[0005] (2) When checking in and consigning baggages, take pictures and weigh the passengers' baggages;

[0006] (3) After taking pictures and weighing, upload the pictures of the flight to the cloud server; load the pre-trained baggage target detection model (YoloV5 target detection model) in step (1), perform target detection on the baggages in the baggage pictures, and after detecting the bounding boxes of the baggages, then crop out the pictures containing only the baggages; the pictures are sent into a pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into a baggage picture autoencoder model for binary coding, and then the generated binary coding of the baggage pictures is stored in a cache database; the said baggage picture autoencoder model is an autoencoder-decoder model;

[0007] (4) Finally, after the flight arrives, if the identification tag of the luggage falls off, the luggage needs to be photographed and weighed; the taken pictures are sent to the cloud server; the pre-trained luggage object detection model in step (1) is loaded to perform object detection on the luggage in the luggage pictures. After the bounding box of the luggage is detected, then the picture containing only the luggage is cropped; the picture is sent into the pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into the luggage picture autoencoder model for binary encoding, and the obtained binary encoding is used to query in the cache database in step (3). Ten pictures with the highest similarity are queried, the weight information corresponding to the luggage photo is matched, and finally 5 pictures with both high picture similarity and weight matching degree are found, and the corresponding passenger information of these 5 pictures is found and returned to the flight arrival client.

[0008] Based on the above solution, the construction method of the YoloV5 object detection model in step (1) is as follows:

[0009] First, a large number of pictures with luggage are collected (the pictures are selected from common flight luggage photos, covering various types of luggage as much as possible, such as suitcase type, bag type, and various colors and styles), about 5000 pictures. The LabelMe software is used to draw the bounding box of the luggage with a rectangle and save it as a json format file. The json file contains the coordinate information of the rectangle box of the luggage target, which are (xmin, ymin) representing the upper left corner coordinates and (xmax, ymax) representing the lower right corner coordinates. Together with the marked json file and the original picture file, they are used as training data. At the same time, the training data is sampled and grouped at a ratio of 5:1. The part with a ratio of 5 continues to be used as the training data set, and the part with a ratio of 1 is used as the validation data set. The training data and the validation data are sent into the YoloV5 object detection model. After 500 generations of training, the model converges, and a trained YoloV5 model with the highest accuracy for luggage detection is obtained, which is in the binary file format and can be used for the detection of luggage targets in the luggage pictures in the subsequent steps. It should be particularly noted that the training of the object detection model is usually a one-time task and does not require model training for each flight's data set.

[0010] Based on the above solution, the luggage picture autoencoder model generates a binary encoded output for the input data as the hash value for picture search.

[0011] Based on the above solution, the luggage picture autoencoder model consists of an encoder, an intermediate layer, and a decoder.

[0012] Based on the above solution, the specific method of step (4) is as follows: After the flight arrives, the unclaimed baggage is sent to the unclaimed baggage retrieval area, where photos are taken and the weight is measured. Then, both the photos and the weight information are uploaded to the cloud server. After receiving the photos, the pre-trained baggage object detection model (YoloV5 pre-trained baggage detection model) in step (1) is loaded to perform object detection on the baggage in the baggage pictures. After detecting the bounding box of the baggage, the picture containing only the baggage is then cropped; the picture is sent into the pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into the baggage picture autoencoder model for binary encoding, and the obtained binary encoding is used to query in the cache database in step (3). Ten pictures with the highest similarity are queried, the weight information corresponding to the baggage photo is matched, and finally, five pictures with the highest rankings in terms of both picture similarity and weight matching degree are found, and the corresponding passenger information of these five pictures is found for manual identification.

[0013] The present invention also provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned baggage matching method based on image retrieval are implemented.

[0014] The present invention also provides a server and a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned baggage matching method based on image retrieval are implemented.

[0015] The baggage matching system based on image retrieval provided by the present invention has the following effects:

[0016] This method can effectively help find the owners of lost unclaimed baggage, reducing the passivity of customer service staff when looking for baggage. It can save a large amount of abnormal expenses for airlines.

[0017] The present invention adopts a method combining image retrieval and weight matching, greatly improving the accuracy.

[0018] The pictures are uploaded to the cloud server, and the query service is provided by the cloud server, which is convenient for the deployment and implementation of the entire technical solution and horizontal capacity expansion.

[0019] Before feature extraction and encoding of the baggage pictures, a classic object detection model YoloV5 is trained. This model is used to perform object detection on the baggage pictures and crop out sub-pictures containing only the baggage, removing the noise of the conveyor belt and airport background. For the encoding, query, etc. of the baggage pictures, noise interference can be excluded, improving the accuracy.

[0020] When encoding the features extracted from the luggage pictures, a pre-trained auto-encoding model based on the encoder-decoder architecture is used. The encoder part of this model outputs binary latent variables and continuous latent variables. The binary latent variables are converted into an adjacency matrix, which, together with the continuous latent variables, serves as the input to the graph neural network (GCN). The output of the graph directly goes to the decoding layer, and after passing through a fully connected layer and an activation function, the output is obtained. The internal structure of the binary latent variables of this structure is fully utilized and discovered during the process of constructing the adjacency matrix to form a graph, and the constructed graph is encoding-driven. The continuous variables focus more on the detailed information of the pictures, and they will feedback the updates of the updated network variables to the encoder through backpropagation, so as to learn more discriminative binary encodings. The final effect is that when the entire model serves as an encoder, its binary encoding output can consider both high-resolution picture information and discriminative details.

[0021] On the cloud server side, after the picture features are extracted and encoded, they are stored in the cache server for subsequent high-performance search, making the overall system performance extremely high and enabling a second-level response to be obtained on the query client. Brief Description of the Drawings

[0022] The following further describes the system and method of the present invention with reference to the accompanying drawings:

[0023] Figure 1 It is the overall structure diagram of the system;

[0024] Figure 2 It is the flow chart of check-in luggage consignment;

[0025] Figure 3 It is the flow chart of the training of the luggage target detection YoloV5 model;

[0026] Figure 4 It is the flow chart of the training of the auto-encoding model of luggage pictures based on the encoder-decoder architecture;

[0027] Figure 5 It is the structure diagram of ResNet-50 feature extraction;

[0028] Figure 6 It is the structure diagram of the auto-encoder model of luggage pictures based on the encoder-decoder architecture;

[0029] Figure 7 It is the flow chart of the feature extraction and encoding storage of luggage pictures on the cloud server;

[0030] Figure 8 It is the effect diagram of the target detection and cropping of luggage pictures;

[0031] Figure 9 It is the flow chart of the retrieval service of the luggage pictures with tags missing on the cloud server;

[0032] Figure 10 Flow chart for finding the owner of lost luggage that arrives at the client for a flight. Detailed implementation manner

[0033] The following uses specific embodiments to further describe the technical solution of the invention, but the protection scope of the invention is not limited to the following embodiments.

[0034] As Figure 1 shown, a luggage matching method based on image retrieval provided by the present invention includes the following steps:

[0035] (1) Construction of a luggage target detection model (YoloV5 target detection model); the constructed model is stored on a cloud server;

[0036] The training process of the luggage target detection model (YoloV5 target detection model) is as Figure 3 shown. First, a large number of pictures with luggage are collected (these pictures are selected from common flight luggage photos, covering various types of luggage as much as possible, such as suitcase-type, bag-type, and various colors and styles), about 5,000 pictures. The bounding box of the luggage is drawn as a rectangle using LabelMe software and saved as a json-format file. The json file contains the coordinate information of the rectangle box of the luggage target, which are (xmin, ymin) representing the upper left corner coordinates and (xmax, ymax) representing the lower right corner coordinates. The marked json file and the original picture file are used together as training data. At the same time, the training data is sampled and grouped at a ratio of 5:1. The part with a ratio of 5 continues to be used as the training data set, and the part with a ratio of 1 is used as the validation data set. The training data and the validation data are fed into the YoloV5 target detection model. After 500 generations of training, the model converges to obtain a trained YoloV5 model with the highest accuracy for luggage detection, which is in binary file format and can be used for detecting the luggage target in the luggage pictures in the subsequent steps. It should be particularly noted that the training of the target detection model is usually a one-time task and does not require training a model for each flight's data set.

[0037] (2) When checking in and checking in luggage, take pictures of the passenger's luggage and weigh the luggage;

[0038] Specifically, the specific description of checking in and checking in luggage is as Figure 2 shown. When a passenger checks in, the checked luggage is submitted to the airline staff. When the luggage is placed on the weighing platform, the luggage is weighed. At the same time, the camera takes pictures of the luggage from four different angles (such as directly in front of the luggage placement direction, from a top view, and from the left and right side view directions). The obtained photo data, weight data, and the corresponding passenger information (flight number, name, ID number, etc.) are simultaneously transmitted to the cloud server and stored on the cloud server.

[0039] (3) After the photographing and weighing are completed, the photos of the flight are uploaded to the cloud server; on the cloud server, the application loads a pre-trained luggage object detection model (YoloV5 object detection model) to perform object detection on the luggage in the luggage picture. After the bounding box of the luggage is detected, a picture containing only the luggage is then cropped. Since the original luggage picture will have some noise interference, especially the conveyor belt and the airport background, these interferences are not conducive to subsequent image retrieval and will cause the retrieval to match these interference information. After cropping, the luggage picture information is more concentrated, excluding interference information such as the background. This picture will be fed into the pre-trained ResNet-50 model (ResNet( https: / / arxiv.org / abs / 1512.03385v1 ), and the output of the penultimate layer is extracted as the output of feature extraction. In order to encode the extracted features, we need to load a pre-trained autoencoder model for luggage pictures. The output of image feature extraction is used as the input of the autoencoder model for luggage pictures, and the binary latent variable of the autoencoder model for luggage pictures is extracted as the binary encoding output. This output is also the binary encoding of the luggage picture and is used for subsequent retrieval. In order to perform more efficient large-scale picture retrieval, the present invention stores the generated binary encoding of the luggage picture in a cache database, such as Redis, for subsequent rapid query.

[0040] The ResNet-50 model is a model based on residual learning, which solves the performance degradation of deep networks. The above-mentioned pre-trained ResNet-50 model takes the images in the ImageNet( https: / / image-net.org / ) image dataset as input to train the ResNet-50 model, and a pre-trained ResNet-50 model will be obtained. This model has learned a large amount of ImageNet picture information, so it can be used as a pre-trained model for other image-related deep learning, which is also called transfer learning. This method uses the ResNet-50 pre-trained model to extract features, enabling the model to fully consider the features of the massive ImageNet picture data. ImageNet is a large visual database for visual object recognition software research, with more than 14 million images.

[0041] The process of training the autoencoder model for luggage pictures is as Figure 4As shown, first, to prepare the training data, it is necessary to extract the features of the luggage picture data, and the output of the feature extraction for each picture is a vector. To construct the training data for the autoencoder, the input of the training data is the result of the feature extraction, and the label of the training data is the same as the input. Then, the training data is grouped at a ratio of 5:1. The part with a ratio of 5 continues to be used as the training data set, and the part with a ratio of 1 is used as the validation data set. After the training set and the validation set are ready, they are fed into the loaded autoencoder-decoder model. First, calculate the forward propagation result, and then perform backpropagation to check whether the number of iterations has reached the maximum number of iterations. If the maximum number of iterations of 500 times has been reached, stop the process and save the trained binary model. Otherwise, return to the forward propagation and continue the iteration.

[0042] Among them, the feature extraction of the luggage data is as Figure 5 shown. The luggage picture data set of this flight is used as the input of the pre-trained ResNet-50 convolutional neural network model. After passing through 5 convolutional layers, the output of the second-to-last layer is taken as the output. This is the extracted picture feature vector.

[0043] The structure of the autoencoder model based on the encoder-decoder architecture is as Figure 6 shown. The purpose of this model is to generate a binary encoded output for the input data as the hash value for picture search. For example, for the data set the purpose is to learn an encoding function

[0044] Here, N refers to the size of the data set, D refers to the dimension of the input, M refers to the length of the output binary, X represents the data set vector, i represents the subscript of the vector, and R represents the set of real numbers.

[0045] As Figure 6 shown, the model consists of an encoder, an intermediate layer, and a decoder.

[0046] In the encoder, after accepting the feature extraction vector of ResNet-50 as the input, it passes through a fully connected layer and the Relu activation function to obtain the output.

[0047] This output is divided into two paths. The first path first passes through a fully connected layer and the Sigmoid activation function, and then is input to the discrete random element. The output of the discrete random element is encoded by the encoding function to obtain the binary latent variable b. The encoding function can be expressed by the following formula:

[0048] b = α(f1(x; θ1), ∈) ∈ {0, 1} M ,

[0049] Here, \(x\) is the input vector, \(\theta_1\) refers to the parameters of the network, \(f_1\) refers to the fully connected layer and the non-linear sigmoid activation function, and \(M\) refers to the length of the output binary. \(\alpha(\cdot,\in)\) here refers to the element-wise discrete random element function, which contains random variables ranging from 0 to 1 and is used as Figure 4 the backpropagation in

[0050]

[0051] where \(i\) represents the \(i\)-th element of the vector, and \(b\) is the output binary vector of 0 and 1.

[0052] The second path of the output passes through a fully connected layer and a Relu activation function to obtain the continuous latent variable \(z\), which is expressed by the formula as follows:

[0053]

[0054] where \(x\) represents the input vector, \(\theta_2\) represents the parameters of the network, \(f_2\) represents the fully connected and Relu activation functions, and \(L\) represents the dimension of the output \(z\).

[0055] In the middle layer, first, the binary latent variable \(b\) and the continuous latent variable \(z\) are received as inputs. For simplicity, denoted as a batch size of binary latent variables, where \(b\) i represents each element of the binary latent variable, \(N\) B represents the maximum length, and \(M\) represents the dimension.

[0056] Denoted as a batch size of continuous latent variables, where \(z\) i represents each element of the binary latent variable, \(N\) B represents the maximum length, and \(L\) represents the dimension.

[0057] To construct a graph based on all the training data, each data is a vertex of the graph, and the edges are determined by the Hamming distance between the binary latent variables. The normalized graph adjacency matrix is calculated as follows:

[0058]

[0059] where is a matrix of all 1s (i.e., all elements in the row and column are 1), \(M\) represents the dimension, and \(B\) B is the binary latent variable.

[0060] For each element in \(A\), this formula is equivalent to the following formula:

[0061] A​​​​ij = 1 - Hamming(b i , b j ) / M

[0062] Here, b i represents the elements of B B , M represents the dimension, and the Hamming code is a coding method for linear debugging codes.

[0063] The processed adjacency matrix A, together with the continuous latent variable z B , serves as the input to the GCN (Graph Convolutional Network), and its output is a vector of dimension L, with each value being (0, 1), denoted as Z'.

[0064] In the decoder, first, it passes through a fully connected layer and a Relu activation function, and then the output is passed to a fully connected layer and an equal activation function.

[0065] The auto - encoding loss function of this model is as follows:

[0066]

[0067] Here, λ is a hyperparameter, b is the binary latent variable, x represents the true value, is the output from the decoder, represents the expectation of the binary latent variable. and are two difference functions, and represent the network parameters, N B represents the maximum length.

[0068] The encoder-decoder architecture is based on an autoencoder model. In the encoder part of this model, binary latent variables and continuous latent variables are output. The binary latent variables are transformed into an adjacency matrix and, together with the continuous latent variables, serve as the input to a graph neural network (GCN). The output of the graph directly goes to the decoding layer. After passing through a fully connected layer and a ReLU activation function, the output is obtained. For pictures of luggage, the internal structure of the binary latent variables of this structure is fully utilized and discovered in the process of constructing the adjacency matrix to form a graph. The constructed graph is encoding-driven, which is conducive to fast encoding indexing and searching during the luggage search process. The continuous variables focus more on the detailed information of the picture. They will feedback the updates of the updated network variables to the encoder in a backpropagation manner, so as to learn more discriminative binary codes. Some luggage has concentrated colors (such as red boxes, black boxes, etc.), and it is necessary to search and find from the detailed parts. By applying the encoder-decoder architecture based on the autoencoder model, the final effect is that when the entire model acts as an encoder, its binary code output can consider both high-resolution picture information and discriminative details, so as to quickly and accurately find the lost luggage in practical applications.

[0069] (4) Finally, after the flight arrives, if the identification tag of the luggage falls off, the luggage cannot be located to the correct flight arrival luggage carousel and retrieved. At this time, pictures of the tagless luggage are taken (for example, 4 pictures are taken from the front, top-down, and two left and right side views of the luggage placement direction), and then the luggage is weighed. The taken pictures are sent to the cloud server, and the cloud server calls the tagless luggage picture retrieval service to query the top 5 most similar pictures and the corresponding passenger information, and returns them to the client (such as an iPad or a self-service query terminal set at the airport arrival area). The tagless luggage retrieval client receives the similar pictures and the corresponding passenger information, and then further confirms with the weight to find the correct passenger information; the identification tag of the luggage can be a QR code tag.

[0070] Specifically, for the process of finding the owner of tagless luggage arriving at the client by flight, as Figure 10As shown in the figure, after the flight arrives, the unclaimed baggage cannot be identified and sent to the baggage carousel, so it is sent to the unclaimed baggage retrieval area. First, 4 photos are taken from multiple angles (for example, from the front directly in the direction of the baggage placement, from the top-down view, and from the left and right side views respectively), and then it is weighed to obtain the weight of the baggage. The photos and the weight are both uploaded to the cloud server. After receiving the photos, the unclaimed baggage image retrieval service on the cloud server calls the retrieval service and returns 10 most similar images of the query image. At the same time, the cloud server also stores the passenger information corresponding to each image and the weight information of the baggage corresponding to this image. Then, through the weight data uploaded by the client, weight matching is performed. The error of weight matching is 10 grams, and the higher the matching degree, the higher the ranking. Thus, by combining image retrieval and weight matching, 5 images with high similarity in pictures and high matching degree in weight are found, and the passenger information corresponding to these 5 images is found and returned to the flight arrival client. Then, the client can query the passenger information with matching similarity in pictures and weight. After manual identification, the owner of the unclaimed baggage can be quickly found. Specifically, after querying the 10 images with the highest similarity, they are sorted in ascending order of approximation as a, b, c, d, e, f, g, h, i, j, and each photo is given a score in this order. The scores are 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 in ascending order of approximation; the weights of the baggage corresponding to the 10 photos are w a , w b , w c , … w j , first calculate the sorting of the difference in baggage weight, calculate abs(w - w k ), where abs represents the absolute value, w represents the weight of the baggage to be searched, and the subscript k represents 10 baggage photos from a to j. Calculate the absolute value of the difference respectively, and then sort them in ascending order of the difference and give scores from 10 to 1 respectively; finally. Add the similarity scores (from 10 to 1) of each photo and the corresponding weight scores (from 10 to 1), and sort the comprehensive scores from high to low. Finally, the 5 photos with the highest comprehensive scores considering both scores and weights can be obtained as the selected baggage photos.

[0071] Specifically, for the retrieval service of unclaimed baggage images on the cloud server, its working process is as follows Figure 9As shown, when the cloud server application service receives the lost-tag luggage picture to be queried, it first loads the pre-trained YoloV5 luggage detection model to perform object detection on the picture. If a luggage target can be detected, a sub-picture containing only the luggage target is cropped from the original picture as the input for the next feature extraction; if no luggage target is detected, this does not necessarily mean that there is no luggage in the picture. It is possible that there is a luggage with a completely different shape and pattern from before, such as a special-shaped bag, which cannot be recognized by the pre-trained luggage target detection model. In this case, the original uncropped picture is used as the input for the next feature extraction. The picture is input into the pre-trained ResNet-50 model, and the output of the penultimate layer is taken as the result of feature extraction. To encode the extracted features, we need to load the pre-trained autoencoder model for luggage pictures. The result of feature extraction by the ResNet-50 model is used as the input for the autoencoder model for luggage pictures, and the binary latent variable of the autoencoder model for luggage pictures is extracted as the binary coding output, which is also the binary coding of the luggage picture. For the binary coding of this picture, search in the cache database of image feature coding (i.e., the saved binary codings that have been trained by the autoencoder model for luggage pictures) to find the binary coding that matches the closest, so as to return the corresponding 10 most similar luggage pictures.

[0072] To verify the effectiveness of this method, about 5000 luggage photos were collected as a picture set on a domestic flight, and 100 of the luggage photos were simulated as lost-tag luggage. To better compare the results, several classic methods of existing image search were selected for comparison. The selected classic methods are:

[0073] siaMAC(https: / / arxiv.org / pdf / 1604.02426v3.pdf)

[0074] LIFT(https: / / arxiv.org / pdf / 1603.09114v2.pdf)

[0075] The experimental environment is the Ubuntu system, the python3.8 environment, and a deep learning platform based on pytorch. To measure the effect, precision (the percentage of the number of correct pictures obtained in the search / the number of actually correct pictures) is used for comparison. The results are shown in Table 1:

[0076] Table 1 Luggage Search Precision

[0077] model Accuracy (%) siaMAC 87.2 LIFT 70.7 the model of the present invention 97.8

[0078] As can be seen from the results, the model of the present invention has higher accuracy, can more accurately search for and match the luggage with missing tags, thereby helping airlines improve the efficiency of lost luggage search and reduce costs, and improve passenger satisfaction.

[0079] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A luggage matching method based on image retrieval, characterized in that, The steps are as follows: (1) Construct a luggage target detection model; construct a YoloV5 target detection model, and the constructed model is stored on the cloud server; the specific construction method is as follows: First, collect a large number of pictures with luggage, which are selected from common flight luggage photos, covering various types of luggage such as suitcase type, bag type, and various colors and styles. Use LabelMe software to draw the bounding boxes of the luggage with rectangles and save them as json format files. The json files contain the coordinate information of the rectangular boxes of the luggage targets, which are (xmin, ymin) representing the upper left corner coordinates and (xmax, ymax) representing the lower right corner coordinates respectively; Together with the marked json files and the original picture files, they are used as training data. At the same time, the training data is sampled and grouped at a ratio of 5:

1. The part with a ratio of 5 continues to be used as the training data set, and the part with a ratio of 1 is used as the validation data set; the training data and the validation data are sent into the YoloV5 target detection model. After 500 generations of training, the model converges, and a trained YoloV5 model with the highest accuracy for luggage detection is obtained, which is in binary file format and can be used for the detection of luggage targets in luggage pictures in the subsequent steps; (2) When checking in and checking luggage, take pictures and weigh the luggage of the passengers; (3) After taking pictures and weighing, upload the pictures of the flight to the cloud server; Load the pre-trained luggage target detection model in step (1). The luggage target detection model is a YoloV5 target detection model. Perform target detection on the luggage in the luggage picture. After detecting the bounding box of the luggage, crop out the picture containing only the luggage; the picture is sent into the pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into the luggage picture autoencoder model for binary encoding, and then the generated binary encoding of the luggage picture is stored in the cache database; the luggage picture autoencoder model is an autoencoder-decoder model; (4) Finally, after the flight arrives, if the identification tag of the luggage falls off, it is necessary to take pictures and weigh the luggage; the taken pictures are sent to the cloud server; load the pre-trained luggage target detection model in step (1), perform target detection on the luggage in the luggage picture, and after detecting the bounding box of the luggage, crop out the picture containing only the luggage; the picture is sent into the pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into the luggage picture autoencoder model for binary encoding, use the obtained binary encoding to query in the cache database in step (3), query out 10 pictures with the highest similarity, match the weight information corresponding to the luggage pictures, and finally find the top 5 pictures with both high picture similarity and weight matching degree, find the corresponding passenger information of these 5 pictures, and return it to the flight arrival client.

2. The baggage matching method based on image retrieval according to claim 1, wherein In the construction method of the YoloV5 target detection model in step (1), the number of pictures with luggage collected is 5000.

3. The baggage matching method based on image retrieval according to claim 1, characterized in that, The described self-encoding model for luggage pictures generates a binary encoded output for the input data to serve as the hash value for picture search.

4. The method for matching luggage based on image retrieval according to claim 3, characterized in that, The described self-encoding model for luggage pictures consists of an encoder, an intermediate layer, and a decoder; After the encoder is used to accept the feature extraction vector of ResNet-50 as input, it passes through a fully connected layer and a Relu activation function to obtain an output; this output is divided into two paths. The first path first passes through a fully connected layer and a Sigmoid activation function, and then is input to a discrete random element. The output of the discrete random element is encoded by an encoding function to obtain a binary latent variable b; the encoding function can be expressed by the following formula: b = α(f1(x; θ1), ∈) ∈ {0, 1} M , Here, \(x\) is the input vector, \(\theta_1\) refers to the parameters of the network, \(f_1\) refers to the fully connected layer and the non-linear sigmoid activation function, \(M\) refers to the length of the output binary; \(\alpha(\cdot,\in)\) here refers to the element-wise discrete random element function, which contains random variables ranging from 0 to 1; The second path of this output passes through a fully connected layer and a Relu activation function to obtain a continuous latent variable z, which is expressed by the following formula: Here, x represents the input vector, θ2 represents the parameters of the network, f2 represents the fully connected and Relu activation functions, and L represents the dimension of the output z; The middle layer first receives the binary latent variable b and the continuous latent variable z as inputs; Denote the binary latent variables with a batch size, where b i represents each element of the binary latent variable, N B represents the maximum length, and M represents the dimension; Denote the continuous latent variables with a batch size, where z i represents each element of the binary latent variable, N B represents the maximum length, and L represents the dimension; To construct a graph based on all the training data, each data point is a vertex of the graph, and the edges are determined by the Hamming distance between binary latent variables; the normalized graph adjacency matrix is calculated as follows: Here is a matrix of all 1s (i.e., all elements in rows and columns are 1), M represents the dimension, and B B is a binary latent variable; for each element in A, this formula is equivalent to the following formula: A ij = 1 - Hamming(b i , b j ) / M Here b i represents an element of B B , M represents the dimension, and the Hamming code is a coding method for linear debugging codes; The processed adjacency matrix A and the continuous latent variable Z B are used together as the input to the GCN (Graph Convolutional Network), and its output is a vector of dimension L, each value of which is in the range (0, 1), denoted as Z'; In the decoder, first pass through a fully connected and Relu activation function, and then the output is passed to a fully connected and equal activation function; The self-encoding loss function of this model is as follows: Here, λ is a hyperparameter, b is a binary latent variable, and x represents the true value. is the output from the decoder, representing the expectation of the binary latent variable; and are two discriminative functions, and represent the network parameters, and N B represents the maximum length.

5. The baggage matching method based on image retrieval according to claim 1, characterized in that The specific method of step (4) is: after the flight arrives, the tagged luggage is sent to the tagged luggage retrieval area, photographed and weighed. Then, the photo and weight information are both uploaded to the cloud server. After receiving the photo, load the pre-trained luggage target detection model in step (1) to perform target detection on the luggage in the luggage picture. After detecting the bounding box of the luggage, crop out the picture containing only the luggage; the picture is sent into the pre-trained ResNet-50 convolutional neural network model for feature extraction, and the output of the penultimate layer is extracted; the output of the feature extraction is sent into the self-encoding model for luggage pictures for binary encoding, and the obtained binary encoding is used to query in the cache database in step (3). Query out 10 pictures with the highest similarity, match the weight information corresponding to the luggage photo, and finally find the top 5 pictures with both high picture similarity and weight matching degree, and find the corresponding passenger information of these 5 pictures for manual identification.

6. The method for matching luggage based on image retrieval according to claim 5, wherein, If the luggage target detection model fails to detect the luggage target in the photographed picture, the original uncropped picture is used as the input for the next feature extraction.

7. The method for luggage matching based on image retrieval according to claim 5, wherein After querying out the 10 images with the highest similarity, sort them in descending order of approximation as a, b, c, d, e, f, g, h, i, j, and assign a score to each photo in this order. The scores are 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 in descending order of approximation; the weights of the luggage corresponding to the 10 photos are w a , w b , w c ,…w j , first calculate the sort of the luggage weight difference, calculate abs(w - w k ), where abs represents the absolute value, w represents the weight of the luggage to be searched, and the subscript k represents the 10 luggage photos from a to j. Calculate the absolute value of the difference respectively, then sort them in ascending order of the difference, and assign scores from 10 to 1 respectively; finally, add the similarity score of each photo and the score of the corresponding weight, sort the comprehensive scores from high to low, and finally 5 photos with the highest comprehensive scores considering both scores and weights can be obtained as the selected luggage photos. Among them, both the similarity score of the photo and the score of the corresponding weight are from 10 to 1.

8. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the luggage matching method based on image retrieval according to any one of claims 1-7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the luggage matching method based on image retrieval according to any one of claims 1-7.

Citation Information

Patent Citations

  • Auxiliary retrieval method for lost luggage of civil aviation

    CN110147749A