A method for detecting foreign objects in the passage of a mobile rack based on prompt extension continuous learning
By introducing a continuous learning method based on prompt extension in the dense rack foreign object detection network model, the problem of poor performance of the existing technology in environmental changes and new scenario migration is solved, the model is achieved with high generalization and universality, and the new scenario deployment process is simplified.
Patent Information
- Application Number
- CN202410374031.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-03-29
AI Technical Summary
The existing dense rack foreign object detection network model based on deep learning is poor in environmental changes and new scenario migration, and it is difficult to adapt to the changes in new data, resulting in cumbersome system deployment process and long cycles.
Using a continuous learning method based on prompt extension, the foreign object detection network model is designed including pre-trained feature extractor, prompt pool and classification head pool. Through dynamic updates of prompt pool and classification head pool, the continuous learning of the model on different data sets is achieved without retraining.
The model is highly generalized and versatile, and can adapt to different site conditions, accelerate the deployment of dense rack systems in new scenarios, reduce deployment complexity and improve efficiency.
Smart Images

Figure CN118172729B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of intelligent image processing, deep learning, and domain incremental continuous learning, and specifically relates to a method for detecting foreign objects in the passage of a compactus based on prompt expansion continuous learning. Background Art
[0002] The intelligent compactus is an actual application of mechatronics technology and is also one of the main development directions of the green, intelligent, and smart modern storage management. Its simple operation and powerful functions improve work efficiency and are widely used in units such as archives and libraries. In the compactus system, the safety of the passage is particularly crucial.
[0003] Currently, to meet the requirements of detecting foreign objects in the passage of the compactus with full coverage and low cost, the new generation of intelligent compactus uses cameras to collect images of the passage inside the compactus, and cooperates with deep learning algorithms to classify the images for the presence or absence of foreign objects, thereby realizing real-time monitoring of foreign objects in the passage. This method has the advantages of strong detection ability, high speed, and low cost. However, as the system is deployed, the existing problems are also exposed. Since deep learning requires a large amount of data for training and static models are difficult to adapt to the changes of new data, the performance of the foreign object detection network model based on deep learning is poor under environmental changes such as lighting, floor patterns, and shadows, and it is difficult to adapt to new scene migrations. This forces technicians to frequently expand data and retrain the network model, resulting in a cumbersome and long deployment process for the compactus system in new scenarios or even the same scenario, which is not conducive to the popularization of the intelligent compactus system. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting foreign objects in the passage of a compactus based on prompt expansion continuous learning, which can realize the continuous learning of the model on different data sets, without the need for retraining and with a small increase in the parameters of the continuous learning model, thereby having high generalization and versatility, being able to adapt to different site conditions, and accelerating the deployment of the archive compactus system in new scenarios.
[0005] The technical method adopted by the present invention is: a method for detecting foreign objects in the passage of a compactus based on prompt expansion continuous learning, including the following steps:
[0006] S1: Deploy network cameras on the compactus to collect images of various foreign objects in the passage, and scale the images to h×h pixels to match the model size, and construct a foreign object image data set. Each task corresponds to 1 foreign object image data set;
[0007] S2: Design a foreign object detection network model, where the foreign object detection network model includes a pre-trained feature extractor, a prompt pool, and a classification head pool; the pre-trained feature extractor includes a pre-trained embedding layer and a pre-trained Transformer encoder; the parameters of the pre-trained embedding layer and the pre-trained Transformer encoder are not updated with model training, and the prompt pool and the classification head pool are updated with model training as trainable parameters;
[0008] S3: Use the foreign object image dataset obtained in step S1 to train the foreign object detection network model;
[0009] S4: Use the foreign object detection network model trained in step S3 to perform foreign object detection on the images collected by the camera, and detect the foreign object status in the channel in real time;
[0010] S5: When the foreign object detection network model learns a new task or is migrated to a new scenario for application, continuously learn the new task or new scenario.
[0011] Further, the specific steps of step S1 are as follows:
[0012] S101: Shoot the picture in the channel through the network camera deployed on the compact shelf; collect the internal images of the channel with different foreign object types, different illuminations, or different shelf spacings, store them according to the presence or absence of foreign object types, and generate corresponding labels according to the image types to form an image dataset;
[0013] S102: Use the random sampling method to divide the image dataset into a training set and a test set;
[0014] S103: Perform data augmentation operations on the training set;
[0015] S104: Scale the image to h×h pixels to match the model size, and each task corresponds to one foreign object image dataset.
[0016] Further, the data augmentation operations include random rotation, flipping mirror, grayscaling, saturation adjustment, and contrast adjustment.
[0017] Further, the pre-trained embedding layer includes a linear embedding and a positional embedding. Both the linear embedding and the positional embedding are learnable parameters. The linear embedding is used to divide the input image into L small blocks and convert it into an image sequence; the positional embedding is used to generate position information for the image blocks and append it to the image sequence to add position features to the image sequence; the input image finally obtains an image sequence of L×D through the pre-trained embedding layer, where L is the number of divided input images, which is also equal to the length of the image sequence, and D is the dimension of the image embedding.
[0018] The pre-trained Transformer encoder includes N blocks in cascade, each block including two regularization layers, a multi-head self-attention layer (MSA), a feed-forward neural network layer, and a residual connection; the image sequence output by the pre-trained embedding layer is linearly mapped into three components, namely, a query component Q, a key component K, and a value component V, after passing through the first regularization layer, and then fed into the multi-head self-attention layer MSA to extract attention features. Subsequently, the image feature representation is obtained through the second regularization layer and the feed-forward neural network layer, and the deep feature representation of the image is finally obtained through N-level blocks.
[0019] The prompt pool consists of t + 1 prompts, and the classification head pool consists of t + 1 classification heads, where t is the number of tasks learned before this learning; the prompts and classification heads are created in pairs, and the quantity grows dynamically with the number of learned tasks, and each task is assigned 1 prompt and classification head; a single prompt is a trainable parameter of N×2×L p ×D, where N is the number of blocks, and L p is the prompt sequence length; a single classification head is a fully connected layer of D×2.
[0020] Furthermore, the specific steps of the training process of the foreign object detection network model are as follows:
[0021] S301: Extract the features of all images in the training set, use the K-means algorithm to cluster the features of all images, and save the obtained K clustering centers to obtain the set C a ={C 1 , C 2 ,..., C i ,..., C K} of the clustering centers of the training set features of task a, where i = 1, 2,..., K;
[0022] S302: Input the training set into the foreign object detection network for training. After the training is completed, save the model parameters. The saved model parameters include the prompt pool and classification head pool parameters, the clustering centers of the learned tasks, and the number of learned tasks; for the prompt and classification head parameters of the current task in the prompt pool and classification head pool, use the Adam optimizer for iterative optimization, freeze the remaining parameters in the network, the prompt pool, and the classification head pool, calculate the classification loss using the cross-entropy function, and use the Step learning rate scheduler to dynamically change the learning rate.
[0023] Furthermore, the specific expression for dynamically changing the learning rate using the Step learning rate scheduler is:
[0024]
[0025] where lr(n) is the learning rate of the nth round, bs is the batch size, λ is the decay coefficient, and N1 represents the total number of training rounds.
[0026] Furthermore, the specific steps for the foreign object detection network model to perform real-time channel foreign object detection are as follows:
[0027] S401: Use the pre-trained feature extractor to extract the features of the input image. By calculating the Manhattan distance between the input image features and all the clustering centers obtained in step S301, find the clustering center in all datasets that is closest to the input image features, so as to obtain the task number m of the closest clustering center. Furthermore, select the corresponding task prompt and classification head in the prompt pool and classification head pool;
[0028] S402: Use the prompt and classification head determined in step S401. Append the selected prompt to the input of the multi-head self-attention layer MSA of each block in the pre-trained Transformer encoder to form a feature extractor after prompt fine-tuning, and connect the selected classification head to form a foreign object detection network model for a specific task;
[0029] S403: Input the image into the foreign object detection network model for a specific task in S402 to obtain the classification result.
[0030] Furthermore, the method for obtaining the task number m in step S401 is as follows:
[0031]
[0032] where m is the task number of the clustering center closest to the input image features, f is the input image features, C a is the set of clustering centers of the dataset features for the a-th task, γ(f, C a ) represents the Manhattan distance, f i represents the i-th element in the input image features, represents the i-th clustering center in C a and D represents the dimension of the image embedding.
[0033] The process of appending the prompt is described using a prompt function, and the specific expression is:
[0034] f(p, h) = MSA(h Q , [p k ; h K , [p v ; h V )
[0035] where f(p, h) is the prompt function, h is the original input of the multi-head self-attention layer MSA, h Q represents the query component of the original input h of the multi-head self-attention layer MSA, h K represents the key component of the original input h of the multi-head self-attention layer MSA, hV Represents the value component of the original input h of the multi-head self-attention layer MSA, p is the prompt, p k is the key component of the prompt p, p v is the value component of the prompt p, [·;·] represents the concatenation operation.
[0036] Furthermore, the specific steps for the foreign object detection network model to continuously learn for new tasks or new scenarios are as follows:
[0037] S501: Construct a dataset for the new task based on the channel images in the new scenario or new task captured by the network camera deployed on the compact rack according to steps S101 - 104;
[0038] S502: Append the prompts and classification heads for the new task to the prompt pool and classification head pool;
[0039] S503: Use the dataset for the new task constructed in step S501 to train the foreign object detection network model with the newly added parameters in step S502 by adopting steps S301 - S302;
[0040] S504: In the manner of steps S401 - S403, use the foreign object detection network model trained in step S503 to perform real-time detection of foreign objects in the compact rack channel under the new scenario or new task.
[0041] The beneficial effects of the present invention are as follows:
[0042] (1) The foreign object detection network model of the present invention learns and expands the prompt parameters for different tasks, thus having the ability of continuous learning. It can adapt to new scenarios while retaining the detection ability for old scenarios, and as the number of training scenarios increases, the generalization ability of the model is further enhanced. With the powerful generalization ability, it can achieve new scenario migration with only a small amount of task training or even without training, so it can reduce the deployment complexity and improve the speed and efficiency;
[0043] (2) For the foreign object detection network model of the present invention, only a small number of prompt parameters are trained for fine-tuning the pre-trained model. As learning progresses, the number of parameters increases slightly, realizing the support for long-term continuous learning;
[0044] (3) The present invention effectively enhances the robustness of the foreign object detection network model by guiding the image acquisition method of the dataset and using data augmentation processing, and by adopting the Step learning rate scheduler to optimize the training of the network model. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0046] Figure 1 It is the flowchart of the method for the embodiment of the present invention;
[0047] Figure 2 It is the structural schematic diagram of the foreign object detection network model in the embodiment of the present invention;
[0048] Figure 3 It is the structural schematic diagram of the block of the pre-trained Transformer encoder in the embodiment of the present invention;
[0049] Figure 4 It is the schematic diagram of the method for selecting trainable parameters in the embodiment of the present invention;
[0050] Figure 5 It is the diagram of the foreign object detection results in the dense rack channel in the embodiment of the present invention under different scenarios. Detailed implementation manners
[0051] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0052] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field of the present application. The "first", "second" and similar terms used in the specification and claims of this patent application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a" or "one" do not indicate a quantity limitation, but mean that there is at least one. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship also changes accordingly.
[0053] As Figure 1 shown, the embodiment of the present invention proposes a method for detecting foreign objects in the dense rack channel based on prompt expansion and continual learning, including the following steps:
[0054] S1: Deploy network cameras on the compact shelves to collect images of various foreign objects in the channel, scale the images to h×h pixels to match the model size, and construct a foreign object image dataset. Each task corresponds to one foreign object image dataset. The specific steps are as follows:
[0055] S101: Use the network cameras deployed on the compact shelves to capture the images in the channel; collect the internal images of the channel with different types of foreign objects, different illuminations or different shelf spacings, store them according to the presence or absence of foreign object types, and generate corresponding labels according to the image types to form an image dataset;
[0056] S102: Use the random sampling method to divide the image dataset into a training set and a test set;
[0057] S103: Perform data augmentation operations on the training set;
[0058] S104: Scale the images to h×h pixels to match the model size. Each task corresponds to one foreign object image dataset.
[0059] In the embodiments of the present invention, the task includes a new compact shelf deployment scenario or a scenario where new compact shelf channels need to be added in the same scenario. The data augmentation operations include random rotation, flipping mirror, grayscale conversion, saturation adjustment and contrast adjustment. After the data augmentation is completed, the images are uniformly scaled to 224×224 pixels to match the model size.
[0060] S2: Design a foreign object detection network model. The structure of the foreign object detection network model is as Figure 2 shown, including a pre-trained feature extractor, a prompt pool and a classification head pool; the pre-trained feature extractor includes a pre-trained embedding layer and a pre-trained Transformer encoder; the parameters of the pre-trained embedding layer and the pre-trained Transformer encoder do not change with the model training, and the prompt pool and the classification head pool are used as trainable parameters and change with the model training.
[0061] The pre-trained embedding layer includes a linear embedding and a positional embedding. Both the linear embedding and the positional embedding are learnable parameters. The linear embedding is used to divide the input image into L small blocks and convert it into an image sequence; the positional embedding is used to generate position information for the image blocks and append it to the image sequence to add position features to the image sequence; the input image finally obtains an image sequence of L×D through the pre-trained embedding layer, where L is the number of divided input images, which is also equal to the length of the image sequence, and D is the dimension of the image embedding;
[0062] The pre-trained Transformer encoder includes N blocks connected in series. The structure of each block is as Figure 3As shown, it includes two regularization layers, a multi-head self-attention layer MSA, a feed-forward neural network layer, and residual connections; the image sequence output by the pre-trained embedding layer is linearly mapped into three components, namely, a query component Q, a key component K, and a value component V, after passing through the first regularization layer, and then sent to the multi-head self-attention layer MSA to extract attention features. Subsequently, through the second regularization layer and the feed-forward neural network layer, an image feature representation is obtained, and through the N-level block, a deep feature representation of the image is finally obtained.
[0063] The prompt pool consists of t + 1 prompts, and the classification head pool consists of t + 1 classification heads, where t is the number of tasks learned before this learning; the prompts and classification heads are created in pairs, and the number dynamically increases with the number of learned tasks, and 1 prompt and classification head are assigned to each task; a single prompt is a trainable parameter of N×2×L p ×D, where N is the number of blocks, and L p is the prompt sequence length; a single classification head is a fully connected layer of D×2.
[0064] In the embodiment of the present invention, the linear embedding divides the input image into 16×16 pixels, a total of 196 image patches, and the input image finally obtains an image sequence of 196×768 through the pre-trained embedding layer. The number of blocks in the pre-trained Transformer encoder is 12, that is, the value of N is 12. The sizes of the prompt pool and the classification head pool are initially set to 1, and a single prompt is a trainable parameter of 12×2×L p ×768, where L p is the prompt sequence length, and in the embodiment of the present invention, L p is set to 10. A single classification head is a fully connected layer of 768×2.
[0065] S3: Use the foreign object image dataset obtained in step S1 to train the foreign object detection network model, and the specific steps are as follows:
[0066] S301: Extract the features of all images in the training set, use the K-means algorithm to cluster the features of all images, and save the obtained K clustering centers to obtain a set C of clustering centers of the training set features of task a a ={C 1 , C 2 ,..., C i ,..., C K}, where i = 1, 2,..., K. In the embodiment of the present invention, the value of K is 5;
[0067] S302: Input the training set into the foreign object detection network for training. After the training is completed, save the model parameters. The saved model parameters include the prompt pool and classification head pool parameters, the learned task clustering centers, and the number of learned tasks. For the prompts and classification head parameters of the current task in the prompt pool and classification head pool, use the Adam optimizer for iterative optimization. Freeze the remaining parameters in the network, the prompt pool, and the classification head pool. Use the cross-entropy function to calculate the classification loss, and use the Step learning rate scheduler to dynamically change the learning rate. The specific expression is as follows:
[0068]
[0069] where lr(n) is the learning rate of the nth round, bs is the batch size, λ is the decay coefficient, and N1 represents the total number of training rounds. In the embodiments of the present invention, the value of the batch size is 256, the total number of training rounds N1 is 200, and the decay coefficient λ is 0.1.
[0070] S4: Use the foreign object detection network model trained in step S3 to detect foreign objects in the images collected by the camera, and real-time detect the foreign object status of the channel. The specific steps are as follows:
[0071] S401: The process of selecting trainable parameters is as Figure 4 shown. Use the pre-trained feature extractor to extract the features of the input image. By calculating the Manhattan distance between the input image features and all the clustering centers obtained in step S301, find the clustering center in all datasets that is closest to the input image features, so as to obtain the task number m of the closest clustering center. Then, select the prompts and classification heads corresponding to the task in the prompt pool and classification head pool;
[0072] The method for obtaining the task number m is:
[0073]
[0074] where m is the task number of the clustering center closest to the input image features, f is the input image features, C a is the set of clustering centers of the dataset features of the ath task, γ(f, C a ) represents the Manhattan distance, f i represents the ith element in the input image features, represents the ith clustering center in C a , and D represents the dimension of the image embedding.
[0075] S402: Use the prompts and classification heads determined in step S401. Append the selected prompts to the input of the multi-head self-attention layer MSA of each block in the pre-trained Transformer encoder to form a feature extractor after prompt fine-tuning, and connect the selected classification head to form a foreign object detection network model for a specific task;
[0076] The hint for the additional process is described using a hint function, and the specific expression is:
[0077] f(p,h) = MSA(h Q ,[p k ;h K ,[p v ;h V );
[0078] Among them, f(p,h) is the hint function, h is the original input of the multi-head self-attention layer MSA, h Q represents the query component of the original input h of the multi-head self-attention layer MSA, h K represents the key component of the original input h of the multi-head self-attention layer MSA, h V represents the value component of the original input h of the multi-head self-attention layer MSA, p is the hint, p k is the key component of the hint p, p v is the value component of the hint p, [·;·] represents the concatenation operation.
[0079] S403: Input the image into the foreign object detection network model for a specific task in S402 to obtain a classification result. The detection results of foreign objects in the dense rack channel in different scenarios in the embodiments of the present invention are as Figure 5 shown.
[0080] S5: When the foreign object detection network model learns a new task or is migrated to a new scenario for application, continuous learning is performed on the new task or new scenario, and the specific steps are as follows:
[0081] S501: Construct a dataset for the new task based on the images captured by the network camera on the dense rack in the new scenario or new task according to steps S101 - 104;
[0082] S502: Append the hint and classification head for the new task to the hint pool and classification head pool;
[0083] S503: Use the dataset for the new task constructed in step S501 to train the foreign object detection network model with the additional new parameters formed in step S502 by adopting steps S301 - S302;
[0084] S504: In the manner of steps S401 - S403, use the foreign object detection network model trained in step S503 to perform real-time detection of foreign objects in the dense rack channel in the new scenario or new task.
[0085] In the hardware environment where the central processing unit model is "12th Gen Intel(R) Core(TM) i7-12700F" and the graphics card model is "NVIDIA GeForce RTX 3060", the following test results can be obtained for the embodiments of the present invention:
[0086] (1) Accuracy test
[0087] The embodiments of the present invention continuously learned 6 different task datasets of foreign object scenarios in the dense rack channels. The number of images in the dense rack channels in each dataset is shown in Table 1.
[0088] Table 1 Number of images in the dense rack channels of different task datasets
[0089] Data set Number of images in the dense rack channel / piece A 1,853 B 1,924 C 2,075 D 1,204 E 1,346 F 1,138
[0090] The average accuracy and average forgetting rate are used to evaluate the continuous learning performance of the foreign object detection network model. The average accuracy is the average of the accuracies of each learned task. The larger the value, the better the continuous learning performance of the network; the average forgetting rate is the average of the forgetting degrees of the model for all old tasks. The smaller the value, the better the continuous learning performance of the network. The specific calculation formulas are as follows:
[0091]
[0092] Among them, A is the average accuracy, F is the average forgetting rate, T represents the total number of tasks, A T,j represents the accuracy of task j after training task T, and A j,j represents the accuracy of task j after training task j.
[0093] Table 2 shows the accuracy, average accuracy, and average forgetting rate of the dataset in each incremental stage. One new dataset is learned in each incremental stage. After continuously learning 6 different datasets, the average accuracy of the network model can reach more than 95%, and the average forgetting rate is less than 2%, showing good continuous learning performance.
[0094] Table 2 Continuous learning performance of the network model
[0095]
[0096] (2) Parameter growth test
[0097] Since the embodiments of the present invention only train the prompt and classification head, the parameters increased during the continuous learning process of the foreign object detection network model are only the parameters of the prompt and classification head. When the prompt length L p = 10 is set in the embodiments of the present invention, the parameter growth amount of the foreign object detection network model for each learned dataset is 185,856. It can be seen that the embodiments of the present invention achieve continuous learning with a small increase in model parameters.
[0098] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting foreign objects in a dense shelving channel based on prompt extension and continuous learning, characterized in that: The steps include: S1: Deploy network cameras on the dense shelves to collect images of various foreign objects in the channel, and scale the images to h×h pixels to match the model size to build a foreign object image dataset. Each task corresponds to one foreign object image dataset; S2: Design a foreign body detection network model, which includes a pre-trained feature extractor, a prompt pool and a classification head pool; the pre-trained feature extractor includes a pre-trained embedding layer and a pre-trained Transformer encoder; the parameters of the pre-trained embedding layer and the pre-trained Transformer encoder are not updated with model training, and the prompt pool and the classification head pool are updated with model training as trainable parameters; S3: Use the foreign body image dataset obtained in step S1 to train the foreign body detection network model. The specific training process is as follows: S301: Extract the features of all images in the training set, cluster the features of all images using the K-means algorithm, and save the obtained K cluster centers to obtain the cluster center set C of the training set features of task a. a ={C1,C2,...,C i ,...,C K }, where i = 1, 2, ..., K; S302: input the training set into the foreign object detection network for training, and save the model parameters after the training is completed. The saved model parameters include the prompt pool and the classification head pool parameters, the cluster center of the learned tasks, and the number of learned tasks; the prompt and classification head parameters of the current task in the prompt pool and the classification head pool are iteratively optimized using the Adam optimizer, and the remaining parameters in the network, the prompt pool, and the classification head pool are frozen. The classification loss is calculated using the cross entropy function, and the learning rate is dynamically changed using the Step learning rate scheduler; S4: Use the foreign object detection network model trained in step S3 to detect foreign objects in the images collected by the camera, and detect the foreign object status of the channel in real time. The specific steps are as follows: S401: extract features of the input image using a pre-trained feature extractor, find the cluster center closest to the input image feature in all data sets by calculating the Manhattan distance between the input image feature and all cluster centers obtained in step S301, thereby obtaining the number m of the task to which the nearest cluster center belongs, and then select the prompt and classification head of the corresponding task from the prompt pool and classification head pool; S402: Using the prompts and classification heads determined in step S401, append the selected prompts to the input of the multi-head self-attention layer MSA of each block in the pre-trained Transformer encoder to form a feature extractor after prompt fine-tuning, and connect the selected classification head to form a foreign object detection network model for a specific task; S403: Input the image to the foreign body detection network model of the specific task in S402 to obtain the classification result; S5: When the foreign object detection network model learns a new task or migrates to a new scene for application, the new task or new scene is continuously learned. The specific steps are as follows: S501: Building a data set for a new task using images of the channel in a new scene captured by a network camera deployed on a dense shelving system according to step S1; S502: Adding the prompt and classification header of the new task to the prompt pool and classification header pool; S503: using the data set of the new task constructed in step S501, and using steps S301 to S302 to train the foreign body detection network model after adding new parameters constructed in step S502; S504: Using the method in steps S401 to S403, use the foreign object detection network model trained in step S503 to perform real-time foreign object detection in the dense shelving channel under new scenarios or new tasks.
2. According to the method of foreign body detection in dense shelving aisles based on prompt extension and continuous learning according to claim 1, it is characterized in that: The prompt pool is composed of t+1 prompts, and the classification head pool is composed of t+1 classification heads, where t is the number of tasks learned before this learning; prompts and classification heads are created in pairs, and the number is dynamically expanded with new learning tasks. For each new task learned, one prompt and one classification head are respectively expanded in the prompt pool and the classification head pool; The structure of a single prompt is a trainable parameter of N×2×Lp×D, where N is the number of blocks and Lp is the length of the prompt sequence; The structure of a single classification head is a D×2 fully connected layer.
3. According to the method of foreign body detection in dense shelving aisles based on prompt extension and continuous learning according to claim 1, it is characterized in that: The method for obtaining the task number m in step S401 is: Among them, m is the task number of the cluster center closest to the input image feature, f is the input image feature, C a is the cluster center set of the dataset features of the ath task, γ(f,C a ) represents the Manhattan distance, f i represents the i-th element in the input image feature, Represents C a The i-th cluster center in , D represents the dimension of image embedding.
Citation Information
Patent Citations
Power transmission multi-scale target detection method and system
CN117292119A
Compact shelving channel foreign matter detection method based on DER incremental learning
CN117372777A