Intelligent storage identification method based on open vocabulary target detection

By adopting an identification method based on open vocabulary object detection in the intelligent warehousing management system, the shortcomings of traditional systems in identifying and classifying inventory product variants and new categories are solved, and efficient, flexible and accurate inventory management is achieved.

CN119990987AInactive Publication Date: 2025-05-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510444212.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for traditional intelligent warehousing management systems to effectively identify and classify variants of inventory products, especially when there are diverse product types and frequent updates of new products, and the classification accuracy of deep learning models under small sample conditions is limited.

Method used

An intelligent warehousing recognition method based on open vocabulary object detection is adopted, and efficient identification and classification of inventory products is achieved through steps such as data preprocessing, multi-scale visual feature extraction, visual-voice deep fusion, bounding box and semantic prediction, target recognition and label matching, and detection result post-processing.

Benefits of technology

It improves the flexibility, accuracy and efficiency of inventory management, can effectively respond to dynamic changes in inventory product variants and new product categories, improves the ability to adapt to new products and dynamic environments, and realizes efficient fine-grained variant recognition and real-time attribute report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990987A_ABST
    Figure CN119990987A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and deep learning, solves the technical problems of frequent updating of stock product variants and new categories, open vocabulary query and fine-grained classification of a traditional method, and particularly relates to an intelligent storage recognition method based on open vocabulary target detection. The method comprises the steps of data preprocessing, multi-scale visual feature extraction, vision-voice deep fusion, bounding box and semantic prediction, target identification and label matching, and detection result post-processing and output. The method shows significant advantages in intelligent warehousing, not only can flexibly adapt to dynamic inventory management requirements, but also realizes efficient fine-grained variant identification, real-time attribute report generation and dynamic anomaly detection. The efficient reasoning process supports rapid processing of large-scale inventory, diversified user query requirements such as inventory checking, product positioning and abnormal state monitoring can be responded in real time, and the warehousing operation process is further optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and specifically to open vocabulary object detection, intelligent warehouse management, and model architecture application based on open vocabulary target detection, which is mainly used for real-time recognition and classification of products in warehouse shelves. Background Art

[0002] Intelligent warehouse management has important applications in modern logistics and supply chain systems, especially in high-efficiency fields such as e-commerce, retail, and manufacturing. With the diversification of product types and dynamic changes in inventory on storage shelves, traditional manual management and automated systems based on fixed categories are difficult to meet the needs of speed and accuracy. Especially in product identification and classification, there are many product variants on the shelves and their characteristics are similar. Traditional detection methods are difficult to cope with efficient and accurate classification needs. In addition, in the open vocabulary detection scenario, the system needs to flexibly identify and classify undefined product categories, while deep learning models often face challenges such as data scarcity and adaptability to new products.

[0003] In recent years, open vocabulary object detection technology based on deep learning has made some progress in intelligent warehouse management. However, this field still faces the following two major technical challenges: Complex product types with similar variant features. Due to the wide variety of product types on the shelves and the possible subtle feature differences between different products, traditional methods are prone to cause classification confusion; New products are frequently introduced and lack adaptability. In a dynamic warehousing environment, the rapid addition of new products requires the model to have the ability to learn and adapt quickly, but the classification accuracy of existing methods is limited under small sample conditions. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides an intelligent warehouse identification method based on open vocabulary target detection, which solves the technical problems that traditional methods have shortcomings in dealing with frequent updates of inventory product variants and new categories, as well as in open vocabulary query and fine-grained classification.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: an intelligent warehouse identification method based on open vocabulary target detection, the method comprising the following steps: Data preprocessing, input image Preprocessing to obtain standardized images , and prompt the user for the text input Convert to text embedding ; Multi-scale visual feature extraction from standardized images Extract feature maps of multiple scales, and optimize and integrate the feature maps of multiple scales to obtain multi-scale feature maps from shallow image details to high-level semantic information ; Deep vision-speech fusion, using the RepVL-PAN module to fuse multi-scale feature maps and embed with text Cooperate with each other to guide optimization and adjustment to generate enhanced multi-scale feature maps And text embedding ; Bounding box and semantic prediction based on multi-scale feature maps Predict multiple bounding boxes of the target containing spatial parameters , and the semantic embedding vector used to capture the semantic features of the object within the bounding box ; Target recognition and label matching, calculation of semantic embedding vector With text embedding The similarity score between , and for each bounding box Match the best text label; Post-processing and output of detection results: Through non-maximum suppression and threshold filtering, high-quality detection results are optimized and output.

[0006] Furthermore, in the data preprocessing, the specific process includes the following steps: Image scaling and resizing, using bilinear interpolation method for input image Perform scaling operation to get the image , the scaled image It is expressed as: ; In the formula, Represents an image scaling operation, which is a processing function that adjusts an image from its original size to its target size; Image normalization processing, the image The pixel value range is mapped from [0,255] to [0,1] to obtain the image , and for the image Each channel is normalized to generate a standardized image , the standardization formula is: ; In the formula, is the mean; is the standard deviation; Prompt the user for text input Converted to 512-dimensional text embedding through CLIP's text encoder , is the number of prompt words, is a 512-dimensional text embedding vector; Input sequence integration and normalization, text embedding through linear projection layer Projected to the normalized image Text embedding adapted to the feature space , and normalize the image channels and text embeddings for text channels Integrate into standardized input forms across modalities ; The projection formula is: ; After projection, the final standardized input is represented as: ; In the formula, , is the projection weight matrix; , is the bias vector.

[0007] Furthermore, in the multi-scale visual feature extraction, the specific process includes the following steps: Initial extraction of image features, extracting standardized images through the initial convolutional layer of the Darknet backbone network Generate a preliminary feature map containing basic visual information such as edge, texture and color changes , the expression is: ; In the formula, Represents batch normalization operation; is the SiLU activation function; Indicates a convolution operation using a convolution kernel of 3×3; Downsampling and multi-scale feature generation, preliminary feature map Perform three downsampling to generate shallow features , middle-level features And deep features Multi-scale features , the expression is: ; in, Indicates The input feature map of the layer; Represents a downsampling convolution operation with a kernel size of 3×3 and a stride of 2. Each operation halves the spatial size of the feature map and gradually increases the number of channels. Feature optimization of CSP module, using CSP module to process multi-scale features , multi-scale features of the input It is divided into two parts, namely: Directly transferred features , and features processed by dense convolution ; Finally, the optimized feature map is generated through the splicing operation , the expression is: ; In the formula, Represents the splicing operation along the channel dimension, combining the two parts of features and merge; The pyramid structure of multi-scale features converts the feature map Shallow features in , middle-level features And deep features The feature pyramid structure is used to integrate the image information from shallow details to high-level semantics. The feature pyramid structure is expressed as: ; in, It is a feature pyramid structure, that is, a multi-scale feature set; Output and transmission of feature maps, multi-scale feature maps Passed to the RepVL-PAN module to support further processing of vision-language, the output form is: ; in, To output the feature set, the subsequent links send it to the cross-modal fusion module to support vision-language fusion.

[0008] Furthermore, in the vision-speech deep fusion, the specific process includes the following steps: Input of multi-scale features, receiving multi-scale feature maps and text embedding As the input of the RepVL-PAN module, the image features are initialized in the form of a feature pyramid structure; Top-down feature fusion, the RepVL-PAN module fuses scale feature maps through a top-down path , from the deep features Start by gradually upsampling and combining with the middle-level features and shallow features Fusion, the fusion formula is: ; in, For upsampling operation, nearest neighbor interpolation is used to increase the feature map resolution by 2 times; is the middle-level feature after fusion; is the shallow feature after fusion; Text-guided optimization, T-CSPL Ayer module uses text embedding The image features are enhanced by 1×1 convolution, the attention map is generated and the features are optimized to highlight the areas related to the query. The optimization formula is: ; in, is the input feature map; Embedded by text The 1×1 convolution weights generated by reparameterization; To take the maximum value along the text dimension, generate an attention map; For the optimized feature graph, highlight the query-related information; I-Pooling Attention’s image-guided adjustment adjusts text embedding by pooling image features , so that the text is more adapted to the current shelf image content, the pooling and attention update formula is: ; ; in, It is the patch mark after pooling; Embed for the adjusted text; is the deep feature after fusion; For maximum pooling, the stride is adjusted to output a 3×3 area; It is a multi-head attention mechanism; Bottom-up feature enhancement and output, further enhance the features through the bottom-up path, from the shallow features after fusion Start downsampling and fusion to ensure the integrity of multi-scale features. The enhancement formula is: ; The final output enhanced multi-scale feature map set is: ; in, For downsampling operations, a convolution with a stride of 2 is used, and the resolution is halved; Enhance shallow features; Enhance mid-level features; Enhance deep features; To enhance the feature set.

[0009] Furthermore, in the bounding box and semantic prediction, the specific process includes the following steps: The detection head receives the multi-scale feature map enhanced by the RepVL-PAN module. , contains shallow, middle and deep information, corresponding to the target detection requirements of different scales; Bounding box regression branch, the regression branch uses convolutional layers to perform multi-scale feature maps Processing is performed to predict the spatial parameters of the bounding box for each feature map grid location, including the center coordinates , Width and Height and confidence , the prediction result of each bounding box is expressed as: ; in, is the predicted value of the center coordinate of the bounding box, indicating the offset relative to the feature network; is the predicted value of the width and height of the bounding box, normalized based on the feature map scale; is the confidence level, ranging from [0,1], indicating the probability of the existence of the target in the bounding box; is the complete prediction parameter of the kth candidate bounding box; The semantic embedding branch uses parallel convolutional layers to embed each bounding box. Generate a 512-dimensional semantic embedding vector , used to capture the semantic features of the target within the bounding box, the formula is: ; in, is the input feature map, from , or ; The number of output channels is 512, and the feature map is converted into a semantic embedding vector; is the semantic embedding vector of the kth bounding box, representing its semantic features; The fusion of bounding boxes and embeddings, each bounding box Its corresponding semantic embedding vector Integrate into a complete prediction unit for subsequent classification and screening. The integrated prediction is expressed as: ; in, Represents the complete prediction information of the kth candidate target, combining spatial position and semantic features; Multi-scale integration of prediction output and detection head for multi-scale feature maps Generate predictions separately and combine predictions of all scales into a unified set of candidate bounding boxes , the integration formula is: ; in, For the Layer feature map The generated set of predictions; The final prediction set contains bounding boxes and semantic embedding information of all scales.

[0010] Furthermore, in the target recognition and label matching, the specific process includes the following steps: Input of candidate targets and text embedding, receiving a set of candidate bounding boxes and text embedding ; Normalization of semantic embedding vector and text embedding, for each bounding box The semantic embedding vector Embedded with user input text Perform normalization processing, the normalization formula is: ; ; in, and Represents a normalized vector with a fixed norm of 1; Similarity calculation, normalized vector calculated by dot product and The similarity score between , and apply a scaling factor and bias ,Right now: ; in, is the similarity score, k is the index of the kth bounding box, and j is the index of the jth text label; is the scaling factor, used to amplify the similarity difference; Label assignment, for each bounding box Assign the text label with the highest similarity score, that is: ; in, Indicates traversing all text label indexes j; Indicates the final result found, making the similarity score The maximum j value; If the similarity score , then it is the bounding box Assigning text labels ; otherwise it is marked as background, that is: ; Where k is the bounding box The corresponding label; For background; Classification result output: Output the classification results C of all bounding boxes, including position, label and confidence. The expression is: ; in, Represents the spatial position, coordinates and size of the kth candidate bounding box; is the confidence, indicating the reliability of the classification; For label.

[0011] Furthermore, in the post-processing and output of the detection results, the specific process includes the following steps: The input of the classification result receives the classification result, including the bounding box, label and confidence information of the candidate object in the shelf image; Non-maximum suppression is performed on each text label, retaining the bounding box with the highest confidence. , while suppressing the overlap with other frames Bounding boxes exceeding the threshold , the calculation formula for non-maximum suppression is: ; like and , then suppress the bounding box ; in, The threshold is 0.5; is the confidence level; Threshold filtering, filter high-quality detection frames through double threshold filtering, namely: Confidence Threshold: Keep The bounding box of ,remove low-confidence false positives; Similarity threshold: Keep The bounding box ensures accurate label assignment; in, represents the maximum similarity score calculated between the k-th bounding box and all text prompts, and is used to determine whether the k-th bounding box should be assigned the most matching text prompt. Integration and optimization of detection results: After non-maximum suppression and threshold filtering, the remaining bounding boxes are integrated to generate an optimized set of optimization results. ,Right now: ; In the formula, For labels; The output and application of test results will output the optimization results in two forms to support the automated management of smart warehousing, namely: Visual output: draw bounding boxes and labels on shelf images to intuitively display detection results for easy viewing by managers; Structured output generates detection reports in JSON or CSV format, including bounding box coordinates, labels, and confidence levels, making it easy to integrate with the system.

[0012] Furthermore, the intelligent warehouse identification method supports the rapid processing of large-scale inventory, can respond to user query requirements including inventory counting, product positioning and abnormal status monitoring in real time, and further optimize the warehouse operation process.

[0013] By means of the above technical solution, the present invention provides an intelligent warehouse identification method based on open vocabulary target detection, which has at least the following beneficial effects: 1. The present invention aims to improve the flexibility, accuracy and efficiency of inventory management, cope with the frequent updates of inventory product variants and new categories, and the shortcomings of traditional methods in open vocabulary query and fine-grained classification. Based on open vocabulary detection technology, efficient product identification and classification in intelligent warehousing can be achieved, which can effectively cope with the challenges of dynamically changing inventory and the introduction of new products.

[0014] 2. The present invention realizes efficient product recognition and classification in intelligent warehousing through multi-scale feature extraction and cross-modal similarity calculation. This method not only improves the accuracy of classification, but also enhances the adaptability to new products and dynamic environments, providing an innovative solution for intelligent warehousing management.

[0015] 3. The intelligent warehousing identification method proposed in this invention has shown significant advantages in intelligent warehousing. It can not only flexibly adapt to dynamic inventory management needs, but also realize efficient fine-grained variant identification, real-time attribute report generation and dynamic anomaly detection, providing innovative and practical solutions for warehousing management. Its efficient reasoning process supports the rapid processing of large-scale inventory, and can respond to diverse user query needs in real time, such as inventory counting, product positioning and abnormal status monitoring, further optimizing the warehousing operation process.

[0016] 4. The present invention effectively improves the recognition and classification accuracy of products in storage shelves by introducing multi-scale feature extraction and cross-modal information fusion mechanisms, especially in scenarios with diverse product types, similar variants, and frequent introduction of new products, overcoming the problems of recognition confusion and insufficient adaptability of new products faced by traditional methods. At the same time, the similarity scoring mechanism further enhances the model's ability to distinguish between different product variants, and has strong robustness and broad application prospects.

[0017] 5. The intelligent warehousing identification method proposed in the present invention shows excellent flexibility and accuracy in open vocabulary detection of intelligent warehousing, supports dynamic product identification and variant classification, and significantly improves the intelligence level of inventory management. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 It is a flow chart of the intelligent storage identification method in the present invention; Figure 2 This is a network structure diagram of the RepVL-PAN module in the present invention. DETAILED DESCRIPTION

[0019] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods, so that the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0020] In order to address the technical defects of open vocabulary object detection technology based on deep learning in intelligent warehouse management, open vocabulary object detection technology has gradually become a development trend. Through visual-language fusion capabilities, combined with multi-scale feature extraction and cross-modal information modeling, it shows strong flexibility and adaptability. This embodiment proposes an open vocabulary object detection method based on open vocabulary target detection, which realizes efficient product recognition and classification in intelligent warehousing through multi-scale feature extraction and cross-modal similarity calculation. This method not only improves the accuracy of classification, but also enhances the system's adaptability to new products and dynamic environments, providing an innovative solution for intelligent warehouse management.

[0021] This example uses the Darknet backbone network and the RepVL-PAN module to generate a multi-scale feature map containing semantic information, which can integrate the global and local visual information of the shelf image and help the model better mine product details and semantic features. At the same time, the similarity score is used to quantify the degree of semantic matching between the candidate target and the text query, making the classification of features with strong consistency more accurate. Figure 1 As shown, the method comprises the following steps: S1, data preprocessing, input image Preprocessing to obtain standardized images , and prompt the user for the text input Convert to text embedding , and the standardized image and text embedding Standardized input that is unified across modal input forms The specific process includes the following steps: S11, image scaling and resizing, using bilinear interpolation method to input image Perform scaling operation to get the image Since the input image The original size is usually (H, W, 3), where H represents the height, W represents the width, and 3 represents the three RGB channels. In order to adapt to the input requirements of the Darknet backbone network, the input image It needs to be scaled to a fixed resolution (640,640,3). The scaling process uses bilinear interpolation to preserve the smoothness of the image content and the integrity of the spatial structure. The scaled image It can be expressed as: ; In the formula, Represents an image scaling operation, which is a processing function that adjusts an image from its original size to its target size; S12, image normalization processing, the image The pixel value range is mapped from [0,255] to [0,1] to obtain the image , and for the image Each channel is normalized to generate a standardized image . The scaled image Further normalization is required to standardize the pixel values ​​and adapt the distribution of the model during training. First, the pixel value range is mapped from [0,255] to [0,1], that is: ; in, It is an image with normalized pixel values. As a large dataset containing millions of diverse images, ImageNet has highly representative statistical values ​​that can be applied to a variety of visual tasks, ensuring the performance consistency of the model in different scenarios. Therefore, based on the statistical values ​​of the ImageNet dataset, each channel is normalized to generate a standardized image. , the standardization formula is: ; In the formula, is the mean; is the standard deviation. The mean used in this embodiment and standard deviation They are: ; Normalized image after normalization The dimensional differences between pixel values ​​are eliminated, making the model more robust to external factors such as lighting and color changes, while remaining consistent with the distribution of training data.

[0022] S13, prompt the text input by the user Converted to 512-dimensional text embedding through CLIP's text encoder , used to specify the detection target. is the number of prompt words, is a 512-dimensional text embedding vector. The specific process is: ; In this embodiment, the text encoder of CLIP is a Transformer structure based on ViT-B / 32. This encoding conversion process extracts text prompts. The semantic information of the image provides a language input basis for the subsequent visual-language cross-modal fusion.

[0023] S14. Input sequence integration and normalization, text embedding through linear projection layer Projected to the normalized image Text embedding adapted to the feature space , and normalize the image channels and text embeddings for text channels Standardized input that integrates into cross-modal input forms .

[0024] Since the spaces of image features and text features are different, text embedding A linear projection layer is required to adapt the feature space of the standardized image, and the output dimension is still 512 text embedding The projection formula is: ; In the formula, , is the projection weight matrix; , is the bias vector.

[0025] After projection, the final standardized input is represented as: ; This embodiment unifies image and text data into a cross-modal input form, laying a foundation for subsequent steps.

[0026] S2, multi-scale visual feature extraction, from the standardized image through the Darknet backbone network Extract feature maps of multiple scales, and optimize and integrate the feature maps of multiple scales to obtain multi-scale feature maps from shallow image details to high-level semantic information The specific process includes the following steps: S21. Initial extraction of image features, extracting standardized images through the initial convolutional layer of the Darknet backbone network Generate a preliminary feature map containing basic visual information such as edge, texture and color changes .

[0027] Standardized image after data preprocessing The image (size 640×640×3) is first fed into the initial convolutional layer of the Darknet backbone network. To extract the basic visual information of the image, a 3×3 convolution kernel (stride 1), SiLU activation function and batch normalization operation are used, with 64 output channels to capture the basic visual information such as edges, textures and color changes in the image to generate a preliminary feature map. , the expression is: ; In the formula, Represents batch normalization operation; is the SiLU activation function; Indicates a convolution operation using a convolution kernel of 3×3.

[0028] S22, downsampling and multi-scale feature generation, preliminary feature map Perform three downsampling to generate shallow features , middle-level features And deep features Multi-scale features . Specifically, the preliminary feature map Multi-scale features are extracted through three downsampling operations, each downsampling uses a 3×3 convolutional layer with a stride of 2, expressed as: ; in, Indicates The input feature map of the layer is initially a preliminary feature map (Size is 640×640×64); Represents a downsampling convolution operation with a kernel size of 3×3 and a stride of 2. Each operation halves the spatial size of the feature map and gradually increases the number of channels.

[0029] Preliminary feature map Three downsamplings are performed to generate features of three scales, namely: Shallow Features (size 80×80×256, stride 8), capturing detailed information of shelf images, such as product edges; Mid-level features (size 40×40×512, stride 16), extracting mid-level semantic information, such as product shape; Deep Features (size is 20×20×1024, stride is 32), providing high-level semantic information such as product category.

[0030] S23, CSP module feature optimization, using CSP (Cross Stage Partial) module to process multi-scale features , multi-scale features of the input It is divided into two parts, namely: Directly transferred features , and features processed by dense convolution ; Finally, the optimized feature map is generated through the splicing operation , the expression is: ; In the formula, It indicates the feature part that is directly passed without additional processing, and the original information is retained; It represents the feature part processed by three 3×3 convolutional layers, which contains three layers of 3×3 convolution to extract deep features; Represents the splicing operation along the channel dimension, combining the two parts of features and merge; Represents the optimized feature map, the number of channels is two parts of the feature and The sum of the number of channels improves the feature expression capability. This process enhances the richness of features and improves the computational efficiency of the model.

[0031] S24, the pyramid structure of multi-scale features, the feature map Shallow features in , middle-level features And deep features The feature pyramid structure is constructed to integrate the image information from shallow details to high-level semantics, providing input for the subsequent cross-modal fusion module. The feature pyramid structure is expressed as: ; in, It is a feature pyramid structure, that is, a multi-scale feature set, which provides rich visual information for subsequent modules.

[0032] S25, output and transmission of feature maps, multi-scale feature maps It is passed to the cross-modal fusion module (i.e., the RepVL-PAN module) to support further processing of vision-language. Its output is in the form of: ; in, To output the feature set, the subsequent links send it to the cross-modal fusion module to support vision-language fusion.

[0033] S3, deep vision-speech fusion, using the RepVL-PAN module to fuse multi-scale feature maps and embed with text Cooperate with each other to guide optimization and adjustment to generate enhanced multi-scale feature maps And text embedding .like Figure 2 As shown, the specific process includes the following steps: S31, multi-scale feature input, receiving multi-scale feature map and text embedding As the input of the RepVL-PAN module, the image features are initialized in the form of a feature pyramid structure to provide a visual basis for subsequent fusion, as follows: ; in, is a text embedding with a size of N×512, where N is the number of query words; As an input feature set, it integrates multi-scale visual information.

[0034] S32, top-down feature fusion, the RepVL-PAN module fuses the scale feature map through a top-down path , from the deep features Start by gradually upsampling and combining with the middle-level features and shallow features Fusion. This process transfers high-level semantic information to lower layers, enhancing the context-awareness of features, and is applicable to the recognition of overall shelf layout and local products in smart warehousing. The fusion formula is: ; in, For upsampling operation, nearest neighbor interpolation is used to increase the feature map resolution by 2 times; is the middle-level feature after fusion; It is the shallow feature after fusion, with a size of 80×80×256; S33, text-guided optimization, T-CSPL Ayer module uses text embedding The image features are enhanced through 1×1 convolution, attention maps are generated and features are optimized to highlight the areas related to the query. In smart warehousing, this sub-step ensures that the model focuses on the target specified by the text and improves the recognition targeting. The optimization formula is: ; in, is the input feature map; Embedded by text The 1×1 convolution weights generated by re-parameterization have a size of 512×N×1×1, where N is the number of query words; To take the maximum value along the text dimension, generate an attention map; It is an optimized feature map with a size of 40×40×512, highlighting query-related information; S34, I-Pooling Attention image-guided adjustment, adjusting text embedding by pooling image features , so that the text is more adapted to the current shelf image content, the pooling and attention update formula is: ; ; in, is the patch mark after pooling, size 27×512; is the adjusted text embedding, size N×512; It is the deep feature after fusion, with a size of 20×20×1024; It is the middle-level feature after fusion, with a size of 40×40×512; It is the shallow feature after fusion, with a size of 80×80×256; For maximum pooling, the stride is adjusted to output a 3×3 area, generating a total of 27 patches (three scales × 9 patches); It is a multi-head attention mechanism. In this embodiment, 8 heads are used, each with a dimension of 64, to update the text embedding.

[0035] S35, bottom-up feature enhancement and output, further enhance the features through the bottom-up path, from the shallow features after fusion Start downsampling and fusion to ensure the integrity of multi-scale features. This process integrates global and local information in smart warehousing and supports the recognition of complex scenes. The enhancement formula is: ; The final output enhanced multi-scale feature map set is: ; in, For downsampling operations, a convolution with a stride of 2 is used, and the resolution is halved; Enhance shallow features, size is 80×80×256; Enhanced mid-level features, size 40×40×512; Enhanced deep features, size is 20×20×1024; To enhance the feature set, it is passed to the detection head to support subsequent recognition.

[0036] S4, Bounding box and semantic prediction based on multi-scale feature maps Predict multiple bounding boxes of the target containing spatial parameters , and the semantic embedding vector used to capture the semantic features of the object within the bounding box , and each bounding box Its corresponding semantic embedding vector Integrate into a unified set of candidate bounding boxes The specific process includes the following steps: S41, the detection head receives input, and the detection head receives the multi-scale feature map enhanced by the RepVL-PAN module , contains shallow, middle and deep information, corresponding to the target detection requirements of different scales. The input feature set is defined as: ; in, It is the input feature set of the detection head, which contains multi-scale information and supports the detection of objects of different sizes.

[0037] S42, bounding box regression branch, the regression branch uses convolutional layers to perform multi-scale feature maps Processing is performed to predict the spatial parameters of the bounding box for each feature map grid location, including the center coordinates , Width and Height and confidence The anchor-free design is adopted to directly predict the offset and size of the bounding box, avoiding the complexity of the traditional anchor-based method. The prediction result of each bounding box is expressed as: ; in, is the predicted value of the center coordinate of the bounding box, indicating the offset relative to the feature network; is the predicted value of the width and height of the bounding box, normalized based on the feature map scale; is the confidence level, ranging from [0,1], indicating the probability of the existence of the target in the bounding box; is the complete prediction parameter of the kth candidate bounding box; Among them, the number of bounding boxes is: 3 bounding boxes are predicted for each feature map grid position, and the total number is N=(802+402+202)×3=25200 bounding boxes, covering multi-scale feature maps .

[0038] S43, semantic embedding branch, through parallel convolutional layers for each bounding box Generate a 512-dimensional semantic embedding vector , which is used to capture the semantic features of the object within the bounding box. This vector supports open vocabulary classification and achieves target matching by calculating similarity with the embedding vector of the text query. The generation formula is: ; in, is the input feature map, from , or ; The number of output channels is 512, and the feature map is converted into a semantic embedding vector; is the semantic embedding vector of the kth bounding box, representing its semantic features.

[0039] S44, fusion of bounding box and embedding, each bounding box Its corresponding semantic embedding vector Integrate into a complete prediction unit for subsequent classification and screening. The integrated prediction is expressed as: ; in, Represents the complete prediction information of the kth candidate target, combining spatial position and semantic features.

[0040] S45, multi-scale integration of prediction output, detection head for multi-scale feature maps Generate predictions separately and combine predictions of all scales into a unified set of candidate bounding boxes , where N=25200. The integration formula is: ; in, For the Layer feature map The generated prediction set contains all ; is the final prediction set, which contains bounding boxes and semantic embedding information of all scales; N=25200 is the total number of candidate boxes, which is obtained by accumulating grid predictions of three scales.

[0041] S5. Target recognition and label matching, calculation of semantic embedding vector Embedded with user input text The similarity score between , and for each bounding box Match the best text label. The specific process includes the following steps: S51: Input of candidate targets and text embedding, receiving candidate bounding box set and text embedding In the smart warehousing scenario, the candidate bounding box represents potential products in shelf images, such as individual backpacks or boxes, while text embeddings Corresponds to specific queries from warehouse managers, such as "blue backpack" or "damaged packaging". The input is clearly defined as: ; in, Represents the bounding box and semantic embedding of the kth candidate object; A collection of embeddings representing a text query.

[0042] S52, normalization of semantic embedding vector and text embedding, for each bounding box The semantic embedding vector Embedded with user input text To ensure the stability of similarity calculation, this embodiment performs normalization on the semantic embedding vector and text embedding Perform L2 normalization, the normalization formula is: ; ; in, and Represents a normalized vector with a fixed norm of 1. In smart warehousing, this normalization process ensures that the product embeddings and query embeddings in different shelf images are compared at a uniform scale to avoid misjudgment due to differences in feature vector size. Normalization ensures that all feature vectors have the same scale to avoid similarity calculations due to size differences. By adjusting the vector norm to 1, a more accurate comparison between the target and the query is ensured.

[0043] S53, similarity calculation, calculate the normalized vector through dot product and The similarity score between , and apply a scaling factor and bias ,Right now: ; in, is the similarity score, k is the index of the kth bounding box, and j is the index of the jth text label; is the scaling factor, used to amplify the similarity difference; S54. Label assignment for each bounding box Assign the text label with the highest similarity score, that is: ; in, Indicates traversing all text label indexes j; Indicates the final result found, making the similarity score The maximum j value; If the similarity score , then it is the bounding box Assigning text labels ; otherwise it is marked as background, that is: ; Where k is the bounding box The corresponding label; In smart warehousing, this step directly assigns the "blue backpack" tag to eligible shelf items, or removes non-matching areas as background, improving the accuracy of inventory counting.

[0044] S55, classification result output, output the classification results C of all bounding boxes, including location, label and confidence. This information can be used for intelligent warehousing tasks, such as generating inventory reports or marking abnormal items, and the expression is: ; in, Represents the spatial position, coordinates and size of the kth candidate bounding box; is the confidence, indicating the reliability of the classification; In smart warehousing, this output provides bounding boxes, labels, and confidence levels for candidate products on each shelf, such as labeling all “blue backpacks” with their locations and confidence levels, allowing managers to respond quickly.

[0045] Where k is the bounding box In smart warehousing, this step directly assigns the "blue backpack" label to eligible shelf items, or removes non-matching areas as background, improving the accuracy of inventory counting.

[0046] S6, post-processing and output of detection results, through non-maximum suppression and threshold filtering, optimize and output high-quality detection results. The specific process includes the following steps: S61, input of classification results, receiving the classification results, including the bounding box, label and confidence information of the candidate target in the shelf image, providing a data basis for subsequent processing; S62, non-maximum suppression (NMS), to eliminate duplicate detections, perform non-maximum suppression on each text label and retain the bounding box with the highest confidence , while suppressing the overlap with other frames Bounding boxes exceeding the threshold , the calculation formula for non-maximum suppression is: ; like and , then suppress the bounding box ; in, The threshold is 0.5, which is used to measure the critical value of the overlap degree; is the confidence level, used to decide which bounding box to keep ; Non-maximum suppression ensures that each product on the shelf is represented by only one bounding box, avoiding duplicate counting. When multiple bounding boxes detect the same item at the same time, non-maximum suppression retains the bounding box with the highest confidence, thereby improving the accuracy of inventory counting.

[0047] S63, threshold filtering, screening high-quality detection frames through double threshold filtering, namely: Confidence Threshold: Keep The bounding box of ,remove low-confidence false positives; Similarity threshold: Keep The bounding box ensures accurate label assignment; in, represents the maximum similarity score calculated between the k-th bounding box and all text prompts, and is used to determine whether the k-th bounding box should be assigned the most matching text prompt. By filtering bounding boxes through confidence and similarity thresholds, false positives and low-quality labels can be eliminated to ensure the reliability of the results.

[0048] S64, detection result integration and optimization, after non-maximum suppression and threshold filtering, the remaining bounding boxes are integrated to generate an optimized optimization result set ,Right now: ; In the formula, For labels; This collection is significantly reduced, retaining only high-quality detection results. In smart warehousing, a concise and high-quality optimized result collection is generated by integrating the filtered bounding boxes, ready for output. And the optimized results can be directly used to generate inventory reports or abnormal alerts.

[0049] S65, output and application of test results, output optimization results in two forms, support the automated management of intelligent warehousing, and help intelligent warehousing automation by providing visual and structured output. That is: Visual output: draw bounding boxes and labels on shelf images to intuitively display detection results for easy viewing by managers; Structured output generates detection reports in JSON or CSV format, including bounding box coordinates, labels, and confidence levels, making it easy to integrate with the system.

[0050] This embodiment aims to improve the flexibility, accuracy and efficiency of inventory management to cope with the frequent updates of inventory product variants and new categories and the shortcomings of traditional methods in open vocabulary query and fine-grained classification.

[0051] This method uses innovative visual-language alignment technology, introduces an open vocabulary query mechanism, and uses the RepVL-PAN module to achieve deep fusion of image and text features. It supports zero-sample detection of variants and new products in the same category, which greatly reduces the burden of retraining and annotation required by traditional methods due to data updates, significantly shortens the time for new products to be put on the shelves, and reduces the cost of manual intervention. At the same time, through the cross-modal similarity score calculation mechanism, the semantic consistency of image and text queries is compared, which greatly improves the accuracy of fine-grained classification, ensures accurate recognition of targets in complex warehousing environments, and is suitable for processing product categories with similar appearance but different functions on shelves.

[0052] This method has shown significant advantages in intelligent warehousing. It can not only flexibly adapt to dynamic inventory management needs, but also achieve efficient fine-grained variant recognition, real-time attribute report generation, and dynamic anomaly detection, providing an innovative and practical solution for warehouse management. Its efficient reasoning process supports the rapid processing of large-scale inventory and can respond to diverse user query needs in real time, such as inventory counting, product positioning, and abnormal status monitoring, further optimizing warehouse operation processes.

[0053] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, so the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0054] The above implementation methods have been described in detail. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. An intelligent warehouse identification method based on open vocabulary object detection, characterized in that: The method comprises the following steps: Data preprocessing, input image Preprocessing to obtain standardized images , and prompt the user for the text input Convert to text embedding ; Multi-scale visual feature extraction from standardized images Extract feature maps of multiple scales, and optimize and integrate the feature maps of multiple scales to obtain multi-scale feature maps from shallow image details to high-level semantic information ; Deep vision-speech fusion, using the RepVL-PAN module to fuse multi-scale feature maps and embed with text Cooperate with each other to guide optimization and adjustment to generate enhanced multi-scale feature maps And text embedding ; Bounding box and semantic prediction based on multi-scale feature maps Predict multiple bounding boxes of the target containing spatial parameters , and the semantic embedding vector used to capture the semantic features of the object within the bounding box ; Target recognition and label matching, calculation of semantic embedding vector With text embedding The similarity score between , and for each bounding box Match the best text label; Post-processing and output of detection results: Through non-maximum suppression and threshold filtering, high-quality detection results are optimized and output.

2. The intelligent warehouse identification method according to claim 1, characterized in that: In the data preprocessing, the specific process includes the following steps: Image scaling and resizing, using bilinear interpolation method for input image Perform scaling operation to get the image , the scaled image It is expressed as: ; In the formula, Represents an image scaling operation, which is a processing function that adjusts an image from its original size to its target size; Image normalization processing, the image The pixel value range is mapped from [0,255] to [0,1] to obtain the image , and for the image Each channel is normalized to generate a standardized image , the standardization formula is: ; In the formula, is the mean; is the standard deviation; Prompt the user for text input Converted to 512-dimensional text embedding through CLIP's text encoder , is the number of prompt words, is a 512-dimensional text embedding vector; Input sequence integration and normalization, text embedding through linear projection layer Projected to the normalized image Text embedding adapted to the feature space , and normalize the image channels and text embeddings for text channels Integrate into standardized input forms across modalities ; The projection formula is: ; After projection, the final standardized input is represented as: ; In the formula, , is the projection weight matrix; , is the bias vector.

3. The intelligent warehouse identification method according to claim 1, characterized in that: In the multi-scale visual feature extraction, the specific process includes the following steps: Initial extraction of image features, extracting standardized images through the initial convolutional layer of the Darknet backbone network Generate a preliminary feature map containing basic visual information such as edge, texture and color changes , the expression is: ; In the formula, Represents batch normalization operation; is the SiLU activation function; Indicates a convolution operation using a convolution kernel of 3×3; Downsampling and multi-scale feature generation, preliminary feature map Perform three downsampling to generate shallow features , middle-level features And deep features Multi-scale features , the expression is: ; in, Indicates The input feature map of the layer; Represents a downsampling convolution operation with a kernel size of 3×3 and a stride of 2. Each operation halves the spatial size of the feature map and gradually increases the number of channels. Feature optimization of CSP module, using CSP module to process multi-scale features , multi-scale features of the input It is divided into two parts, namely: Directly transferred features , and features processed by dense convolution ; Finally, the optimized feature map is generated through the splicing operation , the expression is: ; In the formula, Represents the splicing operation along the channel dimension, combining the two parts of features and merge; The pyramid structure of multi-scale features converts the feature map Shallow features in , middle-level features And deep features The feature pyramid structure is used to integrate the image information from shallow details to high-level semantics. The feature pyramid structure is expressed as: ; in, It is a feature pyramid structure, that is, a multi-scale feature set; Output and transmission of feature maps, multi-scale feature maps Passed to the RepVL-PAN module to support further processing of vision-language, the output form is: ; in, To output the feature set, the subsequent links send it to the cross-modal fusion module to support vision-language fusion.

4. The intelligent warehouse identification method according to claim 1, characterized in that: In the vision-speech deep fusion, the specific process includes the following steps: Input of multi-scale features, receiving multi-scale feature maps and text embedding As the input of RepVL-PAN module; Top-down feature fusion, the RepVL-PAN module fuses scale feature maps through a top-down path , from the deep features Start by gradually upsampling and combining with the middle-level features and shallow features Fusion, the fusion formula is: ; in, For upsampling operation, nearest neighbor interpolation is used to increase the feature map resolution by 2 times; is the middle-level feature after fusion; is the shallow feature after fusion; Text-guided optimization, T-CSPL Ayer module uses text embedding The image features are enhanced by 1×1 convolution, the attention map is generated and the features are optimized to highlight the areas related to the query. The optimization formula is: ; in, is the input feature map; Embedded by text The 1×1 convolution weights generated by reparameterization; To take the maximum value along the text dimension, generate an attention map; For the optimized feature graph, highlight the query-related information; I-Pooling Attention’s image-guided adjustment adjusts text embedding by pooling image features , so that the text is more adapted to the current shelf image content, the pooling and attention update formula is: ; ; in, It is the patch mark after pooling; Embed the adjusted text; is the deep feature after fusion; For maximum pooling, the stride is adjusted to output a 3×3 area; It is a multi-head attention mechanism; Bottom-up feature enhancement and output, further enhance the features through the bottom-up path, from the shallow features after fusion Start downsampling and fusion to ensure the integrity of multi-scale features. The enhancement formula is: ; The final output enhanced multi-scale feature map set is: ; in, For downsampling operations, a convolution with a stride of 2 is used, and the resolution is halved; Enhance shallow features; Enhance mid-level features; Enhance deep features; To enhance the feature set.

5. The intelligent warehouse identification method according to claim 1, characterized in that: In the bounding box and semantic prediction, the specific process includes the following steps: The detection head receives the multi-scale feature map enhanced by the RepVL-PAN module. , contains shallow, middle and deep information, corresponding to the target detection requirements of different scales; Bounding box regression branch, the regression branch uses convolutional layers to perform multi-scale feature maps Processing is performed to predict the spatial parameters of the bounding box for each feature map grid location, including the center coordinates , Width and Height and confidence , the prediction result of each bounding box is expressed as: ; in, is the predicted value of the center coordinate of the bounding box; Predicted values ​​for the width and height of the bounding box; is the confidence level; is the complete prediction parameter of the kth candidate bounding box; The semantic embedding branch uses parallel convolutional layers to embed each bounding box. Generate a 512-dimensional semantic embedding vector , used to capture the semantic features of the target within the bounding box, the formula is: ; in, is the input feature map, from , or ; is the semantic embedding vector of the kth bounding box; The fusion of bounding boxes and embeddings converts each bounding box Its corresponding semantic embedding vector Integrate into a complete prediction unit for subsequent classification and screening. The integrated prediction is expressed as: ; in, Represents the complete prediction information of the kth candidate target; Multi-scale integration of prediction output and detection head for multi-scale feature maps Generate predictions separately and combine predictions of all scales into a unified set of candidate bounding boxes , the integration formula is: ; in, For the Layer feature map The generated set of predictions; The final prediction set contains bounding boxes and semantic embedding information of all scales.

6. The intelligent warehouse identification method according to claim 1, characterized in that: In the target recognition and label matching, the specific process includes the following steps: Input of candidate targets and text embedding, receiving a set of candidate bounding boxes and text embedding ; Normalization of semantic embedding vector and text embedding, for each bounding box The semantic embedding vector Embedded with user input text Perform normalization processing, the normalization formula is: ; ; in, and Represents a normalized vector with a fixed norm of 1; Similarity calculation, normalized vector calculated by dot product and The similarity score between , and apply a scaling factor and bias ,Right now: ; in, is the similarity score, k is the index of the kth bounding box, and j is the index of the jth text label; Label assignment, for each bounding box Assign the text label with the highest similarity score, that is: ; in, Indicates traversing all text label indexes j; Indicates the final result found, making the similarity score The largest j value; If the similarity score , then it is the bounding box Assigning text labels ; otherwise it is marked as background, that is: ; Where k is the bounding box The corresponding label; For background; Classification result output: Output the classification results C of all bounding boxes, including position, label and confidence. The expression is: ; in, Represents the spatial position, coordinates and size of the kth candidate bounding box; is the confidence, indicating the reliability of the classification; For label.

7. The intelligent warehouse identification method according to claim 1, characterized in that: In the post-processing and output of the detection results, the specific process includes the following steps: The input of the classification result receives the classification result, including the bounding box, label and confidence information of the candidate object in the shelf image; Non-maximum suppression is performed on each text label, retaining the bounding box with the highest confidence. , while suppressing the overlap with other frames Bounding boxes exceeding the threshold , the calculation formula for non-maximum suppression is: ; like and , then suppress the bounding box ; in, The threshold is 0.5; is the confidence level; Threshold filtering, filter high-quality detection frames through double threshold filtering, namely: Confidence Threshold: Keep The bounding box of ,remove low-confidence false positives; Similarity threshold: Keep The bounding box ensures accurate label assignment; in, represents the maximum similarity score calculated between the k-th bounding box and all text prompts, and is used to determine whether the k-th bounding box should be assigned the most matching text prompt. Integration and optimization of detection results: After non-maximum suppression and threshold filtering, the remaining bounding boxes are integrated to generate an optimized set of optimization results. ,Right now: ; In the formula, For labels; The output and application of test results will output the optimization results in two forms to support the automated management of smart warehousing, namely: Visual output: draw bounding boxes and labels on shelf images to intuitively display detection results for easy viewing by managers; Structured output generates detection reports in JSON or CSV format, including bounding box coordinates, labels, and confidence levels, making it easy to integrate with the system.

8. The intelligent warehouse identification method according to any one of claims 1 to 7, characterized in that: The intelligent warehouse identification method supports the rapid processing of large-scale inventory, can respond to user query requirements including inventory counting, product positioning and abnormal status monitoring in real time, and further optimize the warehouse operation process.

Citation Information

Patent Citations

  • Multi-modal multi-view target detection and matching method and system

    CN118506102A

  • Virtual-real fusion chemical experiment platform for real-time feedback based on multi-modal perception

    CN119516852A

Cited By

  • Polycrystalline silicon material unmanned warehouse intelligent sorting and inventory management system

    CN122067236A