Multi-target commodity identification method, device and system based on multi-modal data processing
Through the multimodal data processing method, combined with the preprocessing of real-time video data, label extraction, instance segmentation and multimodal model fine-tuning, the occlusion and overlapping problems in multi-target product recognition are solved, and efficient and accurate recognition in complex scenarios are achieved.
Patent Information
- Application Number
- CN202510728593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The prior art is difficult to accurately identify products in multi-objective scenarios, especially in cases of occlusion and overlap, and single modal processing lacks comprehensive analysis of visual, spatial and semantic information.
The multimodal data processing method is adopted to obtain real-time video data, pre-process and tag information extraction, instance segmentation and feature extraction, and fine-tune the multimodal visual language model by combining multi-source privatized data to identify product images and text information.
Accurately identify multiple product targets in complex environments, improve recognition accuracy and stability, and meet the multi-objective recognition needs in dynamic scenarios.
Smart Images

Figure CN120236155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent vending, and particularly to a multi-target commodity recognition method, device and system based on multi-modal data processing. Background Art
[0002] As a common unmanned vending device, intelligent vending cabinets are widely used in multiple fields, such as shopping, food, beverages and daily necessities. Traditional intelligent vending cabinets mainly rely on bar code or QR code scanning to identify commodities. This method usually depends on manual input of commodity information and label matching. However, this method has limitations. For example, it cannot effectively identify situations such as occlusion, overlap or position change of commodities. In addition, the manually labeled tag information is prone to errors, especially in the case of frequent commodity updates, resulting in the vending cabinet being unable to adapt to the display and identification of new commodities in a timely manner. Therefore, how to efficiently and accurately identify commodities in a dynamic environment, especially to handle complex scenarios such as multiple commodity targets appearing simultaneously, occlusion and overlap, has become the core problem in the intelligent upgrade of intelligent vending cabinets.
[0003] Existing technologies usually rely on single visual information or text information in the recognition of multiple commodity targets, and it is difficult to handle complex situations in dynamic transaction scenarios. For example, based on traditional image recognition methods, the model is prone to misrecognition or missed recognition when dealing with occluded and overlapping commodities. At the same time, text information extraction also faces challenges such as unclear, blurred or position-changed labels, and existing methods are mostly single-modal processing, lacking comprehensive analysis of visual information, spatial information and semantic information. Therefore, existing technologies cannot meet the real-time and accurate commodity recognition requirements of intelligent vending cabinets in complex scenarios.
[0004] Existing Chinese Patent CN114445201A discloses a combined commodity retrieval method and system based on a multi-modal pre-training model, including: dividing a commodity image into single-commodity images and combined-commodity images; training a combined-commodity image detector; obtaining and combining the feature encoding, position encoding and segment encoding of the text modality and the picture module in the combined-commodity image, learning the embedding representation, and inputting it into the constructed multi-modal pre-training model; using the multi-modal pre-training model to extract the retrieval features of the picture modality and the text modality of the single-commodity image; the multi-modal pre-training model extracts the retrieval features of the text-picture fusion of the combined-commodity image according to the bounding box and the bounding box features of each target commodity in the combined-commodity image, calculates the pre-distance between the combined-commodity features and the single-commodity features in the retrieval library as the commodity similarity, and selects the most similar single commodity as the result to return. The above patent solution cannot accurately handle situations such as occlusion and overlap between commodities, and cannot cope with the differences between different commodity features and text descriptions. Therefore, it is difficult to ensure accurate commodity recognition in actual scenarios.
[0005] Therefore, how to accurately identify products in a multi-target scenario is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a multi-target product recognition method, device and system based on multi-modal data processing to solve the problem in the prior art that products cannot be accurately identified in a multi-target scenario.
[0007] The technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a multi-target product recognition method based on multi-modal data processing, and the method includes: Obtain real-time video data in a product transaction scenario, and decompose the real-time video data into multiple frames of real-time images; Perform preprocessing and label information extraction on the real-time images to determine the preprocessed target images and the text information corresponding to the product labels; Perform instance segmentation on the target images to determine the product position information; According to the product position information, extract features from the target images to determine the product image feature information; According to the pre-collected multi-source privatized data in an intelligent vending scenario, fine-tune and optimize an open-source multi-modal vision-language model to obtain a multi-modal large model for product recognition; Input the product image feature information and the text information into the multi-modal large model for information fusion, and determine the product target recognition result according to the fused feature information.
[0008] Preferably, the performing preprocessing and label information extraction on the real-time images to determine the preprocessed target images and the text information corresponding to the product labels includes: Adjust the size and perform noise reduction processing on the real-time images to determine the target images; Perform object detection on the target images to determine the product area position information; According to the product area position information, process the product labels in the product area through optical character recognition technology to determine the text information.
[0009] Preferably, the performing instance segmentation on the target images to determine the product position information includes: According to the product area position information, extract the feature information of the product area through a convolutional neural network, and determine candidate areas according to the extracted feature information; Process the candidate areas through an instance segmentation network to determine the binary images corresponding to each product target; Use post-processing technology to process the binary images to determine the product position information.
[0010] Preferably, the multi-modal large model for commodity recognition obtained by fine-tuning and optimizing the open-source multi-modal vision-language model according to the pre-collected multi-source privatized data in the intelligent vending scenario includes: Clean and structure the multi-source raw data according to the pre-collected multi-source raw data in the intelligent vending scenario to obtain an annotated data set; Pair and construct the image information and text labels in the annotated data set, and perform format conversion and unified preprocessing on it to obtain a multi-modal input sample set for training; Load weight parameters for the open-source vision-language pre-trained model according to the multi-modal input sample set, and construct a vision coding and language coding network structure that supports joint optimization to obtain an initial multi-modal model structure for fine-tuning; Perform fine-tuning training on the initial multi-modal model structure according to the characteristics of the vending scenario and the recognition accuracy requirements, and optimize the hyperparameter configuration through a cross-validation strategy to obtain multiple candidate multi-modal models; Evaluate and compare each of the candidate multi-modal models according to the preset accuracy rate, recall rate, and response time to obtain the multi-modal large model.
[0011] Preferably, the inputting the commodity image feature information and the text information into the multi-modal large model for information fusion, and determining the commodity target recognition result according to the fusion feature information includes: Input the commodity image feature information and the text information into the multi-modal large model to obtain fusion feature information that fuses the image features and text semantics; Input the fusion feature information into a pre-trained commodity classification model to obtain an initial commodity category; Judge whether there are similar commodities in the current initial commodity category according to the initial commodity category; When there are similar commodities, obtain the local area of the feature to be extracted and the target feature to be extracted according to the initial commodity category; Extract features from the target image according to the local area and the target feature to obtain local area feature information; Classify the initial commodity category according to the local area feature information to obtain the target commodity category as the commodity target recognition result.
[0012] Preferably, when there are similar commodities, obtaining the local area of the feature to be extracted and the target feature to be extracted according to the initial commodity category includes: Select sample images corresponding to multiple sub-categories under this category from a preset commodity image database according to the initial commodity category; Input each of the sample images into a pre-trained saliency detection model to obtain a saliency heatmap, where the saliency heatmap is used to represent the region in the sample image with the most concentrated visual feature attention; Perform threshold segmentation on the saliency heatmap to obtain multiple candidate regions; Perform a comprehensive scoring on each of the candidate regions, and according to the scoring results, screen and obtain the local region from each of the candidate regions; Perform candidate feature extraction and feature evaluation on the local region, and according to the feature evaluation results, screen and obtain the target feature from the extracted candidate features.
[0013] Preferably, the performing a comprehensive scoring on each of the candidate regions and obtaining the local region according to the scoring results includes: Obtain the significant values in the saliency heatmaps corresponding to each candidate region; According to each of the significant values, calculate the average significant value of each candidate region as the saliency scoring value; According to the initial commodity category, input each of the sample images into a pre-trained image classification model to obtain a class activation map; According to the class activation map, obtain the response heatmap of each sample image under the current initial commodity category; Count the pixels at the corresponding positions of each candidate region in the response heatmap, and calculate the average activation intensity of each candidate region as the class-related scoring value; Perform weighted fusion on the saliency scoring value and the class-related scoring value of each candidate region to obtain the scoring result of each candidate region; Compare each of the scoring results with a preset scoring threshold, and according to the comparison results, select at least one region from each candidate region as the local region.
[0014] Preferably, the performing candidate feature extraction and feature evaluation on the local region and screening and obtaining the target feature from the extracted candidate features according to the feature evaluation results includes: According to the initial commodity category, obtain the feature extraction strategies corresponding to each candidate feature of this category; According to each of the feature extraction strategies, perform multi-path feature extraction on the local region to obtain multiple candidate feature information; Perform distribution consistency analysis on each of the candidate feature information in each of the sample images, and obtain the frequency and position deviation of each candidate feature information appearing in different sample images as the distribution consistency index; Respectively input each of the candidate feature information into a pre-trained product recognition model to obtain a recognition result, and obtain the classification confidence corresponding to the recognition result as the classification response intensity of each candidate feature information; Evaluate each of the candidate feature information according to the distribution consistency index and the classification response intensity to obtain a feature evaluation result; Screen out the target feature from each of the candidate feature information according to the feature evaluation result.
[0015] In a second aspect, the present invention provides a multi-target product recognition device based on multi-modal data processing, and the device includes: A real-time image acquisition module, configured to acquire real-time video data in a product transaction scenario, and decompose the real-time video data into multiple frames of real-time images; A preprocessing and label information extraction module, configured to perform preprocessing and label information extraction on the real-time image to determine a preprocessed target image and the text information corresponding to the product label; An instance segmentation module, configured to perform instance segmentation on the target image to determine product position information; A feature extraction module, configured to extract features from the target image according to the product position information to determine product image feature information; A multi-modal large model training module, configured to fine-tune and optimize an open-source multi-modal vision-language model according to pre-collected multi-source privatized data in a smart vending scenario to obtain a multi-modal large model for product recognition; A product recognition module, configured to input the product image feature information and the text information into the multi-modal large model for information fusion, and determine a product target recognition result according to the fused feature information.
[0016] In a third aspect, an embodiment of the present invention further provides a multi-target product recognition system based on multi-modal data processing, including: an image acquisition device, at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, the method as described above is implemented.
[0017] In summary, the beneficial effects of the present invention are as follows: The multi-object commodity recognition method, device and system based on multi-modal data processing provided by the present invention include: obtaining real-time video data in a commodity trading scenario, and decomposing the real-time video data into multiple frames of real-time images; preprocessing the real-time images and extracting label information to determine the preprocessed target images and the corresponding text information of the commodity labels; performing instance segmentation on the target images to determine the commodity position information; extracting features from the target images according to the commodity position information to determine the commodity image feature information; fine-tuning and optimizing an open-source multi-modal vision-language model according to pre-collected multi-source private data in an intelligent vending scenario to obtain a multi-modal large model for commodity recognition; inputting the commodity image feature information and the text information into the multi-modal large model for information fusion, and determining the commodity target recognition result according to the fusion feature information. The present invention extracts multiple frames of images from real-time video data, and through preprocessing and label information extraction, preliminarily recognizes the commodities in each frame of image, extracts the commodity labels and the corresponding text information. Then, the target images are processed by instance segmentation technology to accurately locate the bounding boxes of each commodity and solve the occlusion problem between commodities. The image features, including visual features, spatial features and semantic features, are further extracted by using the commodity position information. Information fusion is performed through a multi-modal large model, and the commodity image features are combined with the extracted text information to improve the recognition ability of complex commodity targets. Finally, based on the fusion of image and text information, the model can accurately distinguish and recognize multiple commodity targets. Even in a complex environment where multiple commodities coexist and there is partial overlap or occlusion, it can still maintain high recognition performance, not only improving the recognition accuracy, but also being able to operate stably in a dynamic scenario to meet the multi-target recognition requirements. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings, and all of these are within the protection scope of the present invention.
[0019] Figure 1 It is a schematic flow chart of the overall work of the multi-object commodity recognition method based on multi-modal data processing in Embodiment 1 of the present invention; Figure 2 It is a schematic flow chart of preprocessing the real-time images and extracting label information in Embodiment 1 of the present invention; Figure 3 It is a schematic flow chart of performing instance segmentation on the target images in Embodiment 1 of the present invention; Figure 4 It is a schematic flow chart of extracting features from the target images in Embodiment 1 of the present invention; Figure 5 It is a schematic flowchart for determining the target recognition result of a commodity in Embodiment 1 of the present invention; Figure 6 It is a schematic diagram of similar commodities in Embodiment 1 of the present invention; Figure 7 It is a structural block diagram of a multi-target commodity recognition device based on multi-modal data processing in Embodiment 2 of the present invention; Figure 8 It is a schematic structural diagram of a multi-target commodity recognition system based on multi-modal data processing in Embodiment 3 of the present invention. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or sequence between these entities or operations. In the description of the present invention, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements. If there is no conflict, the embodiments of the present invention and the various features in the embodiments can be combined with each other, and all are within the protection scope of the present invention.
[0021] Embodiment 1
[0022] Please refer to Figure 1 , Embodiment 1 of the present invention discloses a multi-target commodity recognition method based on multi-modal data processing, and the method includes: Obtain real-time video data in a commodity trading scenario, and decompose the real-time video data into multiple frames of real-time images; Specifically, first, real-time video data of the transaction scenario is collected in real-time through a camera or other video capture devices. The real-time video data contains multiple commodities and video data from different perspectives. Next, the video stream is decomposed into a series of consecutive image frames, and each frame of the image will be used as the input for subsequent object detection and analysis, laying a foundation for subsequent steps such as commodity recognition, label extraction, instance segmentation, and feature extraction. By effectively decomposing the real-time video data into multiple frames of images, the visual information of the commodities can be accurately captured in a dynamic environment, providing efficient data support for subsequent processing.
[0023] Preprocess the real-time image and extract label information to determine the preprocessed target image and the text information corresponding to the commodity label. Specifically, perform image enhancement processing on each frame of the image, including denoising, brightness and contrast adjustment, etc., to improve the quality and clarity of the image. This process ensures that the image can maintain high readability under different environmental lighting conditions, facilitating subsequent processing. Next, use image processing algorithms, such as edge detection and image smoothing, to preprocess the image, removing unnecessary background noise and highlighting the commodity area to ensure that the commodity target is clearly distinguishable. After the image preprocessing is completed, use OCR (Optical Character Recognition) technology to extract the text information in the commodity label from the image, including commodity name, price, SKU code, etc. OCR technology converts the characters on the commodity label into a machine-readable text format by identifying the text area in the image, and effectively processes problems such as font, size, blur, and rotation to ensure the accurate extraction of text information. Through this process, relevant text information can be assigned to each commodity target, providing comprehensive input data for subsequent commodity recognition, object detection, and feature matching.
[0024] In one embodiment, please refer to Figure 2 , the preprocessing of the real-time image and the extraction of label information to determine the preprocessed target image and the text information corresponding to the commodity label include: Adjust the size and perform noise reduction processing on the real-time image to determine the target image. Specifically, scale and adjust the real-time image according to the preset input size requirements. The size adjustment ensures that the image has a consistent size on different devices and in different environments and facilitates the efficiency of subsequent processing. Subsequently, perform noise reduction processing on the scaled and adjusted image. Noises in the image, such as brightness fluctuations and image blurring, will be removed. Common noise reduction algorithms, such as Gaussian blur and median filtering, are applied to the scaled and adjusted image to enhance the commodity edges and details and improve the image quality. After processing these images, the focus will be on the commodity area to form a clear target image, laying a solid foundation for subsequent object detection and text recognition steps.
[0025] Perform object detection on the target image to determine the position information of the commodity area; Specifically, use object detection algorithms such as YOLO and Faster R-CNN to process the preprocessed target image, identify and locate all commodity objects in the image, and through a convolutional neural network (CNN) and a region proposal network (RPN), effectively classify different objects in the image and calibrate the commodity bounding boxes. These bounding boxes accurately determine the position of the area where the commodity is located, providing position information for subsequent commodity label extraction and feature analysis. Object detection algorithms can efficiently identify commodity objects in real-time images, and can handle occlusions and overlaps well even in complex commodity placement scenarios, paving the way for the accurate identification of commodity targets.
[0026] According to the position information of the commodity area, process the commodity labels in the commodity area through optical character recognition technology to determine the text information.
[0027] Specifically, after obtaining the position information of the commodity area, according to the position information of the commodity area, crop the image of the area where the commodity is located from the original target image, and then apply OCR (optical character recognition) technology to process the cropped commodity label area. OCR technology analyzes the character area in the image and identifies the text information therein, such as important information such as commodity name, price, and production date. To ensure the accuracy of text information extraction, OCR technology will optimize and adjust according to factors such as the font, size, rotation angle, and blur of the text, and at the same time use an adaptive text recognition algorithm to process different label formats and languages. Finally, the text information on the commodity label will be combined with the image features of the commodity, providing reliable multi-modal data support for subsequent commodity identification and processing.
[0028] Perform instance segmentation on the target image to determine the commodity position information; Specifically, an instance segmentation algorithm (such as Mask R-CNN) is applied to the target image. The aim is to accurately segment each independent commodity object in the image and generate an accurate mask for each object. Different from traditional object detection, instance segmentation can not only determine the bounding box of the commodity, but also precisely distinguish the specific contour and shape of the commodity. Therefore, first, the algorithm extracts the features of the image through a convolutional neural network and generates potential target regions through a Region Proposal Network (RPN). Next, a more refined segmentation network is used to perform pixel-level classification on these regions, and finally, a detailed mask for each commodity object is generated. In this way, instance segmentation can not only locate an accurate bounding box for each commodity, but also depict the shape information of the commodity in detail, providing more accurate position information for subsequent commodity feature extraction and localization. This process is particularly important for commodity target recognition in complex backgrounds and can effectively handle problems such as occlusion, overlap, and similar backgrounds of commodities.
[0029] In one embodiment, please refer to Figure 3 , the instance segmentation of the target image to determine the commodity position information includes: According to the commodity region position information, extract the feature information of the commodity region through a convolutional neural network, and determine candidate regions based on the extracted feature information; Specifically, using the commodity region position information (i.e., the bounding box in the target image), extract the commodity regions of interest. These commodity regions will be used as the input of a convolutional neural network (CNN). Through multiple convolutional layers, feature extraction is performed on the image. The convolutional neural network can identify high-level feature information such as the shape, color, and texture of the commodity. These features are crucial for distinguishing different commodities. After feature extraction, the Region Proposal Network algorithm is used to analyze the extracted features and generate multiple candidate regions. These regions represent the parts that may contain commodities. These candidate regions will be used as the input for further processing by the instance segmentation network.
[0030] Process the candidate regions through the instance segmentation network to determine the binary images corresponding to each commodity target; Specifically, input the previously generated candidate regions into the instance segmentation network (such as Mask R-CNN). The instance segmentation network will analyze each candidate region one by one and generate a pixel-level binary mask image for each commodity target. Each pixel value of this binary image indicates whether the position belongs to the target commodity region. Through this pixel-level segmentation, not only can the position of the commodity be accurately identified, but also its specific contour and shape information can be clarified, which helps to distinguish multiple overlapping or adjacent commodity targets.
[0031] Use post-processing techniques to process the binary image to determine the commodity position information.
[0032] Specifically, in the binary image generated by instance segmentation, there may be some noise or incomplete regions. At this time, post-processing techniques (such as morphological operations, contour extraction, region merging, etc.) are used to further optimize the segmentation result. Morphological operations (such as dilation and erosion) remove small noises and fill in the gaps in the commodity region to ensure that the commodity boundary is more coherent and accurate. Then, through the contour extraction algorithm, the boundary shape of each commodity target can be determined, and the position of the commodity in the image can be accurately located according to the segmented region. Finally, the commodity position information obtained after post-processing will be used as the basic data for subsequent steps (such as feature extraction and target recognition).
[0033] According to the commodity position information, feature extraction is performed on the target image to determine the commodity image feature information; Specifically, according to the commodity position information extracted in the previous step, the target image is cropped to obtain a local region containing only the commodity. Then, a convolutional neural network (CNN) is used to extract features from this commodity region, extracting deep features, including information such as the color, texture, shape, and edges of the commodity. These features are processed through multiple convolutional layers, pooling layers, and activation functions to transform the local details in the image into a more abstract high-level feature representation. At the same time, in addition to conventional visual features, more advanced network structures (such as ResNet, VGG, etc.) can be used to extract more complex feature information. During the feature extraction process, position encoding information, such as the relative position of the commodity in the image, can also be combined, so that the extracted features not only reflect the appearance features of the commodity but also take into account its spatial position relationship. In this way, the finally obtained commodity image feature information contains the unique visual and spatial representations of the commodity, providing support for subsequent commodity recognition and matching.
[0034] In one embodiment, please refer to Figure 4 , the performing feature extraction on the target image according to the commodity position information to determine the commodity image feature information includes: According to the commodity position information, determine the position information of the region to be processed related to the commodity in the target image; Specifically, according to the commodity position information obtained in the previous step (such as the bounding box coordinates), the regions related to the commodity in the image are determined. These regions are usually the commodity itself or regions closely related to the commodity, including commodity labels, trademarks, packaging, and other parts related to the commodity. By cropping or marking out the part of the target image containing the commodity, the exact position of the region to be processed is obtained. This region position information provides a clear spatial reference for subsequent feature extraction, analysis, and recognition.
[0035] Extract visual features from the local image of the area to be processed according to the location information of the area to be processed, and determine the visual feature information; Specifically, once the areas to be processed are determined, visual feature information that helps in product recognition is then extracted from these areas. By applying deep learning models such as convolutional neural networks (CNNs) to these local images, basic visual features of the products are extracted, such as color, texture, shape, edges, details, etc. These features are typically gradually extracted and abstracted through multiple convolutional layers and pooling layers, thereby generating feature vectors for product recognition. These visual features provide the necessary information for subsequent product classification, matching, and further analysis.
[0036] Analyze the relative positions between products according to the product location information, and determine the spatial feature information; Specifically, at this step, in addition to considering the visual features of each product area itself, the relative position relationships between products are also analyzed. For example, products are interrelated in the image in terms of distance, arrangement, or aggregation. Using the product location information (such as the coordinates of the bounding boxes), the spatial layout between products is analyzed to identify their position patterns in the image. Such spatial features can include information such as the relative distance, arrangement order, and up-down / left-right relationships between products, helping to better understand the relationships between products in space and providing more context information for subsequent product recognition and sorting.
[0037] Perform weighted feature fusion on the visual feature information and the spatial feature information to determine the product image feature information.
[0038] Specifically, the visual feature information extracted from the area to be processed is weighted and fused with the analyzed spatial feature information. By designing a weighting mechanism, values are assigned to each feature according to its importance, integrating the visual information and spatial information into a unified product image feature representation. This fusion process can be achieved by simple weighted averaging or through feature fusion modules in deep learning (such as attention mechanisms, weighted summation, feature concatenation, etc.); through this weighted feature fusion, the obtained product image feature information is more comprehensive and accurate, providing strong support for product recognition, classification, and subsequent decision-making.
[0039] Fine-tune and optimize an open-source multimodal vision-language model according to pre-collected multi-source private data in an intelligent vending scenario to obtain a multimodal large model for product recognition; In one embodiment, the fine-tuning and optimizing the open-source multimodal vision-language model according to pre-collected multi-source private data in an intelligent vending scenario to obtain a multimodal large model for product recognition includes: Based on the pre-collected multi-source raw data in the intelligent vending scenario, clean and structure the multi-source raw data to obtain an annotated dataset; Specifically, collect multi-modal raw data including transaction records, surveillance videos, and static pictures of in-cabinet goods from multiple actually deployed vending cabinet terminals. Through a unified data interface and format standard, extract the key content of each type of data, such as the product area in the image frame, the product identification code and purchase timestamp in the transaction record, etc. Subsequently, use a predefined label system to perform manual or semi-automatic annotation on the image data, including information such as product category, quantity, and packaging appearance characteristics; at the same time, perform semantic normalization processing on text data (such as product names, barcodes, etc.). Remove invalid, blurred, or incorrect samples through data cleaning, complete the structuring process, and generate a graph-text annotation alignment dataset for model training, laying a foundation for subsequent model input.
[0040] According to the image information and text labels in the annotated dataset, pair and construct the graph-text data, and perform format conversion and unified preprocessing on it to obtain a multi-modal input sample set for training; Specifically, for the already annotated image and text information, pair the image and text data one by one through a graph-text semantic alignment algorithm, and ensure that each pair of graph-text samples has a clear semantic correlation (such as a product image and its product name or description text). To meet the input requirements of the multi-modal pre-training model, perform visual preprocessing operations on the image data such as size normalization, pixel value normalization, and data augmentation (such as random cropping, rotation); for text information, perform standardization processing such as word segmentation, encoding, truncation, and padding according to the language encoder used by the model (such as the Transformer or BERT architecture). Finally, generate a graph-text sample set in a unified format, which can be directly used for the input training of the multi-modal model.
[0041] According to the multi-modal input sample set, load the weight parameters for the open-source vision-language pre-training model, and construct a vision coding and language coding network structure that supports joint optimization to obtain the initial structure of the multi-modal model for fine-tuning; Specifically, select open-source multi-modal pre-trained models with image-text understanding capabilities (such as CLIP, BLIP, etc.) as the base models, load the pre-trained parameter weights released by their official sources, and freeze some underlying modules to retain their general representation capabilities. At the same time, in combination with the characteristics of the vending scenario, design an alignment strategy between a customized visual encoder (such as ResNet or Vision Transformer) and a language encoder (such as BERT or T5), and introduce a cross-modal attention mechanism or a contrastive learning structure to enhance the interactive expression ability between image and text semantics. The constructed model structure supports the joint optimization of the parameters of the visual feature extraction module and the text semantic modeling module, and has good fine-tuning performance, laying a structural foundation for subsequent multi-round training on the private dataset.
[0042] According to the characteristics of the vending scenario and the requirements of recognition accuracy, perform fine-tuning training on the initial structure of the multi-modal model, and optimize the hyperparameter configuration through a cross-validation strategy to obtain multiple candidate multi-modal models. Specifically, to adapt to complex scenario factors such as occlusion, different angles, and low light in the commodity images in the vending cabinet environment, based on the initial structure of the multi-modal model constructed in the third step, design a fine-tuning process that includes multiple rounds of training. In each round of training, introduce the strategy of freezing some backbone parameters and unfreezing specific high-level modules to improve the model's learning ability for domain-specific features; at the same time, introduce a custom loss function to enhance the ability to distinguish between similar commodities of different classes. Through the K-fold cross-validation method, repeat the training experiment on different subsets of the private dataset, and systematically search and adjust hyperparameters such as learning rate, batch size, optimizer type, and regularization weight. Finally, output multiple candidate multi-modal model versions with different performance emphases.
[0043] Evaluate and compare each of the candidate multi-modal models according to the preset accuracy, recall rate, and response time to obtain the multi-modal large model.
[0044] Specifically, to ensure the commodity recognition effect and system response efficiency of the model in actual deployment, construct a multi-dimensional performance evaluation framework, and conduct a systematic comparison of the candidate models output by each round of fine-tuning on the test set. The accuracy and recall rate are used to measure the comprehensiveness and accuracy of the model in the commodity recognition task, and the F1 value is used as a comprehensive indicator to evaluate the recognition stability; the response time evaluates whether it can meet the real-time recognition requirements after being embedded in the vending cabinet terminal. In addition, further test the recognition performance of the model for similar commodities, occluded commodities, and combined commodities. Based on the above multiple indicators, rank the candidate models according to the preset weight scoring system, and select the model version with the best overall performance and deployment performance meeting the requirements as the finally determined multi-modal large model for commodity recognition.
[0045] Input the commodity image feature information and the text information into the multi-modal large model for information fusion, and determine the commodity target recognition result according to the fused feature information.
[0046] Specifically, combine the previously extracted commodity image feature information with the corresponding text information as multi-modal input data. The image features include visual elements (such as color, shape, texture, etc.), while the text information provides a semantic description of the commodity. The two complement each other. Then, input these image features and text information into a pre-trained multi-modal large model. This model is usually based on the Transformer architecture or other deep learning structures, such as the CLIP model, VisualBERT, which can process image and text data simultaneously and utilize the correlation between them for information fusion. In the multi-modal large model, image and text features are complementary and aligned through shared representation learning. The model can understand the relationship between the two, thus forming a deep understanding of the commodity target. Through the fusion of this multi-modal information, the model can perform more accurate commodity target recognition according to the fused information, such as commodity classification, brand recognition, specification matching, etc., and can efficiently and accurately identify and locate commodities in complex commodity trading scenarios.
[0047] In one embodiment, see Figure 5 , the step of inputting the commodity image feature information and the text information into the multi-modal large model for information fusion and determining the commodity target recognition result according to the fused feature information includes: Input the commodity image feature information and the text information into the multi-modal large model to obtain fused feature information that combines fused image features and text semantics; Specifically, use a multi-modal large model, such as a pre-trained CLIP large model or BLIP large model, to jointly model the commodity image feature information and the text information. Extract the alignment information between the commodity image and the text description through a multi-layer attention mechanism, and output a set of fused feature vectors as the fused information, which can more comprehensively represent the multi-dimensional feature attributes of the commodity. The fused information can solve the semantic ambiguity caused by OCR misrecognition, fuzziness, and missing characters, and at the same time make up for the deficiencies of image information in single-modal classification.
[0048] Input the fused feature information into a pre-trained commodity classification model to obtain an initial commodity category; Specifically, the fused feature information that combines image features and text semantics is input into a pre-trained product classification model to obtain an initial product category. The product classification model is trained based on a large amount of product image-text data, has a strong understanding ability of fused semantics, and can output a product category prediction result with high confidence, including a category label and its corresponding confidence score. By leveraging the synergistic effect of multi-modal semantics, the accuracy of the first-level product classification is improved, providing a clear category reference for the subsequent judgment of whether there are similar products.
[0049] Based on the initial product category, determine whether there are similar products in the current initial product category; Specifically, obtain a pre-set similar product mapping database. According to the initial product category, record whether there are sub-categories of products with highly similar visuals / functions under this initial product category. For example, Figure 6 as shown, similar products refer to products in the same series that are highly consistent in appearance elements such as core packaging design, bottle shape structure, material, and main color, and are only differentiated by local identifying differences (such as flavor text, pattern symbols, or color block embellishments). Take Figure 6 as an example, both bottles of beverages have white bottle bodies, black bottle caps, and a minimalist symmetric design, belonging to a series of products under the same brand. Their similarity lies in the unified bottle shape, main color, and standardized layout of the brand logo; while the difference is concentrated in the flavor labels at the lower part of the bottle body: on the left is a light gray peach pattern and white peach flavor text, and on the right is a dark gray apple pattern and green apple flavor text. This design strategy not only maintains the unity of the brand's visual appearance but also achieves precise differentiation of the product line through subtle pattern and text differences. It belongs to a typical multi-flavor similar product in the same series and requires refined identification and classification relying on local features (such as specific patterns, text descriptions). The process of determining whether there are similar products can be based on knowledge graphs, product SKU attribute structures, or Euclidean distances (or cosine similarities) of image-text features between products for clustering analysis. If the model confidence is low or there are multiple highly similar sub-categories under this initial product category, it is determined that there are similar products in the current initial product category, and the next refined identification process needs to be initiated. By automatically triggering a more accurate analysis process in uncertain or ambiguous classification scenarios, the judgment robustness and user experience in complex product identification tasks are improved.
[0050] When there are similar products, according to the initial product category, obtain the local area of the feature to be extracted and the target feature to be extracted; Specifically, when there are similar products, according to the initial product category, salient areas are extracted from multiple sample images under the initial product category, and salient heat maps are generated through salient detection models such as SAM, UNet-Saliency, GradCAM, and CLIP Attention to extract the areas with the highest visual attention, such as packaging edges, logos, special patterns, and color blocks. By analyzing the commonalities and differences of these salient areas in multiple sample images of similar products, a set of target features to be extracted (such as bottle cap shape, edge lines, background texture, etc.) is generated, and the coordinates of the local area are determined for subsequent detailed analysis. By focusing on the detail areas that are truly valuable for identification between easily confused products, the ability of image details to supplement the classification model can be effectively improved, the misclassification rate can be reduced, and significant recognition accuracy can be improved in key products such as different flavors and different capacities of the same brand.
[0051] In one embodiment, when similar commodities exist, obtaining the local area of the feature to be extracted and the target feature to be extracted according to the initial commodity category includes: According to the initial commodity category, sample images corresponding to a plurality of subcategories under the category are selected from a preset commodity image database; Specifically, after the initial product category of the product is preliminarily determined based on the fused feature information, multiple subcategories corresponding to the initial product category will be searched in the pre-established product image database, and sample images corresponding to each subcategory will be selected as the basis for subsequent analysis. These sample images cover common variants under the initial product category, such as product versions with different flavors, packaging styles, or specifications and capacities. By selecting these representative sample images, rich and diverse visual sample resources can be provided for the subsequent differentiation of similar products, which helps to deeply explore the difference characteristics of these products in image details.
[0052] Inputting each of the sample images into a pre-trained saliency detection model to obtain a saliency heat map, wherein the saliency heat map is used to characterize the area in the sample image where the attention to the visual feature is most focused; Specifically, the sample images of each subcategory obtained in the previous step are input into the trained saliency detection model to automatically generate corresponding saliency heatmaps. The above saliency detection model is trained on a large image dataset with manually annotated salient regions. By learning which regions in the image are most likely to extract features, the model establishes a mapping relationship between the input image and the saliency distribution. During the training process, the model usually adopts an encoder-decoder structure, jointly models the global semantics and local texture of the image through a deep convolutional neural network, extracts and restores the spatial features of the image layer by layer to predict the importance of each pixel in terms of saliency. During training, real fixation points, boundary-preserving constraints, or contrast loss functions can be combined to improve the model's ability to perceive key regions in complex product images. Through such a training process, the model is capable of automatically determining the salient regions on unseen product images, thus providing a robust foundation for subsequent local region screening in product recognition. The saliency heatmap can highlight the regions with the most visual attention in the sample image. Common high-saliency regions usually include brand logos, product names, color borders, pattern elements, etc. These regions are often the first parts that machine vision notices when identifying products and are also important references for judging the detailed differences between products. Therefore, through saliency detection, a precise candidate range can be provided for subsequent local feature extraction and comparison.
[0053] Perform threshold segmentation on the saliency heatmap to obtain multiple candidate regions; Specifically, after obtaining the saliency heatmap, a pre-set threshold strategy is used to perform segmentation operations on the saliency heatmap. Regions with saliency values exceeding the pre-set segmentation threshold are extracted from the heatmap to generate several candidate regions. In summary, the saliency value is a numerical value in the saliency heatmap used to represent the degree of attention attraction of each pixel or region in the image. The numerical range is usually between zero and one, and the higher the value, the more obvious the feature at that position. In product images, regions with high saliency values often concentrate on brand logos, text content, pattern details, etc. Therefore, the saliency value can be used as an important indicator to measure the visual focus in the image, helping to locate the local regions that need to be analyzed in depth subsequently. These candidate regions represent information segments with high saliency and likely to contain key visual features in the sample image. In this way, the analysis scope can be effectively narrowed, focusing on the regions with the most differences and recognition value in the product image, providing targeted image segments for subsequent local feature analysis and scoring, thereby improving the recognition efficiency and accuracy.
[0054] Comprehensively score each of the candidate regions, and based on the scoring results, screen and obtain the local regions from each of the candidate regions; Specifically, after the extraction of candidate regions is completed, comprehensive scoring of each candidate region is performed from multiple dimensions. The scoring metrics include the average significance value of the region in the saliency heatmap, and the response intensity in terms of the relevance to the commodity category, etc. Specifically, category activation information is further obtained through the initial commodity category, the performance of the candidate region in the response heatmap is evaluated, and the saliency score value and the category-related score value are weighted and fused to obtain the final score of each candidate region. Finally, according to the scoring results, candidate regions with high discrimination ability are selected as local regions for subsequent extraction of target features and further identification of similar commodities.
[0055] In one embodiment, the comprehensive scoring of each candidate region and obtaining the local region according to the scoring results include: Obtain the significance values in the saliency heatmaps corresponding to each candidate region; Specifically, candidate regions are multiple regions of interest preliminarily selected on the heatmap through methods such as threshold segmentation. The significance values of all pixels included in each candidate region in the saliency heatmap are extracted. These significance values represent the degree of visual attention of the region in the entire image and are the core basis for measuring the saliency of the region.
[0056] According to each of the significance values, calculate the average significance value of each candidate region as the saliency score value; Specifically, the significance values within each candidate region are averaged to obtain a unified value as the saliency score value of the candidate region. This score value reflects the overall visual attractiveness of the region. The higher the value, the more likely the region is to be the first region of attention in most sample images, and thus has higher analysis and recognition value. This score value will be one of the important indicators for subsequent region screening.
[0057] According to the initial commodity category, input each of the sample images into a pre-trained image classification model to obtain a category activation map; Specifically, to further evaluate the relevance between candidate regions and product categories, each of the sample images is input into a pre-trained image classification model. During the recognition process, the model generates a class activation map. The training process of the image classification model is based on a large-scale labeled product image dataset and adopts a deep convolutional neural network structure, such as a residual network or an efficient network structure. By continuously adjusting the weight parameters, it can accurately map the input image to a predefined product category label. During the training process, the model not only learns the correspondence between the overall image and the category but also implicitly learns which regions in the image best represent the category. The class activation map is generated based on this mechanism. When a sample image is input, the feature channels related to the target category are extracted from the last convolutional feature map of the image classification model, and these channels are weighted and summed and then reflected onto the image space to obtain a class activation map. This class activation map shows the image regions that have a strong response to the current initial product category, usually covering key recognition elements such as brand names, product structures, and main patterns, which can further verify and reinforce the regional information provided by the saliency heat map.
[0058] According to the class activation map, obtain the response heat map of each sample image under the current initial product category; Specifically, refine the class activation map into a response heat map related to the sub-categories corresponding to the initial product category. This response heat map focuses on the regions that can most strongly stimulate the response of the classification model under the current category. By extracting the corresponding pixel values of the candidate regions in this response heat map, the correlation between the candidate regions and the current product category can be measured, laying a foundation for calculating the category-related score value in the subsequent steps. This step helps to confirm which candidate regions are not only significant but also highly relevant to the product recognition task itself, improving the final recognition accuracy.
[0059] Statistically analyze the pixels at the corresponding positions of each candidate region in the response heat map, and calculate the average activation intensity of each candidate region as the category-related score value; Specifically, align the class activation map with the position of each candidate region, and extract the corresponding pixel values of the candidate regions in the response heat map. These pixel values reflect the response degree of the region to the initial product category in the image classification model. By averaging all the response values within the candidate region, the category-related score value of the candidate region is obtained, which represents the importance degree of the region to the target category. The advantage of this score is that it does not rely on manual judgment but comes from the internal feature responses of the deep model, effectively quantifying the true association between each region in the image and the category, which helps to accurately select local regions with category discriminative power in the subsequent steps.
[0060] Perform weighted fusion on the saliency score value and the category-related score value of each candidate region to obtain the score result of each candidate region; Specifically, in order to comprehensively evaluate the saliency and category relevance of each candidate region, the saliency score value and the category relevance score value are weighted and fused. The weighting strategy can be set according to the importance of the actual task. For example, a higher weight can be given to the category relevance score to highlight its role in commodity differentiation. The final fused score result represents the comprehensive recognition value of each candidate region, taking into account both the attention tendency (saliency) and the classification focus of the model (category relevance). This fusion mechanism realizes dual guidance of perception and semantics, effectively improving the accuracy and practicality of local feature extraction.
[0061] Compare each of the said score results with a preset score threshold. According to the comparison results, select at least one region from each candidate region as the said local region.
[0062] Specifically, obtain a preset score threshold, which is used to filter out candidate regions with low scores and no recognition significance. Compare the fused score results with this score threshold one by one, and only retain the regions with scores higher than the threshold as valid local regions. If multiple candidate regions meet the conditions, multiple local regions can be selected for subsequent processing. This step can avoid interference from noise regions through the threshold screening mechanism and improve the pertinence of feature extraction; while selecting multiple high-score regions can retain richer commodity detail information, providing a more complete feature basis for subsequent refined comparison and difference analysis.
[0063] Perform candidate feature extraction and feature evaluation on the said local region. According to the feature evaluation results, screen out the said target feature from the extracted candidate features.
[0064] Specifically, use a trained feature extraction network (such as based on a convolutional or vision Transformer architecture) to perform in-depth image feature extraction on each local region. The extracted feature vectors contain multi-dimensional information such as color, shape, texture, and pattern, forming multiple candidate features. Subsequently, evaluate these candidate features, for example, by comparing their similarity with the standard features of the same category of commodities, performing difference analysis, or using a pre-trained feature quality scoring model to score, and screen out the target feature that is the most representative and can best distinguish this commodity. The advantage of this process is that it can automatically extract the key features that are most useful for commodity recognition from complex images, not only improving the accuracy of classification but also supporting efficient applications in subsequent scenarios such as retrieval, comparison, and recommendation.
[0065] In one embodiment, the performing candidate feature extraction and feature evaluation on the said local region, and screening out the said target feature from the extracted candidate features according to the feature evaluation results includes: According to the initial product category, obtain the feature extraction strategies corresponding to each candidate feature of this category; Specifically, after identifying the initial product category to which the current product image belongs, such as carbonated beverages, fruit juice beverages, functional beverages, lactic acid bacteria beverages, etc., the feature extraction scheme specifically configured for this category will be automatically retrieved from the preset feature extraction strategy library. This strategy scheme summarizes the visual structure differences of beverage products. For example, in the category of carbonated beverages, the strategy will first extract the color block boundaries of the bottle body pattern, font style features (such as prominent brand names), bubble textures or bottle cap shapes; while in fruit juice beverages, the feature extraction strategy emphasizes more on the color saturation of the fruit pattern area, the juice content words on the bottle label, and the degree of realism corresponding to the pattern image. Each feature extraction strategy defines parameters such as the spatial position, scale size, priority channel, receptive field range (which can be achieved by adjusting the convolution kernel size or network layer) of feature extraction, guiding subsequent targeted multi-path feature extraction within the local area. The advantage of this strategy design is that it can automatically adapt to the most discriminative image features in different beverage categories, thereby improving the system's ability to distinguish similar products within a large category and reducing the risk of misclassification caused by similar styles.
[0066] According to each of the feature extraction strategies, perform multi-path feature extraction on the local area to obtain multiple candidate feature information; Specifically, according to the selected feature extraction strategy, perform multi-path feature extraction on the local area screened in the previous step. By multi-path, it means that through different neural network branches or parameter paths (such as multi-scale convolution, different channel attention paths, shape feature paths, etc.), the image content of the local area is analyzed in parallel from multiple angles and scales. Each path outputs a set of feature representations, and finally multiple candidate feature information is obtained. The candidate feature information includes multi-dimensional data such as local texture, edge contour, color distribution pattern, and pattern arrangement. This multi-path method can capture the detailed changes in the image that may be used for classification or comparison to the greatest extent, improving the system's representation ability for complex product images and the robustness of subsequent analysis.
[0067] Perform distribution consistency analysis on each of the candidate feature information in each of the sample images, and obtain the frequency and position deviation of each candidate feature information appearing in different sample images as the distribution consistency index; Specifically, to further determine the stability and representativeness of candidate features, distribution consistency analysis is performed on the same candidate feature information extracted from multiple sample images. The distribution consistency analysis includes: statistically analyzing the frequency of occurrence of the candidate feature in all sample images, i.e., whether it can stably appear in most samples; calculating the position deviation of the candidate in the image, such as the degree of difference in its center point coordinates in each image. Features with a high frequency and a small position deviation usually indicate strong commonality and stability in this product category, and thus are given a higher distribution consistency index score. This index not only helps to filter out noise or accidentally occurring invalid features, but also improves the accuracy and interpretability of subsequent target feature screening.
[0068] Respectively input each of the candidate feature information into a pre-trained product recognition model to obtain a recognition result, and obtain the classification confidence corresponding to the recognition result as the classification response intensity of each candidate feature information. Specifically, each of the candidate feature information is input one by one into a product recognition model pre-trained with a large-scale product dataset. This model is usually based on a deep convolutional neural network architecture and can automatically extract multi-level semantic representations from the input local features and perform fine-grained category discrimination. When the model outputs a recognition result, it will give the corresponding classification confidence as the classification response intensity of each candidate feature information. This confidence reflects the degree of certainty of the model's judgment on the product category corresponding to the candidate feature. The classification confidence not only reflects the identification ability of the feature, but also quantifies the contribution of the feature to the overall classification decision, providing a scientific basis for subsequent feature screening.
[0069] Evaluate each of the candidate feature information according to the distribution consistency index and the classification response intensity to obtain a feature evaluation result. Specifically, a comprehensive evaluation is carried out by combining the distribution consistency index and the classification response intensity of the candidate feature. The distribution consistency index measures the frequency of occurrence and spatial position stability of the feature in different samples of the same category, ensuring that the feature is not accidental or noise-induced sporadic information, but a common feature that truly represents this category; while the classification response intensity measures the influence and importance of the feature on the classification result. By designing a weighted fusion or multi-dimensional scoring mechanism, these two indexes are comprehensively considered to obtain a feature evaluation result, avoiding misjudgment caused by a single dimension, so as to obtain a more objective and comprehensive feature quality evaluation result, ensuring that the selected features are both stable and effective.
[0070] Screen out the target feature from each of the candidate feature information according to the feature evaluation result.
[0071] Specifically, based on the above-mentioned feature evaluation results, target features are screened out from all candidate feature information. During the screening process, features with high distribution consistency and high classification response intensity are preferentially selected. These target features exhibit good stability and discriminability in multiple samples. The screening of target features not only improves the accuracy of commodity category differentiation but also effectively supports subsequent similar commodity recognition and fine-grained analysis. Through this screening step, the ability to capture the detailed differences of beverage commodities is greatly improved, and the generalization performance and practical value of the overall recognition model are enhanced.
[0072] Feature extraction is performed on the target image according to the local region and the target feature to obtain local region feature information; Specifically, according to the local region and target feature obtained from the previous screening, a fine feature extraction operation is performed on the target image. By performing feature extraction on the target image, multi-dimensional information including target features is focused on and extracted within the local region of the target image. The multi-dimensional information includes color distribution, texture details, shape structure, and the spatial relationship of feature points, etc. Through deep learning techniques such as multi-path convolutional neural networks or feature pyramid networks, the discriminative detail features in the local region can be fully captured, ensuring that the extracted local region feature information is both rich and representative, thereby providing accurate and high-quality input data for subsequent commodity classification.
[0073] The initial commodity category is classified according to the local region feature information to obtain the target commodity category as the commodity target recognition result.
[0074] Specifically, the extracted local region feature information is input into a fine-grained classification model pre-trained for the initial commodity category for classification determination. This classification model is specifically designed to process local features and can fully exploit the subtle feature differences between different beverage categories, thereby accurately distinguishing the sub-categories within the initial commodity category. The model enhances the perception ability of key local features through multi-layer non-linear transformation and attention mechanisms, and finally outputs the target commodity category as the commodity target recognition result. This process not only improves the accuracy and robustness of classification but also effectively reduces the misjudgment risk caused by the complex or diverse background of commodity images, realizing the efficiency and accuracy of commodity target recognition.
[0075] Embodiment 2
[0076] Please refer to Figure 7 , Embodiment 2 of the present invention also provides a multi-target commodity recognition device based on multi-modal data processing. The device includes: A real-time image acquisition module, configured to acquire real-time video data in a commodity trading scenario and decompose the real-time video data into multiple frames of real-time images; A preprocessing and label information extraction module, which is used to preprocess the real-time image and extract label information, and determine the preprocessed target image and the text information corresponding to the product label; An instance segmentation module, which is used to perform instance segmentation on the target image to determine the product position information; A feature extraction module, which is used to extract features from the target image according to the product position information to determine the product image feature information; A multi-modal large model training module, which is used to fine-tune and optimize an open-source multi-modal vision-language model according to pre-collected multi-source privatized data in the intelligent vending scenario to obtain a multi-modal large model for product recognition; A product recognition module, which is used to input the product image feature information and the text information into the multi-modal large model for information fusion, and determine the product target recognition result according to the fused feature information.
[0077] Specifically, the multi-object commodity recognition device based on multi-modal data processing provided in Embodiment 2 of the present invention is used. The device includes: a real-time image acquisition module for acquiring real-time video data in a commodity trading scenario and decomposing the real-time video data into multiple frames of real-time images; a preprocessing and label information extraction module for preprocessing the real-time images and extracting label information to determine the preprocessed target images and the text information corresponding to the commodity labels; an instance segmentation module for performing instance segmentation on the target images to determine the commodity position information; a feature extraction module for extracting features from the target images according to the commodity position information to determine the commodity image feature information; a multi-modal large model training module for fine-tuning and optimizing an open-source multi-modal vision-language model based on pre-collected multi-source privatized data in a smart vending scenario to obtain a multi-modal large model for commodity recognition; and a commodity recognition module for inputting the commodity image feature information and the text information into the multi-modal large model for information fusion and determining the commodity target recognition result according to the fused feature information. This device extracts multiple frames of images from real-time video data, and through preprocessing and label information extraction, it preliminarily recognizes the commodities in each frame of image, extracts the commodity labels and the corresponding text information. Then, it processes the target images through instance segmentation technology to accurately locate the bounding boxes of each commodity and solve the occlusion problem between commodities. It further extracts image features, including visual features, spatial features, and semantic features, using the commodity position information. Through information fusion with the multi-modal large model, it combines the commodity image features with the extracted text information to enhance the recognition ability for complex commodity targets. Finally, based on the fusion of image and text information, the model can accurately distinguish and recognize multiple commodity targets. Even in a complex environment where multiple commodities coexist and there are partial overlaps or occlusions, it can still maintain high recognition performance, not only improving the recognition accuracy but also being able to operate stably in a dynamic scenario to meet the multi-target recognition requirements.
[0078] Embodiment 3
[0079] In addition, Embodiment 3 of the present invention also provides a multi-object commodity recognition system based on multi-modal data processing, including: an image acquisition device, at least one processor, at least one memory, and computer program instructions stored in the memory. When the computer program instructions are executed by the processor, the method described in Embodiment 1 is implemented.
[0080] Figure 8 The hardware structure diagram of the electronic device provided in Embodiment 3 of the present invention is shown.
[0081] The electronic device may include a processor and a memory storing computer program instructions.
[0082] Specifically, the above-mentioned processor may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.
[0083] The memory may include a mass storage for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to the data processing device. In a particular embodiment, the memory is a non-volatile solid state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0084] The processor reads and executes the computer program instructions stored in the memory to implement any one of the multi-object commodity recognition methods based on multi-modal data processing in the above embodiments.
[0085] In one example, the electronic device may further include a communication interface and a bus. Among them, as Figure 8 shown, the processor, the memory, and the communication interface are connected through the bus and complete communication with each other.
[0086] The communication interface is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present invention.
[0087] A bus includes hardware, software, or both, and couples components of the device to each other. By way of example and not limitation, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus can include one or more buses. Although embodiments of the invention describe and illustrate a particular bus, the invention contemplates any suitable bus or interconnect.
[0088] In summary, embodiments of the present invention provide a multi-target commodity recognition method, apparatus, and system based on multi-modal data processing.
[0089] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0090] The functional blocks shown in the above-described block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an Application Specific Integrated Circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0091] The user information involved in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant location, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0092] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps. That is to say, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0093] As mentioned above, the above are only specific embodiments of the present invention. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A multi-object commodity recognition method based on multi-modal data processing, characterized in that, The method includes: Obtain real-time video data in a commodity trading scenario, and decompose the real-time video data into multiple frames of real-time images; Preprocess the real-time images and extract label information to determine the preprocessed target images and the text information corresponding to the commodity labels; Perform instance segmentation on the target images to determine the commodity position information; According to the commodity position information, extract features from the target images to determine the commodity image feature information; According to the pre-collected multi-source private data in an intelligent vending scenario, fine-tune and optimize an open-source multi-modal vision-language model to obtain a multi-modal large model for commodity recognition; Input the commodity image feature information and the text information into the multi-modal large model for information fusion, and determine the commodity target recognition result according to the fused feature information.
2. The multi-object commodity recognition method based on multi-modal data processing according to claim 1, wherein The preprocessing the real-time images and extracting label information to determine the preprocessed target images and the text information corresponding to the commodity labels includes: Adjust the size and perform noise reduction processing on the real-time images to determine the target images; Perform object detection on the target images to determine the commodity area position information; According to the commodity area position information, process the commodity labels in the commodity area through optical character recognition technology to determine the text information.
3. The multi-object commodity recognition method based on multi-modal data processing according to claim 2, characterized in that, The performing instance segmentation on the target images to determine the commodity position information includes: According to the commodity area position information, extract the feature information of the commodity area through a convolutional neural network, and determine the candidate areas according to the extracted feature information; Process the candidate areas through an instance segmentation network to determine the binary images corresponding to each commodity target; Use post-processing technology to process the binary images to determine the commodity position information.
4. The multi-object commodity recognition method based on multi-modal data processing according to claim 1, characterized in that The according to the pre-collected multi-source private data in an intelligent vending scenario, fine-tune and optimize an open-source multi-modal vision-language model to obtain a multi-modal large model for commodity recognition includes: According to the pre-collected multi-source raw data in an intelligent vending scenario, clean and structure the multi-source raw data to obtain an annotated data set; According to the image information and text labels in the annotated data set, pair and construct the graphic and text data, and perform format conversion and unified preprocessing on it to obtain a multi-modal input sample set for training; According to the multi-modal input sample set, load the weight parameters of an open-source vision-language pre-training model, and construct a vision coding and language coding network structure that supports joint optimization to obtain an initial structure of a multi-modal model for fine-tuning; According to the characteristics of the vending scenario and the recognition accuracy requirements, perform fine-tuning training on the initial structure of the multi-modal model, and optimize the hyperparameter configuration through a cross-validation strategy to obtain multiple candidate multi-modal models; According to the preset accuracy rate, recall rate, and response time, evaluate and compare each candidate multi-modal model to obtain the multi-modal large model.
5. The multi-object commodity recognition method based on multi-modal data processing according to any one of claims 1 to 4, characterized in that, The inputting the commodity image feature information and the text information into the multi-modal large model for information fusion, and determining the commodity target recognition result according to the fused feature information includes: Input the commodity image feature information and the text information into the multi-modal large model to obtain fused feature information that combines image features and text semantics; Input the fused feature information into a pre-trained commodity classification model to obtain an initial commodity category; Based on the initial commodity category, determine whether there are similar commodities in the current initial commodity category; When there are similar commodities, according to the initial commodity category, obtain the local area of the feature to be extracted and the target feature to be extracted; According to the local area and the target feature, perform feature extraction on the target image to obtain local area feature information; According to the local area feature information, classify the initial commodity category to obtain the target commodity category as the commodity target recognition result.
6. The multi-object commodity recognition method based on multi-modal data processing according to claim 5, wherein The step of when there are similar commodities, according to the initial commodity category, obtaining the local area of the feature to be extracted and the target feature to be extracted includes: According to the initial commodity category, select the sample images corresponding to multiple sub-categories under this category from a preset commodity image database; Input each of the sample images into a pre-trained saliency detection model to obtain a saliency heat map, where the saliency heat map is used to represent the area in the sample image with the most concentrated visual feature attention; Perform threshold segmentation on the saliency heat map to obtain multiple candidate regions; Perform a comprehensive score on each of the candidate regions, and according to the scoring result, screen the local area from each of the candidate regions; Perform candidate feature extraction and feature evaluation on the local area, and according to the feature evaluation result, screen the target feature from the extracted candidate features.
7. The multi-object commodity recognition method based on multi-modal data processing according to claim 6, wherein, The step of performing a comprehensive score on each of the candidate regions and obtaining the local area according to the scoring result includes: Obtain the significant value in the saliency heat map corresponding to each candidate region; According to each of the significant values, calculate the average significant value of each candidate region as the saliency scoring value; According to the initial commodity category, input each of the sample images into a pre-trained image classification model to obtain a class activation map; According to the class activation map, obtain the response heat map of each sample image under the current initial commodity category; Count the pixels at the corresponding positions of each candidate region in the response heat map, and calculate the average activation intensity of each candidate region as the class-related scoring value; Perform weighted fusion on the saliency scoring value and the class-related scoring value of each candidate region to obtain the scoring result of each candidate region; Compare each of the scoring results with a preset scoring threshold, and according to the comparison result, select at least one region from each of the candidate regions as the local area.
8. The multi-object commodity recognition method based on multi-modal data processing according to claim 6, characterized in that, The step of performing candidate feature extraction and feature evaluation on the local area and screening the target feature from the extracted candidate features according to the feature evaluation result includes: According to the initial commodity category, obtain the feature extraction strategies corresponding to each candidate feature of this category; According to each of the feature extraction strategies, perform multi-path feature extraction on the local area to obtain multiple candidate feature information; Perform distribution consistency analysis on each of the candidate feature information in each of the sample images, and obtain the frequency and position deviation of each candidate feature information appearing in different sample images as the distribution consistency index; Input each of the candidate feature information into a pre-trained commodity recognition model respectively to obtain recognition results, and obtain the classification confidence corresponding to the recognition results as the classification response intensity of each candidate feature information; Evaluate each of the candidate feature information according to the distribution consistency index and the classification response intensity to obtain a feature evaluation result; Screen out the target features from each of the candidate feature information according to the feature evaluation result.
9. A multi-object commodity recognition device based on multi-modal data processing, characterized in that, The device includes: A real-time image acquisition module, configured to acquire real-time video data in a commodity trading scenario, and decompose the real-time video data into multiple frames of real-time images; A preprocessing and label information extraction module, configured to perform preprocessing and label information extraction on the real-time images to determine the preprocessed target images and the text information corresponding to the commodity labels; An instance segmentation module, configured to perform instance segmentation on the target images to determine commodity position information; A feature extraction module, configured to extract features from the target images according to the commodity position information to determine commodity image feature information; A multi-modal large model training module, configured to fine-tune and optimize an open-source multi-modal vision-language model according to pre-collected multi-source privatized data in a smart vending scenario to obtain a multi-modal large model for commodity recognition; A commodity recognition module, configured to input the commodity image feature information and the text information into the multi-modal large model for information fusion, and determine a commodity target recognition result according to the fused feature information.
10. A multi-target commodity recognition system based on multi-modal data processing, characterized in that, Including: An image acquisition device, at least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1-8 when the computer program instructions are executed by the processor.
Citation Information
Patent Citations
Commodity identification method and system based on multi-modal data
CN115601582A
Goods shelf commodity identification method and system
CN116109992A
Target identification method, commodity identification method, equipment and storage medium
CN118279705A
Retail terminal inventory replenishment method, device and system based on multi-modal data
CN119919060A
Method, device, computer equipment and storage medium for identifying illegal commodity
US20240331425A1
Cited By
Commodity identification method and device
CN120564194A
A merchandise identification method and apparatus
CN120564194B
Clothing data set construction method based on multiple views and structured attributes and application thereof
CN120994646A
A method for constructing a clothing dataset based on multi-view and structured attributes and application thereof
CN120994646B
Target detection method, system and device based on hierarchical collaborative reasoning
CN121330440A