Dinner plate identification method, device and equipment, storage medium and computer program product
Through the combination of instance segmentation model, feature extraction model and vector search library, a lightweight meal plate recognition method is realized, solving the problems of high hardware dependence and strong data dependence in the existing technology, and improving the versatility of meal plate recognition and the ability of new meal plate recognition.
Patent Information
- Application Number
- CN202510284805.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-06
Smart Images

Figure CN120107598A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a method, device, equipment, storage medium and computer program product for plate recognition. Background Art
[0002] Plate recognition refers to the precise identification of the plate's appearance features (such as size, shape, color, pattern), material (such as ceramic, plastic, stainless steel), status (such as completeness, cleanliness) and functional attributes (such as purpose or exclusive marking) through intelligent means. It is widely used in scenarios such as smart catering, inventory management, and smart tableware classification. Existing plate recognition methods are mainly divided into two categories: one is based on sensor technologies such as weight sensors, infrared sensors, and wireless radio frequency identification technology RFID (Radio Frequency Identification) tags to identify the type of plate; the other is to use computer vision technologies such as edge detection, contour analysis, and convolutional neural networks (CNN) to classify plate images.
[0003] On the one hand, sensor-based plate recognition methods have a strong dependence on hardware. The deployment cost of sensors and intelligent assembly lines is high, making it difficult to adapt to small catering scenarios; and hardware equipment requires regular calibration and maintenance, which increases operating costs. On the other hand, computer vision-based plate recognition methods have a strong dependence on data. The accuracy of the model is highly dependent on a large amount of labeled data, and different scenarios require separate data collection and labeling; and the algorithm has poor generalization ability for new plates or plates that are not in the training set. Therefore, it is urgent to propose a more universal plate recognition method.
[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the present application is to provide a method, device, equipment, storage medium and computer program product for plate recognition, aiming to solve the technical problem of poor versatility of plate recognition technology.
[0006] To achieve the above objectives, the present application proposes a plate recognition method, the method comprising:
[0007] Acquire an initial dinner plate image, and perform instance segmentation on the initial dinner plate image using an instance segmentation model to obtain a mask image of a target dinner plate in the initial dinner plate image;
[0008] Acquiring edge pixels of the mask image of the target dinner plate to generate an edge pixel image of the target dinner plate;
[0009] Performing feature extraction on the edge pixel image of the target plate by using a feature extraction model to generate a target feature vector of the target plate;
[0010] A target feature vector of the target plate is subjected to feature retrieval through a vector retrieval library to obtain a plate recognition result of the target plate.
[0011] In one embodiment, the instance segmentation model includes a feature extraction layer, a feature fusion layer, and a detection segmentation layer. The step of performing instance segmentation on the initial plate image by using the instance segmentation model to obtain a mask image of the target plate in the initial plate image includes:
[0012] Preprocessing the initial plate image;
[0013] Performing multi-scale feature extraction on the preprocessed initial plate image through the feature extraction layer to obtain a first plate feature image of the initial plate image;
[0014] Inputting the first dinner plate feature image into the feature fusion layer for global feature fusion to obtain a second dinner plate feature image of the initial dinner plate image;
[0015] The second dinner plate feature image is input into the detection segmentation layer for detection and classification, so as to obtain a mask image of the target dinner plate in the initial dinner plate image.
[0016] In one embodiment, the step of acquiring edge pixels of the mask image of the target plate to generate the edge pixel image of the target plate comprises:
[0017] Preprocessing the mask image of the target dinner plate;
[0018] Performing edge pixel detection on the preprocessed mask image using an edge detection algorithm to obtain an edge pixel detection result of the target dinner plate;
[0019] According to the edge pixel detection result, the area where the edge pixel detection result is located is cropped to generate an edge pixel image of the target plate.
[0020] In one embodiment, the step of extracting features from the edge pixel image of the target plate using a feature extraction model to generate a target feature vector of the target plate includes:
[0021] Performing feature extraction on the edge pixel image by using a feature extraction model to generate an initial feature image of the target plate;
[0022] Performing global average pooling processing on the initial feature image to generate an initial feature vector of the target plate;
[0023] The initial feature vector of the target plate is normalized to generate a target feature vector of the target plate.
[0024] In one embodiment, the step of performing feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate comprises:
[0025] Initialize a vector index based on a standard feature vector, and build a vector search library according to the vector index;
[0026] Calculating the similarity between the target feature vector and the standard feature vector to obtain a similarity ranking table of the target feature vector;
[0027] Based on the similarity ranking table and a preset similarity threshold, a plate recognition result of the target plate is determined from the vector retrieval library.
[0028] In one embodiment, before the step of performing instance segmentation on the initial plate image using an instance segmentation model to obtain a mask image of the target plate in the initial plate image, the step further includes:
[0029] Acquire several groups of first training images;
[0030] Segmenting and annotating the target plate in each of the first training images using an annotation tool to generate an annotated training image;
[0031] Performing preprocessing of adjustment and normalization on each of the labeled training images, and inputting the preprocessed labeled training images into the instance segmentation model for forward propagation to obtain a predicted segmentation result;
[0032] Based on the predicted segmentation result and the standard training image, the segmentation loss of the predicted segmentation result is calculated, and the segmentation loss is back-propagated to the instance segmentation model to update the model parameters to obtain a trained instance segmentation model.
[0033] In one embodiment, before the step of extracting features from the edge pixel image of the target plate using a feature extraction model to generate a target feature vector of the target plate, the step further includes:
[0034] Acquire several groups of second training images and reference labels of the second training images;
[0035] Performing preprocessing of adjustment and normalization on each of the second training images, inputting the preprocessed second training images into a feature extraction model for forward propagation, and obtaining a classification label probability of each of the second training images;
[0036] Based on the classification label probability and the reference label, the prediction loss of the classification label probability of the second training image is calculated, and the prediction loss is back-propagated to the feature extraction model to update the model parameters to obtain a trained feature extraction model.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a plate identification device, which includes:
[0038] An instance segmentation module is used to obtain an initial plate image, perform instance segmentation on the initial plate image through an instance segmentation model, and obtain a mask image of a target plate in the initial plate image;
[0039] A pixel extraction module, used for acquiring edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate;
[0040] A feature extraction module, used for extracting features from the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate;
[0041] The feature retrieval module is used to perform feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
[0042] In addition, to achieve the above-mentioned purpose, the present application also proposes a plate recognition device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the plate recognition method described above.
[0043] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the plate identification method described above are implemented.
[0044] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the plate recognition method described above are implemented.
[0045] One or more technical solutions proposed in this application have at least the following technical effects:
[0046] The present application discloses a plate recognition method, device, equipment, storage medium and computer program product, including: obtaining an initial plate image, performing instance segmentation on the initial plate image through an instance segmentation model to obtain a mask image of a target plate in the initial plate image; performing edge pixel acquisition on the mask image of the target plate to generate an edge pixel image of the target plate; performing feature extraction on the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate; performing feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate. The present application constructs a lightweight plate recognition method through an instance segmentation model, a feature extraction model and a vector retrieval library, reduces the hardware dependence of plate recognition and improves the versatility of plate recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A schematic diagram of a process flow provided for the first embodiment of the plate recognition method of the present application;
[0050] Figure 2 A schematic diagram of the structure of an instance segmentation model provided in an embodiment of the plate recognition method of the present application;
[0051] Figure 3 A schematic diagram of the structure of the feature fusion module provided for the plate recognition method of this application;
[0052] Figure 4 A schematic diagram of the structure of the spatial pyramid pooling module provided for the plate recognition method of this application;
[0053] Figure 5 This is a schematic diagram of the module structure of the dinner plate recognition device according to an embodiment of the present application;
[0054] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the plate recognition method in the embodiment of the present application.
[0055] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0056] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0057] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0058] The main solution of the embodiment of the present application is: obtaining an initial plate image, performing instance segmentation on the initial plate image through an instance segmentation model, and obtaining a mask image of a target plate in the initial plate image; acquiring edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate; performing feature extraction on the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate; performing feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
[0059] In this embodiment, for the convenience of description, the following description is made by taking the dinner plate recognition device as the execution subject.
[0060] Plate recognition refers to the precise identification of the plate's appearance features (such as size, shape, color, pattern), material (such as ceramic, plastic, stainless steel), status (such as completeness, cleanliness) and functional attributes (such as purpose or exclusive marking) through intelligent means. It is widely used in scenarios such as smart catering, inventory management, and smart tableware classification. Existing plate recognition methods are mainly divided into two categories: one is based on sensor technologies such as weight sensors, infrared sensors, and wireless radio frequency identification technology RFID (Radio Frequency Identification) tags to identify the type of plate; the other is to use computer vision technologies such as edge detection, contour analysis, and convolutional neural networks (CNN) to classify plate images.
[0061] On the one hand, sensor-based plate recognition methods have a strong dependence on hardware. The deployment cost of sensors and intelligent assembly lines is high, making it difficult to adapt to small catering scenarios; and hardware equipment requires regular calibration and maintenance, which increases operating costs. On the other hand, computer vision-based plate recognition methods have a strong dependence on data. The accuracy of the model is highly dependent on a large amount of labeled data, and different scenarios require separate data collection and labeling; and the algorithm has poor generalization ability for new plates or plates that are not in the training set. Therefore, it is urgent to propose a more universal plate recognition method.
[0062] The present application provides a solution, which constructs a lightweight plate recognition method through an instance segmentation model, a feature extraction model and a vector retrieval library, reduces the hardware dependence of plate recognition, and improves the versatility of plate recognition.
[0063] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a plate recognition device, etc. The following takes the plate recognition device as an example to illustrate this embodiment and the following embodiments.
[0064] Based on this, the present application embodiment provides a method for plate recognition, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the plate recognition method of the present application.
[0065] In this embodiment, the plate recognition method includes steps S11 to S14:
[0066] Step S11, obtaining an initial plate image, performing instance segmentation on the initial plate image through an instance segmentation model, and obtaining a mask image of a target plate in the initial plate image.
[0067] It should be noted that the initial dinner plate image refers to an original image captured by an image capturing device (such as a camera, a scanner, etc.), which includes the dinner plate and objects that may be attached to it (such as food, tableware, etc.).
[0068] In addition, it should be noted that the instance segmentation model in this application refers to the YOLOv8-seg model, which is used for instance segmentation tasks. It inherits the efficient target detection capability of YOLOv8 and adds the function of semantic segmentation, which can simultaneously complete target detection and pixel-level segmentation. The core idea of YOLOv8-seg is to combine target detection and instance segmentation. It first locates the target object in the image through the detection module of YOLOv8, and then uses the segmentation module to generate a corresponding segmentation mask for each detected object. Among them, the segmentation mask is the final output in the instance segmentation task, which indicates whether each pixel in the image belongs to a specific target object, and indicates the pixel-level boundary of each target object in the image. In the YOLOv8-seg model, the segmentation mask is generated by combining the prototype mask with the features of the detected target object. Specifically, the model selects a suitable mask from the prototype mask for linear combination according to the position and features of the detected target object, and generates the final segmentation mask through a series of subsequent processing steps (such as upsampling, convolution, etc.). The segmentation mask is usually presented in the form of a binary image, where the white area (or pixel value 1) represents the target object and the black area (or pixel value 0) represents the background.
[0069] To better understand the instance segmentation model of this application, please refer to Figure 2 The figure is a schematic diagram of the structure of the YOLOv8-seg model. The YOLOv8-seg model consists of three parts: the backbone network, the neck network, and the head network. The backbone network is mainly composed of three modules: the convolution module CBS (Conv Block with SiLU), the feature fusion module C2f, and the spatial pyramid pooling module SPPF (Spatial Pyramid Pooling Fast). The neck network is mainly composed of the upsampling module UpSample, the concatenation module Concate, the convolution module, and the feature fusion module. The head network consists of a convolution module*2, a two-dimensional convolution module, and a prototype head Proto.
[0070] In addition, it should be noted that the mask image of the target plate refers to the mask image obtained by processing the initial plate image through the instance segmentation model, which only contains pixels in the target plate area, while the pixels in other non-plate areas are set to 0 or a certain background value. The mask image is used to highlight the target plate to facilitate subsequent edge detection and feature extraction. In one embodiment of the present application, the mask image is presented in the form of a binary image, in which the white area (or pixel point with a value of 1) represents the target object, and the black area (or pixel point with a value of 0) represents the background.
[0071] Specifically, first, an initial image containing a target plate is acquired through an image acquisition device such as a camera or a scanner or a network; the initial plate image is input into a trained instance segmentation model for instance segmentation to obtain a mask image of the target plate.
[0072] Step S12, acquiring edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate.
[0073] It should be noted that the edge pixel image of the target plate refers to the image generated by performing an edge pixel acquisition operation on the mask image of the target plate. The edge pixel image only contains the edge pixels of the target plate, that is, those pixels located on the outline of the plate. The edge pixel image is used to represent the shape and outline of the plate and is an important basis for subsequent feature extraction.
[0074] Specifically, edge pixels of the target plate in the mask image are extracted by an edge detection algorithm to generate an edge pixel image of the target plate.
[0075] Step S13, extracting features from the edge pixel image of the target plate using a feature extraction model to generate a target feature vector of the target plate.
[0076] It should be noted that the feature extraction model adopts the PP-LCNetV2 model. PP-LCNetV2 uses a re-parameterization strategy to combine deep convolutions with convolution kernels of different sizes to improve the perception ability of the neural network; optimizing the deep separable convolution not only enhances the model fitting ability, but also improves the model efficiency.
[0077] In addition, it should be noted that the target feature vector of the target plate is a feature value obtained by processing the edge pixel image of the target plate through a feature extraction model, including but not limited to the shape, size, edge smoothness, etc. of the target plate.
[0078] Specifically, the edge pixel image of the target plate is input into a pre-trained feature extraction model for feature extraction to generate a target feature vector of the target plate, which contains key feature information of the edge of the plate.
[0079] Step S14, performing feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
[0080] It should be noted that the vector retrieval library is a database that stores vector indexes, allowing users to retrieve the most similar vectors by calculating the similarity between the query vector and the vectors in the retrieval library. In one embodiment of the present application, the vector retrieval library uses the FAISS vector retrieval library, which is a vector retrieval library designed specifically for processing large-scale data. It focuses on vector similarity search and can quickly and efficiently compare a large number of vectors. Since vectors can represent various data features, FAISS is suitable for a variety of scenarios, such as music recommendation, image recognition, text retrieval, etc.
[0081] In addition, it should be noted that the plate recognition result refers to the result obtained after feature retrieval of the target feature vector of the target plate through the vector retrieval library, which indicates the classification of the target plate determined based on the similarity between the target feature vector of the target plate and the standard feature vector in the vector retrieval library.
[0082] Specifically, a vector retrieval library is pre-constructed, and the target feature vector is input into the vector retrieval library. The retrieval library calculates the similarity between the target feature vector and the feature vectors in the library, finds the known feature vector that is most similar to the target feature vector, and obtains the plate recognition result of the target plate.
[0083] Through the above scheme, this embodiment constructs a lightweight plate recognition method through an instance segmentation model, a feature extraction model, and a vector retrieval library, which reduces the hardware dependence of plate recognition and improves the versatility of plate recognition.
[0084] Based on the above implementation scheme, in a feasible implementation manner, the instance segmentation model includes a feature extraction layer, a feature fusion layer, and a detection segmentation layer. The step of performing instance segmentation on the initial plate image by the instance segmentation model to obtain a mask image of the target plate in the initial plate image includes S21 to S24:
[0085] Step S21, pre-processing the initial plate image.
[0086] Specifically, the initial plate image is preprocessed, including image resizing and normalization, to ensure that the input image meets the input requirements of the instance segmentation model. For example, image resizing can be scaling the input initial plate image to the input size specified by the model (e.g., 640×640); normalization operation refers to normalizing the image pixel values to the range of [-1,1] or [0,1] and performing standardization.
[0087] Step S22, performing multi-scale feature extraction on the preprocessed initial plate image through the feature extraction layer to obtain a first plate feature image of the initial plate image.
[0088] It should be noted that the feature extraction layer of the instance segmentation model refers to the backbone network of the YOLOv8-seg model, namely the Backbone network. The backbone network is mainly composed of three modules: the convolution module CBS, the feature fusion module C2f, and the spatial pyramid pooling module SPPF, which are used to obtain high-quality multi-scale features from the input image. The convolution module is composed of a convolution layer, a batch normalization layer, and an activation layer. The entire backbone network contains 5 convolution modules. Figure 3 The figure shows the network structure of the feature fusion module C2f. In each feature fusion module, the input features pass through two parallel paths. One part of the features is directly connected through the data block Chunk across layers to retain the original features; the other part of the features passes through multiple bottleneck modules BottleNeck. The bottleneck module includes two groups of convolution modules, which are used to extract high-order semantic information. Finally, the two parts of the features are fused. The network structure of the spatial pyramid pooling module SPPF is shown in Figure 1. Figure 4 As shown in the figure, the input features are sequentially passed through a convolution module, three 2D maximum pooling layers, and then the four output features are concatenated according to the channel dimension. Finally, the merged features are input into a convolution module to obtain the final output. The spatial pyramid pooling module integrates the information of the input feature map by operating on different receptive fields, enhancing the model's robustness to changes in target size. Without significantly increasing the computational overhead, the spatial dimension of the feature map is reduced through pooling operations while retaining global context information.
[0089] Specifically, the preprocessed initial plate image is input into the feature extraction layer of the instance segmentation model, and the convolution module in the feature extraction layer performs a preliminary convolution operation on the image to extract basic features; the feature fusion module extracts high-order semantic information through parallel paths and bottleneck modules, and retains the original features for feature fusion; the spatial pyramid pooling module operates on different receptive fields, integrates the information of the input feature map, and enhances the model's robustness to changes in target size; after processing by the entire backbone network, the first plate feature image of the initial plate image is obtained, which contains high-quality features at multiple scales.
[0090] Step S23, inputting the first dinner plate feature image into the feature fusion layer for global feature fusion to obtain a second dinner plate feature image of the initial dinner plate image.
[0091] It should be noted that the feature fusion layer refers to the neck network of the YOLOv8-seg model, namely the Neck network. The neck network contains three modules: upsampling module, downsampling fusion module, and feature fusion module C2f, which are mainly used for global feature fusion. The upsampling fusion module uses upsampling operations to increase the resolution of high-level features to match the resolution of lower-level features, and then splices the upsampled feature map with the feature map from the corresponding layer of the backbone network. The downsampling fusion module performs downsampling through convolution operations to reduce feature resolution and improve semantic expression capabilities. The downsampled features are then spliced with the corresponding features in the upsampling fusion module. The feature fusion module further fuses the spliced features to enhance the expressiveness of the features.
[0092] Specifically, the first plate feature image is input into the feature fusion layer, and the resolution of high-level features is improved through the upsampling fusion module to match the resolution of lower-level features and then spliced; the downsampling fusion module reduces the feature resolution through convolution operation to improve the semantic expression ability, and then splices it with the features in the upsampling fusion module; the feature fusion module further fuses the spliced features to enhance the expression ability of the features; after being processed by the feature fusion layer, the second plate feature image of the initial plate image is obtained, which contains the globally fused features.
[0093] Step S24: input the second plate feature image into the detection and segmentation layer for detection and classification, and obtain a mask image of the target plate in the initial plate image.
[0094] It should be noted that the detection and segmentation layer refers to the head network of the YOLOv8-seg model, namely the Head network, which consists of a detection head and a segmentation head. The detection head contains multiple convolutional layers and fully connected layers, processes the features output by the neck network, predicts the category probability of the target through the classification branch, and predicts the bounding box coordinates of the target through the regression branch. The anchor-free mechanism is adopted to directly predict the target at each position, which reduces the number of hyperparameters and improves the generalization ability of the model. The segmentation head adopts a structure similar to U-Net, including multiple upsampling layers and convolutional layers, upsamples and fuses the feature maps output by the neck network, gradually restores the resolution of the image, and finally generates a segmentation mask of the same size as the input image. Each pixel corresponds to a category label, which is used to indicate whether the pixel belongs to the target instance and which target instance it belongs to. At the same time, the attention mechanism is combined to enable the model to pay more attention to the boundary and detail information of the target and improve the accuracy of segmentation.
[0095] Specifically, the second plate feature image is input to the detection and segmentation layer, the features in the second plate feature image are detected and processed by the detection head, the category probability of the target feature is predicted by the classification branch, and the bounding box coordinates of the target feature are predicted by the regression branch. The segmentation head upsamples and fuses the features of the second plate feature image, gradually restores the resolution of the image, and generates a mask image of the same size as the preprocessed initial plate image. Each pixel in the mask image corresponds to a category label, which is used to indicate whether the pixel belongs to the target instance and which target instance it belongs to. At the same time, the attention mechanism is combined to enable the model to pay more attention to the boundary and detail information of the target, thereby improving the accuracy of segmentation. Finally, the detection and segmentation layer outputs the mask image of the target plate in the initial plate image, and each pixel corresponds to a category label, indicating whether the pixel belongs to the target instance and the corresponding specific target instance.
[0096] Based on the above implementation scheme, in a feasible implementation manner, the step of acquiring edge pixels of the mask image of the target plate to generate the edge pixel image of the target plate includes S31 to S33:
[0097] Step S31, preprocessing the mask image of the target dinner plate.
[0098] Specifically, the mask image of the target plate is subjected to preprocessing operations including image resizing and normalization to ensure that the input image meets the input requirements of the subsequent feature extraction model.
[0099] Step S32, performing edge pixel detection on the preprocessed mask image using an edge detection algorithm to obtain an edge pixel detection result of the target plate.
[0100] It should be noted that edge detection algorithm refers to a technique used in image processing to identify areas in an image where brightness or color changes rapidly. These areas usually correspond to the outline of an object, texture, or structural changes in the scene.
[0101] In addition, it should be noted that the edge pixel detection result is obtained by processing the image with the edge detection algorithm, and includes the position information of the edge pixels in the image. This information can be the coordinates of the edge pixels or the edge image after binarization, in which each edge line is continuous and refined, and the noise is removed to the greatest extent.
[0102] Specifically, select a suitable edge detection algorithm, such as Canny edge detection, Sobel edge detection, Laplacian edge detection, etc. These algorithms can detect grayscale changes in an image, thereby identifying edge pixels. Apply the selected edge detection algorithm to process the preprocessed mask image to obtain edge pixel detection results, and output a binary image containing edge pixels, where edge pixels are marked as white (or a specific color / value), and other pixels are marked as black (or another specific color / value).
[0103] Step S33: According to the edge pixel detection result, the area where the edge pixel detection result is located is cropped to generate an edge pixel image of the target plate.
[0104] It should be noted that the edge pixel image is cropped according to the edge pixel detection result, and it only contains the edge pixel area in the image.
[0105] Specifically, the edge pixel detection results are analyzed to determine the area where the edge pixels are located, determine all connected areas marked as edge pixels, and determine the circumscribed rectangles or irregular shape boundaries of these areas. According to the determined boundaries, the mask image is cropped. The cropping operation will only retain the area containing the edge pixels, thereby generating an edge pixel image of the target plate.
[0106] Through the above solution, this embodiment processes the preprocessed mask image through an edge detection algorithm, which can efficiently identify the edge pixels of the target plate and provide a reliable basis for subsequent image cropping.
[0107] Based on the above implementation scheme, in a feasible implementation manner, the step of extracting features from the edge pixel image of the target plate by using a feature extraction model to generate a target feature vector of the target plate includes S41 to S43:
[0108] Step S41 , extracting features from the edge pixel image using a feature extraction model to generate an initial feature image of the target plate.
[0109] It should be noted that the initial feature image refers to the feature image obtained after feature extraction of the edge pixel image through the feature extraction model PP-LCNetV2.
[0110] Specifically, the edge pixel image is input into the PP-LCNetV2 model, and the image passes through multiple convolutional layers and depth-wise separable convolutional layers in the model to extract features from the edge pixel image and generate the initial feature image of the target plate.
[0111] Step S42, performing global average pooling processing on the initial feature image to generate an initial feature vector of the target plate.
[0112] It should be noted that the initial feature vector is a one-dimensional vector obtained by processing the initial feature image through global average pooling. Global average pooling is a pooling operation that compresses the spatial dimensions of each feature channel (i.e., the height and width of the image) into a single average value, thereby converting the feature map into a feature vector. This vector retains the key information of the image and can be used for subsequent classification or other tasks.
[0113] Specifically, a global average pooling layer is used to perform global average pooling on the feature map processed by the convolution layer, and the feature map of each channel is compressed into a scalar, thereby converting the feature map into a feature vector of a fixed length and generating the initial feature vector of the target plate.
[0114] Step S43, normalizing the initial feature vector of the target plate to generate a target feature vector of the target plate.
[0115] It should be noted that the target feature vector is the vector obtained by normalizing the initial feature vector. Normalization is to scale each element of the vector to a specific range (usually [-1, 1] or [0, 1]) to eliminate the dimensional differences between different features.
[0116] Specifically, the initial feature vector is normalized to make its length 1, and the target feature vector of the target plate is generated. The normalization process eliminates the dimensionality effect between features by mapping the numerical range of each feature to the same interval, making it easier for the model to capture the relationship and pattern between features.
[0117] Through the above scheme, this embodiment uses the heavy parameter strategy through the PP-LCNetV2 model, combines deep convolutions of convolution kernels of different sizes, and improves the perception ability of the neural network; optimizes the deep separable convolution, which not only enhances the fitting ability of the model, but also improves the model efficiency.
[0118] Based on the above implementation scheme, in a feasible implementation manner, the step of performing feature retrieval on the target feature vector of the target plate through the vector retrieval library to obtain the plate recognition result of the target plate includes S51 to S53:
[0119] Step S51, initializing a vector index based on a standard feature vector, and constructing a vector search library according to the vector index.
[0120] It should be noted that standard feature vectors refer to feature vectors extracted from samples of known categories, which are used as references or benchmarks to compare with feature vectors of new samples.
[0121] In addition, it should be noted that vector index is used to store and quickly retrieve vectors in high-dimensional space. It organizes data by mapping vectors to points in multidimensional space, making similarity retrieval efficient. The construction and retrieval process of vector index includes converting words and sentences into vectors, constructing indexes for vectors through clustering algorithms, and retrieving through better retrieval algorithms.
[0122] Specifically, in one embodiment of the present application, the Faiss database is selected as the vector retrieval library, and the inner product similarity index method IndexFlatIP is selected as the vector retrieval method. The feature vectors of each standard plate photo are stored in the memory; these feature vectors are pre-extracted and stored, and represent the standard features of each plate category. Using Faiss's inner product similarity index method, the vector index is initialized based on the stored standard feature vector, which organizes the standard feature vector into a data structure that Faiss can efficiently retrieve. On the basis of the initialized vector index, a vector retrieval library is constructed, which contains the standard feature vectors of all stored plate categories and their corresponding labels, and provides functions such as adding, deleting, checking, and modifying. In the subsequent use process, the vector index in the library can be added and deleted.
[0123] Step S52: Calculate the similarity between the target feature vector and the standard feature vector to obtain a similarity ranking table of the target feature vector.
[0124] It should be noted that the similarity ranking table is a ranking list obtained by calculating the similarity between the query vector and the vectors in the database, wherein each entry contains a vector and its similarity score with the query vector, arranged in ascending or descending order.
[0125] Specifically, using the retrieval function of Faiss, the target feature vector is calculated for similarity with the standard feature vector in the vector retrieval library, and the inner product of the target feature vector and the standard feature vector is calculated as the evaluation index of vector similarity. Based on the similarity calculation results, a similarity ranking table is generated. This table is sorted from high to low or from low to high according to the similarity, and lists the labels and similarity values of the standard feature vectors that are most similar to the target feature vector.
[0126] Step S53: determining the plate recognition result of the target plate from the vector search library based on the similarity ranking table and a preset similarity threshold.
[0127] It should be noted that the similarity threshold is the minimum value set when performing similarity comparison. When the similarity between two feature vectors is greater than or equal to this threshold, they are considered to have sufficient similarity; when the similarity between two feature vectors is less than the similarity threshold, they are considered to have insufficient similarity. The similarity threshold is set according to specific actual conditions.
[0128] Specifically, a similarity threshold is set according to actual needs. In the similarity sorting table, standard feature vector labels with similarity values higher than the preset threshold are screened out. If there is only one label screened out, the plate category corresponding to the label is used as the recognition result of the target plate. If there are multiple labels screened out, the label with the highest similarity can be selected as the recognition result according to actual needs. If the similarity values of all labels in the similarity sorting table are lower than the preset threshold, it may mean that the target plate is not in the preset plate category. At this time, the recognition result of "unmatched" or "unknown category" can be returned.
[0129] Through the above scheme, this embodiment constructs a vector retrieval library and calculates the similarity between the target feature vector and the standard feature vector, which can greatly improve the retrieval efficiency and the recognition accuracy; at the same time, as new standard feature vectors are added to the retrieval library, the recognition ability of the system can be continuously improved, and it has good scalability.
[0130] Based on the above implementation scheme, in a feasible implementation manner, before the step of performing instance segmentation on the initial plate image by using an instance segmentation model to obtain a mask image of the target plate in the initial plate image, the step further includes S61 to S64:
[0131] Step S61, obtaining several groups of first training images.
[0132] It should be noted that the first training image refers to an image used for iterative training of instance segmentation for an instance segmentation model.
[0133] Specifically, the first training images are obtained from various sources, such as the Internet, photography, etc., to ensure that the target plate is clearly visible in the image. The images should be diverse, including different lighting conditions, angles, plate types, and backgrounds.
[0134] Step S62: segment and annotate the target plate in each of the first training images using an annotation tool to generate an annotated training image.
[0135] It should be noted that the labeled training image refers to an image obtained by labeling the target plate on the first training image.
[0136] Specifically, annotate the initial training images using annotation tools, such as Labelimg, etc., to annotate target defects in the images.
[0137] Step S63, performing preprocessing of adjustment and normalization on each of the labeled training images, and inputting the preprocessed labeled training images into the instance segmentation model for forward propagation to obtain a predicted segmentation result.
[0138] Specifically, the annotated training images are preprocessed, including image resizing and pixel normalization, to meet the input requirements of the instance segmentation model. The preprocessed images will be used as input data for model training. The preprocessed annotated training images are input into the YOLOv8-seg model for forward propagation, and the model will output the predicted segmentation results, which include the bounding box and segmentation mask of the target plate.
[0139] Step S64, based on the predicted segmentation result and the standard training image, the segmentation loss of the predicted segmentation result is calculated, and the segmentation loss is back-propagated to the instance segmentation model to update the model parameters to obtain a trained instance segmentation model.
[0140] It should be noted that standard training images refer to images in which the boundaries of the target plate have been accurately annotated by manual or other high-precision methods in the plate recognition or segmentation task. These images are used to provide an accurate reference benchmark for the instance segmentation model so that the model's predicted segmentation results can be evaluated during the training process.
[0141] In addition, it should be noted that the segmentation loss refers to the loss of the predicted segmentation result during the training of the YOLOv8-seg model, which is calculated by the loss function of the YOLOv8-seg model. The loss function of the YOLOv8-seg model includes the detection head loss function and the segmentation head loss function. Among them, the detection head loss function also includes classification loss and regression loss. The classification loss is implemented by the focal loss Focal Loss. Focal Loss solves the problem of uneven positive and negative samples by giving higher weights to complex samples and lower weights to simple samples, so that the model focuses more on complex samples. The regression loss is implemented by the intersection-over-union loss CIoU Loss. When CIoU Loss calculates the distance between the predicted box and the true box, it not only considers the overlapping area and the center point distance, but also the aspect ratio. It can not only better reflect the similarity between the predicted box and the true box, but also accelerate the convergence speed and improve the convergence effect. The segmentation head loss consists of three parts: binary cross entropy loss, similarity loss Dice Loss, and boundary loss. Binary cross entropy loss treats the segmentation prediction of each pixel as a binary classification problem and calculates the binary cross entropy loss between the predicted value and the true label. Dice Loss is used to measure the similarity between the predicted segmentation mask and the true mask, and has good adaptability to the situation of imbalanced foreground and background. Boundary loss calculates the distance or difference between the predicted boundary and the true boundary to constrain the model, allowing the model to pay more attention to the boundary information of the target.
[0142] The calculation expression of classification loss Focal Loss is:
[0143] FL(p t )=-α t (1-p t ) γ log(p t )
[0144] Among them, FL(p t ) is the classification loss; p t is the probability predicted by the instance segmentation model; α t is a hyperparameter that balances the weights of positive and negative samples and is used to adjust the loss contribution of easy-to-classify and difficult-to-classify samples; γ is also a hyperparameter that adjusts the weights of difficult-to-classify samples and is used to further solve the problem of sample imbalance.
[0145] The calculation expression of regression loss CIoU Loss is:
[0146]
[0147] Among them, CIoU is the regression loss; IoU is the intersection over union ratio; ρ is the Euclidean distance, which represents the center point b of the predicted box and the center point b of the real box.gt The distance between them; b is the coordinate of the center point of the prediction box; b gt is the coordinate of the center point of the real box; c is the diagonal length of the minimum enclosing area containing the predicted box and the real box; α and v are parameters related to the aspect ratio, which are used to adjust the impact of the aspect ratio on the loss.
[0148] The calculation expression of binary cross entropy loss is:
[0149] BCE(p,y)=-ylog(p)-(1-y)log(1-p)
[0150] Among them, BCE(p,y) is the binary cross entropy loss; p is the probability predicted by the model that the pixel belongs to the target; y is the true label (0 or 1).
[0151] The calculation expression of Dice Loss is:
[0152]
[0153] Among them, Dice is the Dice loss value; X is the pixel set of the predicted mask; Y is the pixel set of the real mask.
[0154] The calculation expression of boundary loss is:
[0155] BoundaryLoss = H(X, Y)
[0156] Among them, BoundaryLoss is the boundary loss; X is the pixel set of the predicted mask; Y is the pixel set of the real mask; H is the Hausdorff distance between X and Y, which measures the maximum value of the minimum distance from any point in the two sets to the other set.
[0157] In summary, the calculation formula of the YOLOv8-seg model loss function is:
[0158] loss = w 1 FL+w 2 CIoU+w 3 BCE+w 4 Dice+w 5 BoundaryLoss
[0159] Among them, loss is the segmentation loss of the predicted segmentation result; w i is the weight coefficient; FL is the classification loss; CIoU is the regression loss; BCE is the binary cross entropy loss; Dice is the Dice loss value; BoundaryLoss is the boundary loss.
[0160] Specifically, the segmentation loss of the predicted segmentation result is calculated according to the loss function of the segmentation loss, and the gradient of the model parameters in the instance segmentation model is calculated using the back propagation algorithm. The parameters of the model are updated according to the gradient, and the training is iterated to update the parameters until the model converges or reaches the predetermined number of training rounds.
[0161] Through the above scheme, this embodiment trains the instance segmentation model, calculates the segmentation loss of the predicted segmentation result, and updates the parameters of the instance segmentation model, thereby improving the segmentation accuracy of the model for the target plate, and enabling the model to continue to learn and improve to adapt to changing application needs.
[0162] Based on the above implementation scheme, in a feasible implementation manner, the step of extracting features from the edge pixel image of the target plate by using a feature extraction model to generate a target feature vector of the target plate further includes S71 to S73:
[0163] Step S71: Acquire several groups of second training images and reference labels of the second training images.
[0164] It should be noted that the second training image refers to an image used for iterative training of feature extraction for the feature extraction model.
[0165] In addition, it should be noted that the reference label of the second training image refers to the real category label corresponding to the second training image, which is the annotation data used for supervised learning, used to guide the model training process, and help the model learn how to correctly classify the input image.
[0166] Specifically, the first training images are obtained from various sources, such as the Internet, photography, etc., to ensure that the target plate is clearly visible in the image. The images should be diverse, including different lighting conditions, angles, plate types, and backgrounds. A reference label is assigned to each set of second training images, which represents the feature category of the target plate in the image.
[0167] Step S72, performing adjustment and normalization preprocessing on each of the second training images, inputting the preprocessed second training images into a feature extraction model for forward propagation, and obtaining the classification label probability of each of the second training images.
[0168] It should be noted that the classification label probability of the second training image refers to the probability distribution of the target plate in the image belonging to each predefined plate category output by the feature extraction model after the preprocessed second training image is input. It is a probability vector, in which each element represents the confidence that the target plate image belongs to a specific plate category, and the value is between 0 and 1.
[0169] Specifically, the second training image is preprocessed by resizing, cropping, rotating, etc. to ensure the consistency of the input image. The preprocessed image is normalized to scale the pixel values to a suitable range, usually [0, 1] or [-1, 1]. The preprocessed second training image is input into the feature extraction model for forward propagation. The model extracts features from the input image through its network structure, such as optimized depthwise separable convolution. The model outputs a fixed-length feature vector, which is further passed through a classification layer (such as a fully connected layer + softmax) to obtain the classification label probability.
[0170] Step S73, based on the classification label probability and the reference label, calculate the prediction loss of the classification label probability of the second training image, back-propagate the prediction loss to the feature extraction model to update the model parameters, and obtain a trained feature extraction model.
[0171] It should be noted that the prediction loss is the loss of feature extraction during the training process of the PP-LCNetV2 model, which is calculated through the loss function of the PP-LCNetV2 model. The loss function of the PP-LCNetV2 model consists of two parts: cross entropy loss (CrossEntropyLoss) and triplet loss (TripletLoss). Cross entropy loss can speed up the convergence of the network. Triplet loss can learn richer semantic information by comparing the similarities between samples, which helps to improve the generalization ability of the model. Therefore, using cross entropy loss and triplet loss to jointly optimize the plate feature extraction model can not only speed up the convergence of the model, but also improve the generalization ability of the model.
[0172] The calculation formula for cross entropy loss is:
[0173]
[0174] Among them, CrossEntropyLoss is the cross entropy loss; N is the total number of samples of the second training image; C is the number of categories; x i represents the i-th sample input, f(x i ) represents the model output probability distribution corresponding to sample i; y i,j Indicates the true label (0 or 1) of the jth category corresponding to sample i; f(x i ) j Represents the predicted probability of the jth category corresponding to sample i.
[0175] The calculation formula for the triplet loss is:
[0176]
[0177] Where TripletLoss is the triplet loss; N is the total number of samples of the second training image; x i represents the i-th sample input; emb() represents the embedding function; x i represents the i-th sample; Indicates the distance x i The most recent positive example; Indicates the distance x i The nearest negative sample; α is a hyperparameter representing the boundary interval.
[0178] The calculation formula for prediction loss is:
[0179] Loss=CrossEntropyLoss+TripletLoss
[0180] Among them, Loss is the prediction loss; CrossEntropyLoss is the cross entropy loss; TripletLoss is the triplet loss.
[0181] Specifically, the prediction loss of the classification label probability is calculated according to the loss function of the prediction loss, and the gradient of the prediction loss to the model parameters in the feature extraction model is calculated using the back propagation algorithm. The parameters of the model are updated according to the gradient, and the training is iterated to update the parameters until the model converges or reaches the predetermined number of training rounds.
[0182] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the plate recognition method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0183] This application also provides a plate identification device, please refer to Figure 5 , the plate recognition device comprises:
[0184] An instance segmentation module 501 is used to obtain an initial plate image, perform instance segmentation on the initial plate image using an instance segmentation model, and obtain a mask image of a target plate in the initial plate image;
[0185] A pixel extraction module 502 is used to obtain edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate;
[0186] A feature extraction module 503 is used to extract features from the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate;
[0187] The feature retrieval module 504 is used to perform feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
[0188] The plate recognition device provided by the present application adopts the plate recognition method in the above embodiment, which can solve the technical problem of poor versatility of the plate recognition technology. Compared with the prior art, the beneficial effects of the plate recognition device provided by the present application are the same as the beneficial effects of the plate recognition method provided by the above embodiment, and the other technical features of the plate recognition device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0189] The present application provides a plate recognition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the plate recognition method in the above-mentioned embodiment 1.
[0190] Reference below Figure 6 , which shows a schematic diagram of the structure of a plate recognition device suitable for implementing the embodiment of the present application. The plate recognition device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated plate identification device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.
[0191] like Figure 6As shown, the plate recognition device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 to a random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the plate recognition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the plate identification device to communicate with other devices wirelessly or by wire to exchange data. Although the plate identification device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0192] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0193] The plate recognition device provided by the present application adopts the plate recognition method in the above embodiment, which can solve the technical problem that the plate recognition technology has poor versatility. Compared with the prior art, the beneficial effects of the plate recognition device provided by the present application are the same as the beneficial effects of the plate recognition method provided by the above embodiment, and the other technical features in the plate recognition device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0194] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0195] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0196] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the dinner plate recognition method in the above-mentioned embodiment.
[0197] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0198] The computer-readable storage medium may be included in the plate recognition device; or may exist independently without being assembled into the plate recognition device.
[0199] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the plate recognition device, the plate recognition device: obtains an initial plate image, performs instance segmentation on the initial plate image through an instance segmentation model, and obtains a mask image of a target plate in the initial plate image; obtains edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate; extracts features from the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate; and performs feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
[0200] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0201] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0202] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0203] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned plate recognition method, and can solve the technical problem that the plate recognition technology has poor versatility. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the plate recognition method provided in the above-mentioned embodiment, and will not be elaborated here.
[0204] The present application also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the above-mentioned plate identification method are implemented.
[0205] The computer program product provided by the present application can solve the technical problem that the plate recognition technology has poor versatility. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the plate recognition method provided by the above embodiment, which will not be repeated here.
[0206] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for identifying a plate, characterized in that: The plate recognition method comprises: Acquire an initial dinner plate image, and perform instance segmentation on the initial dinner plate image using an instance segmentation model to obtain a mask image of a target dinner plate in the initial dinner plate image; Acquiring edge pixels of the mask image of the target dinner plate to generate an edge pixel image of the target dinner plate; Performing feature extraction on the edge pixel image of the target plate by using a feature extraction model to generate a target feature vector of the target plate; A target feature vector of the target plate is subjected to feature retrieval through a vector retrieval library to obtain a plate recognition result of the target plate.
2. The plate recognition method according to claim 1, characterized in that: The instance segmentation model includes a feature extraction layer, a feature fusion layer, and a detection segmentation layer. The step of performing instance segmentation on the initial plate image through the instance segmentation model to obtain a mask image of the target plate in the initial plate image includes: Preprocessing the initial dinner plate image; Performing multi-scale feature extraction on the preprocessed initial plate image through the feature extraction layer to obtain a first plate feature image of the initial plate image; Inputting the first dinner plate feature image into the feature fusion layer for global feature fusion to obtain a second dinner plate feature image of the initial dinner plate image; The second dinner plate feature image is input into the detection segmentation layer for detection and classification, so as to obtain a mask image of the target dinner plate in the initial dinner plate image.
3. The plate recognition method according to claim 1, characterized in that: The step of acquiring edge pixels of the mask image of the target plate to generate the edge pixel image of the target plate comprises: Preprocessing the mask image of the target dinner plate; Performing edge pixel detection on the preprocessed mask image using an edge detection algorithm to obtain an edge pixel detection result of the target dinner plate; According to the edge pixel detection result, the area where the edge pixel detection result is located is cropped to generate an edge pixel image of the target plate.
4. The method for identifying a plate according to claim 1, wherein: The step of extracting features from the edge pixel image of the target plate by using a feature extraction model to generate a target feature vector of the target plate comprises: Performing feature extraction on the edge pixel image by using a feature extraction model to generate an initial feature image of the target plate; Performing global average pooling processing on the initial feature image to generate an initial feature vector of the target plate; The initial feature vector of the target plate is normalized to generate a target feature vector of the target plate.
5. The method for identifying a plate according to claim 1, wherein: The step of performing feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate comprises: Initialize a vector index based on a standard feature vector, and construct a vector search library according to the vector index; Calculating the similarity between the target feature vector and the standard feature vector to obtain a similarity ranking table of the target feature vector; Based on the similarity ranking table and a preset similarity threshold, a plate recognition result of the target plate is determined from the vector retrieval library.
6. The method for identifying a plate according to claim 1, wherein: Before the step of performing instance segmentation on the initial plate image by using an instance segmentation model to obtain a mask image of the target plate in the initial plate image, the step further includes: Acquire several groups of first training images; Segmenting and annotating the target plate in each of the first training images using an annotation tool to generate an annotated training image; Performing preprocessing of adjustment and normalization on each of the labeled training images, and inputting the preprocessed labeled training images into the instance segmentation model for forward propagation to obtain a predicted segmentation result; Based on the predicted segmentation result and the standard training image, the segmentation loss of the predicted segmentation result is calculated, and the segmentation loss is back-propagated to the instance segmentation model to update the model parameters to obtain a trained instance segmentation model.
7. The method for identifying a plate according to claim 1, wherein: Before the step of extracting features from the edge pixel image of the target plate using a feature extraction model to generate a target feature vector of the target plate, the step further includes: Acquire several groups of second training images and reference labels of the second training images; Performing preprocessing of adjustment and normalization on each of the second training images, inputting the preprocessed second training images into a feature extraction model for forward propagation, and obtaining a classification label probability of each of the second training images; Based on the classification label probability and the reference label, the prediction loss of the classification label probability of the second training image is calculated, and the prediction loss is back-propagated to the feature extraction model to update the model parameters to obtain a trained feature extraction model.
8. A plate recognition device, characterized in that: The plate recognition device comprises: An instance segmentation module is used to obtain an initial plate image, perform instance segmentation on the initial plate image through an instance segmentation model, and obtain a mask image of a target plate in the initial plate image; A pixel extraction module, used for acquiring edge pixels of the mask image of the target plate to generate an edge pixel image of the target plate; A feature extraction module, used for extracting features from the edge pixel image of the target plate through a feature extraction model to generate a target feature vector of the target plate; The feature retrieval module is used to perform feature retrieval on the target feature vector of the target plate through a vector retrieval library to obtain a plate recognition result of the target plate.
9. A plate recognition device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the plate recognition method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the plate identification method according to any one of claims 1 to 7 are implemented.
11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the plate recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Meal delivery detection method and device based on computer vision and storage medium
CN120747559A