Image recognition-based wine bottle type detection system and detection method

By using an image recognition-based wine bottle type detection system, which utilizes the EfficientNet model and channel attention mechanism, combined with a green background layer and a red scale, the system solves the problems of low efficiency, insufficient accuracy, and inconsistent standards in wine bottle type detection, achieving efficient and accurate wine bottle type identification and detection.

CN120708202BActive Publication Date: 2026-01-27SICHUAN XINGWEN DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510790695.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2026-01-27
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency, insufficient accuracy, and inconsistent standards in detecting different types of wine bottles, resulting in packaging defects not being detected in a timely manner and affecting the appearance and quality of the product.

Method used

A bottle type detection system based on image recognition is adopted. It uses the EfficientNet model combined with channel attention mechanism and support vector machine for efficient and accurate recognition. An integrated noise generation device is used to improve the robustness of the model. The bottle image and spatial reference are seamlessly integrated by combining a fixed focal length lens, a green background layer and a red scale.

Benefits of technology

It achieves high-speed and high-accuracy identification of wine bottle types, reduces false detections and missed detections, has flexible design and rapid iteration capabilities, reduces equipment upgrade costs, and improves detection efficiency, accuracy and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708202B_ABST
    Figure CN120708202B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image recognition, and provides a wine bottle type detection system and method based on image recognition. The wine bottle type detection system based on image recognition disclosed by the application comprises an image collecting device used for collecting an overall picture of a bottle to be detected, an image preprocessing device used for preprocessing the overall picture to generate a uniform original picture, and an image recognition device internally provided with an EfficientNet model, wherein the original picture is input into the EfficientNet model to generate a bottle type. In particular, a noise generation device integrated in the system introduces error samples in the model training stage, thereby significantly improving the robustness and anti-interference ability of the model in a complex real scene. Therefore, the system can automatically identify wine bottle types with different shapes, similar designs and even slight differences at a high speed and high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and more specifically, to a wine bottle type detection system and method based on image recognition. Background Technology

[0002] The content in this section provides only background information related to this application and may not constitute prior art.

[0003] Glazed glass bottles are widely used in alcoholic beverage packaging due to their excellent protective properties and display effects. To create a distinctive brand image, different types of alcoholic beverages, and even different series of the same type, often use exclusive bottle designs with varying shapes, sizes, colors, or surface finishes. This results in an extremely diverse range of bottle types, while the quality requirements for packaging materials are becoming increasingly stringent.

[0004] Before wine bottles leave the factory or are filled, it is necessary not only to ensure the correct bottle shape, but also to inspect the outer packaging materials (such as the glaze layer, labels, and printed patterns) for defects to avoid scratches, bubbles, color differences, stains, and other defects affecting product quality. Currently, this inspection process mainly relies on manual visual inspection, which has the following problems:

[0005] Inefficiency: The speed of manual inspection cannot keep up with the bottle feeding rhythm of modern high-speed production lines, creating a production bottleneck;

[0006] Insufficient accuracy: Workers are easily affected by fatigue, lighting and experience, making it difficult to accurately identify minor defects, especially when dealing with different types of wine bottles, resulting in a high rate of false detection and missed detection.

[0007] Inconsistent standards: The subjectivity of human judgment makes it difficult to strictly adhere to consistent testing standards, affecting the stability of product quality.

[0008] These issues may lead to packaging defects going undetected in time, thus affecting the product's appearance quality.

[0009] Therefore, there is an urgent need for an efficient, accurate, stable, and adaptive automated wine packaging material inspection technology to solve this technical problem. Summary of the Invention

[0010] In view of this, the purpose of this application is to provide a wine bottle type detection system based on image recognition. The wine bottle type detection system based on image recognition disclosed in this application can solve the technical problems mentioned in the background art.

[0011] The objective of this application is achieved through the following technical solution:

[0012] A wine bottle type detection system based on image recognition, comprising:

[0013] An image collection device for collecting overall images of the bottle to be inspected;

[0014] An image preprocessing device preprocesses the entire image to generate original images with uniform specifications;

[0015] The image recognition device has a built-in EfficientNet model. The original image is input into the EfficientNet model to generate the bottle type.

[0016] A noise generation device is used to generate erroneous sample data to train the EfficientNet model;

[0017] The EfficientNet model includes:

[0018] An initial feature extraction network extracts features from the original image to generate initial features;

[0019] The core feature learning network extracts initial features based on a channel attention mechanism to output a semantic feature map;

[0020] The classification output network uses a support vector machine to classify the semantic feature map and generate the matching probability for each bottle type.

[0021] This application provides an image recognition-based wine bottle type detection system. It acquires bottle information through an image collection device and generates standardized input data via an image preprocessing device, effectively overcoming the bottlenecks in efficiency, accuracy, and consistency of manual packaging material detection. Its core lies in utilizing a built-in EfficientNet model for efficient and accurate identification: the model first focuses on key visual information through an initial feature extraction layer to reduce interference; then, a core feature learning layer introduces a channel attention mechanism to enhance the in-depth mining of unique bottle details (such as shape outline, surface texture, color, or manufacturing process features), generating highly discriminative semantic feature maps; finally, the classification output layer uses a support vector machine (SVM) to robustly classify the feature maps, outputting the matching probability of each bottle type. Furthermore, an integrated noise generation device introduces error samples during model training, significantly improving the model's robustness and anti-interference ability in complex real-world scenarios. Therefore, this system can automatically identify wine bottle types with varying shapes, similar designs, and even subtle differences with high speed and accuracy, thus accurately completing packaging material detection. Thus, this equipment features a flexible "one machine, multiple inspections" design, enabling intelligent inspection of multiple types of packaging materials of the same kind on a single device, offering high flexibility; well-trained algorithms and unified inspection standards reduce false positives and false negatives; for new product categories, iterative expansion is rapid, reducing overall equipment upgrade costs in the later stages.

[0022] In the process of wine bottle identification, bottle size is a crucial basis for distinguishing types. However, obtaining accurate size information is difficult in existing technologies. Manual measurement is inefficient and cannot meet the cycle time requirements of high-speed production lines; while introducing specialized ranging equipment such as lidar can improve accuracy, it significantly increases system complexity and hardware costs. To address this problem, this application provides the following technical solution:

[0023] In some possible embodiments, the image collection device includes:

[0024] A display stand for placing bottles to be tested;

[0025] A three-dimensional scale is set on the exhibition stand. The three-dimensional scale includes a horizontal axis, a vertical axis, and a height axis with fixed lengths.

[0026] A lens is used to take a full-size image with a fixed focal length;

[0027] A lens mount is used to fix the lens so that the lens can capture an overall image with a fixed focal length and position.

[0028] The booth's surface was covered with a green background layer, and the three-dimensional ruler was coated with a red coating.

[0029] This application utilizes a fixed-focal-length lens and a rigid lens bracket to ensure a constant shooting angle and magnification. Combined with a display stand featuring a built-in 3D scale (horizontal, vertical, and height axes), it achieves simultaneous capture of the bottle image and a precise spatial reference in a single shot. Leveraging the constant image proportions resulting from the fixed focal length and the relative positions of the bottle and scale within the image, the bottle's actual 3D dimensions (especially height) can be directly and non-destructively deduced from the image. Furthermore, the green background layer greatly simplifies the precise segmentation of the bottle's outline, while the red scale coating significantly improves the scale's recognition and positioning accuracy within the image. This solution eliminates the need for additional ranging sensors or time-consuming manual measurements, seamlessly, efficiently, and cost-effectively fusing the bottle's visual features and geometric dimensions during image recognition. This provides a more comprehensive and accurate recognition basis for the subsequent EfficientNet model, significantly enhancing the ability and reliability to identify complex bottle shapes.

[0030] During image acquisition, the light transmittance and high reflectivity of glass bottles are highly susceptible to fluctuations in light source power. This unstable lighting leads to complex and dynamically changing optical noise (such as highlight clipping, glare, and artifacts caused by non-uniform illumination) on the bottle's surface and interior. These noise signals not only severely interfere with the accurate extraction of the bottle's contours but also obscure or distort crucial surface textures, color distributions, and shape details, significantly reducing the accuracy and robustness of subsequent image recognition models (such as EfficientNet) for feature learning and classification. Therefore, this application provides the following technical solution:

[0031] In some possible embodiments, the image preprocessing apparatus includes:

[0032] The grayscale processing module performs grayscale processing on the entire image to generate a grayscale image.

[0033] The gamma value calculation unit assigns a specific gamma value to each pixel of the grayscale image based on a pre-set brightness standard value.

[0034] The gamma transform unit applies a gamma transform to each pixel of the grayscale image to generate the original image.

[0035] The image preprocessing apparatus proposed in this application, after eliminating color information interference through the grayscale processing module, has a core innovation in its adaptive gamma correction mechanism based on a brightness standard. The gamma value calculation unit dynamically calculates and assigns a specific gamma value to each pixel in the grayscale image based on a preset brightness standard value; the gamma transformation unit then applies this value to perform a pixel-by-pixel gamma transformation operation. The core advantage of this scheme lies in its pixel-level adaptive capability: it can intelligently adjust the grayscale response curve of the image according to local or global brightness conditions—non-linearly stretching insufficiently lit areas to enhance details, while compressing overexposed areas to suppress highlight clipping. This targeted processing effectively neutralizes various optical noises induced on the glass bottle by unstable light source power (such as non-uniform illumination and glare), significantly improving the overall clarity and contrast of the image, especially the recognizability of key feature areas (such as bottle edges, label areas, and surface textures). This directly improves the accuracy of bottle type identification.

[0036] In image detection tasks, the original image often contains a large amount of background information unrelated to the target. Directly inputting the complete image into the core feature learning network results in a massive and redundant amount of data to be processed. Especially when the target is a bottle and a 3D ruler, the complex background environment can interfere with the network's effective extraction of the positional relationship features between the two. This not only increases the consumption of computational resources but may also reduce the accuracy of feature extraction and the network's learning efficiency due to the influence of irrelevant data, making it difficult to accurately focus on the key features of the target object.

[0037] In some possible embodiments, the initial feature extraction network extracts the foreground region from the original image based on probability density, and uses the foreground region as the initial feature;

[0038] The foreground area includes the bottle to be inspected and a 3D ruler.

[0039] In the technical solution provided in this application, the detected bottle and the three-dimensional ruler are extracted as the foreground region, thereby reducing the amount of data input to the core feature learning network, reducing data redundancy, and enabling the core feature learning network to accurately extract relevant features based on the positional relationship between the bottle and the three-dimensional ruler.

[0040] This application utilizes an initial feature extraction network to accurately extract the foreground region containing the bottle to be detected and the 3D ruler as initial features based on probability density, effectively filtering redundant background information in the original image. This process significantly reduces the amount of data input to the core feature learning network, avoids interference from irrelevant data in the feature extraction process, and allows the network to concentrate its computational resources on analyzing the positional relationship between the bottle and the 3D ruler. Through targeted feature learning, the core network can more accurately capture the spatial correlation features between the two, improving the accuracy and efficiency of feature extraction. This lays a solid feature foundation for subsequent positional relationship-based detection and measurement tasks, thereby achieving more efficient and accurate target analysis and processing.

[0041] While the initial features generated by shallow convolutions retain low-level details such as the spatial location and edge contours of the targets (bottle and 3D ruler), they lack a deep understanding of the semantic relationships between the targets. Conversely, high-level features generated through deep convolutions, attention mechanisms, and dimensionality reduction, while focusing on semantic associations and abstract features, may lose crucial low-level structural information. If these two processes are handled independently, the information gap between shallow details and deep semantics leads to incomplete feature representations. This weakens the model's ability to accurately locate target positional relationships and affects the richness of semantic features, making it difficult to meet the demands of high-precision detection tasks for multi-dimensional feature fusion.

[0042] In some possible embodiments, the core feature learning network includes:

[0043] The initial convolutional layer performs dimensionality-increasing convolutions on the initial features to generate the first convolutional features;

[0044] A deep convolutional layer is used to perform deep convolution on the first convolutional feature to generate a second convolutional feature.

[0045] The SE module applies an attention mechanism to the second convolutional features to generate attention features;

[0046] The dimension-reduction convolutional layer performs dimension-reduction convolution on the attention features to generate a third convolutional feature;

[0047] The random filtering layer performs random feature filtering on the third convolutional features to generate extracted features;

[0048] The fusion layer fuses the extracted features and the initial features to generate a semantic feature map.

[0049] This solution constructs a bidirectional connection between shallow details and deep semantics through a fusion layer. It organically integrates randomly selected high-level extracted features (including target location-related semantics and attention-focusing features) with the initial features of the original foreground region (including low-level details such as target outline and size proportions), forming a semantic feature map that simultaneously carries low-level spatial positioning information and high-level semantic associations. This design not only avoids the loss of details caused by downsampling in deep networks but also achieves complementary enhancement of spatial coordinate information (such as ruler markings and bottle outlines) and semantic dependencies (such as relative distance and angle mappings) through cross-level feature interaction. This allows the fused feature map to accurately locate target details and understand their semantic associations. Simultaneously, the fusion layer combines randomly selected "sparse strong features" with complete initial features, improving model robustness while avoiding information loss, ultimately forming a feature representation that combines abstract reasoning ability and detail perception.

[0050] In traditional convolutional neural networks (CNNs) for image feature processing involving a bottle and a 3D ruler, the equal processing mechanism of each channel's features leads to the "indiscriminate propagation" of key information and irrelevant noise. In the original second convolutional features, the feature responses of key semantic channels, such as the gradient channel corresponding to the ruler's scale and the color channel corresponding to the bottle's outline, are easily weakened by noise channels such as background texture and irrelevant edges, resulting in high feature redundancy and insufficient discriminative power. This lack of adjustment for the differences in importance between channels makes it difficult for subsequent fusion layers to focus on the target core channels when integrating shallow initial features and high-level semantic features, thus leading to a decrease in the accuracy of the fused semantic feature map in representing the spatial relationship between the bottle and the ruler.

[0051] In some possible embodiments, the SE module includes:

[0052] Compressed layer: The second convolutional features are generated by global average pooling;

[0053] Activation layer: The compressed convolutional features are learned through two fully connected layers to learn the non-linear relationship between channels, and the weight coefficients of each channel are output.

[0054] Calibration layer: The weight coefficient of each channel is multiplied with the second convolutional feature channel by channel to generate attention features.

[0055] This scheme utilizes the compression-excitation-calibration mechanism of the SE module to construct a channel-level adaptive attention mechanism to optimize the quality of high-level features. First, it aggregates spatial information through global average pooling to generate channel-level descriptions. Then, it uses a fully connected layer to learn the non-linear dependencies between channels and generate importance weights. Finally, it multiplies the weights with the original features channel by channel to strengthen key target channels and suppress noisy channels, focusing the attention features on the core semantic information of the bottle and ruler (such as ruler scale gradients and bottle outline colors). The processed high-level features are input to the fusion layer in a "high signal-to-noise ratio" form, ensuring that when combined with shallow initial features (such as target edge details and size ratios), "on-demand fusion" is achieved based on channel weights. This accurately aligns the pixel-level positional information of the ruler scale channels and edge details, while highlighting their crucial role in the measurement task through weights. Ultimately, this achieves a balance between spatial positioning accuracy and semantic abstraction capability in the semantic feature map, providing more discriminative feature representations for target size measurement and positional relationship analysis, significantly improving the fusion layer's efficiency and accuracy in representing complex spatial relationships.

[0056] In bottle type detection tasks, traditional deep learning classification models suffer from a bottleneck in generalization ability in small sample scenarios: relying on large-scale labeled data for training can easily lead to redundant model parameters. When faced with the scarcity of new bottle type samples and the need for fine-grained distinction between multiple categories, they are prone to overfitting the noise of training data and failing to capture subtle differences between categories (such as subtle changes in bottle texture and size ratio). Furthermore, when end-to-end neural networks classify directly, it is difficult to efficiently mine the key role of support vectors in classification decisions, resulting in blurred classification boundaries in complex feature spaces, which seriously affects the detection accuracy in small sample scenarios.

[0057] In some possible embodiments, the classification decision function of the classification output network is f(x), and the type of bottle to be tested is determined based on the value of f(x);

[0058]

[0059] Where x represents the semantic feature map, sign() represents the sign function, and S represents the set of indices of the support vectors. Denotes the optimal Lagrange multiplier, y i Indicates training sample x i The true category label, K(x) i (x) represents the kernel function, b * This represents the bias term, and i represents the index of the training sample.

[0060] This scheme constructs an efficient classification mechanism for small samples by introducing an SVM classification decision function: Utilizing the principle of minimizing the structural risk of SVM, the classification hyperplane is optimized by maximizing the class margin, and the decision boundary can be determined by relying only on sparse support vectors, significantly reducing the dependence on large-scale data and avoiding the problem of overfitting in small samples; by using a kernel function to map the semantic feature map to a high-dimensional nonlinear space, the globally optimal classification hyperplane is solved by convex quadratic programming, effectively handling complex nonlinear relationships such as the correlation between bottle curvature and scale markings and the coupling of material features, and improving the ability to distinguish subtle feature differences.

[0061] When the EfficientNet model identifies bottle types, it is susceptible to interference from complex backgrounds and confusion caused by features of bottles with similar appearances, leading to misidentification of some images containing special noise or difficult-to-distinguish features. However, existing training mechanisms do not effectively utilize these error-prone samples, causing the model to repeatedly learn simple samples but struggling to overcome the feature discrimination bottleneck in error-prone scenes, resulting in slow improvement in the model's recognition accuracy for key error-prone samples.

[0062] In some possible embodiments, the noise generating device includes:

[0063] The sample collection module is used to collect all images that are incorrectly identified and generate a set of error-prone images.

[0064] The sample generation module generates overall training images based on a set of error-prone images;

[0065] The sample input module takes the entire training image as training data and inputs it into the image recognition device to train the EfficientNet model.

[0066] In some possible embodiments, an adversarial learning network is established by using a noise generation device as a generator and an image recognition device as a discriminator.

[0067] This scheme constructs a "targeted reinforcement of error-prone samples" training mechanism through a noise generation device: First, the sample collection module accurately captures images that the model misidentifies, forming an error-prone image set focusing on the model's weaknesses. Then, the sample generation module generates a training set containing typical noise features (such as target occlusion, lighting distortion, and texture confusion) based on this dataset. When this type of data is input into the EfficientNet model for training, it forces the model to focus on learning discrimination rules for difficult-to-distinguish features—for example, strengthening the extraction of correlation features between occluded ruler fragments and bottle outlines, and improving the robustness of color feature representation under uneven lighting, thereby significantly improving the model's ability to remember and generalize from "historical error scenarios". Attached Figure Description

[0068] Figure 1 This is a schematic diagram of an image recognition-based wine bottle type detection system.

[0069] Figure 2 This is a schematic diagram of the image recognition device.

[0070] Figure 3 This is a schematic diagram of the structure of a network for learning core features.

[0071] Figure 4 This is a structural diagram of the SE module.

[0072] Figure 5 This is a schematic diagram of the spatial attention layer. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments. The same reference numerals in the accompanying drawings represent the same components. It should be noted that the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the described embodiments of this application without creative effort are within the scope of protection of this application.

[0074] Compared to the embodiments shown in the accompanying drawings, feasible embodiments within the scope of this application may have fewer components, other components not shown in the drawings, different components, differently arranged components, or components with different connections, etc. Furthermore, two or more components in the drawings may be implemented in a single component, or a single component shown in the drawings may be implemented as multiple separate components.

[0075] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” and similar terms used in this specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not necessarily indicate a quantity limitation. Terms such as “upper” and “lower” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described object changes.

[0076] Example 1:

[0077] refer to Figure 1The first embodiment of this application discloses a bottle type detection system based on image recognition, including: an image collection device, an image preprocessing device, an image recognition device, and a noise generation device. The image preprocessing device preprocesses the entire image to generate a uniform original image; the image recognition device has a built-in EfficientNet model, and the original image is input into the EfficientNet model to generate bottle types; the noise generation device generates erroneous sample data to train the EfficientNet model.

[0078] In this application, bottle type identification relies on an image recognition device. The image collection device is a display stand used to collect overall images of the bottles to be detected. An image preprocessing device preprocesses the overall images, and a noise generation device is used to generate challenging identification images to increase the sample size.

[0079] Specifically, the image acquisition device includes: a display stand, a 3D scale, a lens, and a lens mount. The display stand is a rectangular flat structure with a green background layer on its upper surface to create a uniform, solid-color background. The 3D scale is orthogonally positioned above the display stand, with its horizontal and vertical axes fixed to the edges of the stand along its length and width, respectively. Its height axis is perpendicular to the stand surface and connects to the ends of the horizontal and vertical axes, forming a rectangular coordinate system. The horizontal, vertical, and height axes are all coated with red paint for high-contrast markings. The lens is fixed 50cm directly above the display stand via a lens mount, with its optical axis at a 45° angle to the stand surface. Calibration ensures the lens's focal length is fixed at 35mm, and the shooting angle covers the entire display stand area.

[0080] In actual operation, the bottle to be tested is placed in the center of the display stand. The lens captures an overall image including the bottle, the green background layer, and the red 3D scale at a fixed focal length and position. Since the green background layer and the red scale form a sharp color contrast, subsequent image preprocessing can quickly separate the bottle body from the background through color thresholding. The red markings on the 3D scale provide a positioning reference for the image coordinate system. Combined with the known fixed length parameters of the scale, the position calibration and size normalization of the bottle in the image can be achieved, ensuring that the image coordinates of different samples are consistent and providing standardized input for subsequent inspection processes.

[0081] Image preprocessing devices primarily preprocess the overall image to reduce the impact of lighting variations on image recognition. Bottles are generally made of glass, possessing reflective, translucent, and light-changing capabilities. Therefore, the light and shadow structures cast by the supplementary lighting on the bottle are diverse, resulting in significant noise. This application provides the following technical solution for noise reduction:

[0082] The image preprocessing unit includes: a grayscale processing module, a gamma value calculation unit, and a gamma transformation unit.

[0083] The grayscale processing module performs grayscale processing on the entire image to generate a grayscale image. A grayscale image is created by converting an RGB image into a grayscale value ranging from 0 to 255.

[0084] The gamma value calculation unit assigns a specific gamma value to each pixel of the grayscale image based on a pre-set brightness standard value.

[0085] The calculation process for a specific gamma value γ is as follows:

[0086]

[0087] Where u represents the normalized expected brightness value, u t This represents the normalized grayscale value of the t-th pixel in the grayscale image.

[0088] when u and u t When they are equal, γ equals 1, and no gamma transform is needed. When u t When the value is greater than u, it indicates that the image is too bright. γ > 1 is used to reduce the brightness, and vice versa, γ < 1 is used to increase the brightness.

[0089] The gamma transform unit applies a gamma transform to each pixel of the grayscale image to generate the original image.

[0090]

[0091] Among them, Q t γ represents the normalized grayscale value of pixel t in a grayscale image. t To represent the gamma correction coefficient of pixel t in a grayscale image, I t This represents the normalized grayscale value of a pixel (x, y) in the original image.

[0092] refer to Figure 2 The EfficientNet model consists of an initial feature extraction network, a core feature learning network, and a classification output network. The initial feature extraction network extracts preliminary features, the core feature learning network extracts implicit information from these features, and the classification output network generates the type of bottle to be detected based on this implicit information.

[0093] Specifically: the initial feature extraction network extracts features from the original image to generate initial features; the core feature learning network extracts initial features based on the channel attention mechanism to output a semantic feature map; and the classification output network classifies the semantic feature map based on the support vector machine to generate the matching probability of each bottle type.

[0094] The initial feature extraction network extracts the foreground region from the original image based on probability density, and uses the foreground region as the initial feature; the foreground region includes the bottle to be detected and the 3D ruler.

[0095] The initial feature extraction network extracts features as follows:

[0096] S1: Extract pixels in the original image that have the same color as the background to generate the initial background area;

[0097] S2: Calculate and determine the area of ​​the initial background region. If the area is greater than a preset threshold, it is considered as the background region.

[0098] S3: Use the non-background area in the center of the background area as the initial foreground area;

[0099] S4: Use the pixels in the original image that are the same as those in the red image as the ruler region, use the ruler region and the initial foreground region as the foreground region, and use the foreground region as the initial feature.

[0100] In this application, the initial feature extraction network primarily extracts the foreground region based on color. Compared to image segmentation algorithms, this method is more efficient and accurate. The bottles to be detected need to be placed on a display stand, so the bottles are enveloped by the green stand in the overall image. Therefore, based on the range of green color, it is possible to determine whether a pixel belongs to the foreground or background. Correspondingly, the scale region can be distinguished using the range of red color.

[0101] In practice, to increase efficiency, the RGB colors of the original image can be extracted and used to divide the foreground and background.

[0102] refer to Figure 3 The core feature learning network includes: a preliminary convolutional layer, a deep convolutional layer, an SE module, a dimensionality reduction convolutional layer, a random selection layer, and a fusion layer.

[0103] The initial convolutional layer performs dimensionality-increasing convolution on the initial features to generate the first convolutional features. Specifically, the initial convolutional layer uses a 1×1 convolution kernel, a stride of 1, and padding of the same value to perform dimensionality-increasing processing on the initial features. By increasing the number of feature channels, the diversity of feature representation is expanded.

[0104] The deep convolutional layer performs deep convolution on the first convolutional features to generate the second convolutional features. Specifically, the deep convolutional layer uses a 3×3 depthwise separable convolutional kernel (with 128 groups, meaning each group processes only one channel) and a stride of 1 to refine the spatial features of the first convolutional features and generate the second convolutional features. This step enhances the ability to capture spatial details such as image texture and edges while maintaining computational efficiency.

[0105] The SE module applies an attention mechanism to the second convolutional features to generate attention features. It evaluates and weights the importance of each channel feature using a channel attention mechanism to generate attention features.

[0106] refer to Figure 4 The SE module includes a compression layer, an activation layer, and a calibration layer. The compression layer performs global average pooling on the second convolutional features to generate compressed convolutional features. The activation layer uses two fully connected layers to learn the non-linear relationship between the channels of the compressed convolutional features and outputs the weight coefficients for each channel. The calibration layer multiplies the weight coefficients of each channel with the second convolutional features channel by channel to generate attention features.

[0107] The dimension reduction convolutional layer performs dimension reduction convolution on the attention features to generate a third convolutional feature. The dimension reduction convolutional layer uses a 1×1 convolution kernel and a stride of 1 to reduce the number of channels of the attention features from high to low dimension.

[0108] The random selection layer randomly selects the third convolutional features to generate extracted features; specifically, the random selection layer randomly selects the channels of the third convolutional features with a retention probability of 0.7 to generate extracted features.

[0109] The fusion layer fuses the extracted features and the initial features to generate a semantic feature map.

[0110] The fusion layer adds the extracted features to the initial features element by element to generate the final semantic feature map. This step uses residual connections to fuse the original features with the enhanced features, preserving the low-level information of the initial features while incorporating high-level semantic information, thereby improving the discriminative and generalization capabilities of the features.

[0111] Semantic features generated by the fusion layer Figure 1 Typically, after pooling, the classification layer assigns probabilities to the corresponding samples. However, this method requires a significant amount of time for training and its accuracy is low with small sample sizes. Therefore, this application provides the following technical solution:

[0112] The classification decision function of the classification output network is f(x), and the type of bottle to be tested is determined based on the value of f(x).

[0113]

[0114] Where x represents the semantic feature map, sign() represents the sign function, and S represents the set of indices of the support vectors. Denotes the optimal Lagrange multiplier, y i Indicates training sample x i The true category label, K(x) i (x) represents the kernel function, b * This represents the bias term, and i represents the index of the training sample.

[0115] In this application, the pooling process is replaced by the support vector machine. By leveraging the advantages of the support vector machine in few-sample classification, the training cost can be greatly reduced while ensuring the accuracy of prediction.

[0116] During training, the data input to the classification output network is the training samples, which include samples and labels. The samples are semantic feature maps, and the labels are the types of bottles to be detected.

[0117] The kernel function K(x) in this application i ,x j ) is the Gaussian kernel function;

[0118] K(x i ,x j )=exp(-B||x i -x j || 2 );

[0119] Where i and j represent the indices of the training samples, B represents the regularization parameter, and x i ,x j Let i and j represent the i-th sample and the j-th sample, respectively.

[0120] Example 2: Based on Example 1, Example 2 provides a noise generation device, which is used as a generator and an image recognition device is used as a discriminator to establish an adversarial learning network.

[0121] The noise generation device includes: a sample collection module, used to collect overall images that are prone to misidentification and generate a set of error-prone images.

[0122] The set of error-prone images typically includes overall images that are misidentified by AI models, as well as overall images that are incorrectly labeled by workers when annotating them. These overall images are essentially those that are prone to confusion during the recognition process.

[0123] The sample generation module generates training images based on a set of error-prone images; the sample input module inputs the training images as training data into the image recognition device to train the EfficientNet model.

[0124] The sample generation module essentially mimics these error-prone overall images, automatically generating new sample data to deceive the image recognition device, thereby increasing the recognition performance of the image recognition device through adversarial methods.

[0125] Specifically, the sample generation module includes: a large-scale multiplication convolutional network, a feature enhancement network, and a feature fusion network.

[0126] The large-scale convolutional network consists of six extraction stages, each of which includes a large-scale convolutional block and a feedforward network. The large convolutional block is the Large Kernel ConvolutionBlock, and the feedforward network is the FeedForwardNetwork.

[0127] Based on a lightweight design, large convolutional blocks balance computational efficiency and feature representation capability through multi-scale convolution and residual mechanisms. Feedforward networks focus on feature optimization in the channel dimension and enhance the modeling of dependencies between feature channels through normalization, convolution, and nonlinear activation, thereby improving the overall representation capability.

[0128] The feature enhancement network includes spatial attention units and global attention units. The feature enhancement network is used to add an attention mechanism to the first feature map output by the large-scale multiplication convolutional network.

[0129] The spatial attention unit includes a spatial attention layer and a global information attention layer:

[0130] refer to Figure 5 The spatial attention layer includes four 1×1 convolutional layers and two 3×3 dilated convolutional layers.

[0131] The first feature map Fa is input into the spatial attention layer and undergoes a 1×1 convolution to obtain the hidden feature Fx;

[0132] The first hidden feature Fx is obtained by performing two 3×3 dilated convolutions. b 1 and F b 2, then F b 1 and F b 2. Spatial attention features S are obtained by using the Softmax activation function. ij ;

[0133] The first hidden feature Fx is convolved with a 1×1 layer to obtain the second hidden feature Fs. Then, S... ij The first feature map Fa is multiplied by Fs and fed into a 1×1 convolutional layer for convolution. Then, it is multiplied again by the first feature map Fa and fed into a 1×1 convolutional layer to generate the spatial attention output F. Sa ;

[0134] F Sa =Conv(conv(∑) i,j S ij F s +F x )+Fa); Conv represents a 1×1 convolution.

[0135] The global information attention layer consists of an asymmetric encoding and decoding structure composed of 10 layers of ViTEncoderBlock and 4 layers of TransformerDecoderBlock, which is used to match the spatial attention layer.

[0136] During the encoding stage, the global information attention layer can reasonably divide the feature image, perform multi-dimensional spatial projection on the divided image, and then perform continuous processing by ViTEncoderBlock, thereby generating a more comprehensive image representation.

[0137] The feature fusion layer fuses the outputs of the spatial attention layer and the global information attention layer. The feature fusion layer includes a first fusion unit and a second fusion unit, which have identical structures. The first fusion unit is connected to the spatial attention layer, and the second fusion unit is connected to the global information attention layer. The fusion unit uses 3×3 convolution, ReLU, and BN to calculate similarity. The first and second fusion units calculate the outputs of the fused spatial attention layer and the global information attention layer based on cross-similarity.

[0138] The structures of the noise generation device and the image recognition device are clear, and the loss function during training is the multi-scale loss function LOSS.

[0139] LOSS = λ1LS1 + λ2LS2 + λ3LS3; where λ1, λ2, and λ3 represent weighting coefficients, which can be adjusted as needed.

[0140] LS1 is the adversarial loss function, which is the cross-entropy loss function where the noise generation device acts as the generator and the image recognition device acts as the discriminator.

[0141]

[0142] Where E represents the expectation value operator, x represents the input data (original image), p(x) represents the probability distribution of the input data x, D() represents the discriminator (image recognition device) function, G() represents the generator (noise generation device) function, and G(x) is the image generated based on the input data;

[0143] LS2 is a similarity loss function used to measure the similarity between the generated image and the original image.

[0144]

[0145] W represents the width of the image, H represents the height of the image, i and j represent the pixel indices in the width and height directions of the image, respectively, and G(x) i,j y represents the pixel value of the generated image at (i, j). i,jLet ||·||1 represent the pixel value of the real image at (i, j), and ||·||1 represent the norm.

[0146] LS3 is the loss of detail loss, used to measure the loss of information in the image and ground truth during training;

[0147]

[0148] G(x) i+1,j This represents the pixel value of the generated image at (i+1, j).

[0149] In the technical solution provided in this application, the multi-scale loss function LOSS includes adversarial loss, similarity loss, and loss of detail loss, which can guide the noise generation device to generate more samples that are easy to identify errors during the training process, while increasing the image recognition device's ability to identify difficult samples, thereby increasing the accuracy of the image recognition device.

[0150] Given the loss function, the specific training process will not be provided here.

[0151] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A wine bottle type detection system based on image recognition, characterized in that, include: An image collection device for collecting overall images of the bottle to be inspected; An image preprocessing device preprocesses the entire image to generate original images with uniform specifications; The image recognition device has a built-in EfficientNet model. The original image is input into the EfficientNet model to generate the bottle type. A noise generation device is used to generate erroneous sample data to train the EfficientNet model; The EfficientNet model includes: An initial feature extraction network extracts features from the original image to generate initial features; The core feature learning network extracts initial features based on a channel attention mechanism to output a semantic feature map; The classification output network uses a support vector machine to classify the semantic feature map and generate the matching probability for each bottle type. The classification decision function of the classification output network is f(x), and the type of bottle to be tested is determined based on the value of f(x). ; Where x represents the semantic feature map, sign() represents the sign function, and S represents the set of indices of the support vectors. Denotes the optimal Lagrange multiplier, y i Indicates training sample x i The true category label, Represents the kernel function. This represents the bias term, where i represents the index of the training sample; the pooling process is replaced by a support vector machine.

2. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, The image collection device includes: A display stand for placing bottles to be tested; A three-dimensional scale is set on the exhibition stand. The three-dimensional scale includes a horizontal axis, a vertical axis, and a height axis with fixed lengths. A lens is used to take a full-size image with a fixed focal length; A lens mount is used to fix the lens so that the lens can capture an overall image with a fixed focal length and position. The booth's surface was covered with a green background layer, and the three-dimensional ruler was coated with a red coating.

3. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, The image preprocessing apparatus includes: The grayscale processing module performs grayscale processing on the entire image to generate a grayscale image. The gamma value calculation unit sets a specific gamma value for each pixel in the grayscale image based on a pre-set brightness standard value. The gamma transform unit applies a gamma transform to each pixel of the grayscale image to generate the original image.

4. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, The initial feature extraction network extracts the foreground region from the original image based on probability density, and uses the foreground region as the initial feature. The foreground area includes the bottle to be inspected and a 3D ruler.

5. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, The core feature learning network includes: The initial convolutional layer performs dimensionality-increasing convolutions on the initial features to generate the first convolutional features; A deep convolutional layer is used to perform deep convolution on the first convolutional feature to generate a second convolutional feature. The SE module applies an attention mechanism to the second convolutional features to generate attention features; The dimension-reduction convolutional layer performs dimension-reduction convolution on the attention features to generate a third convolutional feature; The random filtering layer performs random feature filtering on the third convolutional features to generate extracted features; The fusion layer fuses the extracted features and the initial features to generate a semantic feature map.

6. The image recognition-based wine bottle type detection system according to claim 5, characterized in that, The SE module includes: Compressed layer: The second convolutional features are generated by global average pooling; Activation layer: The compressed convolutional features are learned through two fully connected layers to learn the non-linear relationship between channels, and the weight coefficients of each channel are output. Calibration layer: The weight coefficient of each channel is multiplied with the second convolutional feature channel by channel to generate attention features.

7. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, The noise generating device includes: The sample collection module is used to collect all images that are incorrectly identified and generate a set of error-prone images. The sample generation module generates overall training images based on a set of error-prone images; The sample input module takes the entire training image as training data and inputs it into the image recognition device to train the EfficientNet model.

8. The image recognition-based wine bottle type detection system according to claim 1, characterized in that, An adversarial learning network is established by using a noise generation device as the generator and an image recognition device as the discriminator.

9. A method for detecting the type of wine bottle based on image recognition, characterized in that, The bottle type is detected using the image recognition-based bottle type detection system according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Road surface type fusion perception method based on machine vision and deep learning

    CN118072268A

  • Systems and methods for classifying blood cells

    WO2024108104A1