An intelligent control method for foreign matter detection and quality grading of fresh stewed bird's nest production line
Patent Information
- Application Number
- CN202610954262.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]为了解决现有技术中的上述问题,即鲜炖燕窝生产线中异物检测与品质分级依赖人工、检测维度单一且异物检测和品质分级相互独立的问题,本发明提供了一种鲜炖燕窝生产线异物检测与品质分级的智能控制方法,包括:
相比于现有技术,本发明将基于可见光与近红外光谱的多模态图像融合检测方法与异物分级剔除、品质多维度量化分级相结合,构建了从图像采集、异物识别、剔除控制到品质分级的一体化智能控制方法,显著提升了鲜炖燕窝生产线异物检测的准确性和品质分级的客观一致性。
Smart Images

Figure CN122806754A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of food processing and testing technology, specifically relating to an intelligent control method for foreign object detection and quality grading in a fresh stewed bird's nest production line. Background Technology
[0002] As a high-end ready-to-eat health supplement, the control of foreign matter and quality grading during the production process of freshly stewed bird's nest are crucial to ensuring product safety and quality. Currently, freshly stewed bird's nest production lines generally use manual visual inspection for foreign matter screening and quality assessment, which has the following main shortcomings:
[0003] On the one hand, manual inspection relies on the operator's experience and subjective judgment, making it difficult to standardize inspection criteria. Furthermore, prolonged work can easily lead to visual fatigue, resulting in a high rate of missed detection of minute foreign objects. Bird's nest bottles are cylindrical glass bottles, and the transparency of the gel-like material inside is limited. Foreign objects such as hair, fine down, and black plant spots are extremely difficult to identify due to the reflection of the bottle wall and the obstruction of the material.
[0004] On the other hand, quality grading relies heavily on human experience and lacks objective and quantitative evaluation methods for key quality indicators such as the length, transparency, solid content distribution, and uniformity of gel-state moisture in bird's nest strands. This results in poor grading consistency and makes it difficult to meet the quality stability requirements of large-scale production.
[0005] Furthermore, most existing machine vision-based foreign object detection solutions for food use a single visible light imaging method, which is ineffective in detecting transparent foreign objects, suspended foreign objects inside the bottle, and foreign objects with colors similar to the contents of the bottle. A few solutions that introduce near-infrared spectroscopy are only used for detecting a single quality indicator. There is no integrated technical solution that combines visible light and near-infrared spectral images across modes to serve both foreign object detection and quality grading. Summary of the Invention
[0006] To address the aforementioned problems in existing technologies, namely, the reliance on manual labor for foreign object detection and quality grading in fresh-stewed bird's nest production lines, the limited detection dimensions, and the independence of foreign object detection and quality grading, this invention provides an intelligent control method for foreign object detection and quality grading in fresh-stewed bird's nest production lines, comprising: Acquire visible light and near-infrared spectral images of the bird's nest bottle to be tested; The visible light image and the near-infrared spectral image are registered to form a multi-channel image; The multi-channel image is input into a pre-constructed foreign object detection model, which is configured with a cross-modal attention fusion layer. The cross-modal attention fusion layer is used to calculate the correlation between the texture features extracted from the visible light image channel and the spectral features extracted from the near-infrared spectral image channel to obtain cross-modal attention weights. Based on the cross-modal attention weights, the features of the two modalities are weighted and fused, and the foreign object category, location coordinates and confidence level are predicted and output based on the fused features. Foreign object detection results with a confidence level higher than a preset threshold are considered valid detections. When a valid detection includes a foreign object belonging to a preset high-risk foreign object category, or when the size of a non-high-risk foreign object exceeds a preset size threshold based on the location coordinates, a rejection control signal is generated, and the corresponding bottle is rejected based on the rejection control signal. The multi-channel image of the unremoved bottle is input into the pre-constructed quality grading model. The quality grading model is used to extract the morphological and color distribution features of bird's nest strands from the visible light image and the uniformity of gel-state moisture distribution from the near-infrared spectral image. The extracted features are then adaptively weighted and fused to output the quality grade of the bird's nest bottle. The sorting control signal is determined based on the quality grade.
[0007] In some preferred embodiments, after acquiring the visible light image and near-infrared spectral image of the bird's nest bottle to be tested, the method further includes: Acquire polarized light images and backlight images of the bottle; The polarized light image is used as an input channel to introduce the foreign object detection model in order to suppress the reflection interference on the surface of the glass bottle; The backlight image is used as an input channel to introduce the foreign object detection model to extract the outline of the object inside the bottle and the shadow of the foreign object, thereby enhancing the detection capability of transparent and semi-transparent foreign objects.
[0008] In some preferred embodiments, generating a rejection control signal and rejecting the corresponding bottle according to the rejection control signal includes: Access the pre-stored foreign object level mapping table. When a foreign object belonging to the high-risk foreign object set is detected in the effective detection, a rejection control signal is directly generated and the corresponding bottle is rejected. When hair is detected, the hair length is calculated based on the location coordinates. If the hair length exceeds a preset length threshold, a rejection control signal is generated and the corresponding bottle is rejected. When plant-like black spots or fine hairs are detected, the number of foreign objects of this type in a single bottle is accumulated. If the accumulated number exceeds the preset allowable limit, a rejection control signal is generated and the corresponding bottle is rejected.
[0009] In some preferred embodiments, the foreign object detection model performs the following operations on its neck network: Deformable convolutional layers are used to adaptively adjust the receptive field of the slender bird's nest strand region, and a channel attention module is used to enhance the feature channel weights related to foreign objects, thereby improving the prediction accuracy of foreign object location and category.
[0010] In some preferred embodiments, the quality grading model extracts multi-scale feature maps from the multi-channel image through a feature extraction backbone network, and multiple parallel prediction heads predict the solid content score, bird's nest silk integrity score, color uniformity score, and bubble density score respectively. The above four scores are weighted and fused according to preset adaptive weights, and the fusion result is mapped to the quality grade.
[0011] In some preferred embodiments, the prediction process for the color uniformity score includes: The visible light image is converted from the RGB color space to the CIELAB color space, and the pixel value variance of the a channel and the b channel are calculated respectively. A color uniformity score is generated based on the variance. The prediction process for the integrity score of the bird's nest strands includes: Instance segmentation is performed on the visible light image to obtain instance masks for each bird's nest strand. The length distribution and breakage rate of the strands are statistically analyzed based on the instance masks, and a bird's nest strand integrity score is generated accordingly.
[0012] In some preferred embodiments, the method further includes: Determine the foreign object detection rate for the current batch and calculate the index-weighted moving average statistic based on the foreign object detection rate. When the statistic exceeds the preset control limit, the confidence threshold of the foreign object detection model is lowered, and the grade determination boundary of the quality grading model is simultaneously calibrated.
[0013] In some preferred embodiments, the method further includes: Obtain images of bottles that were confirmed as missed or misjudged upon re-inspection; The bottle image is uploaded to a cloud server, where the foreign object detection model and quality grading model are incrementally trained using the bottle image to obtain updated model parameters. The updated model parameters are then deployed to the edge computing devices on the production line to iterate on the current model.
[0014] In some preferred embodiments, acquiring the visible light image and near-infrared spectral image of the bird's nest bottle to be tested includes: When the bottle arrives at the detection station, a trigger signal is generated to start image acquisition, and the bottle is controlled to rotate at least one revolution around its own axis to obtain multi-modal images from multiple angles.
[0015] In some preferred embodiments, the registration of the visible light image and the near-infrared spectral image includes: Based on a pre-calibrated spatial transformation matrix, the near-infrared spectral image is geometrically corrected to achieve pixel-level alignment with the visible light image, forming a registered multi-channel image.
[0016] The beneficial effects of this invention are: Compared with existing technologies, this invention combines a multimodal image fusion detection method based on visible light and near-infrared spectroscopy with foreign object classification and rejection, and multi-dimensional quantitative grading of quality, to construct an integrated intelligent control method from image acquisition, foreign object identification, rejection control to quality grading, which significantly improves the accuracy of foreign object detection and the objective consistency of quality grading in the fresh stewed bird's nest production line.
[0017] On the one hand, the present invention adopts a foreign object detection scheme that combines visible light images and near-infrared spectral images in a dual-modal collaborative manner. Through the cross-modal attention fusion layer configured inside the pre-constructed foreign object detection model, it automatically learns the correlation between visible light texture features and near-infrared spectral features and performs weighted fusion. This effectively overcomes the detection blind spots of a single visible light detection scheme in scenarios such as transparent foreign objects, foreign objects with colors similar to those of the material, and bottle wall reflections, and significantly reduces the foreign object false detection rate.
[0018] On the other hand, this invention introduces a graded rejection strategy based on the hazard level of foreign objects in the foreign object detection process. Foreign objects are divided into high-risk foreign object sets and non-high-risk foreign object sets, and differentiated judgment logic is applied to each set. High-risk foreign objects are rejected with zero tolerance, while non-high-risk foreign objects are conditionally released based on size thresholds and quantity limits. This approach takes into account both the food safety red line and the tolerance requirements for reasonable losses in industrial production, ensuring product safety while avoiding unnecessary costs caused by excessive rejection.
[0019] Furthermore, in some embodiments, this invention integrates foreign object detection and quality grading into the same production line process. Bottles that are not rejected can have their multi-channel images directly reused for quality grading, eliminating the need for separate image acquisition for the quality grading stage and improving the production line's detection efficiency and throughput. The quality grading model extracts morphological, color, and moisture distribution uniformity features from visible light and near-infrared spectral images, respectively, and performs adaptive weighted fusion. This enables objective and quantitative evaluation of multiple quality dimensions, such as the integrity of bird's nest strands, solid content, color uniformity, and bubble density, replacing traditional manual experience-based grading methods. It overcomes the problems of strong subjectivity and poor consistency in manual scoring, providing a stable and reliable quality grading standard for large-scale production. Attached Figure Description
[0020] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating an intelligent control method for foreign object detection and quality grading in a fresh stewed bird's nest production line, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of the framework of an intelligent control system for foreign object detection and quality grading in a fresh stewed bird's nest production line provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a computer system that implements the methods, systems, and electronic devices of the present invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] To more clearly explain the intelligent control method for foreign object detection and quality grading in a freshly stewed bird's nest production line provided by the present invention, the following will be combined with... Figure 1 The steps in the embodiments of the present invention will be described in detail below.
[0024] The first embodiment of the present invention provides an intelligent control method for foreign object detection and quality grading in a fresh stewed bird's nest production line, including steps S10-S60, each step of which is described in detail below: Step S10: Acquire the visible light image and near-infrared spectral image of the bird's nest bottle to be tested; Step S20: Register the visible light image and the near-infrared spectral image to form a multi-channel image; Step S30: Input the multi-channel image into a pre-constructed foreign object detection model. The foreign object detection model is configured with a cross-modal attention fusion layer. The cross-modal attention fusion layer is used to calculate the correlation between the texture features extracted from the visible light image channel and the spectral features extracted from the near-infrared spectral image channel to obtain cross-modal attention weights. Based on the cross-modal attention weights, the features of the two modalities are weighted and fused, and the foreign object category, location coordinates and confidence level are predicted and output based on the fused features. Step S40: Foreign object detection results with confidence levels higher than a preset threshold are taken as valid detections. When the valid detections include foreign objects belonging to a preset high-risk foreign object category, or when the size of a non-high-risk foreign object exceeds a preset size threshold based on the location coordinates, a rejection control signal is generated, and the corresponding bottle is rejected according to the rejection control signal. Step S50: Input the multi-channel image of the unremoved bottle into the pre-constructed quality grading model. Use the quality grading model to extract the morphological and color distribution features of bird's nest strands from the visible light image and extract the uniformity features of gel-state moisture distribution from the near-infrared spectral image. Then, perform adaptive weighted fusion on the extracted features and output the quality grade of the bird's nest bottle. Step S60: Determine the sorting control signal based on the quality grade.
[0025] In this embodiment, the bird's nest bottle to be tested refers to the finished fresh stewed bird's nest bottle that has undergone filling and sealing processes and passes sequentially on the testing conveyor line. The contents of the bottle are a semi-transparent material system formed by mixing bird's nest strands and gel-like liquid. Each bottle is conveyed to the testing station at a uniform speed via a conveyor belt, where multimodal image acquisition is completed.
[0026] Visible light images refer to images of the interior of a bottle obtained through a visible light imaging device. The visible light band is from 380nm to 780nm, which is consistent with the range of human visual perception. It can clearly reflect the morphological characteristics of the bird's nest strands inside the bottle, including the thickness, bending direction, integrity of the strands, as well as the overall color characteristics and the spatial distribution of solids inside the bottle.
[0027] Near-infrared (NIR) spectral images refer to images of the interior of a bottle acquired using a near-infrared imaging device, preferably within one or more characteristic bands in the range of 780 nm to 2500 nm. It is easy to understand that because NIR spectroscopy is sensitive to the overtone and combination frequency absorption of OH bonds in water molecules, and also exhibits characteristic absorption responses to the CH and NH bonds of organic matter such as proteins and sugars in bird's nest, the resulting NIR spectral images can effectively characterize differences in moisture distribution, changes in solid concentration, and areas of abnormal composition within gel-like materials. This provides spectral dimensional information, distinct from visible light texture, for foreign matter identification and quality assessment. For example, in areas of moisture accumulation, the NIR spectral image shows obvious changes in absorption peaks, while dry or high-concentration solid areas exhibit different spectral responses. This difference allows for the quantitative characterization of the spatial distribution of materials within the bottle.
[0028] In the specific implementation of image acquisition, preferably, when the bottle arrives at the inspection station on the conveyor belt, a photoelectric sensor detects the bottle's arrival position and generates a trigger signal. In response to this trigger signal, the multimodal imaging device synchronously performs image acquisition. To further eliminate blind spots caused by material obstruction from a single viewpoint, the inspection station is equipped with a servo rotation mechanism that drives the bottle to rotate at least one revolution around its own axis during image acquisition. During rotation, the imaging device continuously acquires images at a fixed frame rate, obtaining a multi-angle, multimodal image sequence covering the 360-degree inner wall of the bottle. This ensures that all local areas inside the bottle are fully imaged, avoiding blind spots caused by local accumulation of bird's nest fibers or the curvature of the bottle wall.
[0029] In a preferred embodiment, this example acquires both visible light and near-infrared spectral images, as well as polarized light and backlight images of the bottle simultaneously.
[0030] The polarized light image is acquired by adding a polarizer and analyzer to the visible light imaging path. Utilizing the physical property of polarized light to suppress specular reflection from smooth surfaces, it filters out strong reflective patches on the glass bottle surface caused by ambient light, making the internal material area of the bottle, which was previously obscured by reflection, clearly visible in the image. This directly avoids reflective areas becoming interference sources for foreign object detection. The backlit image is acquired using a backlight illumination method, where a surface light source is placed behind the bottle, and the imaging device collects transmitted light passing through the bottle from the front.
[0031] It should be noted that foreign objects and bird's nest materials have fundamentally different light transmittance and absorption characteristics: opaque foreign objects such as metal and glass fragments appear as completely obscured black patches in the backlit image, while transparent foreign objects such as transparent plastic and cellophane, due to their different refractive indices compared to the surrounding gel material, form outline shadows of varying brightness in the backlit image; and semi-transparent foreign objects such as plant black spots appear as areas of grayscale difference. The aforementioned polarized light image and backlit image are used together as input channels in the subsequent foreign object detection model, forming a four-channel or more multimodal input along with the visible light image and near-infrared spectral image.
[0032] Those skilled in the art will understand that, due to unavoidable spatial discrepancies in the physical installation positions of visible light imaging devices and near-infrared spectral imaging devices, the images acquired by the two devices do not perfectly correspond in geometric position. Directly superimposing them will lead to spatial misalignment, manifesting as the same foreign object appearing at different pixel positions in the visible light image and the near-infrared spectral image. This will severely affect the accuracy of subsequent cross-modal feature fusion. Therefore, it is necessary to spatially align the two images beforehand.
[0033] Specifically, this step performs geometric correction on the near-infrared spectral image based on a pre-calibrated spatial transformation matrix. The calibration process for this spatial transformation matrix is completed during the system deployment phase: a precision calibration plate with a known size and checkerboard pattern is placed at the simulated bottle position in the detection station. Images of the calibration plate are acquired using both visible light imaging and near-infrared spectral imaging devices. The checkerboard corner points in the two images are extracted as corresponding feature point pairs. The homography transformation matrix or affine transformation matrix is solved using at least four sets of corresponding point pairs to obtain the precise spatial mapping relationship between the visible light image coordinate system and the near-infrared spectral image coordinate system. This matrix parameter is then stored in the detection system's configuration file.
[0034] During the online detection phase, for each set of synchronously acquired visible light and near-infrared spectral images, a geometric transformation operation is performed frame-by-frame on the near-infrared spectral image using the pre-calibrated spatial transformation matrix. This transformation includes translation correction, rotation correction, and scale correction, ensuring precise pixel-level alignment with the corresponding visible light image. After registration, the visible light and near-infrared spectral image channels are superimposed and merged along the image channel dimension to form a multi-channel image. Each spatial pixel location in this multi-channel image carries at least one visible light information channel and one near-infrared spectral information channel, providing a spatially aligned data foundation for feature extraction and fusion in the subsequent cross-modal attention mechanism of the foreign object detection model.
[0035] In this embodiment, the pre-built foreign object detection model is a deep neural network model that has been trained offline before deployment on the production line. During the training phase, a training set is constructed containing tens of thousands of multi-channel image samples of fresh stewed bird's nest bottles labeled with foreign object categories and detection box locations. The foreign objects cover multiple categories such as metal fragments, glass shards, insect remains, hair, plant black spots, and fine down, covering detection scenarios under different depths, material backgrounds, and lighting conditions inside the bottle. During the training process, the model continuously updates the network weight parameters through the backpropagation algorithm and gradient descent optimizer until it reaches the preset detection accuracy index on the validation set, at which point the training is completed and the model is solidified as the pre-built model.
[0036] In terms of model structure design, the foreign object detection model in this embodiment is internally configured with a cross-modal attention fusion layer. This cross-modal attention fusion layer is the core component connecting the visible light feature extraction branch and the near-infrared spectral feature extraction branch. Its working principle is as follows: First, multi-scale texture feature maps are extracted from the visible light image channels of the multi-channel image through a visible light feature extraction branch. This branch can employ a mature convolutional neural network backbone structure such as ResNet or CSPDarknet, extracting hierarchical feature representations from shallow edge textures to deep semantic information through layer-by-layer convolution and downsampling operations. Simultaneously, multi-scale spectral feature maps are extracted from the near-infrared spectral image channels through a near-infrared feature extraction branch. This branch can employ the same or a separate backbone structure as the visible light branch. The feature maps extracted by the two branches are aligned in the spatial dimension and each encodes discriminative information of different modalities in the channel dimension.
[0037] Then, the cross-modal attention fusion layer calculates the cross-modal correlation of the two types of feature maps. Specifically, using the visible light texture feature map as the query vector and the near-infrared spectral feature map as the key and value vectors, the similarity score matrix of the two modal features at each spatial location is calculated through matrix multiplication. After normalization by the softmax function, the cross-modal attention weight map is obtained. The magnitude of this weight map intuitively reflects the degree of feature correlation between the visible light modal information and the near-infrared spectral modal information at each spatial location—if the feature responses of the two modalities at a certain spatial location are highly consistent and both point to the presence of foreign objects, then that location receives a high attention weight; if the feature responses of the two modalities at a certain location are contradictory or both point to normal materials, then that location receives a low attention weight.
[0038] Finally, based on the obtained cross-modal attention weights, the visible light texture features and near-infrared spectral features are weighted and fused. Specifically, the near-infrared spectral feature vector is weighted and summed with the normalized attention weight matrix, and then residually concatenated with the visible light texture features to obtain the fused cross-modal feature map. Compared to simple channel concatenation (Concat operation) or element-wise feature addition (Add operation), this cross-modal attention fusion mechanism can adaptively learn the complementary relationship between the two modal features, assigning higher weights to spatial locations and feature channels that are highly relevant to foreign object detection, while suppressing background areas or normal material texture areas unrelated to foreign objects.
[0039] Specifically, taking a real-world detection scenario as an example: For transparent plastic fragments with a color highly similar to that of gel-like materials, the edge texture features of this foreign object are extremely weak in the visible light channel, almost blending into the surrounding material. However, the spectral absorption characteristics of the transparent plastic in the near-infrared spectral channel differ significantly from those of bird's nest protein. The cross-modal attention mechanism can automatically detect this spectral anomaly when calculating correlation, assigning high weight to the near-infrared features at this spatial location, thus effectively highlighting the foreign object region in the fused feature map. Conversely, for colored foreign objects (such as dark plant fibers) with similar near-infrared spectral characteristics to the surrounding material but unique visible light texture features, the visible light channel receives high weight. This adaptive weighting mechanism allows the model to flexibly utilize the most effective modal information for different foreign object types.
[0040] After obtaining the cross-modal fused feature map, the foreign object detection model continues to process it further through its neck network and detection head. Preferably, the neck network of the foreign object detection model integrates deformable convolutional layers and channel attention modules.
[0041] Deformable convolutional layers introduce learnable two-dimensional spatial offsets at the regular grid sampling positions of standard convolutional kernels, allowing the sampling shape of the kernels to adaptively deform according to the local geometry of the input feature map. When scanning the bird's nest strand region inside the bottle, the sampling points of the deformable convolutional kernels can be offset and arranged along the curve of the bird's nest strand, thereby completely extracting the length information of the bird's nest strand and perceiving the entire bird's nest strand as a continuous whole, effectively avoiding the missegmentation of continuous long bird's nest strands into foreign object fragments.
[0042] The channel attention module adopts structures such as SENet or ECA-Net to model the interdependencies between channels of the feature map in the channel dimension. By learning the importance weight of each feature channel, it enhances feature channels that are highly related to foreign object discrimination information (such as high-frequency texture channels of metallic reflection, fine edge channels of hair, and high-contrast channels of black dots), and suppresses redundant channels related to background materials and bottle wall textures, thereby further improving the prediction accuracy of foreign object location regression and category classification.
[0043] Based on the multi-scale enhanced feature map output by the neck network, the foreign object detection model's detection head predicts candidate detection boxes on each feature map grid cell, outputting the foreign object category probability distribution, location coordinates, and confidence score for each candidate box. Location coordinates are typically represented by the center point coordinates and width and height dimensions of the detection box, precisely marking the spatial extent of the foreign object in the image. The confidence score, expressed as a probability value between 0 and 1, indicates the model's degree of certainty regarding the detection result.
[0044] Considering that in actual production environments, foreign object detection models inevitably produce a certain number of low-confidence false detections, such as misidentifying normal air bubbles in materials as foreign objects or misidentifying tiny scratches on bottle walls as glass cracks, this embodiment first sets a preset confidence threshold (e.g., 0.5). Of all the detections output by the model, those with a confidence level higher than this threshold are retained as valid detections, while those lower than this threshold are ignored and discarded.
[0045] The setting of this confidence threshold needs to comprehensively consider the risk of missed detection and the cost of false rejection: setting the threshold too high may cause real foreign objects with low confidence to be missed; setting the threshold too low will lead to a large number of false rejections, affecting the production line yield. In practice, this threshold can be calibrated based on the production line's quality control standards and historical detection data statistics.
[0046] For foreign objects detected as valid, a further tiered rejection process is performed based on the object's category and size. The specific rejection process is based on a foreign object hazard level mapping table pre-stored in the control system's memory. This mapping table defines the hazard level classification and handling rules for various types of foreign objects in the form of a data structure. The high-risk foreign object category includes three types: metal, glass, and insects. Experimental data and food safety risk assessments indicate that metal fragments and glass shards may cause physical harm to consumers' mouths and esophagi, while insect remains pose a biological hazard risk; both fall under the category of zero tolerance for food safety issues. Therefore, once any foreign object belonging to this category is detected, regardless of its size, a rejection control signal is immediately generated, and the bottle is marked as a non-compliant product.
[0047] The non-high-risk foreign matter category includes hair, plant-based black spots, and fine down. Although these foreign matter do not meet product quality standards, their threat to consumer health is relatively low when present in trace amounts. Introducing a reasonable tolerance level can effectively reduce production costs and unnecessary waste. For the hair category, an image skeletonization extraction operation is performed on the image region corresponding to its detection location coordinates. The pixel connected regions of the hair region are refined into a center line of single-pixel width. The number of pixels along the center line is accumulated and converted into the actual physical length according to the image spatial resolution. If the calculated hair length exceeds a preset length threshold, a rejection control signal is generated and the product is judged as unqualified; if it does not exceed the threshold, the hair is considered an acceptable trace impurity and is released. For plant-based black spots and fine down, the number of foreign matter belonging to this category among all valid detections in the same bottle is accumulated. If the accumulated number exceeds a preset allowable upper limit, a rejection control signal is generated and the product is judged as unqualified; if it does not exceed the limit, the product is released. The threshold parameters for size and quantity mentioned above can be configured differently according to product grade requirements. For example, the upper limit for premium grade products can be stricter than that for qualified products.
[0048] Once a bottle is determined to be defective, the control system, based on the rejection control signal, drives the actuator located at the rejection position downstream of the inspection station. This actuator can be a pneumatic push rod or a fork mechanism, which pushes the defective bottle from the main conveyor line into the bypass waste collection channel when the bottle reaches the rejection position.
[0049] Based on this, for bottles that successfully pass the foreign object detection stage, the multi-channel images acquired and registered in steps S10 to S20 are further fed into the quality grading model for quality level assessment. This quality grading model was also trained offline before deployment on the production line. Its training data includes tens of thousands of multi-channel image samples where senior quality inspectors labeled the bottles according to the company's quality standards, enabling the model to learn the mapping relationship between multi-channel images and quality levels during the training phase.
[0050] The quality grading model structurally comprises a feature extraction backbone network and multiple parallel task prediction heads. The feature extraction backbone network can employ deep convolutional neural network structures such as ResNet, EfficientNet, or ConvNeXt to extract features layer by layer from the input multi-channel image, ultimately outputting multi-scale feature maps. These feature maps encode multi-level visual and spectral information about the bird's nest inside the bottle, ranging from local texture details (such as the thickness and direction of individual strands) to global distribution patterns (such as the overall settling state of the solids within the bottle).
[0051] Based on the feature maps output by the backbone network, the quality grading model sets up four task prediction heads in parallel: The first prediction head is used to predict the solid content score. Based on the feature map, this prediction head first processes the visible light channel features through a lightweight convolutional subnetwork, using semantic segmentation to identify the solid content regions inside the bottle and calculating the area ratio of the solid content regions in the bottle's cross-sectional projection. Simultaneously, it uses near-infrared spectral channel features to spatially integrate the organic matter concentration inside the bottle; regions with stronger near-infrared spectral responses represent higher solid content concentrations. The visible light area ratio and the near-infrared concentration integral value are fused and output as a solid content score through a fully connected layer. A higher score indicates a higher abundance of bird's nest fibers and a more sufficient total solid content inside the bottle.
[0052] The second prediction head is used to predict the integrity score of the bird's nest strands. This prediction head incorporates a pre-trained instance segmentation sub-network (such as the mask branch of Mask R-CNN) to segment the bird's nest strands in the visible light image one by one, generating an instance mask for each strand. Based on the pixel area and shape parameters of each instance mask, the pixel length distribution of the strands is statistically analyzed—the minimum bounding rectangle of each mask is extracted, and the longer side is used as the estimated length of the strand, and a histogram of the length distribution of all strands is calculated. Simultaneously, morphological analysis is performed on the mask of each strand to detect whether there are breaks within the strand—if a strand's mask shows obvious necking or discontinuity in a local area, it is marked as broken. Based on the above statistics, three sub-indicators are calculated: average strand length, long filament ratio, and breakage rate, which are then combined to generate a bird's nest strand integrity score. Bottles with a high long filament ratio and a low breakage rate receive higher scores, indicating that the bird's nest strands remained intact during processing and were not broken or damaged.
[0053] The third prediction head is used to predict color uniformity scores. This prediction head converts the input visible light image from the RGB color space to the CIELAB color space. The CIELAB color space, defined by the International Commission on Illumination (ICI), is a device-independent, uniform color space where the L channel represents lightness, the a channel represents chromaticity from green to red, and the b channel represents chromaticity from blue to yellow. In this space, the Euclidean distance between two color points is approximately proportional to the degree of color difference perceived by the human eye, making it suitable for the objective quantitative evaluation of color uniformity. This prediction head calculates the spatial variance of pixel values in the a and b channels of the material area inside the bottle: the variance is calculated for all pixel values in the a channel of all material areas, and the variance is calculated for all pixel values in the b channel of all material areas.
[0054] If the material inside the bottle has a highly uniform color, the variance of both channels is small; if there are obvious areas of inconsistent color in the material (such as localized yellowing or localized darkening), the variance increases significantly. Based on the variance, a color uniformity score from 0 to 100 is generated using a preset nonlinear mapping function.
[0055] The fourth prediction head is used to predict the bubble density score. This prediction head performs morphological closing operations on the visible light image to eliminate the interference of bird's nest texture on bubble detection. Then, it uses circular Hough transform or connected component analysis to detect the bubble areas inside the bottle, and counts the number of bubbles and the diameter of each bubble. A weighting coefficient for bubble size is set (large bubbles have a greater impact on visual quality than small bubbles), and a weighted bubble density value is calculated and mapped to a bubble density score. Bottles with fewer and smaller bubbles receive a higher score, reflecting excellent filling and sealing process quality.
[0056] After obtaining the four independent scores, the quality grading model adaptively weights and fuses them using a feature fusion module. The weight parameters of this fusion module can be automatically learned during model training or manually set according to the emphasis placed on each indicator in the company's quality standards. For example, if the market values the integrity of bird's nest strands more, a higher weight coefficient is assigned to the integrity score. The weighted fusion yields a comprehensive quality score ranging from 0 to 100. This comprehensive score is converted into a quality grade using a preset grade mapping rule; for example, a comprehensive score ≥ 90 corresponds to premium grade, 80 to 89 corresponds to first grade, 60 to 79 corresponds to acceptable, and below 60 corresponds to substandard.
[0057] After obtaining the quality grade of the bottles, the control system generates a corresponding sorting control signal based on the quality grade. This sorting control signal is a digitally coded signal, which is sent to a multi-channel sorting mechanism located downstream of the conveyor line via an industrial fieldbus or I / O interface. The sorting mechanism, based on the received grade code, guides bottles of different quality grades to their corresponding grade collection channels, such as premium grade channel, first-grade channel, qualified grade channel, and substandard grade channel, thus completing the quality sorting process.
[0058] In a preferred implementation, during continuous operation of the production line, the control system continuously calculates the ratio of the number of bottles with effectively detected foreign matter in the currently processed batch to the total number of bottles in the batch, which is taken as the foreign matter detection rate for that batch. For the time-series data of foreign matter detection rates across multiple consecutive batches, an exponentially weighted moving average statistic is calculated. For batch t, the formula for calculating the exponentially weighted moving average is: ; Where Yt is the foreign matter detection rate of the t-th batch, and λ is the smoothing coefficient (with a value between 0.1 and 0.3). For the tth The index-weighted moving average statistic for each batch, by assigning higher weights to recent batches and lower weights to older batches, can sensitively capture the long-term slow drift trend of the detection system's performance while suppressing fluctuation noise between random batches.
[0059] This embodiment predefines the upper control limit (UCL) and lower control limit (LCL) of the exponentially weighted moving average statistic. The control limits are determined based on historical data statistics of the system under stable operating conditions. When the statistic exceeds the upper control limit, it indicates that the recent foreign object detection rate is systematically high, possibly due to a low confidence threshold setting in the foreign object detection model, leading to a large number of false rejections. When the statistic exceeds the lower control limit, it indicates that the recent foreign object detection rate is systematically low, possibly due to a high model confidence threshold or deteriorating lighting conditions, leading to an increased risk of missed detections.
[0060] When a statistic exceeds the control limit, the model parameters are automatically adjusted: the confidence threshold of the foreign object detection model is lowered by one step (e.g., 0.05) in the direction of correcting the offset, to restore the detection sensitivity to a normal level. Simultaneously, since fluctuations in the detection rate also affect the distribution of bottle samples entering the quality grading stage, the grade determination boundary of the quality grading model is also recalibrated to ensure that the grade distribution returns to the expected proportion under normal production conditions. After adjustment, the remaining bottles in the current batch that have not yet been sorted are re-executed from step S10, ensuring that all products in the batch are reliably evaluated according to the corrected parameter standards.
[0061] As a preferred implementation method, during the daily operation of the production line, quality inspectors periodically conduct manual re-inspections of bottles that have completed foreign object detection and quality grading through random sampling. For bottles that are confirmed to have missed detection by the foreign object detection model (i.e., samples where foreign objects are detected manually but not by the model), bottles that are confirmed to have misjudged (i.e., samples where the model alarms for foreign objects but are confirmed to be free of foreign objects manually), and bottles that are confirmed to have incorrect quality grading by the model (i.e., samples where the manual rating is inconsistent with the model output rating), their corresponding multi-channel images are marked and collected as difficult samples with high training value.
[0062] The collected images of difficult samples are uploaded to a cloud server via an industrial Ethernet connection on the production line. The cloud server is equipped with a high-performance GPU computing cluster that periodically initiates incremental model training tasks. Incremental training uses the model weights currently deployed on the production line as initial weights. It groups difficult samples with a small number of balanced samples randomly drawn from the historical training set, representing a normal distribution, and updates parameters using a small learning rate (e.g., one-tenth of the original training learning rate) for a limited number of iterations (e.g., 5 to 10 epochs). Simultaneously, a knowledge distillation constraint is added to the loss function, using the original model's output on general samples as a soft label to prevent overfitting on difficult samples and thus forgetting its ability to distinguish normal samples. After training converges, the updated model weight parameters are obtained.
[0063] A model deployment cycle is set (e.g., weekly or bi-weekly), and the updated model weight parameters from the cloud are periodically distributed and deployed to edge computing devices on the production line via the industrial security network. Deployment employs a hot-loading method, meaning that after the edge computing device loads the new model weights into the backup inference engine, the inference request traffic is smoothly switched, achieving model replacement without downtime. Through this continuous update mechanism, the discrimination capabilities of the foreign object detection model and the quality grading model continuously evolve with the accumulation of production line operating time, constantly absorbing discrimination knowledge of newly emerging foreign object types and discrimination experience in quality grading under complex material conditions, achieving long-term stability and gradual improvement in model performance.
[0064] Furthermore, please refer to Figure 2 The second embodiment of the present invention proposes an intelligent control system for foreign object detection and quality grading in a fresh stewed bird's nest production line, the system comprising: Data acquisition module 210 is used to acquire visible light images and near-infrared spectral images of the bird's nest bottle to be tested; Image registration module 220 is used to register the visible light image and the near-infrared spectral image to form a multi-channel image; The foreign object detection module 230 is used to input the multi-channel image into a pre-constructed foreign object detection model. The foreign object detection model is configured with a cross-modal attention fusion layer. The cross-modal attention fusion layer is used to calculate the correlation between the texture features extracted from the visible light image channel and the spectral features extracted from the near-infrared spectral image channel to obtain cross-modal attention weights. Based on the cross-modal attention weights, the features of the two modalities are weighted and fused, and the foreign object category, location coordinates and confidence level are predicted and output based on the fused features. The foreign object detection module 240 is used to take foreign object detection results with a confidence level higher than a preset threshold as valid detections. When the valid detections include foreign objects belonging to a preset high-risk foreign object category, or when the size of a non-high-risk foreign object exceeds a preset size threshold based on the location coordinates, a rejection control signal is generated, and the corresponding bottle is rejected according to the rejection control signal. The quality grading module 250 is used to input the multi-channel image of the unremoved bottle into the pre-constructed quality grading model. The quality grading model is used to extract the morphological and color distribution features of bird's nest strands from the visible light image and the uniformity features of gel-state moisture distribution from the near-infrared spectral image. The extracted features are then adaptively weighted and fused to output the quality grade of the bird's nest bottle. The sorting control module 260 is used to determine the sorting control signal according to the quality grade.
[0065] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the system described above can be found in the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0066] It should be noted that the intelligent control method and system for foreign object detection and quality grading in a fresh-stewed bird's nest production line provided in the above embodiments are only illustrative examples of the above functional module divisions. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.
[0067] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0068] The following is for reference. Figure 3 It shows a schematic diagram of the structure of a computer system for implementing the methods and system embodiments of the present invention. Figure 3 The server shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0069] like Figure 3 As shown, the computer system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read Only Memory (ROM) 302 or programs loaded from storage section 308 into Random Access Memory (RAM) 303. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.
[0070] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0071] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the methods of the present invention. It should be noted that the computer-readable medium described above in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0072] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0074] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0075] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0076] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. An intelligent control method for foreign object detection and quality grading in a fresh-stewed bird's nest production line, characterized in that, include: Acquire visible light and near-infrared spectral images of the bird's nest bottle to be tested; The visible light image and the near-infrared spectral image are registered to form a multi-channel image; The multi-channel image is input into a pre-constructed foreign object detection model, which is configured with a cross-modal attention fusion layer. The cross-modal attention fusion layer is used to calculate the correlation between the texture features extracted from the visible light image channel and the spectral features extracted from the near-infrared spectral image channel to obtain cross-modal attention weights. Based on the cross-modal attention weights, the features of the two modalities are weighted and fused, and the foreign object category, location coordinates and confidence level are predicted and output based on the fused features. Foreign object detection results with a confidence level higher than a preset threshold are considered valid detections. When a valid detection includes a foreign object belonging to a preset high-risk foreign object category, or when the size of a non-high-risk foreign object exceeds a preset size threshold based on the location coordinates, a rejection control signal is generated, and the corresponding bottle is rejected based on the rejection control signal. The multi-channel image of the unremoved bottle is input into the pre-constructed quality grading model. The quality grading model is used to extract the morphological and color distribution features of bird's nest strands from the visible light image and the uniformity of gel-state moisture distribution from the near-infrared spectral image. The extracted features are then adaptively weighted and fused to output the quality grade of the bird's nest bottle. The sorting control signal is determined based on the quality grade.
2. The method according to claim 1, characterized in that, After acquiring the visible light image and near-infrared spectral image of the bird's nest bottle to be tested, the method further includes: Acquire polarized light images and backlight images of the bottle; The polarized light image is used as an input channel to introduce the foreign object detection model in order to suppress the reflection interference on the surface of the glass bottle; The backlight image is used as an input channel to introduce the foreign object detection model to extract the outline of the object inside the bottle and the shadow of the foreign object, thereby enhancing the detection capability of transparent and semi-transparent foreign objects.
3. The method according to claim 1, characterized in that, The process of generating a rejection control signal and rejecting the corresponding bottle according to the rejection control signal includes: Access the pre-stored foreign object level mapping table. When a foreign object belonging to the high-risk foreign object set is detected in the effective detection, a rejection control signal is directly generated and the corresponding bottle is rejected. When hair is detected, the hair length is calculated based on the location coordinates. If the hair length exceeds a preset length threshold, a rejection control signal is generated and the corresponding bottle is rejected. When plant-like black spots or fine hairs are detected, the number of foreign objects of this type in a single bottle is accumulated. If the accumulated number exceeds the preset allowable limit, a rejection control signal is generated and the corresponding bottle is rejected.
4. The method according to claim 1, characterized in that, The foreign object detection model performs the following operations on its neck network: Deformable convolutional layers are used to adaptively adjust the receptive field of the slender bird's nest strand region, and a channel attention module is used to enhance the feature channel weights related to foreign objects, thereby improving the prediction accuracy of foreign object location and category.
5. The method according to claim 1, characterized in that, The quality grading model extracts multi-scale feature maps from the multi-channel image through a feature extraction backbone network. Multiple parallel prediction heads predict the solid content score, bird's nest silk integrity score, color uniformity score, and bubble density score, respectively. The above four scores are weighted and fused according to preset adaptive weights, and the fusion result is mapped to the quality grade.
6. The method according to claim 5, characterized in that, The prediction process for the color uniformity score includes: The visible light image is converted from the RGB color space to the CIELAB color space, and the pixel value variance of the a channel and the b channel are calculated respectively. A color uniformity score is generated based on the variance. The prediction process for the integrity score of the bird's nest strands includes: Instance segmentation is performed on the visible light image to obtain instance masks for each bird's nest strand. Based on the instance masks, the length distribution and breakage rate of the strands are statistically analyzed, and a bird's nest strand integrity score is generated accordingly.
7. The method according to claim 1, characterized in that, The method further includes: Determine the foreign object detection rate for the current batch and calculate the index-weighted moving average statistic based on the foreign object detection rate. When the statistic exceeds the preset control limit, the confidence threshold of the foreign object detection model is lowered, and the grade determination boundary of the quality grading model is simultaneously calibrated.
8. The method according to claim 1, characterized in that, The method further includes: Obtain images of bottles that were confirmed as missed or misjudged upon re-inspection; The bottle image is uploaded to a cloud server, where the foreign object detection model and quality grading model are incrementally trained using the bottle image to obtain updated model parameters. The updated model parameters are then deployed to the edge computing devices on the production line to iterate on the current model.
9. The method according to claim 1, characterized in that, The acquisition of the visible light image and near-infrared spectral image of the bird's nest bottle to be tested includes: When the bottle arrives at the detection station, a trigger signal is generated to start image acquisition, and the bottle is controlled to rotate at least one revolution around its own axis to obtain multi-modal images from multiple angles.
10. The method according to claim 1, characterized in that, The registration of the visible light image and the near-infrared spectral image includes: Based on a pre-calibrated spatial transformation matrix, the near-infrared spectral image is geometrically corrected to achieve pixel-level alignment with the visible light image, forming a registered multi-channel image.