A method and system for detecting the appearance of a penicillin bottle based on multi-view reflection suppression

By employing a multi-view reflection suppression detection method, multi-view video acquisition and unified unfolded domain mapping are performed on vials to decompose reflection and defect characterization. Gating correction and cross-view causal consistency verification are then carried out, solving the problem of defect discrimination stability and consistency in the appearance inspection of transparent curved vials and achieving highly reliable detection results.

CN122335686APending Publication Date: 2026-07-03SHANDONG TAIBANG BIOLOGICAL PROD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG TAIBANG BIOLOGICAL PROD CO LTD
Filing Date
2026-03-23
Publication Date
2026-07-03

Smart Images

  • Figure CN122335686A_ABST
    Figure CN122335686A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of appearance inspection technology, and relates to a method and system for inspecting the appearance of vials based on multi-view reflection suppression. The method includes: acquiring multi-view videos of vials passing through the inspection area and establishing sub-video groups; selecting representative images; locating the effective area of ​​the vial body and establishing a unified unfolded domain to obtain an unfolded image of the vial wall; decomposing the unfolded image of the vial wall into reflection and defect characterizations, and performing gating correction based on reflection risk to obtain single-view candidate defect characterizations; performing propagation matching based on defect prototype maps to obtain single-view category responses; performing cross-view consistency verification and outputting the inspection results. The technical solution of this application, through multi-view video organization analysis and differentiation and correction of reflection interference and defect responses, achieves stable identification and consistent judgment of vial defects, improving defect discrimination stability, detection result consistency and reliability, and meeting the application needs of automated vial inspection in pharmaceutical packaging production lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of appearance inspection technology, and specifically relates to a method and system for appearance inspection of vials based on multi-view reflection suppression. Background Technology

[0002] In existing technologies, the appearance inspection of vials typically relies on manual light inspection or visual inspection methods based on industrial cameras. These methods acquire images of the bottle opening, body, or bottom and identify appearance anomalies such as cracks, scratches, chipping, dirt, and foreign objects to meet the online inspection needs of pharmaceutical packaging production lines for container appearance quality. However, existing methods for inspecting the appearance of vials have some significant shortcomings.

[0003] In practical applications, vials are mostly made of transparent or semi-transparent glass with curved walls. Due to variations in lighting angle, camera view, vial posture, and background brightness, highlights, reflective stripes, and refractive textures are easily observed in the acquired images. These imaging structures are similar to real defects such as fine cracks and shallow scratches in terms of local grayscale, edge response, and texture. At the same time, production line inspection often corresponds to continuous video or multi-view images. The intensity of defects at the same location fluctuates in different frames and from different viewpoints, resulting in weak stability of defect identification criteria and low consistency and reliability of inspection results.

[0004] Therefore, it is evident that existing technologies often suffer from problems such as weak stability in identifying true defects during the appearance inspection of transparent curved vials, and low consistency and reliability of inspection results under different viewing angles. These are the shortcomings of existing technologies.

[0005] In view of this, it is very necessary to provide a method and system for detecting the appearance of vials based on multi-view reflection suppression, so as to solve the above-mentioned defects in the prior art. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies in the appearance inspection of transparent curved vials, such as weak stability in identifying true defects and low consistency and reliability of inspection results under different viewing angles. This invention provides a method and system for inspecting the appearance of vials based on multi-view reflection suppression to solve the aforementioned technical problems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: Firstly, this application provides a method for inspecting the appearance of vials based on multi-view reflection suppression, including: Collect multi-view videos of vials passing through the detection area and establish sub-video groups corresponding to the viewpoints; Based on the stability of the bottle outline, preservation of local texture and blur suppression, representative images were selected from each sub-video group; The effective area of ​​the bottle is located in the representative image, and a unified unfolding domain is established based on the central axis of the bottle, the shoulder boundary of the bottle, and the bottom reference of the bottle to obtain the unfolded image of the bottle wall; The bottle wall unfolding image is decomposed into reflection and defect characterization, and the defect response is gated and corrected according to the reflection risk to obtain single-view candidate defect characterization. Based on the defect prototype map, multi-layer prototype propagation matching is performed on the single-view candidate defect representation to obtain the single-view category response. Based on the corresponding positions in the unified expanded domain, cross-perspective causal consistency verification is performed on the single-perspective category response, reflection risk, and candidate defect topology expression of each perspective, and the results of normal release, abnormal rejection, or re-inspection rollback are output according to the verification results.

[0008] By employing the above technical solution, multi-view videos generated when vials pass through the detection area are organized and utilized. A hierarchical discrimination mechanism is constructed based on the response differences between reflection interference and actual defects. This enables stable extraction, effective identification, and consistent judgment of the appearance information of transparent curved vials. It can reduce the interference of reflection imaging on defect identification conclusions, enhance the stability and reliability of defect discrimination results under different perspectives, and meet the requirements of strong stability in the discrimination of actual defects and high consistency and reliability of detection results under different perspectives during the appearance inspection of transparent curved vials.

[0009] First, establishing sub-video groups corresponding to different viewpoints helps retain dynamic imaging information from each viewpoint during continuous detection, facilitating targeted analysis within a unified viewpoint range. Combining bottle contour stability, local texture preservation, and blur suppression to select representative images helps obtain more effective input suitable for appearance judgment. Locating the effective bottle area in the representative images and establishing a unified unfolded domain based on the bottle's central axis, shoulder boundary, and bottom reference transforms the bottle wall surface image into a more clearly defined unfolded representation. Decomposing reflection and defect representations in the unfolded bottle wall image and implementing gating correction for defect responses based on reflection risk makes the source of abnormal responses clearer. Using a defect prototype map for multi-layer prototype propagation matching of candidate defect representations improves the targeting of category recognition. Combining the corresponding positions in the unified unfolded domain, cross-viewpoint causal consistency verification of category responses, reflection risk, and candidate defect topological representations enhances the reliability of output results for normal release, anomaly rejection, or re-inspection rollback.

[0010] Preferably, the step of selecting representative images from each sub-video group based on bottle outline stability, local texture preservation, and blur suppression includes: The bottle outline is extracted from each frame image in each sub-video group to obtain the bottle region and bottle boundary. The stability of the bottle outline is determined based on the degree of change of the bottle boundary between adjacent frames. Local texture preservation is determined based on the degree of texture gradient preservation between adjacent frames within the bottle area; The blur suppression is determined based on the sharpness of the edges within the bottle area; The stability of the bottle outline, local texture preservation, and blur suppression are jointly scored, and representative images are selected by combining time interval constraints.

[0011] By adopting the above technical solution, and using the joint evaluation of bottle contour stability, local texture preservation and blur suppression, combined with time interval constraints to achieve representative image selection, the comprehensiveness and representativeness of input image quality evaluation can be improved, and the stability of subsequent detection basic data can be enhanced.

[0012] Preferably, the steps of locating the effective area of ​​the bottle in the representative image and establishing a unified unfolded domain based on the bottle's central axis, shoulder boundary, and bottom reference to obtain the unfolded image of the bottle wall include: Perform foreground separation and contour closure filtering on the representative image to obtain the effective area of ​​the bottle; Determine the central axis of the bottle body according to the direction of the main axis of the bottle body, and establish the bottle body reference system based on the upper boundary of the bottle shoulder and the center of the bottle bottom. The bottle wall region is mapped to a unified unfolded domain along the circumferential position around the central axis of the bottle and the height position along the central axis of the bottle to obtain the unfolded image of the bottle wall. Record the corresponding positions of the bottle wall from different perspectives according to the unified unfolded domain.

[0013] By adopting the above technical solution, the bottle wall position can be standardized by using effective area positioning of the bottle body, establishing a bottle body reference system and uniformly expanding domain mapping. This can reduce the impact of differences in bottle surface imaging on detection and analysis, and improve the clarity of positional correspondence between different viewpoints and the consistency of subsequent discrimination.

[0014] Preferably, the steps of decomposing the reflection characterization and defect characterization of the unfolded bottle wall image include: Main defect features were extracted from the unfolded image of the bottle wall to obtain main features that characterize local texture and structural anomalies; Reflection feature encoding is performed on the unfolded image of the bottle wall to obtain a reflection characterization; Defect features are encoded from the unfolded image of the bottle wall to obtain defect characterization; Decorrelation constraints are applied to reflection and defect representations to suppress reflection modes from being mixed into defect representations.

[0015] By adopting the above technical solution, the abnormal information is decoupled by separating and encoding the reflection representation and the defect representation and performing decorrelation constraints. This can reduce the confounding effect of the reflection mode on defect identification and enhance the independence and discriminative specificity of defect feature expression.

[0016] Preferably, the step of gating and correcting the defect response based on reflection risk to obtain a single-view candidate defect characterization includes: Extract the proportion of highlight areas, the distribution pattern of the areas, the local sliding window energy, the local sliding window entropy, and the frequency domain concentration from the unfolded image of the bottle wall to form a reflection risk characterization. By jointly mapping the reflection risk representation and the reflection representation, the reflection gating weights are obtained. Based on the reflection gating weights, the candidate defect responses in the backbone features are corrected channel-by-channel and position-by-position to obtain a single-view candidate defect characterization.

[0017] By adopting the above technical solution, combining reflection risk characterization and reflection characterization to form gating weights and correcting candidate defect responses, the accuracy of abnormal response interpretation in high reflection areas can be improved and the credibility of single-view candidate defect characterization can be enhanced.

[0018] Preferably, the step of performing multi-layer prototype propagation matching on single-view candidate defect representations based on defect prototype maps to obtain single-view category responses includes: For each type of defect, a central prototype, a boundary prototype, and a difficult example prototype are constructed, and the graph relationships between the nodes of each prototype are established to form a defect prototype map. Align the single-view candidate defect representations and defect prototype maps with responses to obtain the initial matching responses of each prototype node. The matching response of each prototype node is propagated and updated based on the graph relationship between prototype nodes; The responses of each prototype node corresponding to the same type of defect are aggregated at the category level to obtain a single-view category response.

[0019] By adopting the above technical solution, the propagation and matching of candidate defect representations can be achieved by utilizing multiple prototypes and their graph relationships in the defect prototype map. This can enhance the adaptability of the category response to intra-class differences and boundary samples, and improve the stability and refinement of single-view defect classification results.

[0020] As a preferred embodiment, the steps for cross-perspective causal consistency verification of single-perspective category responses, reflection risks, and candidate defect topological expressions based on corresponding positions in the unified expansion domain include: The positional correspondence of candidate defect regions from each perspective is determined according to the unified expansion domain. Defect skeletons and topological signatures are extracted for each candidate defect region from each perspective. The topological signature includes skeleton length, main direction, branch features, curvature features and endpoint distribution. Calculate the category consistency, reflection consistency, and topological consistency among each perspective, and obtain the cross-perspective causal consistency verification results based on the category consistency, reflection consistency, and topological consistency.

[0021] By adopting the above technical solution, and combining the positional correspondence in the unified unfolded domain with category consistency, reflection consistency and topological consistency to achieve cross-perspective causal consistency verification, the mutual verification ability between multi-perspective discrimination results can be improved, and the reliability of anomaly judgment conclusions can be enhanced.

[0022] Preferably, the steps for outputting normal release, abnormal rejection, or re-inspection rollback results based on the verification results include: When the cross-perspective causal consistency verification result meets the anomaly establishment condition, the anomaly removal result is output; When the cross-perspective causal consistency verification result meets the normal condition, the normal release result is output. When there are perspective discrepancies, category swings, or unresolved reflection interference in the cross-perspective causal consistency verification results, the re-examination rollback result is output. The execution level corresponding to the main defect category, consistency level, and reflection interference level is given simultaneously.

[0023] By adopting the above technical solution, the normal release, abnormal rejection and re-inspection rollback are output in a graded manner based on the cross-perspective causal consistency verification results, and the execution level information is given at the same time, which can improve the pertinence of the detection result processing and the execution connection.

[0024] Secondly, this application also provides a vial appearance inspection system based on multi-view reflection suppression, comprising: The video grouping module is used to collect multi-view videos of vials passing through the detection area and create sub-video groups corresponding to the viewpoints. The image filtering module is used to filter representative images from each sub-video group based on bottle outline stability, local texture preservation, and blur suppression. The unfolding mapping module is used to locate the effective area of ​​the bottle body in the representative image, and to establish a unified unfolding domain based on the central axis of the bottle body, the shoulder boundary of the bottle and the bottom reference of the bottle to obtain the unfolded image of the bottle wall; The decoupling correction module is used to decompose the reflection characterization and defect characterization of the bottle wall unfolding image, and to perform gated correction on the defect response according to the reflection risk to obtain a single-view candidate defect characterization. The prototype matching module is used to perform multi-level prototype propagation matching on single-view candidate defect representations based on the defect prototype map to obtain single-view category responses. The consistency verification module is used to perform cross-perspective causal consistency verification on single-perspective category responses, reflection risks, and candidate defect topology expressions from various perspectives based on the corresponding positions in the unified expanded domain, and output normal release, anomaly rejection, or re-inspection rollback results according to the verification results.

[0025] Preferably, the decoupling correction module includes: The backbone extraction submodule is used to extract backbone defect features from the unfolded bottle wall image to obtain backbone features that characterize local texture and structural anomalies. The reflection coding submodule is used to encode the reflection features of the unfolded bottle wall image to obtain a reflection representation. The defect coding submodule is used to encode defect features in the unfolded bottle wall image to obtain defect characterization. The decorrelation submodule is used to perform decorrelation constraints on reflection and defect representations to suppress reflection modes from being mixed into defect representations. The gated correction submodule is used to extract the proportion of highlight areas, the distribution pattern of areas, the local sliding window energy, the local sliding window entropy, and the frequency domain concentration from the unfolded image of the bottle wall to form a reflection risk characterization. The reflection risk characterization and the reflection characterization are jointly mapped to obtain the reflection gating weight. Based on the reflection gating weight, the candidate defect response in the main features is corrected channel by channel and position by position to obtain a single-view candidate defect characterization.

[0026] As can be seen from the above technical solutions, the present invention has the following advantages: This application provides a method and system for inspecting the appearance of vials based on multi-view reflection suppression. By organizing and utilizing multi-view videos generated when the vial passes through the inspection area, and constructing a hierarchical discrimination mechanism based on the response differences between reflection interference and real defects, the method achieves stable extraction, effective identification, and consistent judgment of the appearance information of transparent curved vials. This reduces the interference of reflection imaging on defect identification conclusions, enhances the stability and reliability of defect discrimination results under different perspectives, and meets the requirements of strong stability in real defect discrimination and high consistency and reliability of detection results under different perspectives during the appearance inspection of transparent curved vials. Attached Figure Description

[0027] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart of a method for inspecting the appearance of vials based on multi-view reflection suppression provided by the present invention; Figure 2 This is a schematic diagram of a vial appearance inspection system based on multi-view reflection suppression provided by the present invention.

[0029] The modules include: 1. Video grouping module; 2. Image filtering module; 3. Unpacking and mapping module; 4. Decoupling and correction module; 5. Prototype matching module; and 6. Consistency verification module. Detailed Implementation

[0030] Various embodiments of this disclosure are described more fully below with reference to the accompanying drawings. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.

[0031] In the following, the terms “comprising” or “may include”, which may be used in various embodiments of this disclosure, indicate the presence of the disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. Furthermore, as used in various embodiments of this disclosure, the terms “comprising,” “having,” and their cognates are intended only to indicate a particular feature, number, step, operation, element, component, or combination of the foregoing, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more combinations of the foregoing.

[0032] To address the issues of strong reflection interference, weak stability in identifying true defects, and low consistency and reliability of test results under different viewing angles during the appearance inspection of transparent curved vials, existing inspection methods generally lack sufficient support for the stable identification and accurate judgment of abnormal information on the vial surface under complex imaging conditions. This makes it difficult to meet the actual needs of pharmaceutical packaging production lines for stable and reliable inspection of vial appearance quality. This application discloses a vial appearance inspection method and system based on multi-view reflection suppression. By introducing multi-view video organization, unified vial wall unfolding, reflection risk correction, defect prototype propagation matching, and cross-view causal consistency verification mechanisms, it can effectively extract, classify, identify, and consistently judge appearance defect information of transparent curved vials under complex reflection conditions. This enables accurate output of results for normal release, abnormal rejection, or re-inspection regression, thereby effectively improving the stability of identifying true defects during vial appearance inspection, significantly improving the consistency and reliability of test results under different viewing angles, further enhancing the quality of appearance inspection on the production line and the credibility of inspection output, and meeting the application requirements of automated vial appearance inspection in pharmaceutical packaging.

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] like Figure 1 As shown in the figure, this embodiment provides a method for inspecting the appearance of vials based on multi-view reflection suppression, including: Step S1: Collect multi-view videos of the vial passing through the detection area and establish sub-video groups corresponding to the viewpoints; Step S2: Based on the stability of the bottle outline, preservation of local texture, and blur suppression, select representative images from each sub-video group; Step S3: Locate the effective area of ​​the bottle in the representative image, and establish a unified unfolding domain based on the central axis of the bottle, the shoulder boundary and the bottom reference to obtain the unfolded image of the bottle wall; Step S4: Decompose the reflection characterization and defect characterization of the bottle wall unfolded image, and perform gating correction on the defect response according to the reflection risk to obtain single-view candidate defect characterization; Step S5: Perform multi-layer prototype propagation matching on the single-view candidate defect representation based on the defect prototype map to obtain the single-view category response; Step S6: Based on the corresponding positions in the unified expanded domain, perform cross-perspective causal consistency verification on the single-perspective category response, reflection risk, and candidate defect topology expression of each perspective, and output the results of normal release, abnormal rejection, or re-inspection rollback according to the verification results.

[0035] This embodiment employs multi-view video acquisition and viewpoint sub-video organization to enable the vial to generate continuous imaging information facing different observation directions as it passes through the detection area, providing a foundation for stable analysis of appearance anomalies in transparent curved vials. By utilizing the stability of the vial contour, local texture preservation, and blur suppression to representatively screen images from each viewpoint, images entering subsequent discrimination processes have better imaging quality and analytical value, enhancing the stability of appearance detection input. By locating the effective area of ​​the vial in representative images and establishing a unified unfolded domain based on the vial's central axis, shoulder boundary, and bottom reference, the curved vial wall area is transformed into an unfolded representation with clear positional relationships, improving the corresponding analysis capability between different viewpoints. Through the unfolded vial wall diagram... The reflection and defect representations in the image are decomposed, and the defect response is gated and corrected in conjunction with reflection risk, so that reflection interference and real defects can be distinguished and processed, improving the accuracy of defect identification under complex reflection conditions. By performing multi-layer prototype propagation matching based on defect prototype map, and combining the corresponding position in the unified unfolded domain, cross-view causal consistency verification is carried out on the response of each view category, reflection risk and candidate defect topology expression, so that the detection results have stronger category discrimination ability and cross-view consistency judgment ability. Overall, stable detection and reliable output of appearance defects of transparent curved vials are achieved, improving the stability of real defect discrimination and the consistency and credibility of multi-view detection results, meeting the application needs of pharmaceutical packaging production lines for automated inspection of vial appearance.

[0036] The above steps will be specifically described below based on the embodiments of this application.

[0037] In step S1, the core task is to generate multi-view input data that is time-synchronized, visually distinguishable, and consistent in identity around the same target vial. This ensures that subsequent representative image selection, effective region localization of the vial, unified unfolding, single-view discrimination, and cross-view verification are all based on the corresponding imaging of the same target vial. The input to this step is the raw video stream synchronously acquired by multiple industrial cameras within the detection area, and the output is a multi-view sub-video group corresponding one-to-one with a single target vial.

[0038] In this embodiment, a left-view industrial camera, a right-view industrial camera, and a top industrial camera are respectively deployed on both sides and above the conveying direction. The left-view industrial camera covers the left side of the bottle wall, the right-view industrial camera covers the right side of the bottle wall, and the top industrial camera covers the bottle shoulder, bottle neck, and upper bottle wall. The three industrial cameras are driven by the same trigger signal and simultaneously begin acquiring data when the vial enters the detection area, thereby capturing multi-view video of the vial passing through the detection area. For example, the resolution of a single industrial camera can be [missing information]. The frame rate can be The exposure time can be The lighting employs a combination of polarized strip light sources and diffuse surface light sources. The polarized strip light sources provide directional illumination along the circumference of the bottle, while the diffuse surface light sources compensate for the brightness of dark areas, ensuring that the fine textures of the bottle wall can be stably imaged at different heights. For the first... From one perspective, the video sequence acquired within one detection period can be written as: ,in, Indicates the first A video sequence from multiple perspectives, Indicates the first The perspective in the first Images acquired from frame acquisition, This indicates the number of valid frames for that viewpoint within the current detection period. This formula is used to organize continuously acquired results into video inputs that are differentiated by viewpoint and sorted by time, providing a unified data carrier for subsequent frame-by-frame quality evaluation.

[0039] It should be noted that, to avoid video clips from different bottles being mistakenly assigned to the same detection sample, target-level association processing is required for each viewpoint video to establish sub-video groups corresponding to each viewpoint. Specifically, left-view video clips, right-view video clips, and top-view video clips under the same trigger number are bound together, and consistency checks are performed based on the continuity of the bottle's outer frame center displacement, the smoothness of the bottle's area change, and the order in which they enter the detection area. Only frames with continuous position trajectories and complete outlines are retained as valid sub-video groups for the current target vial. For example, each viewpoint can retain... A valid image frame is formed by combining three viewpoints. The original frame detection input. If the distance between two adjacent bottles is too small, resulting in partial overlap of two bottles within the same trigger window, abnormal frame segments can be removed by checking the continuity of the main axis length of the bottle and the integrity of the boundary closure, so as to ensure that the current sub-video group corresponds to only one target vial.

[0040] Thus, step S1 completes the synchronous acquisition and organization of video from multiple fixed observation directions for the same target vial, forming an input data system with trigger numbers as indexes, viewpoint identifiers as distinctions, and single-vial video clips as carriers, providing a reliable video foundation for subsequent representative image selection.

[0041] In step S2, the core task is to select the most suitable representative images for appearance judgment from the continuous videos from each viewpoint, so that the images entering the subsequent bottle effective area positioning and unified unfolding process simultaneously possess the characteristics of stable contours, clear textures, and controlled blur. The input of this step is the sub-video group of each viewpoint, and the output is the representative image of each viewpoint and its quality score.

[0042] In this embodiment, brightness normalization, local contrast enhancement, and noise suppression can be performed frame-by-frame on each sub-video group, and representative images can be selected from each sub-video group based on bottle contour stability, local texture preservation, and blur suppression. Specifically, to form a computable selection criterion, bottle contour extraction can be performed on each frame of the image in each sub-video group to obtain the bottle region and bottle boundary. During this process, edge closure and maximum connected component filtering can be performed on the preprocessed image to obtain the... The first perspective Bottle area of ​​the frame and the set of boundary points of the bottle The purpose of this treatment is to exclude conveyor belt edges, background speckles, and irrelevant bright areas, so that subsequent quality evaluation focuses solely on the bottle itself.

[0043] After obtaining the bottle boundary, the stability of the bottle contour is further determined based on the degree of change in the bottle boundary between adjacent frames. The stability of the bottle contour can be written as:

[0044] in, Indicates the first The first perspective The stability of the bottle's outline in the frame. Indicates the boundary points in the current frame. Indicates the boundary point in the previous frame The nearest corresponding boundary point, Indicates the number of boundary points in the current frame. This represents the length of the diagonal of the rectangle circumscribed around the bottle in the current frame. This formula characterizes the degree of contour fluctuation by calculating the average displacement of boundary points in adjacent frames. The smaller the contour displacement, the more stable the bottle's posture, and the more suitable it is as input for subsequent detection.

[0045] In addition to contour stability, local texture preservation must also be determined based on the degree of texture gradient preservation between adjacent frames within the bottle region. Local texture preservation can be written as:

[0046] in, Indicates local texture preservation. This indicates the number of pixels in the bottle area. This represents the gradient difference normalization benchmark from this perspective. Indicates the first Pixel position of the frame gradient magnitude at that point This represents a very small positive number. This formula is used to measure the degree of preservation of local texture in adjacent frames. If fine cracks, shallow scratches and edge wear in a certain frame can still be stably preserved at the gradient level relative to the previous frame, then the local texture preservation is relatively high.

[0047] Simultaneously, the blur suppression property must be determined based on the sharpness of the edges within the bottle area. The blur suppression property can be written as:

[0048] in, Indicates fuzzy inhibition. Indicates pixel position Laplace response at the location, and These represent the minimum and maximum Laplacian responses within the bottle region of the current frame, respectively. This formula reflects the level of motion blur and trailing suppression by the sharpness of local edges; the sharper the edges, the higher the blur suppression.

[0049] After obtaining the bottle contour stability, local texture preservation, and blur suppression, a joint score is further calculated for these three aspects, and representative images are selected based on time interval constraints. Specifically, the comprehensive score can be written as:

[0050] in, This represents the overall image score. , and These represent the weighting coefficients for the three evaluation items, and can be exemplified by taking... , and The purpose of this formula is to unify the three factors of contour stability, texture preservation, and blurred control into a single scoring framework, avoiding the selection of frames with stable contours but blurred textures, or clear textures but jittery poses, based solely on a single evaluation metric. After sorting the images in descending order of their overall scores, a minimum time interval of 3 frames is set. Within these 3 frames, only the image with the highest overall score is retained as the representative image for that viewpoint, thereby reducing continuous video redundancy and controlling subsequent inference overhead.

[0051] Thus, step S2 completes the selection and conversion from continuous video to high-quality representative images, forming a representative image selection mechanism constrained by contour stability, texture preservation, and blur suppression, providing a clear and stable image foundation for subsequent effective bottle area localization and unified unfolding.

[0052] In step S3, the core task is to convert the bottle wall images with curvature distortion from different viewpoints into a uniformly positioned unfolded representation, so that the same bottle wall position can be mapped to the same coordinate system from different viewpoints. The input of this step is the representative image of each viewpoint, and the output is the unfolded image of the bottle wall and the correspondence between the bottle wall positions of each viewpoint.

[0053] In this embodiment, the effective area of ​​the bottle can be located in the representative image. Specifically, foreground separation and contour closure filtering can be performed on the representative image to obtain the effective area of ​​the bottle. Foreground separation can be performed using a combination of background subtraction and grayscale thresholding, while contour closure filtering removes conveyor belt edges, clamping structures, and scattered bright areas based on aspect ratio, principal axis direction, and boundary integrity. This processing ensures that subsequent calculations for establishing the reference frame and unfolded domain only target the main body area of ​​the bottle, avoiding positional shifts introduced by the background area.

[0054] After determining the effective area of ​​the bottle, a unified unfolding domain is further established based on the bottle's central axis, shoulder boundary, and bottom reference. Specifically, the bottle's central axis can be determined according to the direction of the bottle's principal axis, and a bottle reference system can be established based on the upper boundary of the shoulder and the center of the bottom. For this purpose, the bottle's central axis can be set as... The upper boundary of the bottle shoulder is The center of the bottle bottom is A coordinate system for unfolding the bottle wall surface is constructed, using the central axis of the bottle as the longitudinal reference and the circumferential position around the central axis as the lateral reference. The purpose of establishing this reference system is to unify the spatial aperture of bottle wall images from different perspectives, so that subsequent bottle wall unfolding and cross-perspective position comparisons have a consistent geometric basis.

[0055] After establishing the reference frame, the circumferential position of the bottle wall region along the central axis of the bottle and its height along the central axis can be mapped to a unified unfolded domain, thus obtaining an unfolded image of the bottle wall. For example, for the bottle wall pixels... Its expanded coordinates can be written as:

[0056] in, Represents pixels Circumferential coordinates in the unified expansion domain Represents pixels Height coordinates in the unified unfolded domain This represents the circumferential angle of the pixel relative to the central axis of the bottle. and These represent the minimum and maximum circumferential angles of the visible bottle wall area from the current viewpoint. This indicates the height position of the pixel along the central axis of the bottle. and These represent the upper and lower boundary heights of the effective area of ​​the bottle wall, respectively. The purpose of this formula is to transform the curved bottle wall image, which is originally affected by curvature, into a unified rectangular coordinate system, allowing the same physical location to have a comparable unfolded position expression from different viewpoints. For example, the unfolded bottle wall image size can be uniformly set to 512×256.

[0057] Furthermore, to support subsequent multi-view joint discrimination, the positional correspondence of the bottle wall from each viewpoint can be recorded according to a unified unfolded domain. Specifically, for each viewpoint, the circumferential interval of the visible bottle wall in the unified unfolded domain and the index of the overlapping interval with adjacent viewpoints are recorded, enabling subsequent candidate defect regions from different viewpoints to directly establish correspondences in the unified unfolded domain. The purpose of this processing is to transform cross-viewpoint position comparison from the original image coordinates to unified unfolded coordinates, reducing the difficulty of direct comparison caused by viewpoint differences.

[0058] Thus, step S3 completes the geometric mapping transformation from the representative image to the unfolded image of the bottle wall, forming a unified positional expression mechanism with the central axis of the bottle as the longitudinal reference and the circumferential position as the lateral reference, providing a clear spatial benchmark for subsequent single-view feature decomposition and cross-view consistency verification.

[0059] In step S4, the core task is to establish a single-view detection model that separates reflection and defect representations for the unfolded bottle wall image, and to perform gating correction on the defect response based on reflection risk, thereby suppressing the interference of specular highlights, specular reflections, and refractive pseudostructures on the real defect response at the feature level. The input of this step is the unfolded bottle wall image, and the output is the single-view candidate defect representation.

[0060] In this embodiment, the bottle wall unfolded image can be decomposed into reflection and defect representations. Specifically, in order to form a feature representation that can both preserve local texture anomalies and explicitly distinguish reflection patterns, the bottle wall unfolded image can be extracted to obtain the backbone defect features, thereby obtaining the backbone features representing local texture and structural anomalies. In the specific implementation of this embodiment, the backbone network is composed of a shallow texture coding block, a mid-level structure coding block, and a deep semantic coding block in sequence. The shallow texture coding block uses two 3×3 convolutional layers and one downsampling convolutional layer with a stride of 2, with 32 output channels, to extract fine-grained texture features such as fine cracks, shallow scratches, and small-area edge breaks; the mid-level structure coding block uses three residual convolutional units, with 64 output channels, to extract scratch direction, chipped edge contours, and local damaged structures; the deep semantic coding block uses four dilated residual units, with 128 output channels, to extract high-level semantic differences between the defect area and the background area. The feature maps at the three scales are fused into backbone features after scale alignment and channel mapping.

[0061] in, Indicates the first The first perspective Zhang represents the main features of the image. , and These represent the feature maps of the shallow, middle, and deep layers, respectively. , and These represent projection operators that perform channel mapping and spatial alignment of features at different scales, respectively. The purpose of this formula is to fuse local texture, structural, and semantic information at different scales into a unified backbone feature, so that subsequent reflection coding and defect coding are both based on multi-scale anomaly representation.

[0062] After obtaining the main features, the reflection feature encoding is further performed on the unfolded bottle wall image to obtain the reflection representation; simultaneously, the defect feature encoding is performed on the unfolded bottle wall image to obtain the defect representation. Specifically, two different projection heads can be set on the main features. The reflection encoding projection head mainly consists of two layers of 1×1 convolutions and global average pooling, used to enhance the representation of bright stripes, sheet-like reflections, and continuous refractive textures; the defect encoding projection head mainly consists of 3×3 convolutions, direction-sensitive convolutions, and local pooling layers, used to enhance the representation of local abnormal structures such as fine cracks, shallow scratches, and chipped edges. The two types of representations can be written as:

[0063] in, Represents the reflection representation vector. This represents the defect characterization vector. Represents reflection coding mapping, This represents the defect encoding mapping. The purpose of this formula is to split the backbone features into a reflection subspace and a defect subspace in the feature space, so that subsequent models no longer process specular pseudostructures and real defects in the same representation space.

[0064] Furthermore, to suppress the infiltration of reflection modes into defect representations, decorrelation constraints can be applied to both reflection and defect representations. The decorrelation constraints can be written as:

[0065] in, Indicates the related losses. This represents the inner product of the reflection representation and the defect representation. By constraining the cosine correlation between the two types of representations to be as small as possible, this formula aims to separate the reflection and defect representations in the direction as much as possible, thereby reducing the probability that specular textures are mistaken for real anomalies by defect branches.

[0066] In some embodiments of this application, to further ensure that the reflection representation and the defect representation are not abstractly separated but have an interpretable image mapping relationship, a reflection layer and defect layer reconstruction mechanism can also be introduced.

[0067] Specifically, let the reflective layer be... The defect layer is Then the sparse loss of the defect layer It can be written as:

[0068] in, This represents the total number of pixels. The purpose of this formula is to ensure that the defect layer only responds in localized anomalous regions, rather than forming diffuse activation in a large, uniform background. (Reflection layer smoothing loss) It can be written as:

[0069] in, and These represent the positions of the reflective layer at the pixels. The gradient is defined along both the horizontal and vertical directions. This formula utilizes the characteristic that specular reflection typically exhibits a continuous and smooth variation in space, making the reflective layer more closely approximate the true specular distribution. Reconstruction Loss It can be written as:

[0070] in, Indicates the pixel position of the unfolded image. The original pixel value at that location, This represents the pixel values ​​reconstructed from the reflection layer and the defect layer. The purpose of this formula is to constrain the double decomposition result to still return to the original unfolded image, ensuring that the division of the reflection layer and the defect layer has practical significance at the image level, rather than just a formal separation in the feature space.

[0071] After completing the dual decomposition encoding, gated correction branches can be established around specular and reflection risks. Specifically, the luminance channel can be extracted from the unfolded image of the bottle wall, and its average luminance value can be recorded as... The standard deviation of brightness is The highlight threshold It can be written as:

[0072] in, This represents the specular sensitivity coefficient. The function of this formula is to adaptively delineate candidate specular regions based on the current brightness distribution of the unfolded image, rather than using a fixed grayscale threshold, thus adapting to brightness fluctuations under different bottle surface conditions and lighting conditions. Pixels in the brightness channel above the specular threshold are merged into connected components and morphologically closed to form a specular mask. Subsequently, the proportion of specular regions, their distribution morphology, local sliding window energy, local sliding window entropy, and frequency domain concentration are extracted from the unfolded bottle image to form a reflection risk characterization. Among these, the proportion of specular regions... It can be written as:

[0073] in, This indicates the number of foreground pixels in the specular mask. This represents the total number of pixels in the unfolded image. This formula reflects the overall degree to which the currently unfolded image is covered by highlights. Average area of ​​the highlight region. It can be written as:

[0074] in, This indicates the total number of highlight areas. Indicates the first The pixel area of ​​each specular connected region. This formula is used to characterize whether the specular pseudostructure is scattered or distributed over a large continuous area. The shape factor of each highlight region can be written as:

[0075] in, Indicates the first The shape factor of each highlight area This represents the length of the longer side of the smallest bounding rectangle of the highlight region. This indicates the length of the corresponding short side. This formula is used to distinguish between long, thin strip-shaped specular reflections and clustered, localized bright spots.

[0076] Furthermore, to reflect brightness fluctuations and complex textures, sliding window statistics can be performed on the brightness channel. Local variance of a sliding window It can be written as:

[0077] in, Indicates the first A sliding window area Number of pixels included Indicates the brightness value. This represents the average brightness of the corresponding sliding window. This formula is used to characterize local brightness energy variations; reflective stripes typically cause a significant increase in local variance. Local entropy of a sliding window It can be written as:

[0078] in, Represents the set of brightness values. Indicates brightness value The probability distribution within this sliding window. This formula reflects the complexity of the local brightness distribution; real cracks and highlight bands often exhibit different patterns in local entropy. For the frequency domain concentration, a two-dimensional fast Fourier transform is performed on the unfolded image to obtain the spectrum. The proportion of high-frequency energy can be written as:

[0079] in, Indicates the proportion of high-frequency energy. Indicates the first The first perspective The unfolded image at frequency points Spectral coefficients at that location Represents the set of frequency points in the high-frequency region. This represents the set of frequency points across the entire frequency domain. This formula is used to characterize the energy proportion of the unfolded image in the high-frequency region. Real cracks typically have a more dispersed high-frequency structure, while specular reflections are more likely to exhibit a concentrated spectral pattern.

[0080] After obtaining the above quantities, the highlight ratio will be... Total number of regions Average area Mean shape factor Sliding window variance mean Mean entropy of sliding window High-frequency energy ratio and frequency domain concentration Concatenate into a reflection risk vector :

[0081] This formula encodes various reflection-related statistics into a structured risk representation, which serves as the direct input for subsequent gating correction. Based on this, the reflection risk representation and the reflection representation are further jointly mapped to obtain the reflection gating weights. :

[0082] in, Represents the gated mapping weight matrix. This represents the bias vector. This represents the Sigmoid activation function. Its function is to combine the reflection risk obtained through explicit statistics with the deep reflection characterization to generate the degree of suppression for each channel and location.

[0083] After obtaining the gating weights, the candidate defect responses in the backbone features are further corrected channel-by-channel and position-by-position based on the reflection gating weights, thereby obtaining a single-view candidate defect characterization. :

[0084] in, This indicates element-wise multiplication. The purpose of this formula is to use gating weights to directly suppress pseudo-high responses caused by highlights or reflections in the main features, making the retained candidate anomalous features more similar to the true defect responses.

[0085] Furthermore, to make the gated branches more stable during the training phase, reflection gating constraints can be introduced:

[0086] in, This represents the reflection-gated constraint loss. Indicates the reflection mask at pixel location The value at that location, This indicates the intensity of the defect response at that location before correction. This represents the normalization factor. The purpose of this formula is to further suppress spurious defect responses within high-reflectivity regions during training.

[0087] Thus, step S4 completes the construction of reflection representation, defect representation, and single-view candidate defect representation from the bottle wall unfolded image, forming a single-view feature system with the synergistic effect of reflection-defect dual decomposition coding and reflection risk gating correction, providing a reliable input basis for subsequent single-view category determination.

[0088] In step S5, the core task is to map the single-view candidate defect representation onto the structured defect knowledge space, and then perform multi-layer prototype propagation matching through the defect prototype map to obtain a stable single-view category response. The input of this step is the single-view candidate defect representation, and the output is the single-view category response.

[0089] In this embodiment, multi-layer prototype propagation matching can be performed on single-view candidate defect representations based on a defect prototype map. Specifically, during the training phase, for each type of defect such as cracks, scratches, chipping, fragmentation, dirt, and foreign objects, defect representations are extracted from labeled samples and clustered and hard example mining is performed. Then, central prototypes, boundary prototypes, and hard example prototypes are constructed for each type of defect. For example, for the first... Class defects, whose prototype set can be written as:

[0090] in, Indicates the first The prototype set of class defects Represents the central prototype. Indicates the first A boundary prototype, Indicates the first A difficult prototype, Indicates the number of boundary prototypes. This represents the number of hard example prototypes. The purpose of this formula is to expand each type of defect from a single average feature into a hierarchical set of prototypes, simultaneously preserving intra-class commonalities, inter-class boundaries, and hard example compensation capabilities. The central prototype describes the most stable common expression of the defect type, the boundary prototype describes boundary expressions easily confused with adjacent categories, and the hard example prototype describes anomalous expressions that still hold true even with small areas, weak textures, or strong interference.

[0091] After creating the hierarchical prototypes, the graph relationships between the prototype nodes can be further established, thus forming a defect prototype atlas. All categories of prototypes together constitute the prototype atlas. The similarity weights between nodes in the graph can be written as:

[0092] in, Represents the prototype node With prototype nodes The similarity weights between prototypes. The purpose of this formula is to explicitly express the proximity between prototypes as graph edge weights, so that subsequent matching is no longer just an isolated comparison between the test sample and a single prototype, but can utilize the structural relationships between prototypes to propagate intra-class commonalities and boundary constraints.

[0093] Based on this, during inference on the test sample, the single-view candidate defect representation and defect prototype map can be aligned to obtain the initial matching response of each prototype node:

[0094] in, Indicates the first The first perspective Zhang's representative image and the first Class 1 A prototype The initial matching response between them. The function of this formula is to project the current sample onto various prototype nodes, forming a prototype-level similarity distribution.

[0095] Furthermore, the matching responses of each prototype node are propagated and updated based on the graph relationships between prototype nodes. After the propagation update, the... The graphical response of a defect can be written as:

[0096] in, Indicates the first Spectral response of defects Indicates the number of prototypes for this class. Indicates the first Class 1 The adaptive response coefficients of each prototype after graph propagation. The purpose of this formula is to aggregate prototype-level responses into category-level responses, and to take into account the contributions of central prototypes, boundary prototypes, and hard-case prototypes through propagation coefficients, rather than simply taking the maximum similarity.

[0097] After obtaining various graph responses, further category-level aggregation is performed on the responses of prototype nodes corresponding to the same type of defect to obtain single-view category responses. The single-view prototype graph alignment vector can be written as:

[0098] in, Indicates the first The first perspective Zhang represents the prototype map alignment vector of the image. This represents the total number of defect categories. Based on this, the corrected single-view candidate defect representations, reflection representations, defect representations, and prototype map alignment vectors are fused to obtain the single-view fused feature. :

[0099] in, This represents a compression function that compresses a single-view candidate defect representation into a fixed-length vector. This represents the weight matrix of the fusion layer. This represents the bias vector of the fusion layer. The purpose of this formula is to unify the gated and corrected abnormal responses, reflection information, defect information, and prototype map responses into a single fusion feature, serving as the basis for single-view classification.

[0100] Ultimately, single-view category response It can be written as:

[0101] in, Represents the classification layer weight matrix. This represents the classification layer bias vector. The function of this formula is to output the category probability distribution of various defects from the current perspective, thereby completing the single-view anomaly category determination.

[0102] It should be noted that, in order to keep the prototype graph synchronized with the current feature space, in some embodiments of this application, various prototype nodes can be periodically updated during the training process. For example, after every few rounds of training, defect representations are re-extracted from the training set samples, and the center samples, boundary samples, and hard sample samples are re-clustered and the graph edge weights are updated, thereby avoiding the problem of overly coarse prototypes in the early stages leading to matching distortion in the later stages.

[0103] Furthermore, to enhance the stability of defect characterization of the same sample before and after enhancement, an enhanced consistency distillation loss can be introduced. :

[0104] in, This represents the defect characterization of the original sample. This expression represents the defect representation of the enhanced sample. Its purpose is to maintain the stability of the defect representation under conditions of brightness perturbation, simulated specular enhancement, and slight blur enhancement, and to prevent the model from only remembering the sample features under a single appearance condition.

[0105] Thus, step S5 completes the mapping from single-view candidate defect representation to single-view category response, forming a defect prototype map discrimination mechanism with the synergistic effect of central prototype, boundary prototype, and difficult case prototype, enabling single-view category determination to have stronger intra-class robustness and inter-class discrimination ability.

[0106] In step S6, the core task is to utilize the positional correspondences in the unified unfolded domain to jointly verify the single-view category responses, reflection risks, and candidate defect topology representations from multiple perspectives, and output normal release, anomaly rejection, or re-inspection rollback results based on the verification results. The inputs to this step are the single-view category responses, reflection risks, and unified unfolded positional indices for each perspective, and the outputs are the final execution results and their execution levels.

[0107] In this embodiment, cross-view causal consistency verification can be performed on the single-view category response, reflection risk, and candidate defect topological representation of each viewpoint based on the corresponding positions in the unified unfolded domain. Specifically, the positional correspondence of candidate defect regions for each viewpoint can be determined according to the corresponding positions in the unified unfolded domain. In this process, since step S3 has already recorded the mapping interval of the bottle wall position in the unified unfolded domain for each viewpoint, candidate anomaly regions from different viewpoints can first complete positional alignment in the unified unfolded domain before entering the topological and category consistency calculation. The purpose of this processing is to convert anomaly regions that are inconvenient to directly compare under the original image coordinates into candidate defect regions that can be compared under the same unfolded coordinates.

[0108] After positional alignment, defect skeletons and topological signatures can be extracted for candidate defect regions from each viewpoint. The topological signature includes skeleton length, principal direction, branch features, curvature features, and endpoint distribution. Specifically, for the first... For each candidate defect region from a different perspective, a refinement process can be performed to extract the defect skeleton, and then a topological signature can be constructed.

[0109] in, Indicates the first Topological signature of candidate defects from each perspective. Indicates the length of the skeleton. Indicates the principal direction angle. Indicates the number of branches in the skeleton. Indicates the mean curvature. This indicates the number of endpoints. The purpose of this formula is to transform the slender, single-branch structure commonly seen in crack-type defects, and the multi-endpoint, high-curvature structure commonly seen in edge-breaking and fragmentation defects, into a computable topological expression, thereby providing a structural basis for cross-view verification of true and false defects.

[0110] After obtaining the topological signature, the class consistency, reflection consistency, and topological consistency between each viewpoint are further calculated. The topological consistency between two viewpoints can be written as:

[0111] in, Indicates the first Topological signature of candidate defects from each perspective With the Topological signature of candidate defects from each perspective Topological consistency score between them This represents the topological difference scale parameter. The function of this formula is to transform the structural similarity of candidate defects from different perspectives into... arrive The consistency score is determined by the similarity of the topologies. Category consistency is determined by the similarity between single-view category responses from different perspectives, while reflection consistency is determined by the proximity of reflection risks from different perspectives. For example, when multiple perspectives point to the same defect category and the reflection risks are all low, both category consistency and reflection consistency are high.

[0112] Based on this, cross-perspective causal consistency verification results are obtained according to category consistency, reflection consistency, and topological consistency. Cross-perspective causal consistency score. It can be written as:

[0113] in, Indicates the category consistency score. Indicates the reflectivity consistency score. Represents the topology consistency score. , and This represents the corresponding weight coefficient. The purpose of this formula is to unify category consistency, reflection consistency, and topological consistency into the same cross-perspective verification framework. Only when multiple perspectives are simultaneously valid at the four levels of "same location, same category, non-reflection-dominated, and topologically similar" is the conclusion of the true defect considered stable.

[0114] After obtaining the cross-perspective causal consistency verification results, the system further outputs normal release, anomaly rejection, or re-inspection rollback results based on the verification results. Specifically, when the cross-perspective causal consistency verification results meet the anomaly establishment conditions, anomaly rejection results are output. For example, when at least two perspectives point to the same defect category at their corresponding positions in the unified expansion domain, and the cross-perspective causal consistency score is higher than a preset consistency threshold, the current target bottle can be determined as a genuine anomaly, and a rejection signal is sent to the execution agency. Further, when the cross-perspective causal consistency verification results meet the normal establishment conditions, normal release results are output. For example, when none of the perspectives have formed a stable anomaly category response, and the cross-perspective causal consistency score is insufficient to support the conclusion of a genuine defect, the current target bottle can be determined as normal, and a release signal is sent to the execution agency. For situations in between, when there is perspective divergence, category oscillation, or unresolved reflection interference in the cross-perspective causal consistency verification results, a re-inspection rollback result is output. For example, when an anomaly response is high from a certain perspective but inconsistent with other perspectives, or when multiple perspective categories fluctuate repeatedly between cracks and scratches, or when a high-reflection area highly overlaps with the main anomaly response, the anomaly is not directly rejected. Instead, the process is redirected to a re-inspection path to avoid false defect rejection. Simultaneously, the execution level corresponding to the main defect category, consistency level, and reflection interference level must be provided, allowing the actuator to not only receive the final action signal but also to understand the reliability and interference level of the current conclusion.

[0115] It should be noted that during the training phase, to ensure that reflection representation, defect representation, reflection gating correction, prototype map matching, and cross-view consistency verification converge collaboratively under the same objective, a joint loss function is used for end-to-end optimization. The joint loss function can be written as:

[0116] in, Indicates the total training loss. Indicates the primary classification loss. Indicates the related losses. This represents the sparse loss of the defect layer. This indicates the smoothing loss of the reflective layer. Indicates the losses incurred during reconstruction. This represents the reflection-gated constraint loss. This represents the alignment loss of the prototype map. This indicates cross-perspective consistency loss. This indicates increased consistency in distillation loss. to These are the corresponding loss weights. The purpose of this formula is to unify the model's main task discrimination, reflection-defect decomposition, gating suppression, prototype propagation, and multi-view stability into a single training objective, avoiding the situation where each branch optimizes independently but the overall structure is not closed. For example, the main classification loss can be written as:

[0117] in, Indicates the first The true label of class defects Indicates the first The predicted probability of class defects. This formula serves to constrain the final class determination to be consistent with the true label. The prototype map alignment loss can be written as:

[0118] in, The spectral response value representing the true category. Indicates the first The graphical response values ​​of the defect type, This represents the temperature coefficient. The purpose of this formula is to enhance the relative response advantage of the true category in the prototype map. The cross-view consistency loss can be written as:

[0119] in, and These represent the fused feature vectors of the same target bottle from two different perspectives. This represents the total number of viewpoints. The purpose of this formula is to constrain the fusion features of the same target bottle to remain consistent across different viewpoints, thereby improving the stability of cross-viewpoint joint decision-making. For example, training can use the AdamW optimizer, with an initial learning rate of [value missing]. The batch size is set to 16, and the number of training rounds is set to 120. During training, brightness perturbation, local contrast variation, specular overlay, slight affine transformation, micro-blur perturbation, and local occlusion enhancement can be performed to enable the model to maintain stable discrimination ability in complex reflection scenes.

[0120] Thus far, step S6 has completed the cross-perspective causal consistency verification based on unified expanded domain position correspondence, single-view category response, reflection risk, and candidate defect topology expression, and formed a hierarchical output mechanism of normal release, anomaly rejection, and re-inspection rollback, making the appearance inspection results of transparent curved vials in complex reflection scenarios more stable, more accurate, and more suitable for production line execution.

[0121] In summary, this method achieves effective differentiation between real and false defects under high reflectivity interference on the surface of transparent bottles by uniformly unfolding and representing multi-view images of vials, and combining reflection characterization and defect characterization decomposition, reflection risk gating correction, and cross-view causal consistency verification. It can reliably identify appearance anomalies such as cracks, scratches, chipping, and breakage, reduce the risk of false and missed detections due to reflection, improve detection accuracy, judgment stability, and online sorting reliability in complex optical scenes, and thus improve the automation level of vial appearance inspection and production line quality control capabilities.

[0122] It should be noted that, although the embodiments in this application are based on... Figure 1 The steps are described sequentially, but this does not mean that the steps must be performed in a strict order. The reason this embodiment follows this order is... Figure 1 The order in which each step is described is intended to facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art. In other words, the step numbers are only used to distinguish different steps and do not constitute a limitation on the execution order of the steps; the specific execution order of each step can be appropriately adjusted according to actual needs, functional requirements, and the inherent logic in actual application scenarios.

[0123] In some embodiments of this application, a method for inspecting the appearance of vials based on multi-view reflection suppression is applied to an online inspection production line for lyophilized formulation vials. The capacity of the vials to be inspected is... The height of the bottle is The maximum outer diameter of the bottle is The conveying speed is One industrial camera is installed on each of the three sides of the conveyor belt along the conveying direction in the detection area. All three cameras have a resolution of 2448×2048 and a frame rate of [missing information]. Exposure time was set to Polarized strip light sources are configured on the left and right sides, and a diffuse reflective surface light source is configured on the top. The industrial control computer is configured as follows: Graphics processor, Memory and Processor. The entire implementation process revolves around four typical scenarios: fine crack detection, specular defect suppression, coexistence of complex scratches and dirt, and cross-view inconsistency re-inspection. The main detection process includes stable frame selection, unified bottle wall expansion, reflection-defect dual decomposition coding, reflection risk gating correction, defect prototype map matching, cross-view causal consistency verification, and layered execution output.

[0124] Step 1: Establish multi-view video input for the current target vial.

[0125] When the target vial enters the detection area, the left-view industrial camera, the right-view industrial camera, and the top industrial camera synchronously acquire video under the same trigger signal. Let the... The video sequences corresponding to each viewpoint are:

[0126] in, Indicates the first The perspective in the first Images acquired from frame acquisition, This indicates the number of valid frames for that viewpoint within the current detection period. The purpose of this formula is to organize the continuous acquisition results of the current target bottle into a video sequence with viewpoint identifiers and time order, facilitating subsequent unified frame selection, unified positioning, and unified comparison. In this embodiment, the left viewpoint yields 21 frames, the right viewpoint yields 20 frames, and the top viewpoint yields 22 frames, forming a total of 63 frames of original video data. Subsequently, constrained by the continuity of the bottle's outer frame center displacement, the smoothness of the outer frame area change, and the stability of the principal axis length, residual frames at the tail of the previous bottle and intrusive frames at the head of the next bottle are removed, ultimately retaining the valid multi-view sub-video groups for the current target bottle. After this processing, all subsequent analyses revolve around the same target bottle, preventing adjacent bottle images from being mixed into the same detection conclusion.

[0127] Step two: Select representative images from the multi-view videos.

[0128] For each viewpoint, the video sequence is processed frame-by-frame with luminance normalization, median filtering, and local contrast enhancement. Then, the bottle region, bottle boundary, and local edge responses are extracted. For the first... The first perspective The frame represents the image score calculated by combining bottle contour stability, local texture preservation, and blur suppression.

[0129] in, This indicates the overall score. Indicates the stability of the bottle's outline. Indicates local texture preservation. This indicates blur suppression. The purpose of this formula is to unify the three evaluation criteria—contour stability, texture clarity, and edge sharpness—into a single scoring framework, avoiding the selection of images with unstable poses or obvious local ghosting based solely on a single clarity score. In this embodiment, the score for the 12th frame from the left view is 0.92, the score for the 9th frame from the right view is 0.89, and the score for the 14th frame from the top view is 0.90; therefore, these three frames are selected as representative images. After this processing, the original 63-frame video is compressed into three high-quality representative images, reducing subsequent inference overhead while preserving weak texture features such as fine cracks and shallow scratches to the maximum extent.

[0130] Step 3: Locate and uniformly expand the effective area of ​​the bottle in the representative image.

[0131] Foreground separation, contour closure, and maximum connected component filtering were performed on three representative images to obtain the effective region of the bottle. Then, the central axis of the bottle was determined based on the principal axis fitting results, and the longitudinal reference range was determined based on the upper boundary of the bottle shoulder and the center of the bottle bottom. Finally, the curved surface of the bottle wall was mapped onto a unified unfolded domain. For each pixel of the bottle wall... Its expanded coordinates can be written as:

[0132] in, Represents circumferential coordinates. Represents altitude coordinates. This represents the circumferential angle of a pixel relative to the central axis of the bottle. This represents the height position of a pixel along the central axis of the bottle. The purpose of this formula is to convert the bottle wall image, which is affected by curvature, into a uniform rectangular unfolded image, allowing comparisons of the same bottle wall position observed from different viewpoints using the same coordinates. In this embodiment, the unfolded image size is uniformly 512×256. The current abnormal region's position in the left viewpoint is... The position in the right-hand view is Its position in the top view is The circumferential deviation does not exceed 0.016, and the height deviation does not exceed 0.006, which meets the requirements for subsequent cross-view position correspondence.

[0133] Step four involves expanding the image input reflection suppression detection model to complete backbone feature extraction, reflection-defect dual decomposition encoding, reconstruction constraints, and reflection gating correction.

[0134] The detection model in this embodiment adopts a six-segment structure: "backbone encoder + reflection coding branch + defect coding branch + reconstruction constraint branch + reflection gating branch + prototype map matching branch," specifically adapted to high-reflectivity, small-defect, and multi-view unfolded image scenes of transparent glass vials. The model input is an unfolded image of the vial wall with dimensions of 512×256×3. The backbone encoder consists of three layers. The first layer is a shallow texture coding layer, including two convolutional layers with kernel size of 3×3 and one downsampling convolutional layer with a stride of 2, with an output feature size of 256×128×32, mainly extracting fine cracks, shallow scratches, and local edge abrupt changes. The second layer is a structure coding layer, including three residual convolutional blocks, with an output feature size of 128×64×64, mainly extracting crack direction, chipped edge contours, and local damaged structures. The third layer is a semantic aggregation layer, consisting of four hollow residual blocks with alternating hole ratios of 2 and 4. The output feature size is 64×32×128, primarily extracting the overall semantic differences between the abnormal and background regions. The feature maps from the three layers are upsampled and mapped using a 1×1 channel before being fused into a backbone feature map with a size of 128×64×96. This design is chosen because fine cracks rely on high-resolution shallow textures, edge chipping and fragmentation rely on mid-level structural information, and the distinction between real and fake objects against complex reflective backgrounds requires deep semantic information. The fusion of these three layers is more suitable for subsequent decoupling of reflection and defects.

[0135] Following the backbone feature map, a reflection coding branch and a defect coding branch are set up. The reflection coding branch consists of two 1×1 convolutional layers, one channel attention layer, and one global average pooling layer, outputting a 64-dimensional reflection representation to describe specular stripes, sheet-like reflections, local refraction, and continuous bright bands. The defect coding branch consists of two 3×3 convolutional layers, one orientation-sensitive convolutional layer, and one local pooling layer, also outputting a 64-dimensional defect representation to describe real anomalies such as cracks, scratches, chipping, and fragmentation. The orientation-sensitive convolutional layer contains four sets of orientation convolutional kernels at 0°, 45°, 90°, and 135°, each outputting 16 channels, for a total of 64 orientation-sensitive features, used to enhance the directional representation of slender cracks and strip-shaped scratches.

[0136] To ensure that the reflection and defect representations are not merely formally separated but possess image-level interpretability, a reconstruction constraint branch is established. This branch integrates the reflection layer representation map and the defect layer representation map. Figure 1 The image is then fed into a lightweight reconstructor, which consists of three convolutional layers and two upsampling operations, and outputs a reconstructed image. ,in, Indicated by the reflective layer and defect layer characterization The reconstructed image, This represents the reconstruction mapping function. Its purpose is to constrain the reflection and defect components separated by the network so that they can be recombined back into the original unfolded image, preventing the bibranch from being arbitrarily split without constraint. To further constrain the spatial properties of the two types of components, the training phase keeps the defect layer locally sparse and the reflection layer spatially smooth, thus making the defect layer closer to local anomalous regions and the reflection layer closer to continuous specular distribution.

[0137] After the dual-branch encoding, a reflection-gated branch is constructed. First, the mean brightness is calculated from the brightness channel of the unfolded image. and brightness standard deviation And obtain the highlight threshold:

[0138] in, Indicates the first Each viewpoint represents the highlight threshold of the image. This formula adaptively locates highlight regions based on the current image brightness distribution, rather than using a fixed threshold. In this embodiment, the average brightness of the left viewpoint is 124.6, and the standard deviation is 31.2, resulting in a highlight threshold of 180.76; the average brightness of the right viewpoint is 127.3, and the standard deviation is 29.7, resulting in a highlight threshold of 180.76; the average brightness of the top viewpoint is 118.4, and the standard deviation is 24.1, resulting in a highlight threshold of 161.78. Subsequently, the highlight region proportion, number of connected regions, average area, mean shape factor, mean local variance, mean local entropy, high-frequency energy proportion, and frequency domain concentration are extracted to form an 8-dimensional reflection risk vector. This reflection risk vector is concatenated with a 64-dimensional reflection representation and input into a gating branch. The gating branch consists of two fully connected layers: the first layer has an input dimension of 72 and an output dimension of 128; the second layer has an input dimension of 128 and an output dimension of 96, consistent with the number of trunk feature channels. The gated branch outputs 96-dimensional channel gated weights, which are then broadcast to spatial locations to perform channel-by-channel and position-by-position correction on the main features. Taking the left-view abnormal region in this embodiment as an example, before correction, the true crack response in the main features of this region is 0.83, and the specular pseudo-response is 0.74; after gated branch correction, the true crack response remains at 0.79, while the specular pseudo-response drops to 0.28. This indicates that the gated branch does not weaken the features as a whole, but rather uses the reflection risk vector and reflection characterization to suppress specular pseudo-defects in a targeted manner.

[0139] Step 5: Perform prototype map matching on the single-view candidate defect representation to obtain the single-view category response.

[0140] The prototype map in this embodiment adopts a hierarchical prototype structure consisting of a central prototype, boundary prototypes, and difficult-example prototypes. The crack class has 1 central prototype, 3 boundary prototypes, and 2 difficult-example prototypes; the scratch class has 1 central prototype, 3 boundary prototypes, and 2 difficult-example prototypes; the edge chipping class has 1 central prototype, 2 boundary prototypes, and 2 difficult-example prototypes; the fragmentation class has 1 central prototype, 2 boundary prototypes, and 2 difficult-example prototypes; and the dirt and foreign matter classes each have 1 central prototype, 2 boundary prototypes, and 1 difficult-example prototype, for a total of 25 prototype nodes. Each prototype node has a dimension of 64, consistent with the output dimension of the defect coding branch. After the test sample arrives, the similarity between the single-view candidate defect representation and all prototype nodes is calculated. Then, two rounds of propagation updates are performed within the prototype map, enabling the boundary prototypes to absorb the stable features of the central prototype and the difficult-example prototypes to compensate for the expression gaps of weak textures and small-area anomalies. For the abnormal bottle in this embodiment, the response for cracks from the left perspective is 0.91, for scratches it is 0.34, and for chipped edges it is 0.18; the response for cracks from the right perspective is 0.88, and for scratches it is 0.29; and the response for cracks from the top perspective is 0.74, and for scratches it is 0.26. Therefore, the main defect category in all three perspectives points to the crack category. For specular pseudo-defect samples, although the main features of a single perspective may have a high response locally, after correction by the reflection-gated branch, the matching strength between them and the prototypes of cracks and scratches is insufficient to form a stable prototype alignment relationship, so the abnormal category will not be directly output.

[0141] Step 6: Perform cross-perspective causal consistency verification on results from multiple perspectives.

[0142] First, establish the correspondence between candidate defect regions based on their position coordinates within the unified expanded domain. Then, extract the skeleton length, principal direction, number of branches, and number of endpoints for each candidate region from each viewpoint. Taking the abnormal bottle in this embodiment as an example, the skeleton length of the left viewpoint is 23.4 pixels, the principal direction is 79.6°, the number of branches is 0, and the number of endpoints is 2; the skeleton length of the right viewpoint is 22.1 pixels, the principal direction is 81.3°, the number of branches is 0, and the number of endpoints is 2; the skeleton length of the top viewpoint is 18.7 pixels, the principal direction is 77.9°, the number of branches is 0, and the number of endpoints is 2. The topological consistency between the two views can be calculated using the following formula:

[0143] in, Indicates the first The first perspective and the first Topological consistency score between perspectives and These represent the topological signatures from two different perspectives. This represents the topological difference scale parameter. The function of this formula is to convert the structural similarity of candidate defects from different perspectives into a consistency score between 0 and 1; the more similar the topology, the higher the score. Subsequently, a comprehensive score is calculated based on category consistency, reflection consistency, and topological consistency.

[0144] in, Indicates the cross-perspective causal consistency score. Indicates the category consistency score. Indicates the reflectivity consistency score. This represents the topological consistency score. The purpose of this formula is to simultaneously determine whether they belong to the same type of anomaly, whether neither is dominated by high reflectivity, and whether they have the same geometric structure, avoiding direct judgment of anomaly based solely on high response from a single viewing angle. In the embodiments of this application, , , Therefore, the overall consistency score is 0.89.

[0145] Step 7: Output the execution results and complete the online deployment.

[0146] When the preset consistency threshold is set to 0.78, since the overall consistency score of the current target bottle is 0.89 and the main defect category of all three perspectives is crack, an anomaly rejection result is output. Simultaneously, the main defect category is output as crack, the consistency level as high, and the reflection interference level as medium, and a rejection signal is sent to the air-blowing rejection mechanism. If a bottle has a high response only in one perspective and no corresponding defect expression is formed in other perspectives, it is not directly rejected, but a re-inspection and rollback status is output. If the anomaly responses of all perspectives are low and the prototype map does not form a stable match, a normal release result is output. This hierarchical decision-making logic is consistent with multi-view consistency verification and execution control. The execution mechanism can be a robotic arm, lever, air-blowing mechanism, or guide rail diversion mechanism.

[0147] In this embodiment, the model is deployed using offline training and online inference. The training set consists of 36,800 images collected from 3 production lines and 7 batches, including 21,000 normal samples, 4,200 cracked samples, 3,900 scratched samples, 3,100 chipped samples, 1,800 broken samples, and 2,800 samples with dirt and foreign objects. To enhance the model's adaptability to realistic reflection scenes, brightness perturbation, local contrast perturbation, simulated specular stripe overlay, slight Gaussian blur, random occlusion, and local noise injection are performed during training. The brightness perturbation range is set to ±18%, the simulated specular stripe width is set to 6–18 pixels, and the blur kernel size alternates between 3×3 and 5×5. The AdamW optimizer is used for training, with an initial learning rate of... The batch size is 16, and the training rounds are 120. Starting from round 40, the prototype map is updated every 5 rounds. To ensure that the main defect category determination, reflection representation and defect representation are separated, reflection gating suppression, double decomposition reconstruction constraints, prototype map matching, and cross-view consistency verification converge under the same training objective, a joint loss function is adopted:

[0148] in, Used to constrain the final category determination result. Used to constrain the degree of separation between reflection characterization and defect characterization. Used to keep the defect layer locally sparse in space. Used to keep the reflective layer spatially continuous and smooth. Used to constrain the reflection layer and defect layer to reconstruct the original unfolded image. Used to constrain gated branches to suppress spurious defect responses in high-reflectivity regions. Used to enhance the relative response advantage of the real category in the prototype map. This is used to ensure that the fusion representation of the same target bottle tends to be consistent from different perspectives. This formula is used to constrain the original image input and pseudo-sample input to remain stable in the defect representation space. Its function is to unify the main task discrimination, reflection-defect decoupling, gating suppression, reversible decomposition and reconstruction, prototype propagation, and multi-view stability into a single training objective, avoiding individual branch optimizations that result in an overall non-closed loop. Ultimately, on the independent validation set, the crack class recognition accuracy reached 98.1%, with the overall false positive rate controlled below 1.7%, representing a 43.6% reduction in false positive rate and a 27.4% reduction in false negative rate compared to the basic single-branch model.

[0149] Through the complete implementation process described above, this method can effectively distinguish between real defects such as cracks, scratches, chipping, and breakage and specular reflection artifacts under complex imaging conditions of high reflectivity and high-speed transport of transparent curved vials. It reduces misjudgments caused by high light interference from a single viewpoint, improves the consistency, stability, and online sorting reliability of multi-view detection results, thereby enhancing the automation level of vial appearance inspection and the quality control capability of the production line.

[0150] It should be understood that the step numbers identified by "Step 1, Step 2" and other similar forms in the above embodiments are only used to distinguish different steps and do not limit the steps to be executed in the order of these numbers. The specific execution order of each step can be adjusted according to its functional requirements and the inherent logic in the actual application scenario. The above step numbers should not be interpreted as a limitation on the implementation process of the embodiments of this application.

[0151] like Figure 2As shown, the following is an embodiment of a vial appearance inspection system based on multi-view reflection suppression provided by this disclosure. This vial appearance inspection system based on multi-view reflection suppression belongs to the same inventive concept as the vial appearance inspection method based on multi-view reflection suppression in the above embodiments. For details not described in detail in the embodiments of the vial appearance inspection system based on multi-view reflection suppression, please refer to the embodiments of the vial appearance inspection method based on multi-view reflection suppression described above.

[0152] Based on the same concept, another embodiment of this application provides a vial appearance inspection system based on multi-view reflection suppression, comprising: Video grouping module 1 is used to collect multi-view videos of vials passing through the detection area and to create sub-video groups corresponding to the viewpoints. Image filtering module 2 is used to filter representative images from each sub-video group based on bottle outline stability, local texture preservation, and blur suppression. The unfolding mapping module 3 is used to locate the effective area of ​​the bottle body in the representative image, and to establish a unified unfolding domain based on the central axis of the bottle body, the shoulder boundary of the bottle and the bottom reference of the bottle to obtain the unfolded image of the bottle wall; The decoupling correction module 4 is used to decompose the reflection characterization and defect characterization of the bottle wall unfolding image, and to perform gate correction on the defect response according to the reflection risk to obtain a single-view candidate defect characterization. Prototype matching module 5 is used to perform multi-level prototype propagation matching on single-view candidate defect representations based on defect prototype maps to obtain single-view category responses. The consistency verification module 6 is used to perform cross-perspective causal consistency verification on the single-perspective category response, reflection risk and candidate defect topology expression of each perspective based on the corresponding position in the unified expansion domain, and output the results of normal release, abnormal rejection or re-inspection rollback according to the verification results.

[0153] In some embodiments of this application, the decoupling correction module 4 includes: The backbone extraction submodule is used to extract backbone defect features from the unfolded bottle wall image to obtain backbone features that characterize local texture and structural anomalies. The reflection coding submodule is used to encode the reflection features of the unfolded bottle wall image to obtain a reflection representation. The defect coding submodule is used to encode defect features in the unfolded bottle wall image to obtain defect characterization. The decorrelation submodule is used to perform decorrelation constraints on reflection and defect representations to suppress reflection modes from being mixed into defect representations. The gated correction submodule is used to extract the proportion of highlight areas, the distribution pattern of areas, the local sliding window energy, the local sliding window entropy, and the frequency domain concentration from the unfolded image of the bottle wall to form a reflection risk characterization. The reflection risk characterization and the reflection characterization are jointly mapped to obtain the reflection gating weight. Based on the reflection gating weight, the candidate defect response in the main features is corrected channel by channel and position by position to obtain a single-view candidate defect characterization.

[0154] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.

Claims

1. A method for inspecting the appearance of vials based on multi-view reflection suppression, characterized in that, include: Collect multi-view videos of vials passing through the detection area and establish sub-video groups corresponding to the viewpoints; Based on the stability of the bottle outline, preservation of local texture and blur suppression, representative images were selected from each sub-video group; The effective area of ​​the bottle is located in the representative image, and a unified unfolding domain is established based on the central axis of the bottle, the shoulder boundary and the bottom reference to obtain the unfolded image of the bottle wall; The bottle wall unfolding image is decomposed into reflection and defect characterization, and the defect response is gated and corrected according to the reflection risk to obtain single-view candidate defect characterization. Based on the defect prototype map, multi-layer prototype propagation matching is performed on the single-view candidate defect representation to obtain the single-view category response. Based on the corresponding positions in the unified expanded domain, cross-perspective causal consistency verification is performed on the single-perspective category response, reflection risk, and candidate defect topology expression of each perspective, and the results of normal release, abnormal rejection, or re-inspection rollback are output according to the verification results.

2. The vial appearance inspection method based on multi-view reflection suppression as described in claim 1, characterized in that, Based on the stability of the bottle outline, preservation of local texture, and suppression of blur, the steps for selecting representative images from each sub-video group include: The bottle outline is extracted from each frame image in each sub-video group to obtain the bottle region and bottle boundary. The stability of the bottle outline is determined based on the degree of change of the bottle boundary between adjacent frames. Local texture preservation is determined based on the degree of texture gradient preservation between adjacent frames within the bottle area; The blur suppression is determined based on the sharpness of the edges within the bottle area; The stability of the bottle outline, local texture preservation, and blur suppression are jointly scored, and representative images are selected by combining time interval constraints.

3. The vial appearance inspection method based on multi-view reflection suppression as described in claim 2, characterized in that, The steps for locating the effective area of ​​the bottle in the representative image and establishing a unified unfolded domain based on the bottle's central axis, shoulder boundary, and bottom reference to obtain the unfolded image of the bottle wall include: Perform foreground separation and contour closure filtering on the representative image to obtain the effective area of ​​the bottle; Determine the central axis of the bottle body according to the direction of the main axis of the bottle body, and establish the bottle body reference system based on the upper boundary of the bottle shoulder and the center of the bottle bottom. The bottle wall region is mapped to a unified unfolded domain along the circumferential position around the central axis of the bottle and the height position along the central axis of the bottle to obtain the unfolded image of the bottle wall. Record the corresponding positions of the bottle wall from different perspectives according to the unified unfolded domain.

4. The vial appearance inspection method based on multi-view reflection suppression as described in claim 1, characterized in that, The steps for decomposing reflection and defect characterization on the unfolded image of the bottle wall include: Main defect features were extracted from the unfolded image of the bottle wall to obtain main features that characterize local texture and structural anomalies; Reflection feature encoding is performed on the unfolded image of the bottle wall to obtain a reflection characterization; Defect features are encoded from the unfolded image of the bottle wall to obtain defect characterization; Decorrelation constraints are applied to reflection and defect representations to suppress reflection modes from being mixed into defect representations.

5. The vial appearance inspection method based on multi-view reflection suppression as described in claim 4, characterized in that, The steps for gating and correcting the defect response based on reflection risk to obtain a single-view candidate defect characterization include: Extract the proportion of highlight areas, the distribution pattern of the areas, the local sliding window energy, the local sliding window entropy, and the frequency domain concentration from the unfolded image of the bottle wall to form a reflection risk characterization. By jointly mapping the reflection risk representation and the reflection representation, the reflection gating weights are obtained. Based on the reflection gating weights, the candidate defect responses in the backbone features are corrected channel-by-channel and position-by-position to obtain a single-view candidate defect characterization.

6. The vial appearance inspection method based on multi-view reflection suppression as described in claim 5, characterized in that, The steps for performing multi-layer prototype propagation matching on single-view candidate defect representations based on defect prototype maps to obtain single-view category responses include: For each type of defect, a central prototype, a boundary prototype, and a difficult example prototype are constructed, and the graph relationships between the nodes of each prototype are established to form a defect prototype map. Align the single-view candidate defect representations and defect prototype maps with responses to obtain the initial matching responses of each prototype node. The matching response of each prototype node is propagated and updated based on the graph relationship between prototype nodes; The responses of each prototype node corresponding to the same type of defect are aggregated at the category level to obtain a single-view category response.

7. The method for inspecting the appearance of vials based on multi-view reflection suppression as described in claim 1, characterized in that, Based on the corresponding positions in the unified expanded domain, the steps for cross-perspective causal consistency verification of single-perspective category responses, reflection risks, and candidate defect topological representations from various perspectives include: The positional correspondence of candidate defect regions from each perspective is determined according to the unified expansion domain. Defect skeletons and topological signatures are extracted for each candidate defect region from each perspective. The topological signature includes skeleton length, main direction, branch features, curvature features and endpoint distribution. Calculate the category consistency, reflection consistency, and topological consistency among each perspective, and obtain the cross-perspective causal consistency verification results based on the category consistency, reflection consistency, and topological consistency.

8. The vial appearance inspection method based on multi-view reflection suppression as described in claim 7, characterized in that, The steps for outputting normal release, abnormal rejection, or re-inspection rollback results based on the verification results include: When the cross-perspective causal consistency verification result meets the anomaly establishment condition, the anomaly removal result is output; When the cross-perspective causal consistency verification result meets the normal condition, the normal release result is output. When there are perspective discrepancies, category swings, or unresolved reflection interference in the cross-perspective causal consistency verification results, the re-examination rollback result is output. The execution level corresponding to the main defect category, consistency level, and reflection interference level is given simultaneously.

9. A vial appearance inspection system based on multi-view reflection suppression, characterized in that, include: The video grouping module is used to collect multi-view videos of vials passing through the detection area and create sub-video groups corresponding to the viewpoints. The image filtering module is used to filter representative images from each sub-video group based on bottle outline stability, local texture preservation, and blur suppression. The unfolding mapping module is used to locate the effective area of ​​the bottle body in the representative image, and to establish a unified unfolding domain based on the central axis of the bottle body, the shoulder boundary of the bottle and the bottom reference of the bottle to obtain the unfolded image of the bottle wall; The decoupling correction module is used to decompose the reflection characterization and defect characterization of the bottle wall unfolding image, and to perform gated correction on the defect response according to the reflection risk to obtain a single-view candidate defect characterization. The prototype matching module is used to perform multi-level prototype propagation matching on single-view candidate defect representations based on the defect prototype map to obtain single-view category responses. The consistency verification module is used to perform cross-perspective causal consistency verification on single-perspective category responses, reflection risks, and candidate defect topology expressions from various perspectives based on the corresponding positions in the unified expanded domain, and output normal release, anomaly rejection, or re-inspection rollback results according to the verification results.

10. The vial appearance inspection system based on multi-view reflection suppression as described in claim 9, characterized in that, The decoupling correction module includes: The backbone extraction submodule is used to extract backbone defect features from the unfolded bottle wall image to obtain backbone features that characterize local texture and structural anomalies. The reflection coding submodule is used to encode the reflection features of the unfolded bottle wall image to obtain a reflection representation. The defect coding submodule is used to encode defect features in the unfolded bottle wall image to obtain defect characterization. The decorrelation submodule is used to perform decorrelation constraints on reflection and defect representations to suppress reflection modes from being mixed into defect representations. The gated correction submodule is used to extract the proportion of highlight areas, the distribution pattern of areas, the local sliding window energy, the local sliding window entropy, and the frequency domain concentration from the unfolded image of the bottle wall to form a reflection risk characterization. The reflection risk characterization and the reflection characterization are jointly mapped to obtain the reflection gating weight. Based on the reflection gating weight, the candidate defect response in the main features is corrected channel by channel and position by position to obtain a single-view candidate defect characterization.