A paper cup stacking quantity online identification method and system based on machine vision

By combining machine vision and distance sensors, the problem of inaccurate counting of paper cup stacks on high-speed production lines was solved, achieving high-precision recognition of paper cup stack counts.

CN122265979BActive Publication Date: 2026-08-04CHANGSHA SHITONG PAPER & PLASTIC PRODUCTS AND CHOPSTICKS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGSHA SHITONG PAPER & PLASTIC PRODUCTS AND CHOPSTICKS CO LTD
Filing Date
2026-05-27
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing online paper cup stacking quantity recognition equipment is susceptible to interference from obstructed counting sensors under high-speed production conditions, leading to counting errors and an inability to achieve accurate recognition.

Method used

An online paper cup stacking quantity recognition method based on machine vision is adopted. The paper cup images are acquired by an image sensor, the stacking height of the paper cups is detected by a distance sensor, and data matching is performed using a cup rim template and a pixel height database. Combined with weighted average and difference verification, the accurate counting of paper cups is achieved.

Benefits of technology

It significantly improves the accuracy and stability of online recognition of paper cup stack count, reduces the counting error rate, and meets the counting requirements of high-speed production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265979B_ABST
    Figure CN122265979B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of paper cup large-scale production, and discloses a paper cup stacking quantity online identification method and system based on machine vision, which captures the paper cup after stacking and extrusion through an image sensor, and obtains a single-column stacking image through segmentation; the number of cup rims is extracted based on a cup rim template, the stacking height value is obtained by combining a stacking end pixel position and a pixel height database, and the first paper cup number is calculated by matching stacking feature data; the stacking distance is detected through a distance sensor, and the second paper cup number is obtained through conversion; the visual quantity difference value is calculated, and according to the size of the visual quantity difference value and a visual reference difference value, corresponding data is fused to obtain the final count. The application adopts machine vision and distance sensor dual-source cooperation, effectively solves the counting error problem of traditional methods which are prone to shielding, deformation and high-speed working condition interference, and significantly improves the accuracy and stability of online identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large-scale paper cup production, and in particular to a method and system for online recognition of the number of stacked paper cups based on machine vision. Background Technology

[0002] In the large-scale production of paper cups, accurate identification of the number of stacked paper cups is crucial for ensuring uniform product packaging specifications, meeting downstream customer needs, and achieving closed-loop control of the production process. After forming and finishing, paper cups need to be stacked and packaged according to a preset quantity. If the stacking quantity deviates, it will not only lead to unqualified products and increased production costs, but also affect the company's brand reputation and market competitiveness. Therefore, achieving accurate online identification of the number of stacked paper cups has significant practical importance and application value for paper cup manufacturers.

[0003] Currently, most online paper cup stacking quantity recognition devices adopt non-machine vision recognition methods. These methods are characterized by relatively simple structure, low initial investment cost, and no need for complex image acquisition and processing systems, making them easy to promote and apply in small and medium-sized paper cup manufacturing enterprises. Common non-machine vision recognition methods mainly include weighing counting and infrared through-beam counting. Their core principle is to indirectly count the number of paper cups by detecting their physical parameters, without the need to collect and analyze the appearance features of the paper cups, making them convenient to operate and with low maintenance costs.

[0004] However, there are significant application drawbacks to using non-machine vision recognition methods for online identification of the number of stacked paper cups, especially in high-speed online production environments where counting errors are highly likely. Because non-machine vision methods rely on indirect counting using physical parameters, if a single paper cup obstructs the counting sensor during stacking, it can cause deviations in the counting parameters, resulting in inconsistent numbers of paper cups in the finished packaging. Summary of the Invention

[0005] To improve the accuracy of paper cup stack counting, this application provides a machine vision-based online method and system for recognizing the number of paper cup stacks.

[0006] Firstly, this application provides an online method for recognizing the number of stacked paper cups based on machine vision, employing the following technical solution: A machine vision-based online method for recognizing the number of stacked paper cups includes the following steps: An image sensor is placed next to the stacked paper cups that are about to be transported. The image sensor is used to capture images of the stacked and compressed paper cups and generate images of the paper cups. Multiple rows of paper cups are segmented from the paper cup images to obtain stacked images of a single row of paper cups. The number of cup edges is obtained by extracting cup edges from the stacked image based on the preset cup edge template and accumulating the number. The stacking position of the stacking end of the single column of paper cups in the stacked image is extracted to obtain the stacking end pixel position. The corresponding value is matched from the preset pixel height database according to the stacking end pixel position as the stacking height value. The stacking feature data corresponding to the paper cup type is obtained. The number of the first paper cup is calculated according to the stacking height value and the stacking feature data. A distance sensor is set in the stacking direction of the paper cups. The distance sensor is used to detect the distance after the paper cups are stacked and squeezed and generate distance data. The corresponding value is matched from the preset pixel height database according to the distance data as the stacking height value. The second number of paper cups is calculated according to the stacking height value and the stacking feature data. The absolute value of the difference between the number of cup rims and the number of first paper cups is the visual quantity difference. If the visual quantity difference is less than the preset visual reference difference, the paper cup stack count is calculated based on the number of cup rims and the number of first paper cups. Otherwise, the paper cup stack count is calculated based on the number of first paper cups and the number of second paper cups.

[0007] By adopting the above technical solution, and through the dual detection and fusion counting of machine vision cup rim recognition, height detection and distance sensing, the problem of traditional non-visual recognition methods being easily interfered with by occlusion and having large counting errors under high-speed production conditions is effectively solved, and the accuracy and stability of online recognition of the number of stacked paper cups are greatly improved.

[0008] Furthermore, the step of extracting the cup rims from the stacked images based on a preset cup rim template and accumulating the number of cup rims also includes the following sub-steps: The cup rim template includes paper cup segmentation values, stacking direction, cup body graphic, cup rim segmentation values, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; The cup body shape is removed from the top of the paper cup image along the direction of the paper cup, and the edge shape of the paper cup is removed at a preset inner edge distance to obtain the middle image; Multiple parallel cup rims are extracted from the intermediate image based on the cup rim segmentation value, and the total number of cup rims is obtained by accumulating the number of cup rims; where the paper cup segmentation value is greater than the cup rim segmentation value.

[0009] By adopting the above technical solution, the cup rim extraction method can accurately identify and extract the parallel cup rim features of stacked paper cups through hierarchical threshold segmentation and cup body and edge interference removal processing, effectively avoiding interference such as cup body deformation and edge noise, and significantly reducing the false detection and false detection rate of cup rim counting.

[0010] Furthermore, the step of extracting the cup rims from the stacked images based on a preset cup rim template and accumulating the number of cup rims also includes the following sub-steps: The cup rim template includes paper cup segmentation values, stacking direction, cup rim graphic, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; A first semicircular shape is extracted from one side of the paper cup image along the stacking direction based on the cup rim graphic, and a second semicircular shape is extracted from the other side of the paper cup image along the paper cup direction based on the axisymmetric graphic of the cup rim graphic. The total number of the first semicircle shape and the total number of the second semicircle shape are counted. If the two counts are the same, one of them is taken as the number of the cup rim.

[0011] By adopting the above technical solution, and by symmetrically extracting and bidirectionally verifying the semi-circular edges of stacked paper cups, the problems of missed detection and false detection that are prone to occur in single-sided detection are effectively reduced, and the interference of paper cup stacking offset and partial obstruction on counting accuracy is greatly reduced.

[0012] Furthermore, the step of extracting the cup rims from the stacked images based on a preset cup rim template and accumulating the number of cup rims also includes the following sub-steps: The cup rim template includes paper cup segmentation values, stacking direction, cup rim graphic, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; The paper cup image is folded along the axis of the stacking direction to obtain a folded image, and a semi-circular shape is extracted from the folded image based on the cup rim graphic; The total number of semicircular shapes is used as the number of cup rims.

[0013] By adopting the above technical solution, the paper cup image is folded along the stacking axis, making full use of the symmetrical structure of the paper cup to simplify the extraction of cup rim features, effectively eliminating unilateral image distortion, local occlusion and background noise interference, reducing the complexity of algorithm processing, and significantly improving the accuracy of semi-circular cup rim recognition.

[0014] Furthermore, the step of calculating the number of the first paper cups based on the stacking height value and stacking feature data also includes the following sub-steps: The stacking feature data includes stacking direction, cup height data, and cup rim height data. The cup rim height data is the height of the exposed part of the paper cup located in the middle of the stacked paper cups in the stacking direction. The sum of the cup height data and the cup rim height data is the height of the paper cup. The stack height value, cup body height data, and cup rim height data are all pixel data in the image after correction for imaging distortion. Subtracting the cup body height data from the stack height value yields the temporary value of the middle part, and dividing the temporary value of the middle part by the cup rim height data yields the number of the first paper cups.

[0015] By adopting the above technical solution, and through imaging deformation correction and refined calculation of the pixel height of the cup body and cup rim, the height measurement error caused by paper cup stacking and compression and visual imaging distortion can be effectively reduced.

[0016] Furthermore, the step of calculating the stack count of paper cups based on the number of cup rims and the number of first paper cups also includes the following sub-steps: The number of cups on the rim and the number of the first paper cups are used to calculate the stack count of paper cups by weighted average; Obtain the stack height values ​​corresponding to multiple consecutive first paper cup quantities, and calculate the fluctuation value of multiple stack height values; If the fluctuation value is greater than the preset fluctuation reference value, the weight of the number of first paper cups is controlled according to the fluctuation value. The larger the fluctuation value, the smaller the weight, and the smaller the fluctuation value, the larger the weight.

[0017] By adopting the above technical solution, using weighted average fusion counting, and dynamically and adaptively adjusting the weight of visual height calculation data based on stacking height fluctuations, the counting deviation caused by height fluctuations during the production process can be effectively offset.

[0018] Furthermore, the step of calculating the stack count of paper cups based on the first and second paper cup counts also includes the following sub-steps: The number of paper cups in the first stack and the number of paper cups in the second stack are calculated by weighted average. Obtain the stack height values ​​corresponding to multiple consecutive first paper cup quantities, and calculate the fluctuation value of the multiple stack height values ​​as the first fluctuation value; obtain the stack height values ​​corresponding to multiple consecutive second paper cup quantities, and calculate the fluctuation value of the multiple stack height values ​​as the second fluctuation value; the first fluctuation value and the second fluctuation value are percentage values. The ratio of the first fluctuation value to the second fluctuation value is calculated as the fluctuation ratio. The weight of the number of first paper cups is controlled according to the fluctuation ratio. The larger the fluctuation value, the smaller the weight, and the smaller the fluctuation value, the larger the weight.

[0019] By adopting the above technical solution, and by comparing the fluctuation levels of visual inspection and distance sensing data and adaptively allocating weights according to the fluctuation ratio, the impact of highly volatile data sources can be automatically weakened, the contribution of stable data sources can be strengthened, and production interference can be effectively resisted.

[0020] Furthermore, the method also includes the following steps: The absolute value of the difference between the number of cups on the rim and the number of the first paper cups is the first difference. The absolute value of the difference between the number of cups on the rim and the number of second paper cups is the second difference. The absolute value of the difference between the number of first paper cups and the number of second paper cups is the third difference. If the first difference, the second difference, and the third difference are all greater than the preset reference difference, an alarm for abnormal counting sensor readings will be triggered.

[0021] By adopting the above technical solution, and through the triple difference verification of three sets of counting data, visual recognition, distance detection and sensor malfunctions can be quickly and accurately identified and alarmed in a timely manner.

[0022] Secondly, this application provides an online recognition system for the number of stacked paper cups based on machine vision, employing the following technical solution: A machine vision-based online recognition system for the number of stacked paper cups includes a processor that performs the steps of the machine vision-based online recognition method for the number of stacked paper cups as described in any of the preceding claims. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating the steps of an online method for recognizing the number of stacked paper cups based on machine vision.

[0024] Figure 2 This is a flowchart of the first extraction algorithm.

[0025] Figure 3 This is a flowchart illustrating the steps of the second extraction algorithm. Detailed Implementation

[0026] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0027] This application discloses an online method for recognizing the number of stacked paper cups based on machine vision. This method is mainly applied to high-speed automated production lines for disposable paper cups, such as those with a designed capacity of 80-120 cups / minute, supporting 2-6 columns of parallel stacking. It addresses the industry pain point that traditional weighing and infrared beam counting methods are easily affected by paper cup tilting, compression deformation, and momentary occlusion interference under high-speed conditions, resulting in a counting error rate exceeding 5%. (Refer to...) Figure 1 The specific steps are as follows: An industrial CMOS image sensor is fixedly installed perpendicularly to the stacked end face of the paper cups next to the station where the paper cups have been extruded and stacked by the forming machine and are about to enter the conveyor belt. It is paired with a 16mm low-distortion industrial lens, with the lens optical axis aligned with the stacked paper cup axis. The sensor is installed at a height of 50±5cm from the stack end face to ensure that the imaging range completely covers the multiple rows of stacked paper cups without perspective distortion. The sensor is connected to the host computer vision processing module via Gigabit Ethernet, with a transmission latency of ≤10ms, meeting the requirements for high-speed online inspection.

[0028] Because the production line uses a multi-column parallel stacking design, the original captured images of paper cups contain redundant information such as adjacent columns of paper cups, conveyor belt background, and frame structure, which can easily cause feature interference. Therefore, single-column segmentation is required. Step 1: Convert the colored paper cup image to grayscale using the standard weighted formula: Gray=0.299R+0.587G+0.114B, to remove color interference and reduce data processing load; Step 2: Based on the mechanical positioning parameters of the production line, determine the column pixel range of each column of paper cups. For example, the image column coordinates of the 4 columns of paper cups are 0-320, 320-640, 640-960, and 960-1280 pixels. Use the connected component analysis algorithm to filter out the independent connected components of each column of paper cups. Set the area threshold to 5000-50000 pixels to remove small background noise. Step 3: Extract images by connecting domain boundary coordinates to obtain a single-column stacked paper cup image, with a single image size of 320×720 pixels. Minimize occlusion and edge interference between adjacent columns of paper cups to provide clean image samples for subsequent feature extraction.

[0029] During the system initialization phase, the calibration and storage of three types of basic data must be completed to ensure that the counting logic can be implemented: Cup rim template: For the current paper cup model, such as a 250ml coated paper cup, collect a standard cup rim image and calibrate the core parameters: paper cup segmentation value (grayscale threshold 120-150, used to distinguish the paper cup from the background), stacking direction (vertical direction, i.e., the y-axis direction of the image), cup rim feature contour (a ring protrusion with a width of 3-5 pixels), and extraction algorithm (template matching + contour tracking). Pixel height database: Sensor calibration was performed using Zhang Zhengyou's camera calibration method, establishing a mapping table of "stacked pixel position - actual stacked height - pixel height value". For example, when the y-coordinate of a stacked pixel is 100, the corresponding actual height is 10cm, and the pixel height value is 100 pixels (pixel-cm conversion factor 10:1). All data has been corrected for lens distortion (distortion error ≤0.1%). Stacking feature data: Key parameters of the current paper cup (in pixels) are obtained through actual measurement and image calibration: cup height data (20 pixels, corresponding to an actual height of 2cm, i.e., the height of the complete cup body at the bottom of the stack), cup rim height data (3 pixels, corresponding to an actual height of 0.3cm, i.e., the height of the exposed rim between adjacent paper cups), and total height of a single cup (23 pixels, the sum of the cup body height and the cup rim height).

[0030] Based on a preset cup rim template, cup rim feature recognition and counting are performed on a single-column stacked image: Image preprocessing: Gaussian filtering (kernel size 3×3) is applied to the single-column stacked image to eliminate noise caused by workshop light reflections and stains on the paper cup surface; Template matching: The preprocessed image is matched with the cup rim template to identify all regions that match the cup rim contour features and mark them as candidate cup rims; Deduplication and filtering: By calculating the center distance of the candidate cup rims (less than 5 pixels is considered a duplicate), duplicate recognition results are eliminated; at the same time, continuous contours with a length ≥ 80% of the paper cup diameter are filtered out to eliminate broken or incomplete cup rim noise; Quantity accumulation: The effective cup rims after screening are counted one by one to obtain the number of cup rims, which is denoted as N1. This value directly reflects the physical quantity of paper cups stacked and is a direct visual counting result.

[0031] The number of paper cups is indirectly calculated by measuring the stacking height, serving as a supplementary verification to direct visual counting. Stacked end pixel extraction: The Canny edge detection algorithm (lower threshold 50, upper threshold 150) is used to detect the edge contour of a single-column stacked image and locate the pixel coordinates of the top of the paper cup stack, denoted as y_max, which is the stacked end pixel position; Stack height matching: The corresponding normalized stack height value, denoted as H_vis, is obtained by looking up y_max in the pixel height database, in pixels. For example, when y_max = 115 pixels, the matched value is H_vis = 115 pixels. Quantity conversion: Substitute into the preset calculation formula to obtain the quantity of the first paper cup, denoted as N2: Formula: N2 = round[(H_vis - H_cup body) / H_cup rim]; Where: H_cupbody = 20 pixels (cup body height data), H_cup rim = 3 pixels (cup rim height data), and round() is the rounding function.

[0032] Example: If H_vis=115 pixels, then N2=round[(115-20) / 3]=round(31.67)=32 pixels.

[0033] A laser distance sensor is installed on the side of the paper cup stack (perpendicular to the stack end face), at a position flush with the detection reference plane of the image sensor. The detection end is vertically aligned with the top end face of the paper cup stack, and the actual compression height of the stack is collected in real time to generate continuous distance data, denoted as D, in cm.

[0034] To ensure consistency with the visual inspection data, the same conversion logic as the first paper cup quantity is adopted: Distance-pixel height conversion: Based on the conversion factor of the pixel height database, 10 pixels / cm, the distance data is converted into a stacking height value, denoted as H_dis, in pixels. The formula is: H_dis = D × 10. Quantity conversion: Substituting the same stacking feature data, we obtain the second number of paper cups, denoted as N3: Formula: N3 = round[(H_dis - H_cup body) / H_cup rim]; Example: If the distance sensor detects D=11.6cm, then H_dis=116 pixels, N3=round[(116-20) / 3]=round(32)=32 pixels.

[0035] Consistency between direct counting in computer vision and visual height conversion: Visual quantity difference calculation: Δ_vis = |N1 - N2|; Preset visual reference difference: Based on the counting accuracy requirements of the production line, such as an allowable error of ≤2, set Δ_ref=2.

[0036] Scenario 1: Δ_vis < Δ_ref (good consistency of visual data): This indicates that the image has no severe occlusion, deformation, or noise interference, and the vision system is working stably. The weighted fusion of "number of cup rims + number of first paper cups" fully leverages the advantages of both types of visual data. Weighted fusion formula: N_final = w1 × N1 + w2 × N2, where w1 + w2 = 1; Weighting: Weights are assigned based on data reliability. Counting directly from the cup rim is more intuitive (weight w1=0.6), and height conversion is less affected by deformation (weight w2=0.4). Example: If N1=32, N2=31, and Δ_vis=1<2, then N_final=0.6×32+0.4×31=31.6≈32.

[0037] Scenario 2: Δ_vis ≥ Δ_ref (visual data contains interference): This indicates that cup rim recognition may be affected by factors such as partial occlusion, reflection, and cup offset, leading to a decrease in the reliability of single visual data. Therefore, a dual-source fusion approach of "first cup count + second cup count" is adopted, utilizing the complementarity of visual and non-visual data. Weighted fusion formula: N_final = w3 × N2 + w4 × N3, where w3 + w4 = 1; Weighting: The default is equal weighting (w3=0.5, w4=0.5). If a data source fluctuates significantly, such as a fluctuation value of >5% for 5 consecutive frames, its weight can be dynamically reduced (down to a minimum of 0.2). Example: If N1=35, N2=32, Δ_vis=3≥2, N3=33, then N_final=0.5×32+0.5×33=32.5≈33.

[0038] This method completely solves the core defects of traditional non-visual methods by using a multi-source data fusion architecture of machine vision dual-feature recognition and distance sensing redundancy verification. On the one hand, visual recognition adopts non-contact detection, which reduces counting deviation caused by occlusion. On the other hand, dual-source fusion and dynamic weight adjustment enhance the system's anti-interference ability and reduce the counting error rate under high-speed production conditions.

[0039] The specific steps for obtaining the number of cup rims and the number of the first paper cups based on machine vision are as follows: Before starting the counting process, three types of data need to be calibrated and stored to provide accurate data for cup rim extraction and quantity conversion: Cup rim template library: For the currently produced paper cup model (e.g., 300ml coated paper cup, cup mouth diameter 68mm), 100 sets of standard stacked images are collected for template training, and two types of cup rim templates are stored. The parameters are calibrated as follows: Template 1 (for graded threshold segmentation): Paper cup segmentation value = 130 (grayscale range 0-255, to distinguish paper cup from background), stacking direction = vertical direction (image y-axis), cup body graphic = standard cup body outline (height 65 pixels, cup mouth diameter 136 pixels), cup rim segmentation value = 80 (to distinguish cup rim from cup body), inner edge distance = 5 pixels (safe distance between cup rim and outer edge); Template 2 (for extracting double-sided semicircles): Paper cup segmentation value = 130, stacking direction = vertical, cup rim graphic = standard semicircle outline (radius 4 pixels, arc 180°), axisymmetric graphic = mirrored semicircle outline; Pixel height database: Based on camera calibration results (pixel-millimeter conversion factor = 2 pixels / mm), a mapping table of "stacked end pixel position (y coordinate) - stacked height value (pixel)" is established. For example, when the y coordinate = 200 pixels, the corresponding actual height is 100mm and the stacked height value = 200 pixels. All data have undergone radial distortion correction (correction error ≤ 0.03mm). Stacking Feature Database: Core parameters (in pixels) are calibrated by combining image measurement and physical measurement: Cup height data = 50 pixels (corresponding to an actual height of 25mm, i.e., the non-exposed height of the complete cup body at the bottom of the stack), Cup rim height data = 4 pixels (corresponding to an actual height of 2mm, i.e., the height of the exposed rim between adjacent paper cups), Total height of a single cup = 54 pixels (the sum of cup body height and cup rim height).

[0040] The extraction of the cup rim count includes the following two sub-methods, with the specific steps as follows: Sub-scheme 1: Hierarchical threshold segmentation + cup / edge interference removal method Reference Figure 2 This scheme uses a three-stage process of coarse segmentation → interference removal → fine segmentation to accurately locate the parallel cup edge features. It is suitable for situations where paper cups have slight deformation and a lot of edge noise. The specific sub-steps are as follows: Grayscale conversion: Convert the color image (1920×1080 pixels) of a single stack of paper cups into a single-channel grayscale image using the international standard weighting formula: Gray=0.299R+0.587G+0.114B (R, G, and B are the pixel values ​​of the three channels of the color image), to eliminate color interference and reduce the amount of data processing. Coarse segmentation to extract the paper cup region: Using "paper cup segmentation value = 130" in the cup rim template as the threshold, global binarization is performed: pixels with grayscale values ​​> 130 are marked as the paper cup foreground (pixel value = 255), and pixels with grayscale values ​​≤ 130 are marked as the background (pixel value = 0), resulting in a paper cup image containing only multiple paper cup aggregation areas, thoroughly filtering out background interference such as conveyor belts, racks, and light reflections; Cup shape removal: In the top area of ​​the paper cup image along the stacking direction (vertical direction) (occupying 1 / 5 of the image height), the normalized cross-correlation template matching algorithm is used to slide the image with the pre-stored "cup shape". The matching threshold is 0.85, that is, the similarity is ≥85% and the matching is considered successful. The pixels of the successfully matched cup area are set to 0 (black) to completely remove the interference of the overall cup outline on the extraction of the cup edge; Edge noise removal: According to the preset inner edge distance = 5 pixels, along the left and right outer edges of the paper cup image, the edge area with a width of 5 pixels is cropped and its pixels are set to 0. This removes noise such as rough edges, extrusion deformation, surface stains, and paper scraps from the outer edge of the paper cup, resulting in an intermediate image that retains only the features of the cup rim. Fine segmentation to extract parallel cup rims: Using the cup rim segmentation value = 80 in the cup rim template as the threshold, a second binarization process is performed on the intermediate image to highlight the grayscale difference between the cup rim and the cup body; then, the Hough line detection algorithm is used with parameters: rho = 1 pixel, theta = π / 180 radians, threshold = 30, to extract all horizontally parallel cup rim lines in the image, and those with an angle deviation ≤ 3° are judged as parallel; Cumulative cup rim count: The cup rim lines that are repeatedly identified are removed by using a connected component deduplication algorithm (adjacent lines with a spacing of less than 2 pixels are judged as duplicates); the number of cup rims after filtering is counted one by one in each round, and the cup rim count is recorded as N1.

[0041] Sub-scheme 2: Bilateral symmetrical semicircle extraction + bidirectional quantity verification method Reference Figure 3This scheme utilizes the strictly axially symmetrical physical property of the paper cup's left and right structures. By extracting and cross-validating the features of both semicircles, it solves the problem of single-sided detection being susceptible to occlusion and offset interference. The specific sub-steps are as follows: Grayscale conversion and coarse segmentation: Repeat steps 1-2 of sub-scheme one to complete grayscale conversion and binarization coarse segmentation, and obtain a clean paper cup image containing only the paper cup aggregation area without background interference; Left semicircle feature extraction: Along the stacking direction (vertical direction), in the left region of the paper cup image (occupying 1 / 2 of the image width), the Hough circle transform algorithm (parameters: minimum radius = 3 pixels, maximum radius = 5 pixels, center-to-center distance ≥ 6 pixels, cumulative threshold = 25) is used to extract all first semicircle shapes that meet the size requirements based on the pre-stored "cup rim graphic (standard semicircle)". The first number is obtained by counting the connected components and is denoted as N11. Right semicircle feature extraction: In the right region of the paper cup image (occupying 1 / 2 of the image width), the axisymmetric shape of the cup rim (mirror semicircle) is used as the matching template. The Hough circle transformation operation is repeated to extract all the second semicircle shapes that meet the requirements. The second quantity is counted and denoted as N12. Two-way quantity verification and counting: Calculate the absolute value of the difference between N11 and N12, |ΔN_1|=|N11-N12|. If |ΔN_1|=0, the quantities on both sides are completely consistent, and it is determined that there is no local occlusion, stacking offset or distortion. N11 or N12 is taken as the effective cup edge quantity N1. If |ΔN_1|>0, it is determined that there is interference on one side, triggering the "secondary extraction mechanism". Re-acquire the current stack image and repeat the above steps. If the extraction is still inconsistent after 3 consecutive extractions, the extraction result of sub-scheme 1 is temporarily used as N1.

[0042] After calculating the number of cups at the rim, calculate the number of the first paper cup. The specific steps are as follows: Stacked edge pixel extraction: The Canny edge detection algorithm is used with a lower threshold of 60, an upper threshold of 180, and an apertureSize of 3 to detect the edge contour of a single column of stacked paper cups. The continuous edge at the top of the stack is located, and the center pixel coordinate (y_max) of the edge is taken as the stacked edge pixel. For example, if the detected y coordinate range of the top edge is 198-202 pixels, y_max is set to 200 pixels. Stacking height value matching: Based on y_max=200 pixels, the corresponding standardized stacking height value H_vis=200 pixels is obtained by looking up the table in the preset pixel height database, after deducting the influence of lens distortion; Precise quantity conversion: Retrieve the stacking feature data of the current paper cups, with cup height H_cup_body = 50 pixels and rim height H_rim_4 pixels, and substitute them into the preset calculation formula: N2 = round[(H_vis - H_cup_body) / H_cup_rim]; The `round()` function is used to round to the nearest integer to eliminate calculation errors. Example: If H_vis=200 pixels, then N2=round[(200-50) / 4]=round(37.5)=38, that is, the number of the first paper cups is 38.

[0043] The number of the second paper cup is determined using a distance sensor. The specific steps are as follows: A KeyenceIL-1000 laser distance sensor is installed on the side of the paper cup stack (perpendicular to the stack end face), flush with the detection reference plane of the image sensor. The measurement range is 20-100mm, the measurement accuracy is ±0.05mm, and the response time is ≤5ms. The detection end is vertically aligned with the top end face of the paper cup stack. The actual compression height of the stack is collected in real time at a frequency of 40Hz, generating continuous distance data D in mm, which is synchronously transmitted to the host computer via RS485 interface.

[0044] To ensure consistency with the conversion standard of visual inspection data, the logic is completely consistent with the number of the first paper cups: Distance-pixel height conversion: Based on the conversion factor (2 pixels / mm) of the pixel height database, the distance data is converted into a stacking height value H_dis (unit: pixels), and the formula is: H_dis=D×2; Example: If the distance sensor detects D=98mm, then H_dis=98×2=196 pixels; Quantity conversion: Substituting the same stacking feature data, we obtain the second paper cup quantity N3: N3 = round[(H_dis - H_cup_body) / H_cup_rim]; Example: N3 = round[(196-50) / 4] = round(36.5) = 37.

[0045] The adaptive fusion calculation of the final paper cup stack count follows these steps: 1. Visual quantity difference judgment Calculate the consistency between the number of cup rims N1 and the number of first paper cups N2: Δ_vis=|N1-N2|, preset visual reference difference Δ_ref=2, the specific value is set according to the counting accuracy requirements of the production line, and the allowable error is ≤2.

[0046] 2. Scene-specific fusion strategy Scenario 1: Δ_vis < Δ_ref (good consistency of visual data): This indicates that the image has no severe occlusion, deformation, or noise interference, and the vision system is working stably. A weighted fusion method using "number of cup rims + number of first paper cups" is employed. N_final = w1×N1 + w2×N2, w1 + w2 = 1, where w1 = 0.6 (direct counting along the cup rim is more reliable) and w2 = 0.4 (height conversion provides stronger resistance to deformation). Example: If N1=38, N2=37, and Δ_vis=1<2, then N_final=0.6×38+0.4×37=37.6≈38.

[0047] Scenario 2: Δ_vis ≥ Δ_ref (visual data contains interference): This indicates that cup rim recognition may be affected by factors such as partial occlusion and reflection, so we switched to dual-source fusion of "first paper cup count + second paper cup count": N_final = w3 × N2 + w4 × N3, w3 + w4 = 1, default equal weight allocation (w3 = 0.5, w4 = 0.5). Example: If N1=40, N2=37, Δ_vis=3≥2, N3=38, then N_final=0.5×37+0.5×38=37.5≈38.

[0048] 1. Two cup rim extraction sub-schemes can be adaptively switched according to working conditions: when the paper cup deformation on the production line is small, sub-scheme one is preferred (high computing efficiency, single frame processing time ≤10ms); when there is a lot of stacking offset or occlusion, it automatically switches to sub-scheme two (strong anti-interference ability), balancing efficiency and accuracy. 2. The design of hierarchical threshold segmentation (paper cup segmentation value > cup rim segmentation value) and bilateral symmetrical verification fundamentally solves the pain points of "feature confusion" and "one-sided missed detection" in traditional visual recognition, with a cup rim extraction accuracy of ≥99.5%; 3. The dual-source fusion of visual data and distance sensing data enables a counting error rate of ≤0.3% under high-speed conditions (150 cups / minute), which is far superior to traditional non-visual methods (error rate 8%-12%) and meets the uniform packaging specifications requirements for large-scale paper cup production.

[0049] In this embodiment, the online recognition method for the number of stacked paper cups based on machine vision includes the following specific steps: Utilizing the strictly symmetrical physical structure of paper cups along their stacking axis, this method achieves feature superposition and interference cancellation on both sides of the cup edge through image folding. It is suitable for high-speed production lines (≥150 cups / minute) and applications with unilateral image distortion or slight local occlusion. Its advantages include low algorithm complexity, fast processing speed, and high recognition accuracy. The specific sub-steps are as follows: The core parameters of the "double-sided semicircle extraction template" in the cup rim template library are used, namely, paper cup segmentation value = 130 (grayscale range 0-255), stacking direction = vertical direction (image y-axis), and cup rim graphic = standard semicircle outline (radius 4 pixels, arc 180°), to ensure the consistency of parameters with the previous sub-scheme and reduce recognition fluctuations caused by parameter differences.

[0050] Repeat the grayscale conversion step, using the standard weighted formula Gray=0.299R+0.587G+0.114B, to convert the single-column color stacked image into a single-channel grayscale image and remove color interference; Global binarization is performed with a paper cup segmentation value of 130 as the threshold. Areas with a grayscale value greater than 130 are marked as the paper cup foreground (pixel value = 255), and the rest are the background (pixel value = 0). Morphological closing operations (5×5 circular structuring elements) are used to fill the tiny holes (area < 100 pixels) in the paper cup area. Then, connected component filtering (area ≥ 8000 pixels, aspect ratio 3:1-5:1) is used to obtain a paper cup image with no background interference and a complete outline.

[0051] The stacking direction center axis of the paper cup image is determined by calculating the centroid of the connected components; the effective connected components of the paper cup image are traversed, and the average x-coordinate x_center of all foreground pixels is calculated. The line x=x_center is taken as the axis of symmetry of the stacking direction (perpendicular to the stacking direction and parallel to the cup rim extension direction). Fold the paper cup image symmetrically along the axis x=x_center and merge them. Using the axis as the boundary, mirror the left image area (x<x_center) along the axis and then overlay it with the right image area (x≥x_center) pixel by pixel. The overlay rule is: retain 255 when both sides are 255, and take 255 when either side is 0. This achieves complementary occlusion areas and generates a folded image with the size of 1 / 2 of the original image width × the original image height, such as 300×1080 pixels → 150×1080 pixels.

[0052] Folding effect: Distortions on one side of the image (such as the tilt of the left cup rim) will be canceled out by the normal features after mirroring on the other side. Locally occluded areas (such as 10% occlusion on the right side) will be filled by the superposition on the left side. Background noise is weakened due to the improved signal-to-noise ratio after symmetrical superposition, resulting in a folded image with more concentrated features and less interference.

[0053] Gaussian filtering (kernel size 3×3, standard deviation σ=1.0) is applied to the folded image to further smooth noise and enhance the edges of the semicircular contour; An improved Hough circle transform algorithm (optimizing parameters for superimposed features after folding) is adopted: the minimum radius is set to 3 pixels, the maximum radius to 5 pixels, the center-to-center distance is ≥6 pixels, and the cumulative threshold is set to 28 (2 lower than sub-scheme 2, to adapt to the feature intensity after superposition). Based on the preset cup rim graphic (standard semicircle), all semicircular shapes that meet the size and shape requirements in the folded image are extracted. Feature filtering: Remove discrete semicircles with a center y-coordinate deviation > 5 pixels (excluding non-cup rim noise), retain semicircle features evenly distributed along the stacking direction (y-axis), and ensure that the extraction results correspond only to the rim of the paper cup.

[0054] The effective semicircular features after screening are counted one by one in each round, and the statistical results are directly used as the number of cup rims N1. Special scenario handling: If a semicircular feature is not detected at a certain height position in the folded image, it is determined that there is severe occlusion on both sides at that position, triggering the "single frame re-sampling mechanism"; re-acquire the current stack image and repeat the above steps. If it is still not detected after 2 consecutive attempts, it will automatically switch to sub-scheme 2 (independent extraction on both sides) for complementary recognition to reduce missed detections.

[0055] Adaptive switching logic for cup rim extraction scheme: The system has a built-in scheme switching mechanism that dynamically selects the optimal cup rim extraction sub-scheme based on the real-time operating conditions of the production line, ensuring a balance between efficiency and accuracy across all scenarios. 1. Operating condition judgment indicators: processing time for real-time acquisition of 10 consecutive frames of images, success rate of semicircle extraction (number of effective semicircles / theoretical number of cup edges), and number of occlusion detections; 2. Switching rules: When the processing time is ≤10ms, the extraction success rate is ≥98%, and the number of occlusion detections is 0, sub-scheme 1 (hierarchical threshold segmentation) is preferred to balance efficiency and accuracy. When the extraction success rate is <98% and the number of occlusion detections is ≥3, switch to sub-scheme 2 (double-sided verification) to enhance the anti-occlusion capability. When the production line speed is ≥150 pieces / minute and the processing time requirement is ≤8ms, the system will be forced to switch to sub-scheme 3 (image folding) to prioritize real-time performance. The switching between the three schemes is seamless and does not affect the continuous operation of the production line. During the switching process, the number of valid cup edges from the previous frame is used as a temporary measure to reduce counting interruptions.

[0056] In this embodiment, regardless of which cup rim extraction sub-scheme is used, the calculation logic for the first paper cup quantity N2 remains completely consistent. Only the following steps need to be followed based on the cup rim quantity N1 extracted by sub-scheme three: 1. Pixel extraction at the stacked edge: The Canny edge detection algorithm (lower threshold = 60, upper threshold = 180) is used to locate the center pixel coordinates y_max of the top edge of the stack. 2. Stacking height value matching: H_vis (distortion corrected) is obtained by looking up a table in the pixel height database; 3. Quantity conversion: Substitute into the formula N2=round[(H_vis-H_cupbody) / H_cuprim], where H_cupbody=50 pixels and H_cuprim=4 pixels, and round the result.

[0057] Example: If sub-scheme 3 extracts N1=39, stacked end pixel y_max=202 pixels, and matches H_vis=202 pixels, then N2=round[(202-50) / 4]=round(38)=38, the visual quantity difference Δ_vis=|39-38|=1<Δ_ref=2, and enters scene 1 weighted fusion.

[0058] In this embodiment, the online recognition method for the number of stacked paper cups based on machine vision specifically includes the following steps: Stacking feature data is the core benchmark for quantity conversion, including stacking direction, cup height data, and cup rim height data: Stacking direction: Consistent with the physical direction of paper cup stacking on the production line, calibrated as the vertical direction of the image, positive y-axis direction, from the bottom of the stack to the top, to ensure that the height measurement is consistent with the physical stacking dimension; Cup height data H_cup body: refers to the standardized pixel height of the "non-exposed cup body part" of the paper cup at the bottom when stacked; since the bottom of the paper cup is most squeezed when stacked, the cup body is intact and the cup rim is not exposed, this parameter reflects the core body height of a single paper cup; Cup rim height data H_cup rim: refers to the standardized pixel height of the "exposed cup rim" of the paper cup located in the middle area when stacked; the middle paper cup is uniformly compressed and the exposed cup rim shape is stable, which is the core distinguishing feature between adjacent paper cups, and its height is the cumulative height contributed by a single paper cup; The total height of a single cup, H_single cup: H_single cup = H_cup body + H_cup rim, reflects the actual height occupied by a single paper cup in a stacked state (including the cup body and the exposed rim), and is used to verify the rationality of the conversion logic.

[0059] The calibration process and specific values ​​are illustrated using a 300ml coated paper cup as an example: 1. Physical measurement: Select 100 qualified paper cups, stack them, and use a micrometer to measure the actual height of the bottom paper cup (excluding the exposed rim) as 25mm, the actual height of the exposed rim of the middle paper cup as 2mm, and the total actual height of a single cup as 27mm. 2. Image calibration: Based on the camera calibration results (pixel-millimeter conversion factor = 2 pixels / mm), combined with the imaging deformation correction algorithm, the actual height is converted into standardized pixel data: H_cup body = 25mm × 2 = 50 pixels, H_cup rim = 2mm × 2 = 4 pixels, H_single cup = 27mm × 2 = 54 pixels; 3. Data storage: The calibrated parameters are entered into the stacking feature database and stored according to the paper cup model, such as 300ml-coated-68mm cup mouth, supporting quick switching between multiple models.

[0060] The stacking height, cup height, and cup rim height all require rigorous image distortion correction to eliminate dimensional errors caused by lens radial distortion, tangential distortion, and shooting angle deviation. Specific correction steps are as follows: Correction objects and sources of error: Radial distortion: Image edge stretching caused by lens optical characteristics, such as pixel offset at the edge of a cup, with an error range of ±1.5 pixels; Tangential distortion: The height measurement deviation caused by slight tilt of the camera mounting. For example, if the actual stack height is 100mm, the uncorrected image displays 98mm, with an error range of ±2 pixels. Correction objective: To control the deformation error of all highly correlated pixel data within ±0.3 pixels to ensure conversion accuracy.

[0061] Correction methods and implementation details: 1. Based on the camera intrinsic parameters (focal length f_x=1800 pixels, f_y=1800 pixels, principal point coordinates (u_0=960, v_0=540)) and distortion coefficients (radial distortion coefficients k_1=-0.012, k_2=0.003, tangential distortion coefficients p_1=0.001, p_2=0.0008) obtained by Zhang Zhengyou's planar calibration method, a distortion correction model is established. 2. Perform distortion correction on all foreground pixels (paper cup area) in a single-column stacked image: calculate the mapping relationship between the original coordinates and the corrected coordinates of each pixel through the correction model, and use bilinear interpolation to supplement the corrected pixel values ​​to obtain a distortion-free standardized image; 3. Stacking height value correction: Extract the stacking end pixel y_max of the corrected image, substitute it into the pixel height database, and match to obtain the corrected stacking height value H_vis. For example, if the uncorrected y_max = 203 pixels, the corrected y_max = 200 pixels, and the corresponding H_vis = 200 pixels. 4. Feature data correction: The cup height data and cup rim height data have been corrected for distortion during the calibration stage and can be directly used without repeated processing.

[0062] Based on the corrected height data and stacking feature data, the number of the first paper cups is calculated through a split-cancel-convert logic. The specific steps are as follows: Calculation of temporary value for the middle part: The stacking height value H_vis includes "the complete cup height of the bottom paper cup" and "the sum of the exposed cup rim heights of all paper cups"; the bottom paper cup has no exposed rim (it is covered by the paper cup above), and the exposed cup rims of the middle and top paper cups are stacked to form the core height, so the bottom cup height needs to be split first: H_temporary = H_vis - H_cup body; Example: If H_vis = 200 pixels and H_cup body = 50 pixels, then H_temporary = 200 - 50 = 150 pixels, which is the sum of the exposed cup rim heights of all paper cups.

[0063] First paper cup quantity conversion: Each paper cup contributes 1 exposed rim height H_rim, so divide the temporary value of the middle part by the single rim height to get the first paper cup quantity (rounded to the nearest integer to eliminate decimal errors): N2=round(H_temporary / H_rim); Example: H_temporary=150 pixels, H_rim=4 pixels, then N2=round(150 / 4)=round(37.5)=38 cups.

[0064] Logical verification: Verify the reasonableness of the conversion result using the total height of a single cup: H_verification = H_cup body + N2 × H_cup rim. If |H_verification - H_vis| ≤ 1 pixel, the conversion is considered valid; otherwise, a recalibration prompt is triggered. For example, if H_verification = 50 + 38 × 4 = 202 pixels, and |202 - 200| = 2 pixels, it is necessary to check whether the stacking height value is extracted accurately.

[0065] After this sub-step is refined, the calculation of the number of the first paper cup forms a closed-loop process of "parameter calibration → deformation correction → split conversion → logic verification", which is seamlessly connected with the previous cup rim extraction sub-scheme (three types): no matter which cup rim extraction method is used to obtain the number of cup rims N1, N2 can be obtained through this refined conversion process, and then enter the "visual quantity difference judgment → adaptive fusion counting" stage to ensure the accuracy and stability of the entire process.

[0066] Taking sub-scheme three (image folding extraction) as an example, here is an example: 1. Sub-scheme three extracts N1=38 items; 2. After pixel correction at the stacking end, y_max=200 pixels, matching H_vis=200 pixels; 3. Convert H_temporary = 200 - 50 = 150 pixels, N2 = round(150 / 4) = 38; 4. The visual difference Δ_vis=|38-38|=0<Δ_ref=2, and N_final=0.6×38+0.4×38=38, with a counting error of 0.

[0067] In this embodiment, the adaptive fusion calculation of the final paper cup stack count includes the following sub-steps: Scenario 1: Dynamic weighted fusion of the number of cup rims and the number of first paper cups When the visual quantity difference Δ_vis < Δ_ref, such as Δ_ref = 2, it indicates that the direct recognition of the cup rim and the visual height conversion data are consistent. However, during the production process, there may be fluctuations in the stacking height caused by changes in the tightness of the paper cup compression and slight vibrations of the production line. It is necessary to offset the fluctuation deviation through dynamic weight adjustment. The specific sub-steps are as follows: Basic parameter definition and initial weight setting: Continuous data acquisition frame number K: set to 10 frames, balancing fluctuation detection sensitivity and real-time performance. 10 frames of data can reflect short-term operating condition fluctuations. The processing time of a single frame is ≤9ms, and the total time for 10 frames is ≤90ms, which does not affect the production line cycle time. Fluctuation value calculation method: The coefficient of variation (CV) is used to characterize the degree of fluctuation, eliminating the influence of the stacking height order of magnitude, and reflecting only the relative fluctuation. The formula is as follows: Where σ is the standard deviation of the stacking height values ​​of consecutive K frames. This is the average value; The preset fluctuation reference value CV_ref is calibrated to 1.5% based on the actual measured data of the production line. Under stable working conditions during high-speed production, the stacking height fluctuation is generally ≤1.5%. If it exceeds this value, it is judged as an abnormal fluctuation. Initial weight setting: Based on the reliability priority of the two types of data, the initial weight allocation is as follows: the weight of the number of cup rims w1=0.6, which is a direct and intuitive feature with few sources of error; the weight of the first paper cup number w2=0.4, which is slightly affected by compression deformation in height conversion and satisfies w1+w2=1.

[0068] Stack height fluctuation calculation: Data acquisition: Extract the stacking height values ​​H_{vis,1}, H_{vis,2}, ..., H_{vis,10} corresponding to the number of paper cups in the first 10 consecutive frames. These are all corrected pixel values, such as 200, 201, 199, 202, 198, 201, 200, 199, 201, 200 pixels. Statistical calculations: Calculate the average value _vis = (200 + 201 + 199 + 202 + 198 + 201 + 200 + 199 + 201 + 200) / 10 = 200.1 pixels; Calculate the standard deviation σ_vis = sqrt{[(200-200.1)] 2 +(201-200.1) 2 +...+(200-200.1)2 ] / (10-1)}≈1.05 pixels; The fluctuation value CV_vis is calculated as follows: CV_vis = (1.05 / 200.1) × 100% ≈ 0.52%.

[0069] Judgment criterion: Compare CV_vis with CV_ref = 1.5%; If CV_vis≤CV_ref, such as 0.52%≤1.5% in the example: it indicates that the stacking height is stable, the reliability of the first paper cup quantity is high, and the initial weights w1=0.6 and w2=0.4 are maintained; If CV_vis > CV_ref, such as CV_vis = 2.3%, it indicates that the height fluctuates significantly, and the number of first paper cups is affected by the fluctuation. The weights are adjusted according to the following linear rule: Weight adjustment formula: w2'=w2-(CV_vis-CV_ref)×0.05. For every 1 percentage point exceeding the reference value, the weight is reduced by 0.05 to ensure a smooth adjustment. Weight constraints: w2'≥0.1, the lower limit prevents data invalidation due to excessively low weights; w2'≤0.4, the upper limit does not exceed the initial weights. Synchronously adjust w1' = 1 - w2'; Example: If CV_vis=2.3%, then w2'=0.4-(2.3%-1.5%)×0.05=0.4-0.04=0.36, w1'=0.64.

[0070] The final count is calculated using a weighted fusion formula: N_final = w1 × N1 + w2 × N2, rounded to the nearest integer. Example 1 (Stable operating condition): N1=38, N2=38, w1=0.6, w2=0.4, then N_final=0.6×38+0.4×38=38; Example 2 (fluctuating operating condition): N1=38, N2=37, CV_vis=2.3%, w1'=0.64, w2'=0.36, then N_final=0.64×38+0.36×37=37.64≈38.

[0071] Scenario 2: Weighted fusion of the fluctuations in the number of the first and second paper cups When Δ_vis ≥ Δ_ref, visual data is subject to interference. It is necessary to assign weights by comparing the fluctuations of visual and distance sensing dual-source data to strengthen the role of stable data sources. The specific sub-steps are as follows: Define the number of continuous data acquisition frames K as 10 frames. Fluctuation value types: First fluctuation value CV1 (coefficient of variation of visual stacking height) and second fluctuation value CV2 (coefficient of variation of distance sensor stacking height), both expressed as percentages (%); The fluctuation ratio k = CV1 / CV2 represents the relative stability of the two types of data. The larger the k is, the more unstable the visual data is. Initial weight settings: By default, w3=0.5 (weight of the first paper cup quantity) and w4=0.5 (weight of the second paper cup quantity) are assigned equal weights to ensure that the two sources of data contribute equally in the initial state.

[0072] Calculate the visual side fluctuation value (CV1): Collect the visual stacking height values ​​H_{vis,1}-H_{vis,10} for 10 consecutive frames, such as: 200, 203, 197, 204, 196, 202, 199, 205, 195, 201 pixels; calculate: Pixels, σ1≈3.5 pixels, CV1=(3.5 / 200.2)×100%≈1.75%; Calculate the distance sensor side fluctuation value (CV2): Collect 10 consecutive frames of distance sensor stacking height values ​​H_{dis,1}-H_{dis,10}, corresponding to the pixel values ​​after actual distance conversion, such as: 198, 199, 200, 197, 201, 199, 202, 198, 200, 199 pixels; calculate: For pixels, σ² ≈ 1.5 pixels, CV² = (1.5 / 199.3) × 100% ≈ 0.75%; Fluctuation ratio: k=CV1 / CV2=1.75% / 0.75%≈2.33.

[0073] Adaptive weighting based on volatility ratio: Adjustment rules: Based on the volatility ratio k, w3 and w4 are dynamically allocated, following the principle of "the greater the volatility, the smaller the weight". When k≤1 (CV1≤CV2, visual data is more stable): keep the initial weight w3=0.5, or fine-tune it to w3=0.5+(1-k)×0.1 (maximum not exceeding 0.8). When k > 1 (CV1 > CV2, distance sensing data is more stable): adjust according to the formula w3' = 0.5 - (k - 1) × 0.08, with the lower limit of weight w3' ≥ 0.2 (to reduce the complete failure of visual data). Example calculation: k=2.33>1, then w3'=0.5-(2.33-1)×0.08=0.394≈0.39, w4'=1-0.39=0.61.

[0074] Dual-source weighted fusion computing: The fusion formula is: N_final = w3 × N2 + w4 × N3, rounded to the nearest integer. Example: N2=37 (visual height conversion), N3=38 (distance sensing conversion), w3'=0.39, w4'=0.61, then N_final=0.39×37+0.61×38=37.61≈38.

[0075] Special scene handling: If the fluctuation ratio k≥3 (visual data is extremely unstable, such as CV1=4.5%, CV2=1.5%, k=3): then w3'=0.5-(3-1)×0.08=0.34, further weakening the weight of visual data; If the fluctuation ratio k≤0.5 (distance sensing data is extremely unstable, such as CV1=0.5%, CV2=1.0%, k=0.5): then w3'=0.5+(1-0.5)×0.1=0.55, strengthening the weight of visual data.

[0076] In this embodiment, the online recognition method for the number of stacked paper cups based on machine vision includes the following steps: By cross-checking three independent sets of data—the number of cups along the rim N1, the number of first paper cups N2, and the number of second paper cups N3—problems such as sensor hardware failures, algorithm logic anomalies, and data transmission interruptions can be quickly identified. This reduces batch counting errors caused by equipment failures and ensures packaging quality on the production line. The specific steps are as follows: Define the first difference Δ1: the deviation between direct counting of the cup rim and conversion of visual height, Δ1=|N1-N2|; Define the second difference Δ2: the deviation between direct counting on the cup rim and conversion by distance sensing, Δ2=|N1-N3|; Define the third difference Δ3 as the deviation between visual height conversion and distance sensing conversion, Δ3 = |N2 - N3|.

[0077] The reference difference needs to be calibrated comprehensively based on the production line's counting accuracy requirements and the inherent errors of the equipment to ensure that "normal fluctuations do not trigger alarms, but abnormal faults will trigger alarms": Calibration criteria: Traditional production lines allow a counting error of ≤1, and vision and distance sensing devices have an inherent error of ≤0.5. The overall setting is Δ_alarm=2, meaning that if the difference in all three sets exceeds 2, it is considered abnormal. Calibration process: Through the statistical analysis of 1000 sets of counting data under normal operating conditions, the maximum value of the difference between the three sets is ≤1.8. Therefore, Δ_alarm=2 can effectively distinguish between normal fluctuations and abnormal faults, with a false alarm rate of ≤0.3%.

[0078] For example: Example 1 (Normal operating conditions): N1=38, N2=38, N3=37; Δ1=|38-38|=0, Δ2=|38-37|=1, Δ3=|38-37|=1; All three differences are ≤ Δ_alarm = 2, which is considered normal.

[0079] Example 2 (Image sensor failure): The image sensor lens is dirty, causing the cup rim recognition to fail, N1=33; visual height extraction is affected, N2=34; the distance sensor is normal, N3=38; Δ1=|33-34|=1, Δ2=|33-38|=5, Δ3=|34-38|=4; Although Δ1=1≤2, Δ2 and Δ3 are both>2, so no alarm will be triggered for now. The condition of "both are greater than" must be met.

[0080] Example 3 (simultaneous failure of both sensors): Image sensor shifted (N1=42, N2=43), distance sensor malfunctioned (N3=37). Δ1=|42-43|=1, Δ2=|42-37|=5, Δ3=|43-37|=6; If the condition is still not met ("all greater than"), no alarm will be triggered.

[0081] Example 4 (Algorithm anomaly + sensor failure): Cup rim extraction algorithm crashes (N1=30), visual height conversion error (N2=35), distance sensor data drift (N3=28). Δ1=|30-35|=5, Δ2=|30-28|=2 (equal to the reference difference, does not meet the "greater than"), Δ3=|35-28|=7; They still haven't called the police.

[0082] Example 5 (System-wide anomaly): Image sensor has no image output (N1=0), visual height extraction failed (N2=5), distance sensor disconnected (N3=45); Δ1=|0-5|=5, Δ2=|0-45|=45, Δ3=|5-45|=40; The difference between the three sets of values ​​is greater than Δ_alarm=2, which is considered abnormal and triggers an alarm.

[0083] Strict judgment criteria: An abnormality in the counting sensor is only determined when the first difference is greater than Δ_alarm, the second difference is greater than Δ_alarm, and the third difference is greater than Δ_alarm, thus reducing false alarms caused by fluctuations in a single data source.

[0084] When the system triggers an alarm, it simultaneously notifies operators through multiple channels to ensure a rapid response. Audible and visual alarm: Industrial alarm lights (constantly red) and buzzers (frequency 2Hz, volume ≥85dB) are installed on the production line to continuously alarm until the fault is cleared; The host computer pop-up window shows a red fault message box in the production monitoring system, displaying "Counting sensor abnormality", "Abnormal difference: Δ1=×, Δ2=×, Δ3=×" and "Suggested troubleshooting: image sensor, distance sensor, algorithm parameters". Stop Interlock: If the alarm lasts for more than 10 seconds (customizable), the system will automatically send a signal to the production line PLC to suspend the paper cup transfer and packaging process to prevent batches of unqualified products from being released. Data logging: Automatically stores the time of anomaly occurrence, three sets of count data, differences, and equipment operation logs (such as sensor voltage and data transmission delay) to facilitate subsequent fault tracing.

[0085] After an alarm is triggered, troubleshooting and system reset should be performed, with the following priority: 1. Hardware Troubleshooting: Check if the image sensor lens is clean, if it is misaligned, and if the power supply and network cable are properly connected; check if the distance sensor detection end is blocked, if the signal cable is loose, and if the power supply is stable. 2. Software Troubleshooting: Verify that the cup rim template, pixel height database, and stacking feature data match the current paper cup model; check that the visual processing algorithm is running normally and that there are no program crash messages. Reset Procedure: After troubleshooting, the operator clicks "Reset Alarm" in the monitoring system. The system automatically collects 3 frames of new data and recalculates. If the difference between the three sets of data is ≤ Δ_alarm = 2, the alarm is cleared and the production line resumes operation.

[0086] This application also discloses an online recognition system for the number of stacked paper cups based on machine vision, including a processor, wherein the processor executes the steps of the online recognition method for the number of stacked paper cups based on machine vision as described in any of the above embodiments.

[0087] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A machine vision-based online identification method for the number of stacked paper cups, characterized in that, Includes the following steps: An image sensor is placed next to the stacked paper cups that are about to be transported. The image sensor is used to capture images of the stacked and compressed paper cups and generate images of the paper cups. Multiple rows of paper cups are segmented from the paper cup images to obtain stacked images of a single row of paper cups. The number of cup edges is obtained by extracting cup edges from the stacked image based on the preset cup edge template and accumulating the number. The stacking position of the stacking end of the single column of paper cups in the stacked image is extracted to obtain the stacking end pixel position. The corresponding value is matched from the preset pixel height database according to the stacking end pixel position as the stacking height value. The stacking feature data corresponding to the paper cup type is obtained. The number of the first paper cup is calculated according to the stacking height value and the stacking feature data. A distance sensor is set in the stacking direction of the paper cups. The distance sensor is used to detect the distance after the paper cups are stacked and squeezed and generate distance data. The corresponding value is matched from the preset pixel height database according to the distance data as the stacking height value. The second number of paper cups is calculated according to the stacking height value and the stacking feature data. The absolute value of the difference between the number of cup rims and the number of first paper cups is the visual quantity difference. If the visual quantity difference is less than the preset visual reference difference, the paper cup stack count is calculated based on the number of cup rims and the number of first paper cups. Otherwise, the stack count of paper cups is calculated based on the number of the first paper cup and the number of the second paper cup; The step of extracting the cup rim from the stacked images based on a preset cup rim template and accumulating the number of cup rims further includes the following sub-steps: The cup rim template includes paper cup segmentation values, stacking direction, cup body graphic, cup rim segmentation values, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; The cup body shape is removed from the top of the paper cup image along the direction of the paper cup, and the edge shape of the paper cup is removed at a preset inner edge distance to obtain the middle image; Multiple parallel cup rims are extracted from the intermediate image based on the cup rim segmentation value, and the total number of cup rims is obtained by accumulating the number of cup rims; where the paper cup segmentation value is greater than the cup rim segmentation value.

2. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The step of extracting the cup rims from the stacked images based on a preset cup rim template and accumulating the number of cup rims also includes the following alternative step: The cup rim template includes paper cup segmentation values, stacking direction, cup rim graphic, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; A first semicircular shape is extracted from one side of the paper cup image along the stacking direction based on the cup rim graphic, and a second semicircular shape is extracted from the other side of the paper cup image along the paper cup direction based on the axisymmetric graphic of the cup rim graphic. The total number of the first semicircle shape and the total number of the second semicircle shape are counted. If the two counts are the same, one of them is taken as the number of the cup rim.

3. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The step of extracting the cup rims from the stacked images based on a preset cup rim template and accumulating the number of cup rims also includes the following alternative step: The cup rim template includes paper cup segmentation values, stacking direction, cup rim graphic, and extraction algorithm; The extraction algorithms include: The stacked image is converted to a grayscale image, and the paper cup image is obtained by segmenting multiple aggregated paper cup images from the grayscale image based on the paper cup segmentation value; The paper cup image is folded along the axis of the stacking direction to obtain a folded image, and a semi-circular shape is extracted from the folded image based on the cup rim graphic; The total number of semicircular shapes is used as the number of cup rims.

4. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The step of calculating the number of the first paper cups based on the stacking height value and stacking feature data also includes the following sub-steps: The stacking feature data includes stacking direction, cup height data, and cup rim height data. The cup rim height data is the height of the exposed part of the paper cup located in the middle of the stacked paper cups in the stacking direction. The sum of the cup height data and the cup rim height data is the height of the paper cup. The stack height value, cup body height data, and cup rim height data are all pixel data in the image after correction for imaging distortion. Subtracting the cup body height data from the stack height value yields the temporary value of the middle part, and dividing the temporary value of the middle part by the cup rim height data yields the number of the first paper cups.

5. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The step of calculating the stack count of paper cups based on the number of cup rims and the number of the first paper cups also includes the following sub-steps: The number of cups on the rim and the number of the first paper cups are used to calculate the stack count of paper cups by weighted average; Obtain the stack height values ​​corresponding to multiple consecutive first paper cup quantities, and calculate the fluctuation value of multiple stack height values; If the fluctuation value is greater than the preset fluctuation reference value, the weight of the number of first paper cups is controlled according to the fluctuation value. The larger the fluctuation value, the smaller the weight, and the smaller the fluctuation value, the larger the weight.

6. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The step of calculating the stack count of paper cups based on the number of the first paper cup and the number of the second paper cup also includes the following sub-steps: The number of paper cups in the first stack and the number of paper cups in the second stack are calculated by weighted average. Obtain the stack height value corresponding to multiple consecutive first paper cup quantities, and calculate the fluctuation value of multiple stack height values ​​as the first fluctuation value; Obtain the stack height values ​​corresponding to multiple consecutive second paper cup quantities, and calculate the fluctuation value of multiple stack height values ​​as the second fluctuation value; The first fluctuation value and the second fluctuation value are percentage values; The ratio of the first fluctuation value to the second fluctuation value is calculated as the fluctuation ratio. The weight of the number of first paper cups is controlled according to the fluctuation ratio. The larger the fluctuation value, the smaller the weight, and the smaller the fluctuation value, the larger the weight.

7. The online recognition method for the number of stacked paper cups based on machine vision according to claim 1, characterized in that, The method also includes the following steps: The absolute value of the difference between the number of cups on the rim and the number of the first paper cups is the first difference. The absolute value of the difference between the number of cups on the rim and the number of second paper cups is the second difference. The absolute value of the difference between the number of first paper cups and the number of second paper cups is the third difference. If the first difference, the second difference, and the third difference are all greater than the preset reference difference, an alarm for abnormal counting sensor readings will be triggered.

8. A machine vision-based online recognition system for the number of stacked paper cups, characterized in that, The device includes a processor that performs the steps of the online recognition method for the number of stacked paper cups based on machine vision as described in any one of claims 1-7.