State detection and control method of automatic packing system for food PET bottles

By setting up a multi-dimensional detection module and neural network analysis on the PET bottle packaging equipment, and combining it with acoustic feedback for online learning, the problem of inconsistent quality in PET bottle packaging equipment when facing individual differences and batch fluctuations has been solved, achieving efficient and stable film packaging quality control.

CN122126532APending Publication Date: 2026-06-02JIUSU TECH (SHANGHAI) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIUSU TECH (SHANGHAI) CO LTD
Filing Date
2026-04-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing automatic packaging equipment for heat-shrink film on PET bottles in the food packaging industry cannot effectively cope with individual bottle differences and batch quality fluctuations, resulting in inconsistent film packaging quality and a lack of adaptive adjustment capabilities, often requiring manual intervention.

Method used

By setting up weight, vision, and humidity detection modules on the conveyor chain to acquire multidimensional state data, using neural networks for individual state vector analysis and compensation, and combining acoustic detection feedback for online incremental learning, packaging parameters are dynamically adjusted.

Benefits of technology

This has achieved stability and consistency in the packaging quality of PET bottles, reduced defects, and improved production efficiency and energy management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122126532A_ABST
    Figure CN122126532A_ABST
Patent Text Reader

Abstract

This invention discloses a state detection and control method for an automated packaging system for food-grade PET bottles, belonging to the field of intelligent control of packaging equipment. Addressing the problem that existing packaging systems cannot dynamically adjust parameters based on individual bottle differences and batch quality fluctuations, this method sets up a detection module on the conveyor chain to acquire the weight, bottle outline, ellipticity, cap sealing ring status, bottle spacing, and condensate coverage of each bottle, and correlates these to form an individual state vector. The statistical distribution characteristics of multiple consecutive bottles are calculated to determine the degree of overall batch quality fluctuation, and based on this, a standard packaging mode, a compensating packaging mode, or a batch diversion warning is selected. In the compensating packaging mode, the individual state vector is input into a neural network, and a pre-compensation coefficient is output for advance adjustment. An acoustic detection device is installed at the heat shrink oven outlet, and the knocking spectrum analysis results are fed back to the neural network for online incremental learning. This method is used for quality detection and adaptive control in the automated packaging process of food-grade PET bottles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology for packaging equipment, specifically relating to a status detection and control method for an automatic packaging system for food PET bottles. Background Technology

[0002] In the existing food packaging industry, automatic packaging equipment for PET bottles using heat shrink film typically wraps groups of bottles, but it faces the following problems and drawbacks in actual production.

[0003] On the one hand, PET bottles exhibit individual differences during filling, capping, and transportation. These differences include variations in bottle ovality, inconsistent cap sealing ring conditions, and uneven surface humidity distribution due to condensation. These differences affect the quality of subsequent heat-shrink film packaging. For instance, excessive bottle ovality can lead to poor film adhesion, excessive condensation coverage reduces the adhesion between the heat-shrink film and the bottle, and a loose cap sealing ring may cause air leakage or appearance defects during shrinkage. However, existing packaging systems typically rely on simple photoelectric switches or weight thresholds for control, lacking simultaneous acquisition and correlation analysis of multi-dimensional states (such as weight, profile, humidity, and spacing) for each bottle. This prevents the controller from understanding the overall quality fluctuations of the current batch of bottles and makes it difficult to adjust drying air temperature, shaping clamping force, or heat shrink oven temperature to address individual differences.

[0004] On the other hand, in actual production, PET bottles from the same batch often exhibit certain quality fluctuations, though the degree of fluctuation can vary. Most existing systems use fixed packaging parameters, which may lead to unnecessary energy consumption or over-adjustment when fluctuations are small, and fail to effectively compensate for larger fluctuations, resulting in defects such as loose packaging, loose caps, or even breakage. Operators often have to stop the machine to make adjustments after a defect occurs, leading to inefficiency. Some systems have attempted to introduce single inspection indicators (such as weight or appearance), but due to the coupling effects between multiple state parameters, the accuracy of single-indicator judgment is insufficient, and it cannot distinguish defects caused by different reasons.

[0005] Furthermore, when inspecting the quality of film packaging at the outlet of the heat shrink oven, existing methods mostly rely on manual visual inspection or simple image comparison, making it difficult to detect hidden defects such as loose bottle caps or loose internal arrangement. Some studies have attempted to use acoustic testing by tapping, but due to variations in factors such as PET bottle material and film packaging tightness, acoustic judgments with fixed thresholds are prone to false alarms or false negatives. Moreover, traditional systems cannot feed back acoustic test results to front-end control parameters for self-learning, leading to repeated occurrences of the same problems.

[0006] The difficulty in solving the above problems lies in the following: PET bottles pass through the conveyor belt at high speed, requiring multi-dimensional detection and data correlation to be completed in a very short time; the fluctuation characteristics between batches are random and difficult to describe with a fixed mathematical model; the inverse mapping relationship between acoustic signals and packaging parameters is complex and lacks an effective online learning mechanism. Therefore, existing technologies lack a state detection and control method that can dynamically adjust packaging parameters based on the individual bottle state and the overall batch fluctuation, and continuously optimize using acoustic feedback. Summary of the Invention

[0007] This invention provides a state detection and control method for an automatic packaging system for food-grade PET bottles. Based on the multidimensional state vector of each PET bottle and the overall batch quality fluctuation, it adaptively selects a standard packaging mode, a compensation packaging mode, or a batch diversion warning. In the compensation mode, it uses a neural network to output a pre-compensation coefficient for advance adjustment. At the same time, it uses acoustic detection feedback from the heat shrink oven outlet for online incremental learning, thereby improving the stability and consistency of the film packaging quality and reducing defects.

[0008] To achieve these objectives and other advantages of the present invention, a method for status detection and control of an automatic packaging system for food-grade PET bottles is provided, comprising: S1. A weight detection module, a 3D vision detection module, and a surface humidity detection module are sequentially set on the conveyor chain to obtain the weight value, bottle outline and cap sealing ring status, bottle ellipticity, distance between adjacent bottles, and condensate coverage of each PET bottle. The controller associates the detection data of the same PET bottle by timestamp to form an individual state vector and calculates the statistical distribution characteristics of the state vectors of N consecutive PET bottles. S2. The controller determines the overall quality fluctuation of the current batch of PET bottles based on statistical distribution characteristics: if the fluctuation is less than or equal to the first threshold, the standard packaging mode is executed; if the fluctuation is greater than the first threshold but less than or equal to the second threshold, the compensation packaging mode is executed; if the fluctuation is greater than the second threshold, a batch diversion warning is triggered. S3. In the compensation packaging mode, the individual state vector of each PET bottle is input into a pre-trained neural network, which outputs pre-compensation coefficients for drying air temperature, shaping clamping force, and heat shrink oven zone temperature; the neural network establishes an inverse mapping between individual defects and packaging parameters based on historical failure cases. S4. The controller makes advance adjustments according to the pre-compensation coefficient; when the condensate coverage rate is too high and the ellipticity is also too high, the drying air temperature compensation output by the neural network is positive, the clamping force compensation is negative, and the downstream temperature compensation is negative. S5. An acoustic detection device is installed at the outlet of the heat shrink oven to feed the impact spectrum analysis results back to the neural network for online incremental learning; the actual film packing quality score in the feedback signal is provided by a subsequent quality detection module independent of the acoustic detection.

[0009] Preferably, the statistical distribution characteristics in step S2 are as follows: The controller constructs an N×m matrix from the individual state vectors of N consecutive PET bottles, where m is the dimension of the state vector; Calculate the covariance matrix Σ of the matrix, then calculate the Mahalanobis distance of each individual state vector relative to the mean of the current batch, and define the root mean square value of all Mahalanobis distances as the overall quality fluctuation degree. The first threshold and the second threshold are set as the 70th percentile and the 90th percentile of the root mean square value of the Mahalanobis distance of historical qualified batches, respectively, and these two percentiles are automatically updated after every 100 batches. When the root mean square value of the Mahalanobis distance is less than or equal to the 70th percentile, the fluctuation level is considered low, and the standard packaging mode is executed; when it is greater than the 70th percentile but less than or equal to the 90th percentile, the fluctuation level is considered moderate, and the compensated packaging mode is executed. Package mode; when the value is greater than the 90th percentile, it is determined that the fluctuation level is high, triggering a batch diversion warning.

[0010] Preferably, the knock spectrum analysis in step S5 is obtained through the following steps: S51. The pneumatic impactor in the acoustic detection device is activated by photoelectric trigger only when the membrane package reaches the predetermined impact position. The impact frequency is synchronized with the conveyor chain speed. The impact air pressure is 0.3±0.05MPa. The impact head is made of silicone material with a Shore A hardness of 50. The duration of a single impact is 0.1s. S52: The microphone acquires the tapping sound signal at a sampling rate of 20kHz, extracts the effective signal within a time window of 0.2s-0.5s after the tapping begins, and applies a Hanning window for windowing processing. S53. Perform a fast Fourier transform on the windowed signal to obtain the power spectral density in the 0-10kHz frequency domain. Divide the power spectral density into 15 frequency bands according to 1 / 3 octave bands and calculate the energy proportion of each frequency band. S54. The controller calculates the cosine similarity between the measured energy proportions of the 15 frequency bands and the pre-stored standard membrane package template. The standard membrane package template is established by the average spectrum of 10 consecutive qualified membrane packages. The similarity S = (A·B) / (‖A‖‖B‖), where A is the measured feature vector and B is the template feature vector. S55. Simultaneously calculate the difference ΔE between the energy percentage of the measured spectrum in the 2kHz-5kHz frequency band and the energy percentage of the standard membrane-coated template in the same frequency band. S56: If the similarity S is less than 0.85 and ΔE is greater than 0.1, it is judged as an abnormality of loose bottle cap; if S is less than 0.85 and ΔE is not greater than 0.1, it is judged as an abnormality of foreign matter or loose arrangement inside the film package; otherwise, it is judged as qualified. S57: The abnormal category or qualified status in the judgment result and the measured spectrum feature vector are used as feedback signals and input into the neural network for online incremental learning, wherein the weight coefficient of the abnormal sample is set to twice that of the qualified sample.

[0011] Preferably, the online incremental learning in step S5 specifically includes: The controller maintains a cyclic experience playback buffer, the capacity of which is dynamically set according to the product of the average number of packages per minute in the current batch and the duration of a standard production shift, with an upper limit of no more than 1000 samples; the experience playback adopts confidence weighting, and the sample weight is set comprehensively based on the similarity S of acoustic detection and subsequent quality score. Each sample contains an individual state vector, pre-compensation coefficients, acoustic detection judgment results, and an actual membrane pack quality score provided by a subsequent quality detection module independent of acoustic detection. The subsequent quality detection module includes a membrane pack tensile testing device or a membrane pack contour visual detection device, the output of which serves as a supervision signal. After each new quality score is obtained, the sample is added to a buffer, and a batch of samples is randomly drawn from the buffer with probability p, where p increases linearly from 0.1 to 0.5 with the buffer fill rate. A loss function is calculated for the extracted batch samples. This loss function consists of the sum of a mean squared error prediction error term and an elastic weight consolidation regularization term. The mean squared error prediction error term calculates the mean squared error between the pre-compensation coefficient vector and the expected compensation coefficient vector output by the neural network. The expected compensation coefficient vector is generated in real time by the controller based on the acoustic detection judgment results and subsequent quality scores, according to preset process rules. The regularization term uses the Fisher information matrix to perform a weighted square summation of the changes in each weight parameter in the neural network to limit the drift amplitude of important weights. An adaptive moment estimation optimizer is used to dynamically adjust the learning rate. The initial learning rate is set to a preset value between 0.001 and 0.0001, and the gradient clipping threshold is set to 1.0. After each preset number of incremental updates, the controller selects a predetermined number of samples from the buffer that have the largest Mahalanobis distance from the current batch for replay training.

[0012] Preferably, the neural network in step S3 and the neural network updated by online incremental learning in S5 are the same network. The network parameters are determined by offline training at the initial moment and are asynchronously updated by the feedback signal generated in step S5 after each packaging is completed. The controller sets a network parameter version number, which increments after each incremental update. The controller updates the network parameter version uniformly after each production batch ends, and the network parameters used for inference within the same batch remain unchanged. When the feedback signal in step S5 is qualified for three consecutive batches, the controller automatically reduces the learning rate of incremental learning to 50% of the original value to prevent over-adjustment of the converged model.

[0013] Preferably, step S1 further includes: setting a photoelectric trigger switch at the inlet of the conveyor chain, and generating a start signal when the PET bottle passes through the photoelectric trigger switch; The controller calculates the predicted arrival times of the PET bottles at the weight detection module, 3D vision detection module, and surface moisture detection module based on the real-time running speed of the conveyor chain, and generates the trigger delay for each module based on the predicted time; each module starts sampling according to the corresponding delay and aligns the timestamp of the sampled data to the start signal time of the photoelectric trigger switch; The controller also periodically collects speed fluctuation signals from the conveyor chain. When the speed fluctuation exceeds ±5%, it dynamically corrects the trigger delay of subsequent PET bottles.

[0014] Preferably, the neural network described in step S3 employs a multi-source domain adaptive transfer learning strategy during the offline training phase: Historical packaging data under at least three different bottle types, two different materials, and two different environmental humidity conditions were collected in advance, and the source domain model was trained accordingly. When the actual bottle type or environmental conditions during on-site production change, the controller calculates the maximum mean difference between the distribution characteristics of the state vectors of the first 20 PET bottles in the current batch and the data distribution of each source domain. It selects the source domain model with the smallest difference as the initial model and uses the detection data of the first 50 PET bottles in the current batch to quickly adapt and fine-tune the model. During fine-tuning, only the weight parameters of the last two fully connected layers of the neural network are updated.

[0015] Preferably, the advance adjustment in step S4 further includes: The controller calculates the actual conveying distance from each detection point to the drying nozzle, shaping fixture, and heat shrink oven entrance based on the position coordinates of the weight detection module, 3D vision detection module, and surface humidity detection module arranged sequentially on the conveyor chain. Combined with the real-time speed of the conveyor chain, it obtains the advance adjustment time window corresponding to the drying air temperature, shaping clamping force, and heat shrink oven zone temperature, as well as the advance adjustment time window for the air volume / speed of each zone of the heat shrink oven. When the individual state vector of the same PET bottle triggers multiple compensation coefficients, the controller stores the pre-compensation coefficients in a time-sorted compensation queue according to the spatial order of each actuator, and extracts the compensation coefficients from the queue and applies them within the corresponding time window before each actuator's action. If the time windows in the compensation queue overlap due to sudden changes in conveyor speed, the controller adopts a priority arbitration strategy: clamping force compensation takes precedence over temperature compensation, and the overlapping part is based on clamping force compensation. If the time windows are misaligned, the starting point of the window is recalculated according to the actual arrival time.

[0016] Preferably, in online incremental learning, the controller also maintains a parameter coupling suppression matrix C, the dimension of which is equal to the number of pre-compensation coefficients in the neural network output, and the elements c in the parameter coupling suppression matrix C are... ij This represents the cross-influence coefficient of the i-th compensation parameter on the j-th compensation parameter; Each time, the neural network outputs the original pre-compensation coefficient vector P. raw Then, the controller calculates the corrected compensation coefficient vector P. corr = P raw + C·(P raw - P hist ), where P hist The average compensation coefficient for the 10 most recent qualified bottles in history; When the drying air temperature compensation is positive and the clamping force compensation is negative, the corresponding cross-influence coefficient c in the parameter coupling suppression matrix C 12 Set to 0.3 to suppress the offsetting effect of increased drying air temperature on the actual clamping force; After the controller completes 100 incremental updates, it adaptively adjusts the coupling suppression matrix based on the feedback results of acoustic detection: it counts the occurrence rate of loose and abnormal arrangement in the last 100 membrane packs. If the occurrence rate exceeds 5%, it increases the absolute value of the cross-influence coefficient of clamping force on drying air temperature by 0.05 each time, with an upper limit of 0.5; if the occurrence rate is less than 1%, it gradually decreases the coefficient.

[0017] Preferably, during the asynchronous update process, the controller also sets up a core sample memory, which is used to store the sample with the largest Fisher information matrix trace in the historical failure cases. Each core sample contains its individual state vector, the corresponding pre-compensation coefficient, and the acoustic detection judgment result. When incremental learning is triggered, the controller first extracts a batch of samples from the recurrent experience replay buffer to calculate the first loss function. Then, it randomly extracts 20% of the batch samples from the core sample memory to calculate the second loss function. The second loss function uses the same mean squared error form as the first loss function but does not weight the regularization term. The first and second loss functions are added together with a weight of 0.7:0.3 to obtain the total loss function, which is used to update the neural network parameters. The core sample memory has a fixed capacity of 200 samples and adopts a clustering-based update strategy: the samples in the memory are divided into multiple clusters according to the defect type, and the sample with the largest loss gradient norm is retained in each cluster. When a new sample is added, it is assigned to its own cluster and the sample with the smallest gradient norm in that cluster is replaced. If the new sample represents a new defect type, a new cluster is created and the entire cluster with the smallest average gradient norm in all current clusters is deleted to maintain sample diversity. The calculation of the Fisher information matrix trace is performed only when core sample candidates are added, and a diagonal approximation of the Fisher information matrix is ​​used to reduce the amount of computation. It is calculated at most once per production batch.

[0018] The present invention has at least the following beneficial effects: This invention, by sequentially setting up a weight detection module, a 3D vision detection module, and a surface humidity detection module, can simultaneously acquire multi-dimensional state data for each PET bottle and associate the data of the same bottle by timestamp to form an individual state vector. The controller calculates statistical distribution characteristics based on the state vectors of N consecutive bottles, thereby scientifically judging the overall quality fluctuation of the batch and automatically selecting a standard packaging mode, a compensatory packaging mode, or a batch diversion warning. In the compensatory packaging mode, a neural network is used to establish an inverse mapping based on historical failure cases, outputting pre-compensation coefficients for drying air temperature, shaping clamping force, and heat shrink oven zone temperature, achieving individualized and precise compensation. Simultaneously, acoustic detection results are fed back to the neural network for online incremental learning, enabling the system to continuously learn and improve from actual packaging results, forming a closed-loop optimization. This method improves the stability and consistency of film packaging quality and reduces defect generation.

[0019] This invention uses the root mean square value of Mahalanobis distance as a quantitative indicator of the overall quality fluctuation. This indicator can comprehensively consider the correlation between multi-dimensional state vectors and more accurately reflect the quality consistency of bottles within a batch than traditional one-dimensional variance or mean. By setting the first and second thresholds to the 70th and 90th percentiles of the root mean square value of Mahalanobis distance of historical qualified batches, respectively, and automatically updating them after every 100 batches, the thresholds can adaptively follow changes in production conditions. This dynamic threshold mechanism avoids misjudgments between different batches using fixed thresholds, making the switching between standard packaging mode, compensated packaging mode, and batch diversion warning more reasonable. It can save energy when the fluctuation is small, provide effective compensation when the fluctuation is moderate, and divert problematic batches in a timely manner when the fluctuation is too large, reducing the risk of defective products leaving the batch.

[0020] The impact spectrum analysis method proposed in this invention automatically strikes the membrane package upon arrival using a pneumatic impactor, and acquires the acoustic signal at a 20kHz sampling rate. After windowing using a Hanning window and FFT transformation, the energy proportions of 15 frequency bands in 1 / 3 octave bands are obtained. By calculating the cosine similarity with a standard membrane package template and combining the energy proportion differences in the 2kHz-5kHz frequency band, it can accurately distinguish between bottle cap loosening and abnormalities such as foreign objects or loose arrangement inside the membrane package, with an accuracy rate far exceeding that of the fixed threshold method. More importantly, the abnormality category, qualified status, and measured spectrum feature vector are used as feedback signals input into the neural network for online incremental learning, and the weight of abnormal samples is set to twice that of qualified samples, making the neural network pay more attention to failure cases and continuously improve the detection capability of hidden defects.

[0021] This invention employs a cyclic experience replay buffer, the capacity of which is dynamically set based on the product of the number of packages packed per minute and the shift duration, with an upper limit of no more than 1000 samples. This ensures sufficient historical sample diversity while controlling computational overhead. Experience replay uses confidence weighting, with sample weights combining acoustic detection similarity S and subsequent quality scores, allowing high-confidence samples to contribute more to model updates. The loss function is the sum of the mean squared error prediction error term and the elastic weight consolidation regularization term. The regularization term uses the Fisher information matrix to weight and penalize changes in important weights, effectively preventing catastrophic forgetting. The adaptive moment estimation optimizer dynamically adjusts the learning rate, and a gradient pruning threshold of 1.0 ensures training stability. After each preset number of iterations, the sample with the largest Mahalanobis distance to the current batch is selected for replay training, strengthening the learning of marginally distributed samples and improving the model's generalization ability.

[0022] This invention sets the neural network in step S3 and the network updated by online incremental learning in step S5 to be the same network, but uses an asynchronous update mechanism. The controller sets the network parameter version number, which increments after each incremental update. Step S3 always calls the latest version of the network parameters during inference, ensuring that the front-end control always uses the current optimal model. When the feedback signal is qualified for three consecutive batches, the learning rate of incremental learning is automatically reduced to 50% of the original value, avoiding over-adjustment that would lead to performance degradation when the model has converged. This version control and adaptive learning rate reduction strategy ensures that the model can continuously learn from new samples while preventing excessive perturbation of learned knowledge, achieving a good balance between learning efficiency and stability.

[0023] This invention uses a photoelectric trigger switch at the conveyor belt inlet as a unified starting signal source. The controller calculates the predicted arrival time of PET bottles at each detection module based on the real-time speed of the conveyor belt and generates corresponding trigger delays, ensuring that the timestamps of the sampling data from each module are aligned to the same starting signal. This guarantees that data such as the weight, three-dimensional contour, and surface moisture of the same bottle can be accurately correlated to form an individual state vector. When the conveyor belt speed fluctuates by more than ±5%, the controller dynamically corrects the trigger delay of subsequent PET bottles, compensating for the time deviation caused by speed changes. Compared to the method of each module using independent sensor triggers, this invention reduces the probability of misalignment of data from different dimensions in the individual state vector, thus improving the accuracy of the detection data.

[0024] This invention employs a multi-source domain adaptive transfer learning strategy, pre-collecting historical packaging data under different bottle types, materials, and humidity conditions, and training source domain models separately. When the actual bottle type or environmental conditions during on-site production change, the controller calculates the maximum mean difference between the state vector distribution of the first 20 bottles in the current batch and the data distribution of each source domain, automatically selecting the most similar source domain model as the initial model. Then, the model is rapidly adapted and fine-tuned using the detection data of the first 50 bottles in the current batch, updating only the weight parameters of the last two fully connected layers. This strategy significantly reduces the amount of data and computation time required for model adaptation, enabling rapid switching between different bottle types and environmental conditions without retraining the entire network.

[0025] This invention calculates independent advance adjustment time windows for drying air temperature, shaping clamping force, and heat shrink oven zone temperature based on the actual conveying distance from each detection module to the actuator and the real-time speed of the conveyor belt. When multiple compensation coefficients are triggered for the same bottle, the controller stores the pre-compensation coefficients in the compensation queue according to the spatial sequence of the actuators and extracts and applies the coefficients within the corresponding time window before each actuator's action. When a sudden change in conveyor belt speed causes time windows to overlap, a priority arbitration strategy is adopted: clamping force compensation takes precedence over temperature compensation, and the overlapping part is based on clamping force compensation. If the windows are misaligned, the starting point of the window is recalculated. This mechanism ensures that multiple compensation coefficients are applied in an orderly manner within the correct time window, avoiding compensation timing errors or execution conflicts caused by speed fluctuations, and guaranteeing the compensation effect.

[0026] This invention introduces a parameter coupling suppression matrix C to quantify the cross-influence coefficient between different compensation parameters. When the drying air temperature compensation is positive and the clamping force compensation is negative, the cross-influence coefficient c is... 12 Setting it to 0.3, the original compensation coefficient is decoupled and corrected using a correction formula, effectively suppressing the offsetting effect of increased drying air temperature on the actual clamping force. After every 100 incremental updates, the controller adaptively adjusts the coupling suppression matrix based on the occurrence rate of loose arrangement anomalies in acoustic detection: if the occurrence rate exceeds 5%, the absolute value of the cross-influence coefficient is increased (by 0.05 each time, with a maximum of 0.5); if it is below 1%, it is gradually decreased. This adaptive mechanism allows the coupling suppression matrix to be optimized according to changes in actual production conditions, further improving the accuracy of multi-parameter collaborative compensation.

[0027] This invention establishes a core sample memory specifically for storing samples with the largest loss gradient norm from historical failure cases (a larger gradient norm indicates a greater impact of the sample on model parameter updates). Each core sample contains an individual state vector, pre-compensation coefficients, and acoustic detection judgment results. During incremental learning, the controller simultaneously extracts batches of samples from the recurrent experience replay buffer to calculate the first loss function, and randomly extracts 20% of the core samples from the core sample memory to calculate the second loss function. The two are added together with a weight of 0.7:0.3 to obtain the total loss function. This ensures that high-information failure cases are not forgotten. The calculation of the loss gradient norm uses an approximation method and is calculated at most once per batch, significantly reducing the computational load. This mechanism effectively alleviates the catastrophic forgetting problem, enabling the neural network to maintain its ability to identify and compensate for rare but important defects over a long period.

[0028] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating the status detection and control method of the automatic packaging system for food-grade PET bottles according to the present invention. Detailed Implementation

[0030] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.

[0031] It should be understood that terms such as “having,” “comprising,” and “including” as used herein do not exclude the presence or addition of one or more other elements or combinations thereof.

[0032] like Figure 1 As shown, this embodiment of the invention provides a status detection and control method for an automatic packaging system for food-grade PET bottles, comprising: S1. A weight detection module, a 3D vision detection module, and a surface humidity detection module are sequentially set on the conveyor chain to obtain the weight value, bottle outline and cap sealing ring status, bottle ellipticity, distance between adjacent bottles, and condensate coverage of each PET bottle. The controller associates the detection data of the same PET bottle by timestamp to form an individual state vector and calculates the statistical distribution characteristics of the state vectors of N consecutive PET bottles. S2. The controller determines the overall quality fluctuation of the current batch of PET bottles based on statistical distribution characteristics: if the fluctuation is less than or equal to the first threshold, the standard packaging mode is executed; if the fluctuation is greater than the first threshold but less than or equal to the second threshold, the compensation packaging mode is executed; if the fluctuation is greater than the second threshold, a batch diversion warning is triggered. S3. In the compensation packaging mode, the individual state vector of each PET bottle is input into a pre-trained neural network, which outputs pre-compensation coefficients for drying air temperature, shaping clamping force, and heat shrink oven zone temperature; the neural network establishes an inverse mapping between individual defects and packaging parameters based on historical failure cases. S4. The controller makes advance adjustments according to the pre-compensation coefficient; when the condensate coverage rate is too high and the ellipticity is also too high, the drying air temperature compensation output by the neural network is positive, the clamping force compensation is negative, and the downstream temperature compensation is negative. S5. An acoustic detection device is installed at the outlet of the heat shrink oven to feed the impact spectrum analysis results back to the neural network for online incremental learning; the actual film packing quality score in the feedback signal is provided by a subsequent quality detection module independent of the acoustic detection.

[0033] In the above embodiments, the conveyor chain is the core path for carrying and transporting PET bottles. To improve the accuracy of packaging quality control, this method sequentially sets up three detection modules on the conveyor chain: a weight detection module, a 3D vision detection module, and a surface moisture detection module. The weight detection module typically uses a dynamic electronic scale or a strain gauge weighing sensor to acquire the weight value of each PET bottle in real time, thereby determining whether there are abnormalities such as underfilling, overfilling, or bottle damage. The 3D vision detection module can be composed of a laser profilometer, a structured light camera, or a multi-view stereo vision system, which can reconstruct the 3D contour of the bottle body and then extract the bottle body ellipticity, the state of the bottle cap sealing ring, and the spacing between adjacent bottles. Among them, the bottle body ellipticity reflects the degree to which the cross-section of the bottle is close to a circle; excessive ellipticity will affect the tightness of subsequent film wrapping. The state of the bottle cap sealing ring refers to whether the sealing ring is intact, displaced, or damaged. The spacing between adjacent bottles is used to determine whether the bottles are evenly arranged; uneven spacing will cause misalignment of the bottles when the heat shrink film is wrapped. The surface humidity detection module can use infrared thermal imaging, capacitive humidity sensors, or laser reflectivity analysis to detect the proportion of the bottle surface covered by condensation. In high humidity environments or under conditions of large temperature differences, condensation easily forms on the surface of PET bottles, reducing the adhesion between the heat-shrink film and the bottle. The controller, typically a programmable logic controller, industrial embedded computer, or edge computing-based intelligent controller, receives the detection data from the three modules mentioned above and correlates information such as the weight, bottle outline, cap sealing ring status, bottle ellipticity, distance between adjacent bottles, and condensation coverage of each PET bottle according to a unified timestamp as it passes through the conveyor belt. This correlation forms an individual state vector. This vector is a digital characteristic description of the bottle, containing key parameters affecting subsequent packaging quality. Subsequently, the controller continuously collects the individual state vectors of N PET bottles (N can be 100 or 150, the specific value can be flexibly set according to production speed and batch size), and calculates the statistical distribution characteristics of these vectors, such as the mean, covariance matrix, or Mahalanobis distance, thus providing a basis for judging the quality consistency of the entire batch of bottles.

[0034] The controller assesses the overall quality fluctuation of the current batch of PET bottles based on the aforementioned statistical distribution characteristics. This fluctuation reflects the magnitude of differences in multidimensional states among individual bottles within the batch, including the comprehensive dispersion of multiple aspects such as weight, ellipticity, and humidity. For example, if the statistical distribution shows a small root mean square value of Mahalanobis distance, it indicates a high degree of consistency among the bottles; conversely, a large value indicates significant differences. Depending on the magnitude of the fluctuation, the controller executes three different control modes. When the fluctuation is less than or equal to the first threshold, it indicates stable bottle quality, and the standard packaging mode is executed, i.e., packaging is performed according to preset fixed parameters such as drying air temperature, clamping force, and heat shrink oven temperature to save energy and maintain efficient production. When the fluctuation is greater than the first threshold but less than or equal to the second threshold, it indicates a moderate degree of individual difference, and the system enters the compensation packaging mode, adjusting parameters individually for each bottle's specific state. When the fluctuation is greater than the second threshold, it means that the quality of bottles within the batch is severely inconsistent, and the system triggers a batch diversion warning, automatically guiding the batch of bottles to the diversion channel to prevent them from entering the packaging process and causing a large number of defective products. In actual operation, the first and second thresholds can be dynamically set based on statistical data from historical qualified batches. For example, the root mean square value of the Mahalanobis distance of the most recent 100 qualified batches can be collected, with the 70th percentile used as the first threshold and the 90th percentile as the second threshold. If historical data is lacking, empirical values ​​can also be used, such as setting the first threshold to 0.6 and the second threshold to 0.9. These values ​​are selected within a dimensionless normalized reference range and are automatically recalculated and updated after every 100 batches. This hierarchical control strategy avoids the drawbacks of a single mode handling all situations and achieves a reasonable match between the degree of quality fluctuation and production parameters.

[0035] In compensated packaging mode, the controller inputs the individual state vector of each PET bottle into a pre-trained neural network. This neural network can be a multilayer perceptron, a radial basis function network, or a lightweight convolutional neural network. The number of nodes in its input layer matches the dimension of the individual state vector, for example, six dimensions or more. The output layer corresponds to the pre-compensation coefficients for the packaging parameters that need adjustment, specifically including pre-compensation coefficients for the drying air temperature, the shaping clamping force, and the zone temperature of the heat shrink oven. The drying air temperature is used to regulate the temperature of the hot air blowing onto the bottle to dry surface condensation. The shaping clamping force is used to adjust the mechanical force that holds the bottles in place. The heat shrink oven is typically divided into multiple temperature zones, such as a front zone, a middle zone, and a rear zone. The temperature of each zone can be adjusted independently and each corresponds to a pre-compensation coefficient. The neural network training process utilizes historical failure cases, such as past defects like loose wrapping, loose caps, or broken wrapping. Through inverse mapping learning, it learns from the individual defect characteristics to deduce what packaging parameters should be adjusted to avoid the defect. For example, if historical data shows that bottles with high ellipticity and excessive condensation are better packaged using higher drying air temperature and lower clamping force, the neural network will learn this mapping relationship. After training, the neural network can infer online and quickly output a set of pre-compensation coefficients for the current bottle state, which can be used by subsequent actuators. It should be noted that the so-called inverse mapping means that the neural network learns the relationship between the individual state vector in historical failure cases, the packaging parameters used at that time, and the final defect type to infer "how to adjust the parameters to avoid this type of defect". During the training phase, historical failure cases themselves do not directly provide correct compensation coefficient labels; therefore, this invention adopts a strategy combining contrastive learning and label generation based on expert rules: for each failure case, the system forms a negative sample pair with its individual state vector and the actual packaging parameters that caused the failure; at the same time, the packaging parameters of the failure case are corrected according to process knowledge (for example, if the membrane detachment is caused by excessive condensation, the drying air temperature is corrected upward by one step), generating a "pseudo-optimal compensation coefficient" that can avoid the defect as a positive sample. The neural network learns an inverse mapping from the individual state vector to the optimal compensation coefficient by minimizing the prediction error of positive sample pairs (state, pseudo-optimal compensation) and maximizing the prediction error of negative sample pairs (state, failure parameters). Furthermore, during the online incremental learning phase, the actual scores from subsequent quality control modules serve as reward signals for reinforcement learning, further optimizing the compensation coefficients output by the network. This allows the network to continuously extract correct parameter adjustment patterns from historical failures.

[0036] After obtaining the pre-compensation coefficient, the controller does not apply it immediately. Instead, it makes advance adjustments based on the spatial position of each actuator on the conveyor chain and the predicted arrival time of the bottle. In other words, the controller calculates the time required for the bottle to travel from the detection point to each actuator based on the real-time speed of the conveyor chain, and sends the corresponding compensation coefficient to the actuator just before the bottle arrives, achieving precise time synchronization. This approach avoids compensation failures caused by transmission delays or delayed mechanism responses. A typical compensation logic example is as follows: When the system detects that a PET bottle has a high condensate coverage and a high ellipticity (e.g., a condensate coverage exceeding 40% and an ellipticity greater than 0.8 mm), the neural network will output a positive compensation for drying air temperature (increasing the drying air temperature to accelerate surface moisture evaporation), a negative compensation for shaping clamping force (reducing the clamping force to prevent deformation or damage to the bottle due to excessive ellipticity), and a negative compensation for the temperature at the rear of the heat shrink oven (lowering the temperature at the rear to avoid excessive shrinkage or localized melting due to irregular bottle shape during high-temperature shrinkage). This collaborative compensation strategy fully considers the physical coupling relationship between multiple parameters, avoiding the side effects that may result from adjusting a single parameter.

[0037] At the exit of the heat shrink oven, this method also incorporates an acoustic detection device. This device collects the tapping sound and performs spectral analysis by gently tapping the already packaged membrane bundles, for example, using a pneumatic tapper or electromagnetic hammer. Different internal states produce tapping sound spectra with characteristic differences; for example, loose bottle caps result in more high-frequency components, loose arrangement alters the mid-frequency energy distribution, and the presence of foreign objects produces abnormal spikes. The controller uses the spectral analysis results as feedback signals to determine whether the current membrane bundle is of acceptable quality and what type of defect exists. More importantly, this feedback signal is fed back into a neural network for online incremental learning. That is, the neural network not only relies on a pre-trained offline model but also continuously updates its weight parameters based on the actual acoustic detection results after packaging, thereby adapting to changes in production conditions, such as bottle type changes or seasonal fluctuations in environmental humidity. Furthermore, the actual membrane bundle quality score in the feedback signal is not given by the acoustic detection itself but comes from a subsequent quality inspection module independent of the acoustic detection, such as a membrane bundle strength detection device based on tensile testing or a membrane bundle appearance inspection system based on high-precision vision. This design avoids cyclical verification of the same detection method, improving the objectivity and reliability of the feedback signal. Through the coordinated work of the above five steps, the system achieves a complete closed loop from individual PET bottle status acquisition, batch quality assessment, differential compensation control to acoustic feedback and model self-optimization.

[0038] Compared to existing PET bottle packaging control technologies, most current technologies rely on fixed packaging parameters or single detection indicators such as weight or appearance. These technologies cannot address the multidimensional differences in the individual bottle states, leading to frequent defects such as loose wrapping, loose caps, or surface damage when batch quality fluctuates significantly. Furthermore, they lack adaptive adjustment capabilities, often requiring manual intervention after defects occur. This new method, by sequentially setting up a weight detection module, a 3D vision detection module, and a surface moisture detection module, and constructing an individual state vector, achieves for the first time the simultaneous acquisition and correlation analysis of multiple key quality characteristics for each PET bottle, providing a reliable data foundation for personalized compensation. Based on this, a hierarchical control strategy based on statistical distribution characteristics can intelligently identify the overall batch quality fluctuation level and automatically switch between standard mode, compensation mode, or diversion warning, avoiding unnecessary energy consumption or over-compensation and effectively preventing severely defective batches from flowing into subsequent processes. In the compensated packaging mode, a neural network is used to establish an inverse mapping based on historical failure cases. This allows for the output of optimal pre-compensation coefficients for drying air temperature, clamping force, and heat shrink oven zone temperature for each bottle's specific defect combination. Particularly noteworthy is the differentiated compensation strategy employed when both condensation and ellipticity are excessively high, significantly improving film bonding and bottle integrity. Furthermore, the introduction of an acoustic detection device, combined with feedback from an independent quality scoring module, enables the neural network to continuously perform online incremental learning during production, constantly optimizing the output accuracy of the compensation coefficients and thus drastically reducing the recurrence rate of similar defects. The entire method forms a closed-loop control system from detection, decision-making, execution to feedback self-optimization, effectively improving the stability, adaptability, and finished product qualification rate of the automated packaging process for food-grade PET bottles.

[0039] In one specific embodiment, the statistical distribution characteristics in step S2 are as follows: The controller constructs an N×m matrix from the individual state vectors of N consecutive PET bottles, where m is the dimension of the state vector; Calculate the covariance matrix Σ of the matrix, then calculate the Mahalanobis distance of each individual state vector relative to the mean of the current batch, and define the root mean square value of all Mahalanobis distances as the overall quality fluctuation degree. The first threshold and the second threshold are set as the 70th percentile and the 90th percentile of the root mean square value of the Mahalanobis distance of historical qualified batches, respectively, and these two percentiles are automatically updated after every 100 batches. When the root mean square value of the Mahalanobis distance is less than or equal to the 70th percentile, the fluctuation level is considered low, and the standard packaging mode is executed; when it is greater than the 70th percentile but less than or equal to the 90th percentile, the fluctuation level is considered moderate, and the compensated packaging mode is executed. Package mode; when the value is greater than the 90th percentile, it is determined that the fluctuation level is high, triggering a batch diversion warning.

[0040] In the above implementation, the specific calculation method for the statistical distribution characteristics in step S2 is further defined. The controller arranges the individual state vectors of N consecutive PET bottles into an N x m matrix, where N represents the number of bottles and m represents the dimension of each state vector. For example, m can be six-dimensional, including parameters such as weight, ellipticity, bottle cap sealing ring state, distance between adjacent bottles, and condensate coverage. Based on this matrix, the controller calculates the covariance matrix. The covariance matrix reflects the degree of linear correlation between state parameters of different dimensions, such as whether there is a positive or negative correlation between condensate coverage and ellipticity. Further, the controller calculates the Mahalanobis distance of each individual state vector relative to the mean state of all bottles in the current batch. Mahalanobis distance is a distance metric that considers the correlation between dimensions, and it reflects the degree of deviation between a bottle and the batch average state more accurately than ordinary Euclidean distance. The root mean square value of the Mahalanobis distance is obtained by summing the squares of all Mahalanobis distances, dividing by the number of bottles, and then taking the square root. This value is defined as the overall quality fluctuation. The larger this value, the higher the dispersion of the bottle state within the batch.

[0041] In this implementation, the first and second thresholds are set as the 70th and 90th percentiles of the root mean square distance (RMS) of historical qualified batches, respectively. For example, the system records the RMS of the Mahalanobis distance for the past 100 qualified production batches, sorts these values ​​from smallest to largest, and the value at the 70th percentile is the first threshold, and the value at the 90th percentile is the second threshold. This historical percentile-based method allows the thresholds to adaptively follow changes in production conditions. For example, when equipment aging or changes in raw material batches cause an overall shift in the normal fluctuation range, the thresholds will adjust accordingly. To maintain the timeliness of the thresholds, the system automatically recalculates these two percentiles after every 100 batches, updating them once with data from the most recent 100 qualified batches. In the absence of initial historical data, empirical reference values ​​can be used first, such as setting the first threshold to 0.5 and the second threshold to 0.8, and then switching to dynamic update mode after accumulating enough batches.

[0042] The controller compares the calculated root mean square (RMS) Mahalanobis distance value with the two thresholds mentioned above, performing a three-level judgment. When the RMS value is less than or equal to the 70th percentile, it is judged as low volatility, and the standard packaging mode is executed, i.e., using the system's preset fixed packaging parameters without additional compensation. When the RMS value is greater than the 70th percentile but less than or equal to the 90th percentile, it is judged as medium volatility, and the compensation packaging mode is executed, i.e., the neural network is activated to adjust individual parameters. When the RMS value is greater than the 90th percentile, it is judged as high volatility, triggering a batch diversion warning, automatically guiding the batch of bottles to the defective channel or the inspection area to avoid entering the heat shrink packaging process and causing a large number of defects. This three-level judgment mechanism achieves a smooth transition from normal production, intelligent compensation to abnormal diversion.

[0043] Compared to existing methods that use fixed thresholds or single-index variance judgments, this implementation introduces the root mean square value of Mahalanobis distance as a quantitative indicator of the overall quality fluctuation, which can comprehensively consider the correlation between multi-dimensional state parameters and avoid the problems of misjudgment or ignoring coupling effects in a single dimension. The dynamic threshold update mechanism based on the percentile of historical qualified batches enables the system to adaptively follow changes in production conditions without frequent manual threshold adjustments. The three-level judgment mode achieves refined hierarchical control, saving energy and reducing equipment wear when fluctuations are small, providing precise compensation when fluctuations are moderate, and promptly diverting problematic batches when fluctuations are excessive, thereby reducing the risk of defective products leaving the production line and improving the adaptability and batch consistency of the production line.

[0044] In one specific implementation, the tapping spectrum analysis in step S5 is obtained through the following steps: S51. The pneumatic impactor in the acoustic detection device is activated by photoelectric trigger only when the membrane package reaches the predetermined impact position. The impact frequency is synchronized with the conveyor chain speed. The impact air pressure is 0.3±0.05MPa. The impact head is made of silicone material with a Shore A hardness of 50. The duration of a single impact is 0.1s. S52: The microphone acquires the tapping sound signal at a sampling rate of 20kHz, extracts the effective signal within a time window of 0.2s-0.5s after the tapping begins, and applies a Hanning window for windowing processing. S53. Perform a fast Fourier transform on the windowed signal to obtain the power spectral density in the 0-10kHz frequency domain. Divide the power spectral density into 15 frequency bands according to 1 / 3 octave bands and calculate the energy proportion of each frequency band. S54. The controller calculates the cosine similarity between the measured energy proportions of the 15 frequency bands and the pre-stored standard membrane package template. The standard membrane package template is established by the average spectrum of 10 consecutive qualified membrane packages. The similarity S = (A·B) / (‖A‖‖B‖), where A is the measured feature vector and B is the template feature vector. S55. Simultaneously calculate the difference ΔE between the energy percentage of the measured spectrum in the 2kHz-5kHz frequency band and the energy percentage of the standard membrane-coated template in the same frequency band. S56. If the similarity S is less than 0.85 and ΔE is greater than 0.1, it is judged as an abnormality of loose bottle cap; if S is less than 0.85 and ΔE is not greater than 0.1, it is judged as an abnormality of foreign matter or loose arrangement inside the film package; otherwise, it is judged as qualified. S57. The abnormal category or qualified status in the judgment result and the measured spectrum feature vector are used as feedback signals and input into the neural network for online incremental learning, wherein the weight coefficient of the abnormal sample is set to twice that of the qualified sample.

[0045] In the above embodiment, the specific steps for obtaining the tapping spectrum analysis in step S5 are further defined. The acoustic detection device includes a pneumatic tapper, which is activated by a photoelectric trigger switch only when the membrane package reaches the predetermined tapping position, avoiding noise interference caused by ineffective tapping. The tapping frequency is synchronized with the running speed of the conveyor belt; for example, the tapper operates once for every membrane package spacing the conveyor belt advances. The tapping air pressure is set between 0.25 MPa and 0.35 MPa, with an optimal value of 0.3 MPa. This pressure range produces a sufficiently loud tapping sound without damaging the membrane package or bottle. The tapping head is made of silicone material with a Shore A hardness of 50. This hardness ensures the clarity of the tapping sound while providing sufficient elasticity to reduce impact damage to the bottle cap. The duration of a single tap is 0.1 seconds, which is the total time for the tapping head to contact the membrane package and rebound. The tapping sound is collected by a microphone with a sampling rate set to 20 kHz, covering the frequency range required for human hearing and analysis. After acquiring the signal, the system extracts the effective signal within a time window of 0.2 to 0.5 seconds after the initial impact, encompassing the main vibration decay process following the impact. The extracted signal is then windowed using a Hanning window to reduce spectral leakage at signal boundaries. It should be noted that the above key parameters are determined based on extensive experiments examining the impact acoustic characteristics of food-grade PET bottles and commonly used heat-shrink film materials. A combination of a striking pressure of 0.3±0.05MPa and a silicone striking head with a Shore hardness of A50 can generate a stable and moderately energetic broadband excitation signal without damaging the bottle cap or diaphragm. Experiments show that the sound pressure level variation coefficient is less than 5% within this pressure range, while the A50 hardness avoids high-frequency harmonic distortion. The 0.2s-0.5s time window after the start of striking is selected because the first 0.2s includes nonlinear collision impact and mechanical noise between the striking head and the diaphragm. During this period, the signal-to-noise ratio is low and the correlation with internal defects is weak. Discarding this window highlights the steady-state decay segment after 0.2s, dominated by the inherent vibration of the diaphragm-bottle system, thereby improving the distinguishability of defect features. The similarity threshold is 0.8. The energy difference threshold of 0.1 for 5 kHz and 2-5 kHz is derived from the statistical analysis of the spectral distribution of qualified products and two typical defects (loose bottle caps and loose arrangement). The average cosine similarity of qualified products is 0.92±0.03, while that of defective products is below 0.80. A threshold of 0.85 can balance missed detections and false detections. Loose bottle caps will generate an additional resonance peak in the 2-5 kHz range, making ΔE>0.12. Loose arrangement mainly changes the energy distribution in the mid-to-low frequency range and ΔE<0.05. Therefore, 0.1 is the optimal separation threshold. The above parameters were verified in five common bottle types, two film packaging materials, and environments with relative humidity of 40~90%. The recognition accuracy remained above 92%, indicating that the fixed values ​​have sufficient engineering robustness.

[0046] The windowed signal is subjected to a Fast Fourier Transform (FFT) to obtain the power spectral density in the 0-10 kHz frequency domain. Then, following the commonly used one-third octave band division method in acoustic engineering, the entire frequency domain is divided into 15 bands, with the center frequencies of each band distributed proportionally. The system calculates the percentage of energy in each band relative to the total energy, forming a 15-dimensional feature vector. The controller performs cosine similarity calculations between the measured energy percentages of these 15 bands and a pre-stored standard membrane template. The standard membrane template is established by the average spectrum of 10 consecutive qualified membranes, representing the ideal spectral distribution. The closer the cosine similarity value is to 1, the more similar the measured spectrum is to the template. Simultaneously, the system also calculates the difference ΔE between the energy percentage of the measured spectrum in the sensitive 2-5 kHz band and the energy percentage of the standard membrane template in this band. The judgment rules are as follows: If the similarity is less than 0.85 and ΔE is greater than 0.1, it is judged as an abnormality of loose bottle cap, because a loose bottle cap will generate additional high-frequency vibration components; if the similarity is less than 0.85 and ΔE is not greater than 0.1, it is judged as an abnormality of foreign matter or loose arrangement inside the membrane package. These defects mainly change the distribution of mid-to-low frequency energy rather than the 2 to 5 kHz frequency band; all other cases are judged as qualified. The similarity threshold of 0.85 and the difference threshold of 0.1 are reference values ​​optimized through experiments, and can be fine-tuned according to product requirements in actual applications.

[0047] The abnormal categories from the above judgment results, such as loose bottle caps or loose arrangement, along with the qualified status, are used as feedback signals along with the measured 15-dimensional spectral feature vector, and input into the neural network for online incremental learning. To enhance the neural network's learning effect on defect cases, the weight coefficient of abnormal samples is set to twice that of qualified samples. That is, when the neural network updates its parameters, the error gradient contribution of one abnormal sample is equivalent to that of two qualified samples, making the network pay more attention to those situations that are likely to cause packaging failure. In this way, acoustic detection is not only used for on-the-spot judgment but also continuously provides data support for the optimization of the control model.

[0048] Compared to existing acoustic detection methods that rely on fixed thresholds or manual listening, the impact spectrum analysis method proposed in this embodiment objectively quantifies the acoustic characteristics of the diaphragm by calculating the energy proportion of one-third octave bands and cosine similarity. Combined with the energy difference in the 2kHz to 5kHz frequency band, it accurately distinguishes between bottle cap looseness and internal loose arrangement, significantly improving the recognition accuracy. Furthermore, by doubling the weight of abnormal samples and feeding them back into the neural network for incremental learning, the system continuously improves from each anomaly, resulting in a sustained decrease in both false alarm and false negative rates over time. The trigger synchronization of the pneumatic impactor and the selection of a silicone impact head also ensure consistent detection and non-destructive treatment of the product.

[0049] In one specific implementation, the online incremental learning in step S5 specifically includes: The controller maintains a cyclic experience playback buffer, the capacity of which is dynamically set according to the product of the average number of packages per minute in the current batch and the duration of a standard production shift, with an upper limit of no more than 1000 samples; the experience playback adopts confidence weighting, and the sample weight is set comprehensively based on the similarity S of acoustic detection and subsequent quality score. Each sample contains an individual state vector, pre-compensation coefficients, acoustic detection judgment results, and an actual membrane pack quality score provided by a subsequent quality detection module independent of acoustic detection. The subsequent quality detection module includes a membrane pack tensile testing device or a membrane pack contour visual detection device, the output of which serves as a supervision signal. After each new quality score is obtained, the sample is added to a buffer, and a batch of samples is randomly drawn from the buffer with probability p, where p increases linearly from 0.1 to 0.5 with the buffer fill rate. A loss function is calculated for the extracted batch samples. This loss function consists of the sum of a mean squared error prediction error term and an elastic weight consolidation regularization term. The mean squared error prediction error term calculates the mean squared error between the pre-compensation coefficient vector and the expected compensation coefficient vector output by the neural network. The expected compensation coefficient vector is generated in real time by the controller based on the acoustic detection judgment results and subsequent quality scores, according to preset process rules. The regularization term uses the Fisher information matrix to perform a weighted square summation of the changes in each weight parameter in the neural network to limit the drift amplitude of important weights. An adaptive moment estimation optimizer is used to dynamically adjust the learning rate. The initial learning rate is set to a preset value between 0.001 and 0.0001, and the gradient clipping threshold is set to 1.0. After each preset number of incremental updates, the controller selects a predetermined number of samples from the buffer that have the largest Mahalanobis distance from the current batch for replay training.

[0050] In the above implementation, the specific implementation method of online incremental learning in step S5 is further defined. The controller maintains a cyclic experience replay buffer, the capacity of which is dynamically set according to the product of the average number of packs per minute in the current batch and the duration of a standard production shift. For example, if 120 membrane packs are packed per minute and a shift is 8 hours (480 minutes), the product is 57600, but the upper limit is no more than 1000 samples, so the actual capacity is 1000. This design ensures that there are enough historical samples for training while controlling memory and computational overhead. The experience replay adopts a confidence-weighted mechanism. The weight of each sample is set comprehensively based on the similarity S obtained from acoustic detection and the actual membrane pack quality score given by the subsequent independent quality detection module. The closer the similarity S is to 1, the more reliable the acoustic judgment is; the higher the quality score, the better the packing effect. The weighted sum of the two is used as the importance weight of the sample, so that samples with high confidence and high-quality feedback play a greater role in model updates.

[0051] Each sample stored in the buffer contains the following information: an individual state vector (i.e., features such as the bottle's weight, ellipticity, and humidity), pre-compensation coefficients from the neural network output, acoustic detection results (e.g., pass or specific anomaly category), and an actual membrane pack quality score provided by a subsequent quality inspection module independent of acoustic detection. This subsequent quality inspection module can be a membrane pack tensile testing device or a membrane pack contour visual inspection device. The former tests the tightness of the membrane pack by applying tensile force, while the latter compares the regularity of the membrane pack's shape by taking photos; its output serves as a supervisory signal. Each time a new quality score is obtained, the system adds the sample to the buffer and randomly draws a batch of samples from the buffer with probability p for training. The probability p increases linearly from 0.1 to 0.5 with the buffer's fill rate; for example, the drawing probability is 0.1 when the buffer is empty and 0.5 when the buffer is full. This progressive drawing strategy avoids overfitting to a small number of samples in the early stages of training and makes full use of rich historical data in the later stages.

[0052] A loss function is calculated for the extracted batch samples. This loss function consists of two summed parts. The first part is the mean squared error prediction error term, which measures the mean squared error between the pre-compensation coefficient vector output by the neural network and the expected compensation coefficient vector. The expected compensation coefficient vector is generated in real time by the controller based on the acoustic detection judgment results according to preset process rules (e.g., when the bottle cap is judged to be loose, the historical average compensation coefficient is taken and the clamping force is increased by 0.05 and the subsequent temperature is decreased by 0.05; when the arrangement is judged to be loose, the clamping force is increased by 0.05 and the drying air temperature is increased by 0.025; when it is judged to be qualified, the current network output is directly taken). The second part is the elastic weight consolidation regularization term, which uses the Fisher information matrix to perform a weighted square summation of the changes in each weight parameter to limit the drift of important weights. The system uses an adaptive moment estimation optimizer to dynamically adjust the learning rate, with the initial learning rate set to 0.0005 and the gradient clipping threshold set to 1.0. After every 100 incremental updates, the controller selects the 20 samples with the largest Mahalanobis distance from the current batch from the buffer for replay training to improve the model's generalization ability to marginal distributions. It is important to note that the expected compensation coefficient vector in this step is not generated automatically based on acoustic detection results, but rather employs a two-layer mechanism of "objective quality scoring as the primary factor and acoustic judgment as the auxiliary factor for classification." Specifically, the controller first obtains the actual membrane package quality score provided by a subsequent quality detection module (e.g., a membrane package contour visual inspection device or a membrane package tensile testing device), independent of acoustic detection. This score is an objective, non-cyclic true value and serves as the primary basis for generating the expected compensation coefficient. Simultaneously, the acoustic detection judgment results (such as "abnormally loose bottle cap," "abnormally loose arrangement," or "qualified") are only used as defect type identifiers to select the corresponding adjustment template from a preset process rule library. Their values ​​(e.g., similarity S, energy difference ΔE, etc.) do not directly participate in the numerical calculation of the expected compensation coefficient. The following are exemplary preset process rules: When the actual membrane package quality score is lower than the qualified threshold (e.g., lower than 6 out of 10), and the acoustic judgment result is "abnormally loose bottle cap", the expected compensation coefficient is set as follows: clamping force increases by +0.05 (relative value), downstream temperature decreases by -0.05, and drying air temperature remains unchanged; when the actual membrane package quality score is lower than the qualified threshold, and the acoustic judgment result is "abnormally loose arrangement", the expected compensation coefficient is set as follows: clamping force increases by +0.05, drying air temperature increases by +0.025, and downstream temperature remains unchanged; when the actual membrane package quality score is higher than the excellent threshold (e.g., ≥9 points), the current neural network output is considered to be optimal, and the expected compensation coefficient is directly taken as the pre-compensation coefficient of the current neural network output. In other cases, the expected compensation coefficient is generated by the controller according to the proportional integral rule based on the difference between the quality score and the target value.

[0053] In the above design, independent quality scoring ensures the objectivity of the supervision signal, avoiding the neural network from falling into the trap of "self-verification"; acoustic judgment results only assist in distinguishing defect types, making the compensation rules more targeted. The combination of these two ensures the correct direction of online incremental learning.

[0054] Compared to simple online learning or training on a fixed dataset, this implementation effectively balances the utilization efficiency of new and old samples by combining a recurrent experience replay buffer with confidence weighting and a progressive extraction strategy, preventing the neural network model from forgetting historical knowledge. The introduction of a resilient weight consolidation regularization term alleviates the catastrophic forgetting problem, enabling the neural network model to maintain accurate compensation for old conditions while continuously learning new ones. An adaptive moment estimation optimizer and gradient pruning further improve the stability and convergence speed of the training process. By replaying samples that differ most from the current batch, the robustness of the neural network model to anomalies and marginal cases is enhanced.

[0055] In one specific implementation, the neural network in step S3 and the neural network updated by online incremental learning in S5 are the same network. Its network parameters are determined by offline training at the initial moment and are asynchronously updated by the feedback signal generated in step S5 after each packaging is completed. The controller sets a network parameter version number, which increments after each incremental update. The controller updates the network parameter version uniformly after each production batch ends, and the network parameters used for inference within the same batch remain unchanged. When the feedback signal in step S5 is qualified for three consecutive batches, the controller automatically reduces the learning rate of incremental learning to 50% of the original value to prevent over-adjustment of the converged model.

[0056] In the above implementation, the neural network in step S3 and the neural network updated by online incremental learning in step S5 are the same network. The initial parameters of this network are determined by offline training, that is, basic training is completed using a large amount of historical data before the system is put into use. During production operation, after each packaging is completed, the acoustic detection feedback signal generated in step S5 will trigger an asynchronous update of the network. The so-called asynchronous update means that the inference process (outputting pre-compensation coefficients according to the current bottle state in step S3) and the learning process (adjusting network parameters according to the feedback signal in step S5) do not occur on the same timeline. Inference can continue to use the old version of parameters, while learning is carried out in the background. The network parameter version is updated uniformly after each production batch, and the network parameters used for inference within the same batch remain unchanged. The controller sets a network parameter version number, for example, the initial version is V1.0, and the version number is incremented to V1.1, V1.2, etc. after each incremental update. This strategy of fixing parameters within a batch avoids the risk of continuous defective products due to model mutation caused by a single abnormal update.

[0057] When the feedback signals in step S5 are qualified for three consecutive batches, that is, there are no film package quality defects or only a very small number of acceptable defects in three consecutive production batches, the controller automatically reduces the learning rate of incremental learning to 50% of the original value. For example, if the original learning rate is 0.0005, it is reduced to 0.00025. The function of this mechanism is to prevent the performance from degrading due to overlearning the noise in new samples when the neural network model has converged to a good state. The qualified judgment for three consecutive batches provides sufficient confidence to indicate that the current model is good enough and does not need to be updated significantly. When defective batches appear again later, the learning rate can be restored to the original value or dynamically adjusted according to the defect frequency.

[0058] Through version number management and adaptive learning rate reduction, the system achieves a balance between the stability and adaptability of the neural network model. The version number enables the system to easily roll back to a previous stable version. If the defect rate increases after a certain incremental update, the operator can manually switch back to the old version. The automatic reduction of the learning rate avoids unnecessary parameter perturbations during stable production while retaining the ability to quickly restore a high learning rate when quality fluctuations occur again. This design is particularly suitable for automated packaging production lines that operate continuously for a long time.

[0059] Compared with the traditional methods of separating the inference network and the update network or using a fixed learning rate, in this embodiment, through the asynchronous update and the mechanism of fixing parameters within a batch, the neural network can continuously improve without interrupting production while avoiding the risk of model mutation within a batch. The strategy of automatically reducing the learning rate after consecutive qualified batches effectively prevents the over-adjustment of the converged model and reduces the performance fluctuations caused by noise samples.

[0060] In one specific embodiment, step S1 further includes: setting a photoelectric trigger switch at the entrance of the conveyor track, and generating a start signal when a PET bottle passes through the photoelectric trigger switch; The controller calculates the predicted moments when the PET bottle arrives at the weight detection module, the three-dimensional vision detection module, and the surface humidity detection module respectively according to the real-time running speed of the conveyor track, and generates the trigger delays for each module based on the predicted moments; each module starts sampling according to the corresponding delay, and aligns the timestamps of the sampled data to the start signal moment of the photoelectric trigger switch; The controller also regularly collects the speed fluctuation signals of the conveyor track, and when the speed fluctuation exceeds ±5%, dynamically corrects the trigger delays of subsequent PET bottles.

[0061] In the above embodiment, a photoelectric trigger switch is installed at the entrance of the conveyor belt. When a PET bottle passes through the switch, the switch generates a start signal. This signal serves as the time reference for the bottle. The controller collects the operating speed of the conveyor belt in real time, for example, by obtaining the linear velocity of the belt through an encoder or laser speed sensor. Based on this speed, the controller calculates the predicted time required for the PET bottle to reach the weight detection module, the 3D vision detection module, and the surface humidity detection module from the entrance. Since the three modules are arranged sequentially on the conveyor belt and are spaced at a fixed distance from each other, the arrival time at each module is different. The controller generates trigger delays for each module based on these predicted times; for example, the delay for the weight detection module is 1.2 seconds, the delay for the 3D vision module is 2.5 seconds, and the delay for the humidity detection module is 3.8 seconds. Each module starts sampling according to its corresponding delay, and the timestamps of the sampled data are uniformly aligned to the start signal time of the photoelectric trigger switch. In this way, even if the sampling times of the three modules are different, the final data can be associated with the same bottle using the same time reference.

[0062] The operating speed of the conveyor belt is not completely constant in actual production and may fluctuate due to load changes, motor fluctuations, or mechanical friction. The controller periodically collects speed fluctuation signals from the conveyor belt, for example, sampling the speed value every 0.5 seconds. When a speed fluctuation exceeding ±5% is detected, for example, a drop from 10 meters per minute to 9.4 meters per minute, the controller recalculates the predicted arrival time of subsequent PET bottles at each module based on the new speed and dynamically corrects the corresponding trigger delay. This correction is not done all at once, but is calculated for each newly entering bottle based on the current real-time speed, thus ensuring that each bottle receives an accurate sampling time.

[0063] Through the aforementioned photoelectric trigger alignment and dynamic compensation for speed fluctuations, the weight, 3D contour, and humidity data of the same bottle can still be correctly correlated even if the conveyor speed changes. This is crucial for subsequent individual state vector construction, because if the humidity data of a bottle is mistakenly correlated with that of a previous bottle, it will lead to incorrect input to the neural network and invalid compensation coefficients at the output. This mechanism also reduces the cost of installing multiple independent trigger sensors on the conveyor, as only one entry trigger is needed to provide a reference for all detection modules.

[0064] Compared to methods where each detection module uses its own independent photoelectric sensor for triggering, this implementation combines a single-entry photoelectric trigger with speed prediction and delay generation, reducing the number of sensors and installation complexity, while avoiding data misalignment caused by differences in the response times of different sensors. The dynamic speed fluctuation correction function enables the system to maintain high-precision data alignment even when the chain conveyor speed is unstable, thereby improving the accuracy of individual state vectors and the effectiveness of subsequent compensation control.

[0065] In one specific implementation, the neural network described in step S3 employs a multi-source domain adaptive transfer learning strategy during the offline training phase: Historical packaging data under at least three different bottle types, two different materials, and two different environmental humidity conditions were collected in advance, and the source domain model was trained accordingly. When the actual bottle type or environmental conditions during on-site production change, the controller calculates the maximum mean difference between the distribution characteristics of the state vectors of the first 20 PET bottles in the current batch and the data distribution of each source domain. It selects the source domain model with the smallest difference as the initial model and uses the detection data of the first 50 PET bottles in the current batch to quickly adapt and fine-tune the model. During fine-tuning, only the weight parameters of the last two fully connected layers of the neural network are updated.

[0066] In the above implementation, the neural network in step S3 is specified to employ a multi-source domain adaptive transfer learning strategy during the offline training phase. Before the system is put into actual production, developers collect historical packaging data under at least three different bottle types, two different materials, and two environmental humidity conditions. For example, bottle types may include round, square, and oval bottles; materials may include ordinary PET and high-temperature resistant PET; and environmental humidity conditions may include a low-humidity environment of 30% relative humidity and a high-humidity environment of 80% relative humidity. For each combination, a source domain model is trained, meaning each source domain corresponds to a trained neural network. These source domain models are stored in the controller's model library for subsequent use.

[0067] When the actual bottle type or environmental conditions during on-site production change, such as switching from producing round bottles to square bottles, or from a dry winter to a humid summer, the controller first collects the state vectors of the first 20 PET bottles in the current batch. It calculates the distribution characteristics of these 20 vectors, such as the mean and covariance matrix, and then performs a maximum mean difference calculation with the data distributions of each source domain. Maximum mean difference is a non-parametric statistic that measures the difference between two distributions; a smaller value indicates a more similar distribution. The controller selects the source domain model with the smallest difference as the initial model. For example, if the current distribution of square bottles has the smallest difference from the pre-trained square bottle source domain model, then that model is directly selected without starting training from scratch.

[0068] After selecting the source domain model, the controller uses the detection data of the first 50 PET bottles in the current batch to quickly adapt and fine-tune the model. During fine-tuning, only the weight parameters of the last two fully connected layers of the neural network are updated, while the parameters of all preceding layers remain unchanged. This is because shallow layers of a neural network typically learn general feature extractors (such as the basic shape of the bottle outline), while deeper layers are more task-specific (such as mapping compensation coefficients). By fine-tuning only the last two layers, it is possible to quickly adapt to new bottle shapes or environmental conditions with limited data, while avoiding overfitting. The fine-tuning process typically requires only a few dozen iterations, significantly shortening the model adaptation time.

[0069] Compared to the traditional method of retraining the neural network for each product type change, this implementation uses multi-source domain pre-training and source domain selection guided by the maximum mean difference. This allows the system to quickly find the initial model that best matches the current operating conditions, greatly reducing the workload and time of offline training. The strategy of only fine-tuning the last two layers significantly reduces the need for new sample data while maintaining adaptability; even data from only 50 bottles can achieve effective adaptation. This enables the production line to quickly switch between various bottle types and seasonal changes, improving the equipment's flexible production capabilities.

[0070] In one specific embodiment, the advance adjustment in step S4 further includes: The controller calculates the actual conveying distance from each detection point to the drying nozzle, shaping fixture, and heat shrink oven entrance based on the position coordinates of the weight detection module, 3D vision detection module, and surface humidity detection module arranged sequentially on the conveyor chain. Combined with the real-time speed of the conveyor chain, it obtains the advance adjustment time window corresponding to the drying air temperature, shaping clamping force, and heat shrink oven zone temperature, as well as the advance adjustment time window for the air volume / speed of each zone of the heat shrink oven. When the individual state vector of the same PET bottle triggers multiple compensation coefficients, the controller stores the pre-compensation coefficients in a time-sorted compensation queue according to the spatial order of each actuator, and extracts the compensation coefficients from the queue and applies them within the corresponding time window before each actuator's action. If the time windows in the compensation queue overlap due to sudden changes in conveyor speed, the controller adopts a priority arbitration strategy: clamping force compensation takes precedence over temperature compensation, and the overlapping part is based on clamping force compensation. If the time windows are misaligned, the starting point of the window is recalculated according to the actual arrival time.

[0071] In the above embodiment, the advance adjustment in step S4 is further defined. The controller calculates the actual conveying distance from each detection point to each actuator based on the position coordinates of the weight detection module, 3D vision detection module, and surface moisture detection module arranged sequentially on the conveyor chain. The actuators include drying nozzles, shaping fixtures, and the entrances to each zone of the heat shrink oven. Combining the real-time speed of the conveyor chain, the controller obtains the advance adjustment time windows corresponding to the drying air temperature, shaping clamping force, and heat shrink oven zone temperature. For example, if the distance from the 3D vision detection module to the shaping fixture is 2 meters and the chain speed is 0.5 meters per second, the time window is 4 seconds; that is, the controller needs to issue a clamping force compensation command 4 seconds before the bottle reaches the shaping fixture. It should be noted that heat shrink ovens are typically tunnel-type structures with large thermal inertia of the internal heating elements, making it difficult to directly change the oven air temperature in a short time. Therefore, the advance adjustment of the heat shrink oven zone temperature described in this invention does not refer to directly changing the set temperature of the heating elements, but rather to pre-adjusting the circulating air volume or speed of the corresponding zone. By increasing or decreasing the fan speed, the flow rate of hot air blown onto the membrane pack can be quickly changed, thereby achieving an equivalent temperature regulation effect as the membrane pack passes through the zone. The response time is typically within 0.5 seconds, meeting the time window requirements for proactive adjustment. Furthermore, the controller can also indirectly regulate the heat shrinkage intensity by fine-tuning the instantaneous speed of the conveyor chain to change the residence time of the membrane pack in each temperature zone.

[0072] When the individual state vector of the same PET bottle triggers multiple compensation coefficients, such as when the drying temperature and shaping clamping force need to be adjusted simultaneously, the controller stores the pre-compensation coefficients in a time-sorted compensation queue according to the spatial order of each actuator on the conveyor chain. Each entry in the queue contains the compensation type, compensation value, and planned execution time. Within the corresponding time window before each actuator's action, the controller retrieves the compensation coefficient from the queue and sends it to the actuator. If a sudden change in conveyor speed, such as a sudden acceleration or deceleration, causes the originally staggered time windows to overlap, meaning two compensation commands need to be executed consecutively at the same moment or within a very short time, the controller adopts a priority arbitration strategy. Specifically, clamping force compensation takes precedence over temperature compensation because incorrect adjustment of clamping force may cause bottle deformation or disordered arrangement, while a slight delay in temperature adjustment has a relatively smaller impact. For overlapping time periods, clamping force compensation takes precedence, and temperature compensation can be appropriately delayed until after the overlap ends.

[0073] If a change in the conveyor speed causes a misalignment between the time window in the compensation queue and the actual arrival time of the bottle at the actuator—for example, a decrease in speed causing the bottle to arrive later than planned—the controller will recalculate the start point of each time window based on the new real-time speed and update the execution time in the compensation queue. This dynamic adjustment ensures that the compensation coefficient is applied at the correct time regardless of speed changes. The entire mechanism is continuously monitored by a background thread of the controller, monitoring the conveyor speed and queue status, requiring no manual intervention.

[0074] Compared to traditional compensation methods that use fixed delays or ignore speed fluctuations, this implementation achieves the orderly and precise application of multiple compensation coefficients by dynamically calculating and adjusting the time window and compensation queue in advance. The priority arbitration strategy resolves the problem of command conflicts during sudden speed changes, prioritizing the correct execution of clamping force and reducing the risk of equipment damage or product quality issues. The recalculation function after time window misalignment gives the system good robustness to chain speed changes, ensuring accurate compensation timing even on variable speed production lines.

[0075] In one specific implementation, during online incremental learning, the controller also maintains a parameter coupling suppression matrix C, the dimension of which is equal to the number of pre-compensation coefficients in the neural network output, and the elements c in the parameter coupling suppression matrix C... ij This represents the cross-influence coefficient of the i-th compensation parameter on the j-th compensation parameter; Each time, the neural network outputs the original pre-compensation coefficient vector P. raw Then, the controller calculates the corrected compensation coefficient vector P. corr = P raw + C·(P raw - P hist ), where P hist The average compensation coefficient for the 10 most recent qualified bottles in history; When the drying air temperature compensation is positive and the clamping force compensation is negative, the corresponding cross-influence coefficient c in the parameter coupling suppression matrix C 12 Set to 0.3 to suppress the offsetting effect of increased drying air temperature on the actual clamping force; After the controller completes 100 incremental updates, it adaptively adjusts the coupling suppression matrix based on the feedback results of acoustic detection: it counts the occurrence rate of loose and abnormal arrangement in the last 100 membrane packs. If the occurrence rate exceeds 5%, it increases the absolute value of the cross-influence coefficient of clamping force on drying air temperature by 0.05 each time, with an upper limit of 0.5; if the occurrence rate is less than 1%, it gradually decreases the coefficient.

[0076] In the above implementation, a parameter coupling suppression matrix C is introduced in online incremental learning. The dimension of this matrix is ​​equal to the number of pre-compensation coefficients output by the neural network. For example, if three compensation coefficients are output (drying air temperature, shaping clamping force, and heat shrink oven zone temperature), then the matrix is ​​3 rows and 3 columns. The element c in the matrix... ij This represents the cross-influence coefficient of the i-th compensation parameter on the j-th compensation parameter. For example, c 12 This indicates the cross-effect of changes in drying air temperature on the actual clamping force, c 21 This represents the cross-effect of clamping force changes on the drying air temperature effect. These coefficients are typically positive or negative, reflecting how a change in one compensation parameter enhances or weakens the effect of another. The original pre-compensation coefficient vector P is output by the neural network at each step. raw Then, the controller calculates the corrected compensation coefficient vector P. corr The purpose of this corrected formula is to decouple the mutual interference between parameters.

[0077] When both condensate coverage and ellipticity are high, the neural network outputs positive compensation for drying air temperature and negative compensation for clamping force. In this case, directly applying these two compensations might cause the increased drying air temperature to slightly soften the bottle surface, requiring a greater clamping force to achieve the desired shaping effect; that is, there is a negative coupling between drying air temperature and clamping force. To suppress this offsetting effect, this implementation uses the corresponding cross-influence coefficient c in the parameter coupling suppression matrix C. 12 The value is set to 0.3. This means that when the drying air temperature increases, the correction formula will automatically make additional adjustments to the clamping force compensation to offset the weakening effect of the increased drying air temperature on the actual clamping force. The specific value of this coefficient, 0.3, is the optimal reference value obtained from experiments on the mechanical properties of typical PET bottle materials.

[0078] After every 100 incremental updates, the controller adaptively adjusts the coupling suppression matrix based on feedback from acoustic detection. Specifically, the system calculates the incidence of "loose arrangement anomalies" in the last 100 membrane packs. If this incidence exceeds 5%, it indicates that the current clamping force may be insufficient or the drying air temperature may be too high, causing misalignment of the bottles. In this case, the absolute value of the cross-influence coefficient between clamping force and drying air temperature needs to be increased by 0.05 each time, but the upper limit is no more than 0.5. Conversely, if the incidence is below 1%, it indicates that the current cross-influence coefficient may be too large, leading to overcompensation. In this case, the coefficient is gradually decreased, for example, by 0.02 each time, but not below 0. This adaptive mechanism allows the coupling suppression matrix to automatically optimize with changes in production conditions, without the need for manual calibration.

[0079] Compared to traditional compensation methods that ignore the coupling relationship between parameters, this implementation reduces mutual interference between multiple compensation parameters by introducing a parameter coupling suppression matrix and decoupling the original compensation coefficients based on a modified formula. Especially when the drying air temperature and clamping force are in opposite directions, the offsetting effect of thermal and mechanical effects is effectively suppressed by setting a cross-influence coefficient. Adaptive adjustment based on acoustic feedback further enables the matrix coefficients to self-optimize, and the synergistic effect of compensation improves with increasing operating time. This directly improves the neatness of the membrane pack arrangement and the overall packaging quality.

[0080] In one specific implementation, during the asynchronous update process, the controller also sets up a core sample memory, which is used to store the sample with the largest Fisher information matrix trace in the historical failure cases. Each core sample contains its individual state vector, the corresponding pre-compensation coefficient, and the acoustic detection judgment result. When incremental learning is triggered, the controller first extracts a batch of samples from the recurrent experience replay buffer to calculate the first loss function. Then, it randomly extracts 20% of the batch samples from the core sample memory to calculate the second loss function. The second loss function uses the same mean squared error form as the first loss function but does not weight the regularization term. The first and second loss functions are added together with a weight of 0.7:0.3 to obtain the total loss function, which is used to update the neural network parameters. The core sample memory has a fixed capacity of 200 samples and adopts a clustering-based update strategy: the samples in the memory are divided into multiple clusters according to the defect type, and the sample with the largest loss gradient norm is retained in each cluster. When a new sample is added, it is assigned to its own cluster and the sample with the smallest gradient norm in that cluster is replaced. If the new sample represents a new defect type, a new cluster is created and the entire cluster with the smallest average gradient norm in all current clusters is deleted to maintain sample diversity. The calculation of the Fisher information matrix trace is performed only when core sample candidates are added, and a diagonal approximation of the Fisher information matrix is ​​used to reduce the amount of computation. It is calculated at most once per production batch.

[0081] In the above implementation, a core sample memory is further set up during the asynchronous update process. This memory is specifically used to store samples from historical failure cases that have the largest loss gradient norm under the current model parameters. The loss gradient norm is defined as the L2 norm of the gradient vector of the loss function with respect to the neural network parameters. The larger the value, the greater the impact of the sample on the current model parameter update, and the more critical it is to maintaining model performance. Each core sample contains an individual state vector, the corresponding pre-compensation coefficient, and the acoustic detection judgment result. The capacity of the memory is fixed at 200 samples, and a clustering-based update strategy is adopted to balance importance and diversity: the samples in the memory are divided into multiple clusters according to defect type, and the sample with the largest loss gradient norm calculated based on the latest model parameters is retained in each cluster. Since the model parameters change continuously during online learning, after each production batch (or after every 50 incremental updates), the controller recalculates the loss gradient norm of all samples in the memory based on the network parameters at that time, and updates the sample with the largest norm in each cluster, ensuring that "largest" always reflects the importance of the current model.

[0082] The specific process for adding a new sample is as follows: The controller first determines the defect type of the new sample (based on the acoustic detection result) and assigns it to the corresponding cluster; if there is no corresponding cluster for the defect type, a new cluster is created. Then, the loss gradient norm of the new sample is calculated based on the current model parameters and compared with the gradient norm of the existing samples in the cluster (if the cluster has not been recalculated in this batch, the current gradient norm of all samples in the cluster is temporarily calculated). If the gradient norm of the new sample is greater than the current minimum gradient norm in the cluster, the sample with the minimum gradient norm in the cluster is replaced with the new sample, thus ensuring that the sample with the maximum gradient norm is always retained in the cluster. If the creation of new clusters results in too many clusters (e.g., more than 20 clusters, each cluster retains an average of 10 samples, but each cluster can actually retain multiple samples to maintain quantity balance), the entire cluster with the minimum average gradient norm among all current clusters is deleted (i.e., all samples in the cluster are deleted) to free up space and maintain sample diversity. The average gradient norm of a cluster is defined as the arithmetic mean of the loss gradient norms of all samples in the cluster.

[0083] Through the aforementioned periodic recalculation and comparison mechanism based on the current model parameters, the principle of "retaining the sample with the largest loss gradient norm within each cluster" can be reliably achieved. This not only preserves samples with high information content but also avoids the crowding out of rare defect types by recent high-gradient samples, effectively maintaining sample diversity.

[0084] When incremental learning is triggered, the controller first draws a batch of samples from the recurrent experience replay buffer and calculates the first loss function. Then, it randomly draws 20% of the batch size of core samples from the core sample memory (e.g., if the first batch has 100 samples, then 20 core samples are drawn from the memory). A second loss function is calculated for these core samples. The second loss function uses the same mean squared error form as the first loss function, but does not weight the regularization term; it only focuses on the prediction error. The first and second loss functions are added together with a weighted average of 0.7 to 0.3 to obtain the total loss function, which is used to update the neural network parameters. This dual-loss mechanism ensures that the model does not forget rare but crucial failure cases while learning new samples.

[0085] To reduce computational overhead, the loss gradient norm is not calculated with every training iteration. Instead, it is calculated on demand using a combination of periodic recalculations: the loss gradient norm is calculated at most once per production batch (only when new samples are added or periodic recalculations are performed), and the calculation is further simplified by using a diagonal approximation of the Fisher information matrix. For periodic recalculations of samples already in the memory bank, since there is only one recalculation per batch and the total number of samples is only 200, the computational cost is well within the tolerance of the industrial controller. This approximation method has proven to be sufficiently accurate in practical applications while significantly saving computational resources.

[0086] The number of devices and processing scale described herein are for the purpose of simplifying the description of the invention. Applications, modifications, and variations of the invention will be readily apparent to those skilled in the art.

[0087] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details.

Claims

1. A method for status detection and control of an automatic packaging system for food-grade PET bottles, characterized in that, include: S1. A weight detection module, a 3D vision detection module, and a surface humidity detection module are sequentially set on the conveyor chain to obtain the weight value, bottle outline and cap sealing ring status, bottle ellipticity, distance between adjacent bottles, and condensate coverage of each PET bottle. The controller associates the detection data of the same PET bottle by timestamp to form an individual state vector and calculates the statistical distribution characteristics of the state vectors of N consecutive PET bottles. S2. The controller determines the overall quality fluctuation of the current batch of PET bottles based on statistical distribution characteristics: if the fluctuation is less than or equal to the first threshold, the standard packaging mode is executed; if the fluctuation is greater than the first threshold but less than or equal to the second threshold, the compensation packaging mode is executed; if the fluctuation is greater than the second threshold, a batch diversion warning is triggered. S3. In the compensation packaging mode, the individual state vector of each PET bottle is input into a pre-trained neural network, which outputs pre-compensation coefficients for drying air temperature, shaping clamping force, and heat shrink oven zone temperature; the neural network establishes an inverse mapping between individual defects and packaging parameters based on historical failure cases. S4. The controller makes advance adjustments according to the pre-compensation coefficient; when the condensate coverage rate is too high and the ellipticity is also too high, the drying air temperature compensation output by the neural network is positive, the clamping force compensation is negative, and the downstream temperature compensation is negative. S5. An acoustic detection device is installed at the outlet of the heat shrink oven to feed the impact spectrum analysis results back to the neural network for online incremental learning; The actual membrane pack quality score in the feedback signal is provided by a subsequent quality inspection module independent of the acoustic detection.

2. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that, The statistical distribution characteristics in step S2 are as follows: The controller constructs an N×m matrix from the individual state vectors of N consecutive PET bottles, where m is the dimension of the state vector; Calculate the covariance matrix Σ of the matrix, then calculate the Mahalanobis distance of each individual state vector relative to the mean of the current batch, and define the root mean square value of all Mahalanobis distances as the overall quality fluctuation degree. The first threshold and the second threshold are set as the 70th percentile and the 90th percentile of the root mean square value of the Mahalanobis distance of historical qualified batches, respectively, and these two percentiles are automatically updated after every 100 batches. When the root mean square value of the Mahalanobis distance is less than or equal to the 70th percentile, the fluctuation level is considered low, and the standard packaging mode is executed; when it is greater than the 70th percentile but less than or equal to the 90th percentile, the fluctuation level is considered moderate, and the compensated packaging mode is executed. Package mode; when the value is greater than the 90th percentile, it is determined that the fluctuation level is high, triggering a batch diversion warning.

3. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that, The knock spectrum analysis in step S5 is obtained through the following steps: S51. The pneumatic impactor in the acoustic detection device is activated by photoelectric trigger only when the membrane package reaches the predetermined impact position. The impact frequency is synchronized with the conveyor chain speed. The impact air pressure is 0.3±0.05MPa. The impact head is made of silicone material with a Shore A hardness of 50. The duration of a single impact is 0.1s. S52: The microphone acquires the tapping sound signal at a sampling rate of 20kHz, extracts the effective signal within a time window of 0.2s-0.5s after the tapping begins, and applies a Hanning window for windowing processing. S53. Perform a fast Fourier transform on the windowed signal to obtain the power spectral density in the 0-10kHz frequency domain. Divide the power spectral density into 15 frequency bands according to 1 / 3 octave bands and calculate the energy proportion of each frequency band. S54. The controller calculates the cosine similarity between the measured energy proportions of the 15 frequency bands and the pre-stored standard membrane package template. The standard membrane package template is established by the average spectrum of 10 consecutive qualified membrane packages. The similarity S = (A·B) / (‖A‖‖B‖), where A is the measured feature vector and B is the template feature vector. S55. Simultaneously calculate the difference ΔE between the energy percentage of the measured spectrum in the 2kHz-5kHz frequency band and the energy percentage of the standard membrane-coated template in the same frequency band. S56: If the similarity S is less than 0.85 and ΔE is greater than 0.1, it is judged as an abnormality of loose bottle cap; If S is less than 0.85 and ΔE is not greater than 0.1, it is determined to be foreign matter inside the membrane package or abnormally loose arrangement; otherwise, it is determined to be qualified. S57: The abnormal category or qualified status in the judgment result and the measured spectrum feature vector are used as feedback signals and input into the neural network for online incremental learning, wherein the weight coefficient of the abnormal sample is set to twice that of the qualified sample.

4. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that... It lies in, The online incremental learning in step S5 specifically includes: The controller maintains a cyclic experience playback buffer, the capacity of which is dynamically set according to the product of the average number of packages per minute in the current batch and the duration of a standard production shift, with an upper limit of no more than 1000 samples; the experience playback adopts confidence weighting, and the sample weight is set comprehensively based on the similarity S of acoustic detection and subsequent quality score. Each sample contains an individual state vector, pre-compensation coefficients, acoustic detection judgment results, and an actual membrane pack quality score provided by a subsequent quality detection module independent of acoustic detection. The subsequent quality detection module includes a membrane pack tensile testing device or a membrane pack contour visual detection device, the output of which serves as a supervision signal. After each new quality score is obtained, the sample is added to a buffer, and a batch of samples is randomly drawn from the buffer with probability p, where p increases linearly from 0.1 to 0.5 with the buffer fill rate. A loss function is calculated for the extracted batch samples. This loss function consists of the sum of a mean squared error prediction error term and an elastic weight consolidation regularization term. The mean squared error prediction error term calculates the mean squared error between the pre-compensation coefficient vector and the expected compensation coefficient vector output by the neural network. The expected compensation coefficient vector is generated in real time by the controller based on the acoustic detection judgment results and subsequent quality scores, according to preset process rules. The regularization term uses the Fisher information matrix to perform a weighted square summation of the changes in each weight parameter in the neural network to limit the drift amplitude of important weights. An adaptive moment estimation optimizer is used to dynamically adjust the learning rate. The initial learning rate is set to a preset value between 0.001 and 0.0001, and the gradient clipping threshold is set to 1.

0. After each preset number of incremental updates, the controller selects a predetermined number of samples from the buffer that have the largest Mahalanobis distance from the current batch for replay training.

5. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that... It lies in, The neural network in step S3 is the same neural network updated by online incremental learning in S5. Its network parameters are determined by offline training at the initial moment and are asynchronously updated by the feedback signal generated in step S5 after each packaging is completed. The controller is set to a network parameter version number, which is incremented after each incremental update. The controller updates the network parameter version uniformly after the end of each production batch, and the network parameters used for inference within the same batch remain unchanged. When the feedback signal in step S5 is qualified for three consecutive batches, the controller automatically reduces the learning rate of incremental learning to 50% of the original value to prevent over-adjustment of the converged model.

6. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that... It lies in, Step S1 also includes: setting a photoelectric trigger switch at the entrance of the conveyor chain, and generating a start signal when the PET bottle passes through the photoelectric trigger switch; The controller calculates the predicted arrival times of the PET bottles at the weight detection module, 3D vision detection module, and surface moisture detection module based on the real-time running speed of the conveyor chain, and generates the trigger delay for each module based on the predicted time; each module starts sampling according to the corresponding delay and aligns the timestamp of the sampled data to the start signal time of the photoelectric trigger switch; The controller also periodically collects speed fluctuation signals from the conveyor chain. When the speed fluctuation exceeds ±5%, it dynamically corrects the trigger delay of subsequent PET bottles.

7. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that... It lies in, The neural network described in step S3 employs a multi-source domain adaptive transfer learning strategy during the offline training phase: Historical packaging data under at least three different bottle types, two different materials, and two different environmental humidity conditions were collected in advance, and the source domain model was trained accordingly. When the actual bottle type or environmental conditions during on-site production change, the controller calculates the maximum mean difference between the distribution characteristics of the state vectors of the first 20 PET bottles in the current batch and the data distribution of each source domain. It selects the source domain model with the smallest difference as the initial model and uses the detection data of the first 50 PET bottles in the current batch to quickly adapt and fine-tune the model. During fine-tuning, only the weight parameters of the last two fully connected layers of the neural network are updated.

8. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that, The advance adjustment in step S4 further includes: The controller calculates the actual conveying distance from each detection point to the drying nozzle, shaping fixture, and heat shrink oven entrance based on the position coordinates of the weight detection module, 3D vision detection module, and surface humidity detection module arranged sequentially on the conveyor chain. Combined with the real-time speed of the conveyor chain, it obtains the advance adjustment time window corresponding to the drying air temperature, shaping clamping force, and heat shrink oven zone temperature, as well as the advance adjustment time window for the air volume / speed of each zone of the heat shrink oven. When the individual state vector of the same PET bottle triggers multiple compensation coefficients, the controller stores the pre-compensation coefficients in a time-sorted compensation queue according to the spatial order of each actuator, and extracts the compensation coefficients from the queue and applies them within the corresponding time window before each actuator's action. If the time windows in the compensation queue overlap due to sudden changes in conveyor speed, the controller adopts a priority arbitration strategy: clamping force compensation takes precedence over temperature compensation, and the overlapping part is based on clamping force compensation. If the time windows are misaligned, the starting point of the window is recalculated according to the actual arrival time.

9. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 1, characterized in that... It lies in, In online incremental learning, the controller also maintains a parameter coupling suppression matrix C, the dimension of which is equal to the number of pre-compensation coefficients in the neural network output, and the elements c in the parameter coupling suppression matrix C are... ij This represents the cross-influence coefficient of the i-th compensation parameter on the j-th compensation parameter; Each time, the neural network outputs the original pre-compensation coefficient vector P. raw Then, the controller calculates the corrected compensation coefficient vector P. corr = P raw + C·(P raw - P hist ), where P hist The average compensation coefficient for the 10 most recent qualified bottles in history; When the drying air temperature compensation is positive and the clamping force compensation is negative, the corresponding cross-influence coefficient c in the parameter coupling suppression matrix C 12 Set to 0.3 to suppress the offsetting effect of increased drying air temperature on the actual clamping force; After the controller completes 100 incremental updates, it adaptively adjusts the coupling suppression matrix based on the feedback results of acoustic detection: it counts the occurrence rate of loose and abnormal arrangement in the last 100 membrane packs. If the occurrence rate exceeds 5%, it increases the absolute value of the cross-influence coefficient of clamping force on drying air temperature by 0.05 each time, with an upper limit of 0.5; if the occurrence rate is less than 1%, it gradually decreases the coefficient.

10. The status detection and control method of the automatic packaging system for food-grade PET bottles as described in claim 5, characterized in that, During the asynchronous update process, the controller also sets up a core sample memory, which is used to store the sample with the largest Fisher information matrix trace in the historical failure cases. Each core sample contains its individual state vector, the corresponding pre-compensation coefficient, and the acoustic detection judgment result. When incremental learning is triggered, the controller first extracts a batch of samples from the recurrent experience replay buffer to calculate the first loss function. Then, it randomly extracts 20% of the batch samples from the core sample memory to calculate the second loss function. The second loss function uses the same mean squared error form as the first loss function but does not weight the regularization term. The first and second loss functions are added together with a weight of 0.7:0.3 to obtain the total loss function, which is used to update the neural network parameters. The core sample memory has a fixed capacity of 200 samples and adopts a clustering-based update strategy: the samples in the memory are divided into multiple clusters according to the defect type, and the sample with the largest loss gradient norm is retained in each cluster. When a new sample is added, it is assigned to its own cluster and the sample with the smallest gradient norm in that cluster is replaced. If the new sample represents a new defect type, a new cluster is created and the entire cluster with the smallest average gradient norm in all current clusters is deleted to maintain sample diversity. The calculation of the Fisher information matrix trace is performed only when core sample candidates are added, and a diagonal approximation of the Fisher information matrix is ​​used to reduce the amount of computation. It is calculated at most once per production batch.