Quality inspection method and system for metal content in wet tissue based on multi-modal data fusion
By using a hardware synchronization beacon module and a bidirectional attention fusion model, the problem of misaligned data acquisition timing was solved, enabling high-precision detection and decision transparency of tiny metallic foreign objects in wet wipes, thereby improving production efficiency and quality management.
Patent Information
- Application Number
- CN202511446463.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing intelligent quality inspection technologies suffer from inconsistent data due to misaligned data acquisition timing in high-speed production line environments, insufficient reliability of artificial intelligence models, and opaque decision-making processes, making it difficult to achieve high-precision and high-reliability detection of tiny metal foreign objects in wet wipes.
By deploying hardware synchronization beacon modules to ensure precise spatiotemporal synchronization of multimodal data, and combining bidirectional attention fusion and attribution deep learning models, precise spatiotemporal synchronization of data sources is achieved, and positive prediction and reverse attribution are performed to ensure data consistency and decision transparency.
It achieves high-precision and high-reliability detection of tiny metal foreign objects in wet wipes, makes the decision-making process transparent, shortens the fault diagnosis time, and improves production efficiency and product quality management.
Smart Images

Figure CN120932767A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method and system for quality inspection of metal content in wet wipes based on multimodal data fusion. Background Technology
[0002] In modern industrial manufacturing, especially in industries with stringent requirements for product purity and safety such as hygiene products, food, and pharmaceuticals, efficient and accurate online detection of minute metal foreign objects that may be introduced during production is crucial for ensuring product quality, maintaining brand reputation, and protecting consumer health. As a daily consumer product that comes into direct contact with the human body, wet wipes require particularly stringent quality control during production; any residual metal fragments can pose a serious safety hazard. Traditional metal foreign object detection technologies typically include metal detectors based on electromagnetic induction or X-ray detection systems utilizing density difference imaging, providing fundamental guarantees for quality control in industrial production during a specific historical period. Specifically, metal detectors generate an alternating magnetic field; when a conductive metal foreign object passes through, it causes magnetic field disturbances that are captured by the sensor, triggering an alarm. This method is relatively sensitive to ferromagnetic metals and its cost is relatively controllable. X-ray detection, on the other hand, emits X-rays that penetrate the product, generating grayscale images based on the differences in the absorption rates of different substances, thereby identifying foreign objects such as metals with densities much higher than the product's substrate. It is capable of detecting various metals and high-density non-metallic foreign objects.
[0003] However, as related industries continue to evolve towards higher speeds, greater precision, and greater intelligence, and as consumers' demands for product quality continue to rise, the inherent characteristics of the aforementioned traditional testing methods at the principle level are gradually revealing their limitations in addressing new challenges. On the one hand, for thin and lightweight products moving rapidly on high-speed conveyor belts, the electromagnetic field changes or X-ray absorption differences caused by extremely small or thin-sheet / filament-shaped metallic foreign objects are very weak and easily drowned out by background noise, leading to a higher false negative rate with traditional methods. On the other hand, detection methods relying solely on physical properties cannot provide in-depth information about the source of contaminants. In view of this, the industry has begun to explore intelligent quality inspection technologies using multimodal data fusion, attempting to construct a more comprehensive and three-dimensional product quality profile by integrating data sources from different dimensions. Such technologies typically combine machine vision systems to capture abnormal color spots on the product surface, utilize spectral analysis technology to detect the chemical composition characteristics of materials, and correlate with production process data from the manufacturing execution system. These heterogeneous data are then input into an artificial intelligence model for comprehensive analysis. This paradigm of fusion analysis can theoretically significantly improve the sensitivity and accuracy of detection, representing the development direction of intelligent quality inspection technology.
[0004] Nevertheless, when translating multimodal fusion technology from laboratory concepts into rigorous industrial production practices, a deep-seated and less obvious technical contradiction emerges. This contradiction stems from the conflict between the physical reality of data acquisition in high-speed production line environments and the fundamental data quality requirements of artificial intelligence models. At its core, current multimodal solutions generally follow a technical path of independent acquisition and post-processing fusion. This means that each subsystem, such as visual sensors, spectrometers, and production log systems, is physically deployed and operates independently, generating data streams based on its own clock and processing cycle. Finally, these data streams are correlated in the central processing unit through software-level timestamp alignment. On modern high-speed production lines that process tens or even hundreds of products per second, the inherent flaws of this architecture are amplified dramatically. The inherent millisecond-level processing latency between different sensors and systems, network jitter, and minor clock asynchrony all contribute to a timing misalignment problem. This means that image data, spectral data, and production log data used by the model to describe the same wet wipe product sample may actually originate from physically adjacent but not identical products. This kind of data contamination undermines the authenticity and consistency of input data at its source, making the data foundation upon which AI model training is based distorted and unreliable. Consequently, the black-box problem of the decision-making process of a model trained and reasoned based on contaminated data becomes increasingly intractable. Even if the model makes an unqualified judgment, operators cannot be certain whether the judgment is based on a real defect pointed to by multiple dimensions of features, or merely a logical illusion caused by data misalignment. In this situation, any attempt to explain or trace the model's decisions becomes meaningless, because the object of analysis is a logically invalid set of data from the outset. This not only hinders the effective handling of individual alarms but also fundamentally severs the path of using quality inspection data to feed back into production processes and achieve a closed-loop quality management system.
[0005] Therefore, how to construct a low-level mechanism from the physical source of data collection that can ensure the precise synchronization of multimodal data in the spatiotemporal dimension, and on this basis, develop an intelligent analysis method that can not only make accurate judgments, but also transparently attribute and trace the basis of the judgments, so as to truly overcome the inherent contradiction between data integrity and model interpretability in existing technologies, has become a key challenge and a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0006] This invention discloses a method and system for quality inspection of metal content in wet wipes based on multimodal data fusion. It aims to solve the technical challenges of existing intelligent quality inspection technologies in high-speed production line environments, such as data inconsistency due to misaligned data acquisition timing, insufficient reliability of artificial intelligence models, and opaque decision-making processes. This invention constructs a low-level mechanism that ensures precise spatiotemporal synchronization of multimodal data from the physical source of data acquisition, and combines this with a deep learning model possessing both positive prediction and reverse attribution capabilities. This enables high-precision and high-reliability detection of minute metallic foreign objects in wet wipes, and provides deterministic and transparent traceability of the judgment criteria.
[0007] To achieve the above objectives, the present invention provides a method for quality inspection of metal content in wet wipes based on multimodal data fusion, the method comprising the following steps: First, a hardware synchronization beacon module is deployed at the testing station where the wet wipes under test pass. When a specific physical reference point of the wet wipes under test enters the preset trigger area of the hardware synchronization beacon module, the high-frequency stable clock source and unique identifier generation logic circuit embedded in the hardware synchronization beacon module immediately generate a globally unique synchronization identifier containing a nanosecond-level precision timestamp, and synchronously send a differential trigger signal with high anti-interference characteristics to a visual data acquisition unit, a spectral data acquisition unit, and a production process log acquisition unit.
[0008] Secondly, the visual data acquisition unit, spectral data acquisition unit, and production process log acquisition unit execute their respective data acquisition operations in parallel at the same moment they receive the differential trigger signal. Specifically, the visual data acquisition unit captures a frame of digital image data covering the surface of the wet wipe product under test; the spectral data acquisition unit acquires the reflection or transmission spectral data of predetermined detection points on the surface of the wet wipe product under test; and the production process log acquisition unit initiates a precise query to the manufacturing execution system's backend database or real-time data interface to obtain structured log data associated with the current production status. After completing data acquisition, each acquisition unit immediately rigidly binds the obtained synchronization identifier to the acquired data, forming a data unit containing the synchronization identifier and data payload.
[0009] Next, the data units generated by the three data acquisition units, each bound with the same synchronization identifier, are transmitted to a data aggregation and preprocessing unit. This unit uses the synchronization identifier as a unique key to aggregate the image data, spectral data, and structured log data into a structured multimodal data packet. This data packet follows a predefined, strict data pattern, ensuring an absolute one-to-one correspondence between different modalities.
[0010] The structured multimodal data packets are then fed into a pre-trained bidirectional attention fusion and attribution model deployed in a central processing unit. The model performs a forward reasoning process in which an intermodal collaborative attention mechanism within the model weights and fuses deep features from different modalities, calculating and outputting a quantified risk score characterizing the likelihood of a metallic foreign object in the tested wet wipe product.
[0011] Finally, a threshold judgment is performed on the risk score. When the risk score exceeds a preset risk threshold, the model's reverse attribution function is immediately activated. This function utilizes the attention weight matrix generated during the forward inference process and stored within the model to trace back and calculate the contribution of each input feature to the final high-risk score. Accordingly, a set of key input features that play a decisive role in the high-risk judgment is accurately identified and output. These key input features are further associated with their original modalities and presented in specific forms, including one or more specific pixel coordinate regions in the digital image, one or more specific absorption or reflection peak bands in the spectral data, and one or more specific parameter entries in the structured log data. Furthermore, the risk score and the identified key input features, which are enhanced with highlighted or marked visual processing, are presented together on the human-computer interaction interface, providing operators with direct and clear decision-making basis.
[0012] A further improvement to the aforementioned technical solution is that the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator, which provides a system master clock signal with a frequency stability better than one part per million. The hardware synchronization beacon module also includes a field-programmable gate array (FPGA) chip. The internal logic circuitry of the FPGA chip is responsible for generating a synchronization identifier conforming to a universally unique dictionary-ordered identifier specification based on the system master clock signal, and integrates a low-voltage differential signal transmitter to output the trigger signal in differential pairs, thereby ensuring a high signal-to-noise ratio and transmission integrity in industrial electromagnetic environments.
[0013] A further improvement to the aforementioned technical solution is that the core component of the visual data acquisition unit is an industrial camera employing a CMOS image sensor with a global shutter function. Its pixel resolution is 2048 x 2048 pixels, and it is equipped with a telecentric lens with a focal length of 75mm to eliminate perspective distortion caused by variations in the thickness of the wet wipes being tested. The industrial camera connects to the system via a gigabit Ethernet interface supporting a precise time protocol, ensuring that its internal clock remains synchronized with the clock source of the hardware synchronization beacon module at the sub-microsecond level. The spectral data acquisition unit uses a near-infrared spectrometer built based on Fabry-Perot interferometry technology, with an operating wavelength range covering 900 nm to 1700 nm and a spectral resolution of 10 nm. The production process log acquisition unit is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the relational database of the manufacturing execution system through an open database connection interface. Upon receiving a trigger signal, it executes a structured query statement containing precise timestamp matching conditions to obtain parameters such as the batch number, running speed, cutter number used, and cumulative working time of the cutter in an atomic operation manner.
[0014] A further improvement to the aforementioned technical solution is that the data aggregation and preprocessing unit is deployed on an edge computing device. The structured multimodal data packets are encapsulated using JavaScript object representation. This data packet structure includes the following main fields: a string field storing a universally unique dictionary sorting identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and a nested modal data object field. The modal data object field contains a visual data object field, which further includes metadata such as image format, Base64-encoded image data, resolution, and exposure time. The modal data object field also contains a spectral data object field, which includes an array recording all sampled wavelength points, an array recording the normalized reflection intensity values for each wavelength point, and metadata such as integration time. The modal data object field also contains a process log data object field, which directly stores structured log data obtained from the manufacturing execution system in key-value pairs, including parameters such as batch number, production line speed, cutter number, and cumulative cutter working time. After generating the data packet, the data aggregation and preprocessing unit transmits the data packet asynchronously and reliably to the central processing unit through a message-oriented middleware, specifically a zero-message queue, in a publish-subscribe mode.
[0015] A further improvement to the aforementioned technical solution is that the specific network structure of the bidirectional attention fusion and attribution model is designed as a deep neural network architecture consisting of three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head.
[0016] The three parallel feature extractor branches are as follows: The visual feature extractor uses a pre-trained residual network without a top classification layer as its backbone. The input is a normalized 2048x2048x1 single-channel grayscale image tensor, and the output is a deep feature map with dimensions of 64x64x2048.
[0017] The spectral feature extractor employs a one-dimensional transformer network consisting of six stacked encoder layers. The input is a spectral intensity sequence with a dimension of 81 by 1, corresponding to 81 wavelength points. After position encoding, the sequence is fed into the network, and the output is a sequence feature embedding with a dimension of 81 by 512.
[0018] The log feature extractor employs a multilayer perceptron. For categorical log data, an embedding layer is used to convert it into a dense vector; for numerical log data, standardization is performed. All processed log features are concatenated into a one-dimensional vector, which is then passed through the multilayer perceptron network to output a log feature embedding with dimensions 1 by 512.
[0019] The intermodal collaborative attention fusion module is the core of this model, comprising a four-layer stacked collaborative attention submodule. Within each submodule, bidirectional attention computation is performed: on one hand, the visual feature map is reshaped into a 4096x2048 sequence as the query; the spectral feature embedding and log feature embedding are concatenated to form an 82x512 sequence as the key and value; the attention weights of the visual features on the spectral and log features are calculated, generating a visual context vector modulated by spectral and log information. On the other hand, the spectral and log feature embedding sequence is used as the query, and the visual feature map sequence is used as the key and value; the attention weights of the spectral and log features on the visual features are calculated, generating a vector modulated by visual information, containing spectral and log context information. Subsequently, the original modal features and their corresponding context vectors are concatenated, and information fusion is performed through a feedforward neural network. This process is performed layer by layer in the four submodules, achieving deep, iterative interaction and fusion between multimodal features. Finally, the module outputs a fused feature vector with a dimension of 1x1024, highly condensed with information from all modalities.
[0020] The prediction and attribution head contains two functional paths: The forward prediction path inputs the 1x1024 fused feature vector into a classifier composed of a multilayer perceptron containing two fully connected layers and a sigmoid activation function, and outputs a scalar value between zero and one, which is the risk score.
[0021] The reverse attribution path employs a composite algorithm based on gradient and attention flow when attribution is triggered. First, the gradient of the risk score output value relative to the attention weight matrix of the last layer in the intermodal collaborative attention fusion module is calculated. Then, this gradient is multiplied by the attention weight matrix using a Hadamard product to obtain a saliency score for gradient attention fusion. This score characterizes the contribution of different feature flows to the final decision. Next, this saliency score is backpropagated along the network until the input layer. For visual data, a saliency map of the same size as the original image is generated, with the brightest regions on it representing key pixel regions. For spectral and log data, importance scores for each feature dimension are generated, with the bands and log entries with the highest scores being the key features.
[0022] This invention also provides a multimodal data fusion-based quality inspection system for metal content in wet wipes, the system comprising: A hardware synchronization beacon module is used to generate a globally unique synchronization identifier and send a differential trigger signal when the wet wipes under test enter the preset trigger area.
[0023] A visual data acquisition unit, connected to the hardware synchronization beacon module, is used to receive the differential trigger signal and capture digital image data covering the surface of the wet wipe product under test, and bind the synchronization identifier to the image data.
[0024] A spectral data acquisition unit is connected to the hardware synchronization beacon module, used to receive the differential trigger signal and acquire the spectral data of predetermined detection points on the surface of the wet wipe product to be tested, and to bind the synchronization identifier to the spectral data.
[0025] A production process log acquisition unit is connected to the hardware synchronization beacon module, used to receive the differential trigger signal and obtain structured log data associated with the current production status from the manufacturing execution system, and bind the synchronization identifier to the structured log data.
[0026] A data aggregation and preprocessing unit, connected to the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit, is used to aggregate the received image data, spectral data, and structured log data into a structured multimodal data packet based on the synchronization identifier.
[0027] A central processing unit is connected to the data aggregation and preprocessing unit and is equipped with the bidirectional attention fusion and attribution model. The model receives the multimodal data packets and performs forward inference to calculate a risk score. When the risk score exceeds a preset threshold, the model activates its reverse attribution function to identify key input features and outputs the risk score and the key input features.
[0028] A human-computer interaction interface, connected to the central processing unit, is used to receive and visualize the risk score and the key input features.
[0029] A further improvement to the aforementioned technical solution is that the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator, which provides a system master clock signal with a frequency stability better than one part per million. The hardware synchronization beacon module also includes a field-programmable gate array (FPGA) chip. The internal logic circuitry of the FPGA chip is responsible for generating a synchronization identifier conforming to a universally unique dictionary-ordered identifier specification based on the system master clock signal, and integrates a low-voltage differential signal transmitter to output the trigger signal in differential pairs, thereby ensuring a high signal-to-noise ratio and transmission integrity in industrial electromagnetic environments.
[0030] A further improvement to the aforementioned technical solution is that the core component of the visual data acquisition unit is an industrial camera employing a CMOS image sensor with a global shutter function. Its pixel resolution is 2048 x 2048 pixels, and it is equipped with a telecentric lens with a focal length of 75mm to eliminate perspective distortion caused by variations in the thickness of the wet wipes being tested. The industrial camera connects to the system via a gigabit Ethernet interface supporting a precise time protocol, ensuring that its internal clock remains synchronized with the clock source of the hardware synchronization beacon module at the sub-microsecond level. The spectral data acquisition unit uses a near-infrared spectrometer built based on Fabry-Perot interferometry technology, with an operating wavelength range covering 900 nm to 1700 nm and a spectral resolution of 10 nm. The production process log acquisition unit is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the relational database of the manufacturing execution system through an open database connection interface. Upon receiving a trigger signal, it executes a structured query statement containing precise timestamp matching conditions to obtain parameters such as the batch number, running speed, cutter number used, and cumulative working time of the cutter in an atomic operation manner.
[0031] A further improvement to the aforementioned technical solution is that the data aggregation and preprocessing unit is deployed on an edge computing device. The structured multimodal data packets are encapsulated using JavaScript object representation. This data packet structure includes the following main fields: a string field storing a universally unique dictionary sorting identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and a nested modal data object field. The modal data object field contains a visual data object field, which further includes metadata such as image format, Base64-encoded image data, resolution, and exposure time. The modal data object field also contains a spectral data object field, which includes an array recording all sampled wavelength points, an array recording the normalized reflection intensity values for each wavelength point, and metadata such as integration time. The modal data object field also contains a process log data object field, which directly stores structured log data obtained from the manufacturing execution system in key-value pairs, including parameters such as batch number, production line speed, cutter number, and cumulative cutter working time. After generating the data packet, the data aggregation and preprocessing unit transmits the data packet asynchronously and reliably to the central processing unit through a message-oriented middleware, specifically a zero-message queue, in a publish-subscribe mode.
[0032] Building upon the aforementioned technical solution, a further improvement is made to the network architecture of the bidirectional attention fusion and attribution model deployed within the central processing unit. This architecture includes three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head. The visual feature extractor employs a pre-trained residual network. The spectral feature extractor employs a one-dimensional transformer network. The log feature extractor employs a multilayer perceptron. The intermodal collaborative attention fusion module comprises four stacked collaborative attention sub-modules, each performing bidirectional attention computation to achieve deep interactive fusion of multimodal features. The prediction and attribution head includes a forward prediction path and a reverse attribution path. The forward prediction path is used to calculate a risk score, while the reverse attribution path employs a composite algorithm based on gradient and attention flow to identify key input features.
[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows: 1. Data Source Synchronization and Fidelity: This invention employs a hardware synchronization beacon module based on a highly stable clock source and a highly interference-resistant differential trigger signal to physically ensure absolute spatiotemporal alignment of multiple heterogeneous data streams targeting the same object at the moment of acquisition, even in high-speed dynamic environments. This fundamentally eliminates data contamination caused by timing misalignment, providing a high-quality raw data foundation with inherent logical consistency for subsequent intelligent analysis, and greatly improving the training effect and inference accuracy of the model.
[0034] 2. Transparency and Attributability of the Decision-Making Process: The bidirectional attention fusion and attribution model designed in this invention not only enables high-precision risk assessment through deep intermodal interaction fusion, but its unique reverse attribution mechanism also decomposes the abstract risk scoring decision-making process into a quantitative contribution to specific, observable input features. These features specifically include a blob on an image, a peak in a spectrum, or a record in a log. This opens the model's black box, providing a clear and verifiable chain of evidence for each judgment, enhancing operators' trust in the system's decisions.
[0035] 3. Closed-loop capability for fault diagnosis and process optimization: By visualizing quantified attribution results and risk scores, operators can instantly and intuitively understand the root cause of alarms, determining whether it is a physical defect in the product itself or an anomaly in specific production process parameters. This traceability capability, accurate to specific characteristics, significantly shortens fault diagnosis and handling time, reducing the average fault diagnosis and handling time from over thirty minutes to less than two minutes. It also provides direct and effective data support for continuous improvement of production processes and the development of preventative maintenance strategies, thus constructing a complete closed loop from intelligent quality inspection to intelligent manufacturing, significantly improving production efficiency and product quality management.
[0036] 4. System Robustness and Industrial Applicability: The technical solution of this invention fully considers the harsh environment of industrial sites. From the mechanical structure stability of sensor installation and the gantry structure with vibration damping design, to the low-voltage differential signal used for signal transmission to ensure high electromagnetic interference resistance, and the edge-center collaborative processing mode adopted in the data processing architecture, all reflect a high degree of engineering and systematic consideration, ensuring that the entire method and system can operate stably and reliably in long-term, continuous industrial production, and have good industrial applicability. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the overall technical architecture of the multimodal data fusion-based quality inspection system for metal content in wet wipes proposed in this invention. Figure 2 This is a schematic diagram of the core principle framework of the bidirectional attention fusion and attribution model in this invention.
[0038] Figure 3 This is a logical flowchart of the multimodal data synchronous acquisition in this invention.
[0039] Figure 4 This is a logical flowchart of the multimodal data aggregation and preprocessing in this invention.
[0040] Figure 5 This is a schematic diagram of the overall data flow from acquisition to decision-making in this invention.
[0041] Figure 6 This is a schematic diagram of the core principle framework for identifying key input features using the reverse attribution function in this invention.
[0042] Figure reference numerals: 1-Hardware synchronization beacon module; 2-Visual data acquisition unit; 3-Spectral data acquisition unit; 4-Production process log acquisition unit; 5-Manufacturing execution system backend database; 6-Data aggregation and preprocessing unit; 7-Central processing unit; 8-Human-machine interface. Detailed Implementation
[0043] This invention provides a method and system for quality inspection of metal content in wet wipes based on multimodal data fusion. It aims to solve the technical challenges of existing intelligent quality inspection technologies in high-speed production line environments, such as data inconsistency due to misaligned data acquisition timing, insufficient reliability of artificial intelligence models, and opaque decision-making processes. This invention constructs a low-level mechanism that ensures precise spatiotemporal synchronization of multimodal data from the physical source of data acquisition, and combines it with a deep learning model possessing both positive prediction and reverse attribution capabilities. This enables high-precision and high-reliability detection of minute metallic foreign objects in wet wipes, and provides deterministic and transparent traceability of the judgment criteria. The technical solution of this invention will be described in detail below with specific embodiments.
[0044] See Figure 5 This invention provides a method for quality inspection of metal content in wet wipes based on multimodal data fusion, the method comprising the following steps: First, see Figure 3 S1: Deploy a hardware synchronization beacon module 1 at the detection station through which the wet wipe product under test passes. When a specific physical reference point of the wet wipe product under test enters the preset trigger area of the hardware synchronization beacon module 1, the high-frequency stable clock source and unique identifier generation logic circuit embedded in the hardware synchronization beacon module 1 immediately generate a globally unique synchronization identifier containing a nanosecond-level precision timestamp, and synchronously send a differential trigger signal with high anti-interference characteristics to a visual data acquisition unit 2, a spectral data acquisition unit 3, and a production process log acquisition unit 4.
[0045] Specifically, on the high-speed production line of wet wipes, the hardware synchronization beacon module 1 is precisely installed and deployed at key detection nodes along the product transport path. The physical implementation of the hardware synchronization beacon module 1 includes a temperature-compensated crystal oscillator, which serves as the system's master clock source, providing an ultra-high stability clock signal with a frequency stability better than one part per million. This high-stability clock signal is the physical basis for ensuring nanosecond-level timestamp accuracy, effectively avoiding time deviations caused by clock drift. The hardware synchronization beacon module 1 also includes a field-programmable gate array (FPGA) chip, whose internal dedicated logic circuitry is responsible for performing the following core functions: First, real-time monitoring of the position of the wet wipes on the conveyor belt. This monitoring mechanism is implemented through a high-precision laser beam sensor array or a high-speed photoelectric switch. When the physical reference point of the wet wipes under test, such as its leading edge or a preset marker point, precisely enters the preset trigger area defined by the overlap of multiple sensors, the FPGA chip immediately captures this event. Second, based on the system's master clock signal, the synchronization identifier generation logic circuitry inside the FPGA chip generates a globally unique synchronization identifier in real time, conforming to the universally unique dictionary-ordered identifier specification. This identifier not only includes a Unix timestamp with nanosecond precision but also integrates a serial number or batch information to ensure its uniqueness and traceability throughout the system. The precise time point for generating this identifier is defined as the "zero point" moment when the wet wipe product is at the inspection station. Third, the field-programmable gate array chip integrates a low-voltage differential signal transmitter. Simultaneously with the generation of the synchronization identifier, this transmitter immediately outputs a differential trigger signal with high anti-interference characteristics in differential pairs to the pre-connected visual data acquisition unit 2, spectral data acquisition unit 3, and production process log acquisition unit 4. This differential signal transmission mechanism effectively suppresses electromagnetic noise interference common in industrial environments, ensuring signal integrity and extremely low latency during long-distance transmission of the trigger signal. This guarantees that each acquisition unit can physically receive the trigger command simultaneously with sub-microsecond or even nanosecond synchronization precision. The core of this step lies in eliminating the inconsistency problem caused by timing misalignment of multimodal data from the physical source of data acquisition through a high-precision synchronization mechanism at the hardware level, laying a solid and reliable foundation for subsequent data fusion and intelligent analysis.
[0046] Secondly, S2: At the same moment they receive the differential trigger signal, the visual data acquisition unit 2, the spectral data acquisition unit 3, and the production process log acquisition unit 4 execute their respective data acquisition operations in parallel. Specifically, the visual data acquisition unit 2 captures a frame of digital image data covering the surface of the wet wipe product under test; the spectral data acquisition unit 3 acquires the reflection or transmission spectral data of predetermined detection points on the surface of the wet wipe product under test; and the production process log acquisition unit 4 initiates a precise query to the manufacturing execution system backend database 5 or real-time data interface to obtain structured log data associated with the current production status. After completing data acquisition, each acquisition unit immediately rigidly binds the obtained synchronization identifier to the acquired data, forming a data unit containing the synchronization identifier and data payload.
[0047] Specifically, the core component of the visual data acquisition unit 2 is an industrial camera employing a complementary metal-oxide-semiconductor (CMOS) image sensor with a global shutter function. The industrial camera boasts an ultra-high resolution of 2048 x 2048 pixels and is equipped with a telecentric lens with a focal length of 75mm. The telecentric lens is optically designed to eliminate perspective distortion caused by variations in the thickness of the wet wipes being tested, ensuring that the image size and shape remain consistent regardless of the wipe's location within its depth of field, thus guaranteeing measurement accuracy. The industrial camera connects to the system via a Gigabit Ethernet interface supporting a precise time protocol, and its internal clock maintains sub-microsecond synchronization with the clock source of the hardware synchronization beacon module 1, further ensuring precise alignment of data acquisition time. When the visual data acquisition unit 2 receives a differential trigger signal from the hardware synchronization beacon module 1, the industrial camera immediately performs a single-frame image capture operation. The captured digital image data contains high-resolution visual information about the surface of the wet wipes being tested. After the acquisition is completed, the control logic of the visual data acquisition unit 2 will rigidly bind the synchronization identifier received from the hardware synchronization beacon module 1 to the digital image data in the form of a digital signature or metadata embedding.
[0048] The spectral data acquisition unit 3 employs a near-infrared spectrometer constructed using Fabry-Perot interferometry technology based on microelectromechanical systems (MEMS). This spectrometer operates in the wavelength range of 900 nm to 1700 nm, with a spectral resolution of 10 nm, and features high-speed scanning and high sensitivity. Its optical probe is precisely calibrated and fixed above a predetermined detection point on the surface of the wet wipe product, for example, its central region. When the spectral data acquisition unit 3 receives a differential trigger signal, the near-infrared spectrometer immediately performs a reflection or transmission spectral scan on the predetermined detection point on the surface of the wet wipe product within a preset integration time, acquiring the spectral data for that point. The spectral data is represented as a sequence of 81 spectral intensity values sampled at 10 nm intervals within the wavelength range of 900 nm to 1700 nm. These values reflect the absorption or reflection characteristics of the wet wipe material and any foreign matter present at specific wavelengths of near-infrared light. After data acquisition, the spectral data acquisition unit 3 also rigidly binds the received synchronization identifier to the acquired spectral data.
[0049] The production process log acquisition unit 4 is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the relational database of the manufacturing execution system (MES) backend through an open database connection interface, maintaining uninterrupted communication. When the production process log acquisition unit 4 receives a differential trigger signal, the software agent program immediately executes a predefined structured query statement containing precise timestamp matching conditions. The query statement, in an atomic operation, retrieves structured log data associated with the current production status from the MES backend database 5 in real time. This log data includes, but is not limited to, key process parameters such as the current production line batch number, real-time operating speed, the cutter number used, and the cumulative working time of the cutter. The precise timestamp matching conditions ensure a high degree of consistency in time between the retrieved log data and the wet wipe product indicated by the synchronization identifier. After acquiring the log data, the production process log acquisition unit 4 rigidly binds the synchronization identifier to the acquired structured log data.
[0050] Through the aforementioned parallel and synchronous data acquisition and identifier binding mechanism, heterogeneous data from vision, spectroscopy, and production process logs are ensured to be precisely mapped to the same wet wipe product under test in physical time. Each data unit generates a data unit containing the same synchronization identifier, which serves as the unique key for subsequent data aggregation.
[0051] Again, see Figure 4S3: The data units generated by the three data acquisition units and bound with the same synchronization identifier are transmitted to a data aggregation and preprocessing unit 6. The data aggregation and preprocessing unit 6 uses the synchronization identifier as a unique key to aggregate the image data, the spectral data, and the structured log data into a structured multimodal data packet. This data packet follows a predefined, strict data pattern, ensuring an absolute one-to-one correspondence between different modalities.
[0052] Specifically, the data aggregation and preprocessing unit 6 is deployed on a dedicated edge computing device. This edge computing device is equipped with a high-performance processor and large-capacity memory to meet the real-time processing requirements of high-speed data streams. The data aggregation and preprocessing unit 6 is connected to each data acquisition unit via a high-speed industrial Ethernet, receiving data unit streams bound with synchronization identifiers. Upon receiving a data unit, the data scheduling module within the data aggregation and preprocessing unit 6 first caches and matches these data units according to their internal synchronization identifiers. Once visual data units, spectral data units, and production process log data units with the same synchronization identifier are received, the data scheduling module immediately triggers the data aggregation process.
[0053] The core of the data aggregation process is to accurately match and merge different modal data belonging to the same wet wipes product based on the synchronization identifier as a unique key. The structured multimodal data packet is encapsulated in JavaScript object representation format, a lightweight data exchange format that is easy for machines to parse and generate, and has good scalability. This data packet structure follows a predefined, strict data pattern, ensuring an absolute one-to-one correspondence and data integrity between different modal data. The data packet structure includes the following main fields: a string field storing a universally unique dictionary sorting identifier, which precisely records the synchronization identifier; a 64-bit integer field storing a nanosecond-level Unix timestamp, which indicates the data packet's generation time and is highly consistent with the timestamp in the synchronization identifier; and a nested modal data object field.
[0054] The modal data object field contains a visual data object field. This visual data object field further includes the image format, such as portable web graphics or the Joint Image Experts Group standard; Base64 encoded image data, which converts binary image data into a text-transferable string for easy network transmission and storage; and metadata such as resolution and exposure time, which provide key parameter information during image acquisition.
[0055] The modal data object field also includes a spectral data object field. This spectral data object field contains an array recording all sampled wavelength points, precisely listing 81 wavelength points in 10-nanometer intervals within the range of 900 nm to 1700 nm; an array recording the normalized reflection intensity values for each wavelength point, the normalization typically employing standard normal variable transformation or multivariate scattering correction to eliminate baseline drift and particle size effects; and metadata such as integration time, which records the duration of spectral acquisition.
[0056] The modal data object field also includes a process log data object field. This process log data object field directly stores structured log data obtained from the manufacturing execution system in key-value pairs, including parameters such as batch number, production line speed, cutter number, and cumulative cutter working time. The key of each key-value pair is the parameter name, and the value is the corresponding parameter data.
[0057] After successfully generating the structured multimodal data packets, the data aggregation and preprocessing unit 6 transmits the data packets asynchronously and reliably to the central processing unit 7 via a message-oriented middleware, specifically a zero-message queue, in a publish-subscribe pattern. The zero-message queue provides high-performance, low-latency message transmission capabilities. Its publish-subscribe pattern allows the central processing unit 7 to subscribe to the required data while ensuring reliable delivery of data packets. Even in the event of a momentary network interruption, the message persistence mechanism ensures that data is not lost. This step ensures the complete aggregation and standardization of data at the logical level, providing standardized, high-quality input for subsequent deep learning model inference.
[0058] Then, see Figure 2 S4: The structured multimodal data packets are fed into a pre-trained bidirectional attention fusion and attribution model deployed in the central processing unit 7. The model performs a forward reasoning process in which the intermodal collaborative attention mechanism within the model performs weighted fusion of deep features from different modalities, calculates and outputs a quantitative risk score characterizing the likelihood of the presence of metallic foreign objects in the tested wet wipes product.
[0059] Specifically, the central processing unit 7 is typically a high-performance server equipped with multiple graphics processing units to provide the powerful computing capabilities required for deep learning model inference. The bidirectional attention fusion and attribution model is pre-trained on the central processing unit 7. Its training dataset consists of multimodal data packets containing a large number of normal wet wipes and wet wipes containing tiny metallic foreign objects, and is labeled with ground truth by expert annotations. Upon receiving the structured multimodal data packets as input, the model immediately initiates the forward inference process.
[0060] The specific network structure of the bidirectional attention fusion and attribution model is designed as a deep neural network architecture consisting of three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head.
[0061] The three parallel feature extractor branches are as follows: First, the visual feature extractor employs a pre-trained residual network without a top classification layer as its backbone. This residual network, pre-trained on large-scale image datasets such as ImageNet, possesses powerful image feature extraction capabilities. The input is a normalized 2048x2048x1 single-channel grayscale image tensor. The normalization process typically includes operations such as scaling pixel values to the zero-to-one range, mean normalization, and standard deviation normalization to eliminate the influence of differences in lighting and camera parameters. The visual feature extractor, through processing multiple convolutional layers, pooling layers, and residual blocks, outputs a deep feature map with dimensions of 64x64x2048, which encodes the spatial and semantic information in the image.
[0062] Second, the spectral feature extractor employs a one-dimensional transformer network consisting of six stacked encoder layers. This one-dimensional transformer network is specifically designed for processing sequential data and can capture long-range dependencies within the sequence. The input is a spectral intensity sequence with a dimension of 81 x 1, corresponding to 81 wavelength points. This sequence is first position-encoded to inject relative position information of the wavelength points, and then fed into the transformer network. Each encoder layer of the transformer network contains a multi-head self-attention mechanism and a feedforward network. Through the processing of the six encoder layers, the spectral feature extractor outputs a sequence feature embedding with a dimension of 81 x 512, which is rich in information about the shape, peaks, valleys, and subtle variations of the spectral curve in different wavelength regions.
[0063] Third, the log feature extractor employs a multilayer perceptron. For categorical log data, such as cutter numbers, an embedding layer is used to convert them into dense vectors, thus mapping discrete categorical information into a continuous vector space. For numerical log data, such as production line speed and cumulative cutter working time, standardization is performed, typically using zero-mean unit variance scaling to eliminate dimensional differences. All processed log features, including the embedded vectors and standardized values, are concatenated into a one-dimensional vector, which is then non-linearly transformed through the multilayer perceptron network, outputting a log feature embedding with dimensions 1 x 512, representing the current production process status.
[0064] The intermodal collaborative attention fusion module is the core of this model, comprising a collaborative attention submodule consisting of four stacked layers. Within each submodule, bidirectional attention computation is performed, enabling deep, iterative interaction and fusion between multimodal features. Specifically: On one hand, the visual feature map (64 x 64 x 2048 dimensions) is first reshaped and flattened into a 4096 x 2048 sequence, which serves as the query vector. Simultaneously, the spectral feature embedding (81 x 512) and the log feature embedding (1 x 512) are concatenated to form an 82 x 512 sequence, which serves as the key and value vectors. Then, the attention weights of the visual features on the spectral and log features are calculated. This process can be represented as: In this context, Q represents the visual query, a sequence obtained by reshaping and flattening the visual feature map, representing the visual feature vector used for the query. K is the concatenation key of the spectral and log data, a sequence formed by concatenating the spectral feature embeddings and log feature embeddings, serving as the "key" for matching the query vector. V is the concatenation value of the spectral and log data, a sequence formed by concatenating the spectral feature embeddings and log feature embeddings, representing the "value" corresponding to the "key," which is ultimately weighted by attention weights. It is the dimension of the key K, used to scale the dot product of Q and K to avoid the dot product value being too large due to excessively high dimension.
[0065] This attention mechanism generates a visual context vector modulated by spectral and log information, which enhances the parts of the visual features that are related to spectral and log information.
[0066] On the other hand, the spectral and log feature embedding sequence (82 x 512) is used as the query vector. Simultaneously, the sequence reshaped from the visual feature map (4096 x 2048) is used as the key and value vectors. Attention is also utilized to generate a vector modulated by visual information, containing both spectral and log contextual information.
[0067] Subsequently, the original modal features—namely, the original visual features, original spectral features, and original log features—are concatenated with their corresponding context vectors. The concatenated features are then fused using a feedforward neural network containing activation functions that introduce non-linearity. This bidirectional attention computation and feature fusion process is performed layer by layer in four sub-modules, with each layer using the fused features from the previous layer as input, enabling deep, iterative interaction and fusion between multimodal features. Finally, this module outputs a highly condensed fused feature vector with dimensions 1 x 1024, encapsulating information from all modalities.
[0068] The prediction and attribution head contains two functional paths: The forward prediction path inputs the 1x1024 fused feature vector into a classifier composed of a multilayer perceptron containing two fully connected layers and a sigmoid activation function. The fully connected layers perform a linear transformation on the fused features, and the sigmoid activation function compresses the output value to the range of zero to one, ultimately outputting a scalar value between zero and one, which is the risk score. A higher risk score indicates a greater likelihood of the presence of metallic foreign objects in the tested wet wipe product.
[0069] Finally, see Figure 6 S5: Threshold judgment is performed on the risk score. When the risk score exceeds a preset risk threshold, the reverse attribution function of the model is immediately activated. This function uses the attention weight matrix generated during the forward inference process and stored inside the model to reverse trace and calculate the contribution of each input feature to the final high-risk score. Accordingly, a set of key input features that play a decisive role in the high-risk judgment is accurately identified and output. The key input features are further associated with their original modalities and presented in a specific form, including one or more specific pixel coordinate regions in the digital image, one or more specific absorption or reflection peak bands in the spectral data, and one or more specific parameter entries in the structured log data. Furthermore, the risk score and the identified key input features, which are enhanced in a highlighted or marked form, are presented together on the human-computer interaction interface 8, providing the operator with direct and clear decision-making basis.
[0070] Specifically, the preset risk threshold is set and calibrated through experiments on a large number of samples and expert experience before system deployment; for example, it is set to 0.85. When the risk score output by the model's positive prediction path exceeds this threshold, the system immediately triggers an alarm and activates the model's reverse attribution function.
[0071] When the reverse attribution path is triggered, it employs a composite algorithm based on gradient and attention flow. First, the gradient of the risk score output value relative to the last layer of the attention weight matrix in the intermodal collaborative attention fusion module is calculated. This gradient characterizes the strength and direction of the influence of each element in the attention weight matrix on the final risk score. Then, this gradient is multiplied element-wise by the attention weight matrix to obtain a saliency score matrix for gradient attention fusion. Each element of this score matrix characterizes the contribution of a specific step in the interaction between specific modal features to the final decision. This process can be represented as: in, This represents the Hadamard product. It's an element-wise multiplication operation used to combine the gradient with the attention weight matrix. The elements are multiplied one by one to obtain the fused significance score matrix. Significance (significance score matrix): The final matrix, where each element represents the contribution of a specific link in the interaction between modal features to the final decision (e.g., risk score). Gradient: The gradient of the risk score output value relative to the last layer of attention weight matrix in the inter-modal collaborative attention fusion module. It reflects the strength and direction of the influence of each element in the attention weight matrix on the final risk score. The attention weight matrix, derived from the last layer of the "intermodal collaborative attention fusion module," records the attention weight relationships between different features.
[0072] Next, this saliency score is backpropagated along the network up to the input layer. During backpropagation, the saliency score is inversely mapped back to the original input space through a feature extractor, thereby quantifying the contribution of the original input features. For visual data, this process generates a saliency map of the same size as the original image, where bright areas are the key pixel regions that play a decisive role in high-risk judgments, such as a dark spot with a diameter of three millimeters. For spectral data, the backpropagation results generate importance scores for each feature dimension, with the bands with the highest scores identified as the bands with key absorption or reflection peaks, such as the anomalous absorption peaks in the 1030 nm to 1050 nm band. For log data, importance scores are also generated for each parameter item, with the parameter item with the highest score being the key feature, such as the item that the cutter numbered three has accumulated more than eight hours of working time.
[0073] Finally, the central processing unit 7 transmits the calculated risk score along with the identified key input features, enhanced with highlights or markers, to the human-machine interface 8. The human-machine interface 8 presents this information in an intuitive and graphical manner. For example, in digital images, key pixel areas are highlighted with red borders or semi-transparent overlays; in spectral curves, key bands are marked with special colors or shaded areas; in production process logs, key parameter entries are bolded or highlighted. This integrated and transparent information presentation provides operators with direct and clear decision-making support, enabling them to quickly understand the source of risk and take appropriate corrective measures, such as immediately stopping the machine for inspection, replacing the cutter, or adjusting production process parameters. This step greatly improves decision-making efficiency and system reliability, achieving full transparency and traceability of the intelligent quality inspection process.
[0074] See Figure 1 The present invention also provides a multimodal data fusion-based quality inspection system for metal content in wet wipes, the system comprising: A hardware synchronization beacon module 1 is designed to precisely generate a globally unique synchronization identifier containing a nanosecond-level timestamp when a wet wipe product enters a preset detection area, and synchronously send a highly interference-resistant differential trigger signal to each data acquisition unit. The physical implementation of this module includes a temperature-compensated crystal oscillator that provides a system master clock signal with frequency stability better than one part per million. An internal field-programmable gate array (FPGA) chip is responsible for generating a synchronization identifier conforming to a universally unique dictionary-ordered identifier specification based on the master clock signal, and integrates a low-voltage differential signal transmitter to output the trigger signal in differential pairs, ensuring a high signal-to-noise ratio and transmission integrity in industrial electromagnetic environments.
[0075] A visual data acquisition unit 2 is connected to the hardware synchronization beacon module 1. Upon receiving the differential trigger signal, this unit accurately captures a frame of high-resolution digital image data covering the surface of the wet wipe product under test. After acquisition, the synchronization identifier is rigidly bound to the captured image data. The core component of the visual data acquisition unit 2 is an industrial camera employing a complementary metal-oxide-semiconductor image sensor with a global shutter function. Its pixel resolution is 2048 x 2048 pixels, and it is equipped with a telecentric lens with a focal length of 75mm to eliminate perspective distortion caused by variations in the thickness of the wet wipe product under test. The industrial camera is connected to the system via a gigabit Ethernet interface supporting a precise time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module 1 at the sub-microsecond level.
[0076] A spectral data acquisition unit 3 is connected to the hardware synchronization beacon module 1. Upon receiving the differential trigger signal, this unit accurately acquires the reflection or transmission spectral data of predetermined detection points on the surface of the wet wipe product under test. After acquisition, the synchronization identifier is rigidly bound to the acquired spectral data. The spectral data acquisition unit 3 employs a near-infrared spectrometer built based on Fabry-Perot interferometry technology, with an operating wavelength range covering 900 nm to 1700 nm and a spectral resolution of 10 nm, enabling high-speed and high-sensitivity capture of spectral information.
[0077] A production process log acquisition unit 4 is connected to the hardware synchronization beacon module 1. Upon receiving the differential trigger signal, this unit initiates a precise query to the manufacturing execution system backend database 5 or a real-time data interface to obtain structured log data associated with the current production status. After obtaining the data, the synchronization identifier is rigidly bound to the acquired structured log data. The production process log acquisition unit 4 is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the manufacturing execution system backend relational database through an open database connection interface. Upon receiving the trigger signal, it executes a structured query statement containing precise timestamp matching conditions to obtain key parameters such as the current production line batch number, operating speed, cutter number used, and the cumulative working time of the cutter in an atomic operation manner.
[0078] A data aggregation and preprocessing unit 6 is connected to the visual data acquisition unit 2, the spectral data acquisition unit 3, and the production process log acquisition unit 4. This unit uses the synchronization identifier as a unique key to precisely aggregate the received image data, spectral data, and structured log data into a structured multimodal data packet. The data aggregation and preprocessing unit 6 is deployed on an edge computing device, and the structured multimodal data packet is encapsulated using JavaScript object representation. The data packet structure includes a string field storing a universally unique dictionary sorting identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and nested modal data object fields. The modal data object fields internally include a visual data object field, which contains metadata such as image format, Base64 encoded image data, resolution, and exposure time; a spectral data object field, which contains an array recording all sampled wavelength points, an array recording the normalized reflection intensity values for each wavelength point, and integration time; and a process log data object field, which directly stores the structured log data obtained from the manufacturing execution system in key-value pairs. After generating the data packet, the data aggregation and preprocessing unit 6 transmits the data packet asynchronously and reliably to the central processing unit 7 through a message-oriented middleware, specifically a zero-message queue, in a publish-subscribe mode.
[0079] A central processing unit 7 is connected to the data aggregation and preprocessing unit 6. The central processing unit 7 houses the bidirectional attention fusion and attribution model. This unit receives the multimodal data packets and performs a forward inference process to calculate a quantified risk score. When the risk score exceeds a preset risk threshold, the model's reverse attribution function is immediately activated to trace back and identify key input features that play a decisive role in the high-risk judgment, and finally outputs the risk score and the identified key input features. The network architecture of the bidirectional attention fusion and attribution model deployed in the central processing unit 7 includes three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head. The visual feature extractor uses a pre-trained residual network. The spectral feature extractor uses a one-dimensional transformer network. The log feature extractor uses a multilayer perceptron. The intermodal collaborative attention fusion module contains four stacked collaborative attention sub-modules, each performing bidirectional attention calculation to achieve deep interactive fusion of multimodal features. The prediction and attribution head includes a forward prediction path and a reverse attribution path. The forward prediction path is used to calculate a risk score, and the reverse attribution path uses a composite algorithm based on gradient and attention flow to identify key input features.
[0080] A human-machine interface 8 is connected to the central processing unit 7. The human-machine interface 8 is used to receive and visualize the risk score and the identified key input features in an intuitive and graphical manner, for example, by highlighting or marking key parts of images, spectra, and log data to provide operators with direct and clear decision-making basis.
[0081] The above are merely specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent substitutions, or modifications made based on the present invention to solve essentially the same technical problems and achieve essentially the same technical effects are all covered within the protection scope of the present invention.
Claims
1. A method for quality inspection of metal content in wet wipes based on multimodal data fusion, characterized in that, Includes the following steps: S1: Deploy a hardware synchronization beacon module at the detection station through which the wet wipes under test pass. When a specific physical reference point of the wet wipes under test enters the preset trigger area of the hardware synchronization beacon module, the hardware synchronization beacon module generates a globally unique synchronization identifier and synchronously sends a differential trigger signal to a visual data acquisition unit, a spectral data acquisition unit, and a production process log acquisition unit. S2: At the same time that the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit receive the differential trigger signal, they execute their respective data acquisition operations in parallel. The visual data acquisition unit captures a frame of digital image data, the spectral data acquisition unit acquires the spectral data of the predetermined detection points on the surface of the wet wipe product to be tested, and the production process log acquisition unit acquires the structured log data associated with the current production status from the manufacturing execution system. Each acquisition unit binds the obtained synchronization identifier with the acquired data to form a data unit containing the synchronization identifier and the data payload. S3: The data units generated by the three data acquisition units and bound with the same synchronization identifier are transmitted to a data aggregation and preprocessing unit. The data aggregation and preprocessing unit aggregates the digital image data, the spectral data and the structured log data into a structured multimodal data packet based on the synchronization identifier as a unique key value. S4: The structured multimodal data packet is fed into a pre-trained bidirectional attention fusion and attribution model deployed in a central processing unit. The model performs a forward reasoning process to calculate and output a quantified risk score. S5: Perform threshold judgment on the risk score. When the risk score exceeds a preset risk threshold, activate the reverse attribution function of the model, identify and output a set of key input features that play a decisive role in the high-risk judgment. The key input features are further associated with their original modalities and presented in a specific form. The risk score and the key input features with enhanced visualization processing are presented together on the human-computer interaction interface.
2. The method according to claim 1, characterized in that, In step S1, the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator that provides a system master clock signal with a frequency stability better than one part per million. The hardware synchronization beacon module also includes a field-programmable gate array (FPGA) chip. The internal logic circuit of the FPGA chip is responsible for generating the synchronization identifier that conforms to the universally unique dictionary sorting identifier specification based on the system master clock signal, and integrates a low-voltage differential signal transmitter for outputting the differential trigger signal in the form of differential pairs.
3. The method according to claim 1, characterized in that, In step S2: The core component of the visual data acquisition unit is an industrial camera with a CMOS image sensor that has a global shutter function. The industrial camera is connected to the system through a gigabit Ethernet interface that supports a precise time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module at the sub-microsecond level. The spectral data acquisition unit employs a near-infrared spectrometer constructed using Fabry-Perot interferometer technology based on microelectromechanical systems. The production process log acquisition unit is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the relational database of the manufacturing execution system through an open database connection interface. Upon receiving a trigger signal, it executes a structured query statement containing precise timestamp matching conditions to obtain parameters such as the batch number, running speed, cutter number used, and cumulative working time of the cutter in an atomic operation manner.
4. The method according to claim 1, characterized in that, In step S3, the data aggregation and preprocessing unit is deployed on an edge computing device; the structured multimodal data packet is encapsulated using JavaScript object representation format. The data packet structure includes a string field storing a universally unique dictionary sorting identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and a nested modal data object field; after generating the data packet, the data aggregation and preprocessing unit asynchronously and reliably transmits the data packet to the central processing unit through a message-oriented middleware, specifically a zero message queue, in a publish-subscribe pattern.
5. The method according to claim 1, characterized in that, In steps S4 and S5, the network architecture of the bidirectional attention fusion and attribution model includes three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head. The reverse attribution path employs a composite algorithm based on gradient and attention flow to calculate the gradient of the risk score output value relative to the attention weight matrix of the last layer in the intermodal collaborative attention fusion module. The gradient is then multiplied by the attention weight matrix to obtain a saliency score, which is then propagated back along the network to the input layer to identify the key input features.
6. A multimodal data fusion-based quality inspection system for metal content in wet wipes, characterized in that, include: A hardware synchronization beacon module is used to generate a globally unique synchronization identifier and send a differential trigger signal when the wet wipes under test enter the preset trigger area; A visual data acquisition unit is connected to the hardware synchronization beacon module, used to receive the differential trigger signal and capture digital image data covering the surface of the wet wipe product under test, and bind the synchronization identifier to the digital image data; A spectral data acquisition unit is connected to the hardware synchronization beacon module, used to receive the differential trigger signal and acquire the spectral data of predetermined detection points on the surface of the wet wipe product to be tested, and to bind the synchronization identifier to the spectral data; A production process log acquisition unit is connected to the hardware synchronization beacon module, used to receive the differential trigger signal and obtain structured log data associated with the current production status from the manufacturing execution system, and bind the synchronization identifier to the structured log data; A data aggregation and preprocessing unit, connected to the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit, is used to aggregate the received digital image data, the spectral data, and the structured log data into a structured multimodal data packet based on the synchronization identifier. A central processing unit is connected to the data aggregation and preprocessing unit and is equipped with the bidirectional attention fusion and attribution model. The model receives the multimodal data packets and performs forward inference to calculate the risk score. When the risk score exceeds a preset threshold, the model activates the reverse attribution function to identify key input features and outputs the risk score and the key input features. A human-computer interaction interface, connected to the central processing unit, is used to receive and visualize the risk score and the key input features.
7. The system according to claim 6, characterized in that, The physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator that provides a system master clock signal with a frequency stability better than one part per million. The hardware synchronization beacon module also includes a field-programmable gate array (FPGA) chip. The internal logic circuit of the FPGA chip is responsible for generating the synchronization identifier that conforms to the universally unique dictionary sorting identifier specification based on the system master clock signal, and integrates a low-voltage differential signal transmitter for outputting the differential trigger signal in the form of differential pairs.
8. The system according to claim 6, characterized in that: The core component of the visual data acquisition unit is an industrial camera with a CMOS image sensor that has a global shutter function, a pixel resolution of 2048 by 2048 pixels, and a telecentric lens with a focal length of 75mm. The industrial camera is connected to the system through a gigabit Ethernet interface that supports a precise time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module at the sub-microsecond level. The spectral data acquisition unit is a near-infrared spectrometer built based on Fabry-Perot interferometer technology of microelectromechanical systems. Its working wavelength range covers 900 nm to 1700 nm and the spectral resolution is 10 nm. The production process log acquisition unit is a dedicated software agent program running on an industrial personal computer. This program establishes a persistent connection with the relational database of the manufacturing execution system through an open database connection interface. Upon receiving a trigger signal, it executes a structured query statement containing precise timestamp matching conditions to obtain parameters such as the batch number, running speed, cutter number used, and cumulative working time of the cutter in an atomic operation manner.
9. The system according to claim 6, characterized in that, The data aggregation and preprocessing unit is deployed on an edge computing device. The structured multimodal data packets are encapsulated using JavaScript object representation format. The data packet structure includes a string field storing a universally unique dictionary sorting identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and a nested modal data object field. After generating the data packets, the data aggregation and preprocessing unit asynchronously and reliably transmits the data packets to the central processing unit through a message-oriented middleware, specifically a zero message queue, in a publish-subscribe pattern.
10. The system according to claim 6, characterized in that, The network architecture of the bidirectional attention fusion and attribution model deployed within the central processing unit includes three parallel feature extractor branches, an intermodal collaborative attention fusion module, and a prediction and attribution head. The visual feature extractor employs a pre-trained residual network; the spectral feature extractor employs a one-dimensional transformer network; and the log feature extractor employs a multilayer perceptron. The intermodal collaborative attention fusion module comprises four stacked collaborative attention sub-modules, each performing bidirectional attention computation to achieve deep interactive fusion of multimodal features. The prediction and attribution head includes a forward prediction path and a reverse attribution path. The forward prediction path is used to calculate a risk score, and the reverse attribution path employs a composite algorithm based on gradient and attention flow to identify key input features.
Citation Information
Patent Citations
Metal foreign matter detection method and device and terminal equipment
CN113260882A
Online detection method and system for foreign matters in meat products
CN119619427A
Rapid detection system for medicine quality
CN119959179A
Water quality heavy metal real-time detection method and system based on multi-source data fusion
CN120044202A
Vegetable pesticide content detection process
CN120254199A
Cited By
Foreign matter detection method and device for glass fiber cloth cover and medium
CN121558769A