A multi-modal data fusion wet wipe metal content quality inspection method and system
By using a hardware synchronization beacon module and a bidirectional attention fusion model, the problem of misaligned data acquisition timing was solved, enabling high-precision detection and transparent decision-making for tiny metal foreign objects in wet wipes, thus improving the detection reliability and fault diagnosis efficiency of the production line.
Patent Information
- Application Number
- CN202511446463.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing intelligent quality inspection technologies, in high-speed production line environments, suffer from problems such as data inconsistency due to misaligned data acquisition timing, insufficient reliability of artificial intelligence models, and opaque decision-making processes. These issues hinder accurate judgment of product quality and optimization of production processes.
By deploying hardware synchronization beacon modules to ensure precise spatiotemporal synchronization of multimodal data, and combining bidirectional attention fusion and attribution deep learning models, absolute data alignment and transparent decision-making processes are achieved.
It achieves high-precision and high-reliability detection of tiny metal foreign objects in wet wipes, provides clear judgment criteria and fault tracing capabilities, and improves production efficiency and quality management level.
Smart Images

Figure CN120932767B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a multi-modal data fusion wet tissue metal content quality inspection method and system. BACKGROUND
[0002] In the field of modern industrial manufacturing, especially in the hygiene product, food and pharmaceutical industries which have strict requirements on product purity and safety, efficient and accurate online detection of tiny metal foreign matters mixed in the production process is a key link to guarantee product quality, maintain brand reputation and protect consumer health. Wet tissue, as a type of daily consumer product that directly contacts the human body, has particularly strict quality control in the production process, and any residual metal debris can pose a serious safety hazard. Traditional metal foreign matter detection technologies, typically including metal detection doors based on electromagnetic induction principles or X-ray detection systems using density difference imaging, have provided a fundamental guarantee for quality control in industrial production in a specific historical period. Specifically, the metal detection door generates an alternating magnetic field, and when a conductive metal foreign matter passes through, it causes a disturbance in the magnetic field and is captured by a sensor, triggering an alarm. This method is relatively sensitive to ferromagnetic metals and has a relatively controllable cost. X-ray detection emits X-rays to penetrate the product, generates a grayscale image based on the difference in X-ray absorption rate of different substances, and identifies metal and other foreign matters with a density much higher than the product substrate, and has detection capability for various metals and non-metal high-density foreign matters.
[0003] However, with the continuous evolution of related industries towards high-speed, fine and intelligent directions, and the continuous improvement of consumer requirements for product quality, some inherent characteristics of the above traditional detection schemes in the principle level make them gradually show limitations in response to new challenges. On the one hand, for light and thin products moving rapidly on a high-speed conveyor belt, the electromagnetic field changes or X-ray absorption differences caused by extremely tiny or sheet-shaped and filament-shaped metal foreign matters are very weak and are easily overwhelmed by background noise, leading to an increased missed detection rate of traditional methods. On the other hand, detection methods that simply rely on physical properties cannot provide deep information about the source of contaminants. In view of this, the industry has begun to explore intelligent quality inspection technologies using multi-modal data fusion, trying to integrate data sources from different dimensions to build a more comprehensive and three-dimensional product quality portrait. Such technologies usually combine machine vision systems to capture color abnormal spots on the product surface, use spectral analysis techniques to detect the chemical composition characteristics of the material, and correlate production process data in the manufacturing execution system, and then input these heterogeneous data into an artificial intelligence model for comprehensive research and judgment. This fusion analysis paradigm can theoretically significantly improve the sensitivity and accuracy of detection, and represents the development direction of intelligent quality inspection technology.
[0004] However, when converting the multi-modal fusion technology from a laboratory concept to a rigorous industrial production practice, a deep and non-obvious technical contradiction emerges. This contradiction is rooted in the conflict between the physical reality of data acquisition in a high-speed production line environment and the basic requirements of artificial intelligence models for data quality. At its root, current multi-modal solutions generally follow an independent acquisition, post-fusion technical path, in which vision sensors, spectrometers, production log systems, and other subsystems are physically deployed and operated independently, each generating data streams according to their own clock and processing period, and finally correlating in a central processor through software-level timestamp alignment. In modern high-speed production lines that process tens or even hundreds of products per second, the inherent deficiencies of this architecture are dramatically magnified. The inherent millisecond processing delays, network jitter, and slight clock desynchronization between different sensors and systems collectively result in a timing misalignment problem. This means that image data, spectral data, and production log data that the model considers to describe the same wet wipe product sample may actually come from physically adjacent but not the same product. This data contamination undermines the authenticity and consistency of the input data from the source, making the data basis on which the artificial intelligence model is trained itself distorted and unreliable. Accordingly, a model trained and reasoned based on contaminated data makes the black box problem of its decision-making process even more intractable. Even if the model gives an unqualified judgment, operators cannot be sure whether the judgment is based on a real, multi-dimensional feature pointing to a defect, or is just a logical illusion caused by data misalignment. In such a case, any attempt to explain or trace the model's decision will be meaningless, because the object of its analysis is a set of logically inconsistent data from the start. This not only hinders the effective handling of a single alert, but also fundamentally cuts off the path of using quality inspection data to feed back to the production process and achieve a closed-loop quality management.
[0005] Therefore, how to build a bottom-layer mechanism from the physical source of data acquisition that can ensure the accurate synchronization of multi-modal data in the space-time dimension, and on this basis, develop an intelligent analysis method that not only can make accurate judgments, but also can transparently attribute and trace the basis of the judgment, so as to truly overcome the inherent contradiction between data integrity and model explainability in the existing technology, has become a key challenge and technical problem to be solved for those skilled in the art. SUMMARY
[0006] The application discloses a kind of multi-modal data fusion's wet wipe metal content quality inspection method and system, to solve the technical problems of the existing intelligent quality inspection technology under the environment of high-speed production line, data inconsistency caused by data acquisition time sequence dislocation, artificial intelligence model reliability is insufficient and decision-making process is not transparent.The application constructs a bottom mechanism from data acquisition physical source, i.e.
[0007] To achieve the above object, the application provides a kind of multi-modal data fusion's wet wipe metal content quality inspection method, the method comprises the following steps:
[0008] First, deploy a hardware synchronization beacon module in the detection station of the wet wipe product to be measured, when the specific physical reference point of the wet wipe product to be measured enters the preset trigger area of the hardware synchronization beacon module, the high-frequency stable clock source and the unique identifier generation logic circuit embedded in the hardware synchronization beacon module generate a globally unique synchronization identifier containing nanosecond-level precision timestamp, and a differential trigger signal with high anti-interference characteristic is sent to a visual data acquisition unit, a spectral data acquisition unit and a production process log acquisition unit.
[0009] Secondly, the visual data acquisition unit, the spectral data acquisition unit and the production process log acquisition unit perform parallel data acquisition operations at the same time when receiving the differential trigger signal. Specifically, the visual data acquisition unit captures a frame of digital image data covering the surface of the wet wipe product to be measured;The spectral data acquisition unit obtains the reflection or transmission spectrum data of the predetermined detection point on the surface of the wet wipe product to be measured;The production process log acquisition unit initiates an accurate query to the manufacturing execution system backend database or real-time data interface to obtain the structured log data associated with the current production state. Each acquisition unit immediately binds the obtained synchronization identifier with the acquired data to form a data unit containing synchronization identifier and data payload after completing data acquisition.
[0010] Thirdly, the data units generated by the three data acquisition units, which are bound with the same synchronization identifier, are transmitted to a data aggregation and preprocessing unit. The data aggregation and preprocessing unit aggregates the image data, the spectral data and the structured log data into a structured multi-modal data packet according to the synchronization identifier as the unique key value. The data packet follows a predefined and strict data mode, ensuring the one-to-one correspondence between different modal data.
[0011] Then, the structured multi-modal data packet is fed as input into a pre-trained bidirectional attention fusion and attribution model deployed on a central processing unit.
[0012] Finally, the risk score is thresholded. When the risk score exceeds a pre-set risk threshold, the reverse attribution function of the model is activated. This function utilizes the attention weight matrix generated during the forward inference process and stored inside the model to trace back and calculate the contribution of each input feature to the final high risk score. Accordingly, a set of key input features that play a decisive role in the high risk determination are accurately identified and output. The key input features are further associated with their original modalities to present in specific forms, including one or more specific pixel coordinate regions in the digital image, one or more specific absorption or reflection peak wavebands in the spectral data, and one or more specific parameter entries in the structured log data. Moreover, the risk score and the key input features identified are highlighted or marked for enhanced visualization and presented together on a human-machine interaction interface to provide direct and explicit decision-making basis for the operator.
[0013] Further improvements based on the technical solution are that the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator that provides a system master clock signal with a frequency stability better than one part in a million. The hardware synchronization beacon module also includes a field programmable gate array chip, and the internal logic circuit of the field programmable gate array chip is responsible for generating a synchronization identifier that meets the specification of the universal unique dictionary identifier based on the system master clock signal. A low voltage differential signal transmitter is integrated to output the trigger signal in differential pair form, thereby ensuring high signal-to-noise ratio and transmission integrity of the signal in an industrial electromagnetic environment.
[0014] On the basis of the technical scheme, further improvement is made, the core component of the visual data acquisition unit is an industrial camera using a CMOS image sensor with a global shutter function, the pixel resolution is 2048 by 2048 pixels, and a telecentric lens with a focal length of 75 mm is equipped to eliminate the perspective distortion caused by the thickness fluctuation of the wet wipe product to be measured. The industrial camera is connected with the system through a gigabit Ethernet interface supporting the precise time protocol, ensuring that the internal clock of the camera is synchronized with the clock source of the hardware synchronization beacon module to within sub-microsecond level. The spectral data acquisition unit adopts a near-infrared spectrometer based on a micro-electro-mechanical system Fabry-Perot interferometer technology, which has a working wavelength range of 900 nm to 1700 nm and a spectral resolution of 10 nm. The production process log acquisition unit is a special software agent program running on an industrial personal computer, which establishes a persistent connection with the manufacturing execution system back-end relational database through an open database connection interface, and after receiving a trigger signal, executes a structured query statement containing a precise time stamp matching condition to obtain the batch number, running speed, cutter number and cumulative working time of the cutter of the current production line in an atomic operation.
[0015] On the basis of the technical scheme, further improvement is made, the data aggregation and preprocessing unit is deployed on an edge computing device. The structured multi-modal data packet is packaged in JavaScript Object Notation format. The data packet structure includes the following main fields: a string field storing a universally unique dictionary identifier, a 64-bit integer field storing a nanosecond-level Unix timestamp, and a nested modal data object field. Inside the modal data object field, there is a visual data object field, which further includes image format, Base64 encoded image data, and resolution and exposure time metadata. Inside the modal data object field, there is also a spectral data object field, which includes an array recording all sampling wavelength points, an array recording normalized reflection intensity values corresponding to each wavelength point, and integration time metadata. Inside the modal data object field, there is also a process log data object field, which directly stores structured log data obtained from the manufacturing execution system in the form of key-value pairs, including batch number, production line speed, cutter number and cutter cumulative working time. After generating the data packet, the data aggregation and preprocessing unit transmits the data packet asynchronously and reliably to the central processing unit in a publish-subscribe mode through a message-oriented middleware, specifically a zero message queue.
[0016] On the basis of the technical scheme, further improvement is made, and the specific network structure of the bidirectional attention fusion and attribution model is designed as a deep neural network architecture composed of three parallel feature extractor branches, an inter-modal collaborative attention fusion module and a prediction and attribution head.
[0017] The three parallel feature extractor branches are respectively:
[0018] The visual feature extractor adopts a pre-trained residual network without a top classification layer as a backbone network. The input is a normalized 2048x2048x1 single-channel grayscale image tensor, and the output is a deep feature map with a dimension of 64x64x2048.
[0019] The spectral feature extractor adopts a one-dimensional transformer network composed of six encoder layers. The input is a spectral intensity sequence with a dimension of 81x1, corresponding to 81 wavelength points, which is sent into the network after position encoding, and the output is a sequence feature embedding with a dimension of 81x512.
[0020] The log feature extractor adopts a multilayer perceptron. For categorical log data, an embedding layer is used to convert it into a dense vector; for numerical log data, standardization processing is performed. All processed log features are concatenated into a one-dimensional vector, which is output by the multilayer perceptron network as a log feature embedding with a dimension of 1x512.
[0021] The inter-modal collaborative attention fusion module is the core of the model, which includes a collaborative attention submodule stacked by four layers. In each layer submodule, bidirectional attention calculation is performed: on the one hand, the visual feature map is reshaped into a 4096x2048 sequence as a query, and the spectral and log feature embeddings are concatenated to form an 82x512 sequence as a key and a value, the attention weights of the visual features on the spectral and log features are calculated, and a visual context vector modulated by spectral and log information is generated. On the other hand, the spectral and log feature embedding sequences are used as queries, and the visual feature map sequence is used as a key and a value, the attention weights of the spectral and log features on the visual features are calculated, and a vector modulated by visual information is generated, which contains spectral and log context information. Then, the original modal features and their corresponding context vectors are concatenated, and a feedforward neural network is used for information fusion. This process is performed layer by layer in the four submodules, realizing deep and iterative interaction and fusion among multi-modal features. Finally, the module outputs a fusion feature vector with a dimension of 1x1024, which highly condenses all modal information.
[0022] The prediction and attribution head includes two functional paths:
[0023] a forward prediction path that inputs the 1-by-1024 fused feature vector into a classifier consisting of multiple layers of perceptron with two fully connected layers and sigmoid activation function, outputting a scalar value between zero and one, which is the risk score.
[0024] a backward attribution path that adopts a gradient-and-attention flow based composite algorithm when attribution is triggered. First, the gradient of the risk score output value with respect to the last layer of attention weight matrix in the inter-modal collaborative attention fusion module is calculated. Then, the gradient is Hadamard multiplied with the attention weight matrix to obtain a gradient-attention fused saliency score. This score represents the contribution of different feature flows to the final decision. Next, the saliency score is back-propagated along the network until the input layer. For visual data, a saliency map of the same size as the original image is generated, with high-intensity regions being the key pixel regions. For spectral data and log data, importance scores of each feature dimension are generated, with the highest-scored wavebands and log entries being the key features.
[0025] The application also provides a multi-modal data fusion wet wipe metal content quality inspection system, the system comprising:
[0026] a hardware synchronization beacon module for generating a globally unique synchronization identifier and sending a differential trigger signal when the wet wipe product to be tested enters a preset trigger area.
[0027] a visual data acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and capturing digital image data covering the surface of the wet wipe product to be tested, and binding the synchronization identifier with the image data.
[0028] a spectral data acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and obtaining spectral data of predetermined detection points on the surface of the wet wipe product to be tested, and binding the synchronization identifier with the spectral data.
[0029] a production process log acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and obtaining structured log data associated with the current production state from a manufacturing execution system, and binding the synchronization identifier with the structured log data.
[0030] a data aggregation and preprocessing unit connected to the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit for aggregating the received image data, spectral data, and structured log data into a structured multi-modal data package according to the synchronization identifier.
[0031] A central processing unit connected to the data aggregation and preprocessing unit and deployed with the bidirectional attention fusion and attribution model, configured to receive the multi-modal data package and perform forward inference to calculate a risk score, and activate a reverse attribution function of the model to identify key input features when the risk score exceeds a preset threshold, and output the risk score and the key input features.
[0032] A human-computer interaction interface connected to the central processing unit, configured to receive and visually present the risk score and the key input features.
[0033] Further improvements made on the basis of the technical scheme are that the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator, which provides a system master clock signal with a frequency stability better than one part in a million. The hardware synchronization beacon module also includes a field programmable gate array chip, and the internal logic circuit of the field programmable gate array chip is responsible for generating a synchronization identifier conforming to the specification of the universal unique dictionary identifier based on the system master clock signal, and integrating a low-voltage differential signal transmitter for outputting the trigger signal in the form of a differential pair, thereby ensuring high signal-to-noise ratio and transmission integrity of the signal in an industrial electromagnetic environment.
[0034] Further improvements made on the basis of the technical scheme are that the core component of the visual data acquisition unit is an industrial camera using a CMOS image sensor with a global shutter function, which has a pixel resolution of 2048 by 2048 pixels and is equipped with a telecentric lens with a focal length of 75 mm to eliminate perspective distortion caused by fluctuations in the thickness of the wet wipe product to be measured. The industrial camera is connected to the system through a gigabit Ethernet interface supporting the precision time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module to within sub-microseconds. The spectral data acquisition unit uses a near-infrared spectrometer based on a micro-electromechanical system Fabry-Perot interferometer technology, which has a working wavelength range of 900 nanometers to 1700 nanometers and a spectral resolution of 10 nanometers. The production process log acquisition unit is a special software agent program running on an industrial personal computer. The program establishes a persistent connection with the manufacturing execution system back-end relational database through an open database connection interface, and after receiving the trigger signal, it executes a structured query statement containing a precise time stamp matching condition to atomically obtain parameters such as the batch number, running speed, cutter number used, and cumulative working time of the cutter of the current production line.
[0035] On the basis of the technical scheme, further improvement is made, the data aggregation and preprocessing unit is arranged on an edge computing device, the structured multi-modal data packet is encapsulated in JavaScript Object Notation format, the data packet structure includes the following main fields: a string field for storing a universally unique identifier, a 64-bit integer field for storing a Unix timestamp in nanoseconds, and a nested modal data object field, the modal data object field includes a visual data object field, which further includes an image format, Base64 encoded image data, and resolution and exposure time metadata, the modal data object field also includes a spectral data object field, which includes an array recording all sampling wavelength points, an array recording normalized reflection intensity values corresponding to each wavelength point, and integration time metadata, the modal data object field also includes a process log data object field, which directly stores structured log data obtained from a manufacturing execution system in the form of key-value pairs, including batch number, production line speed, cutter number, and cutter cumulative working time parameters, and the data aggregation and preprocessing unit generates the data packet and transmits it to the central processing unit in a publish-subscribe mode through a message-oriented middleware, specifically a zero message queue, in an asynchronous and reliable manner.
[0036] On the basis of the technical scheme, further improvement is made, the network architecture of the bidirectional attention fusion and attribution model arranged in the central processing unit includes three parallel feature extractor branches, an inter-modal collaborative attention fusion module, and a prediction and attribution head, the visual feature extractor adopts a pre-trained residual network, the spectral feature extractor adopts a one-dimensional transformer network, and the log feature extractor adopts a multi-layer perceptron, the inter-modal collaborative attention fusion module includes four stacked collaborative attention submodules, each submodule performs bidirectional attention calculation to realize deep interaction and fusion of multi-modal features, and the prediction and attribution head includes a forward prediction path and a reverse attribution path, the forward prediction path is used to calculate a risk score, and the reverse attribution path adopts a composite algorithm based on gradient and attention flow to identify key input features.
[0037] Compared with the prior art, the application has the advantages and positive effects that:
[0038] 1. Data source synchronization and fidelity: The invention ensures the absolute alignment of multiple heterogeneous data streams for the same detection object in space and time at the moment of acquisition in a high-speed dynamic environment by using a hardware synchronization beacon module based on a high-stability clock source and a high-anti-interference differential trigger signal from the physical layer. This fundamentally eliminates data pollution caused by timing misalignment, provides high-quality raw data basis with inherent logical consistency for subsequent intelligent analysis, and greatly improves the training effect and inference accuracy of the model.
[0039] 2. Transparency and accountability of decision-making process: The bidirectional attention fusion and attribution model designed by the invention can not only make high-precision risk judgments through inter-modal deep interaction and fusion, but also can analyze the abstract risk score decision-making process into the quantitative contribution of specific and observable input features, including a spot on the image, a peak in the spectrum, or a record in the log. This opens the black box of the model and provides a clear and verifiable evidence chain for each judgment, enhancing the trust of operators in the system's judgment.
[0040] 3. Closed-loop capability of fault diagnosis and process optimization: By visualizing the quantified attribution results together with the risk score, operators can immediately and intuitively understand the root cause of the alarm and determine whether it is a physical defect of the product itself or an abnormality of a specific production process parameter. This traceability to specific features greatly shortens the fault diagnosis and disposal time from more than thirty minutes to within two minutes, and provides direct and effective data support for continuous improvement of production processes and development of preventive maintenance strategies, thereby building a complete closed loop from intelligent quality inspection to intelligent manufacturing, significantly improving production efficiency and product quality management level.
[0041] 4. System robustness and industrial applicability: The technical solution of the invention fully considers the harsh environment of industrial sites, from the mechanical stability of sensor installation, the gantry structure with vibration damping design, to the low-voltage differential signal used in signal transmission to ensure high electromagnetic interference immunity, to the edge-center collaborative processing mode used in data processing architecture, all of which reflect high engineering and systematic considerations, ensuring that the entire method and system can operate stably and reliably in long-term and continuous industrial production, with good industrial applicability. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 is the overall technical scheme architecture diagram of the multi-modal data fusion wet wipe metal content quality inspection system proposed by the invention;
[0043] Figure 2is the core principle framework diagram of bidirectional attention fusion and attribution model in the application.
[0044] Figure 3 is the logical flow framework diagram of multi-modal data synchronous acquisition in the application.
[0045] Figure 4 is the logical flow framework diagram of multi-modal data convergence and preprocessing in the application.
[0046] Figure 5 is the whole data flow diagram of multi-modal data from acquisition to decision in the application.
[0047] Figure 6 is the core principle framework diagram of reverse attribution function identifying key input features in the application.
[0048] The drawing mark: 1-hard ware synchronization beacon module;2-vision data acquisition unit;3-spectrum data acquisition unit;4-production process log acquisition unit;5-manufacturing execution system back-end database;6-data convergence and preprocessing unit;7-central processing unit;8-human-computer interaction interface. DETAILED DESCRIPTION
[0049] The application provides a kind of multi-modal data fusion's metal content quality inspection method and system in wet wipe, to solve the technical problems of the existing intelligent quality inspection technology in high-speed production line environment, data inconsistency caused by data acquisition time sequence dislocation, artificial intelligence model reliability is insufficient and decision-making process is not transparent.The application constructs a bottom mechanism from data acquisition physical source, i.e.to ensure the space-time accurate synchronization of multi-modal data, and combines a kind of deep learning model with forward prediction and reverse attribution dual functions, to realize the high-precision, high-reliability detection of small metal foreign matter in wet wipe, and determine the certainty, transparent traceability of judgment basis.The technical solutions of the application will be described in detail in combination with specific embodiments.
[0050] Referring to Figure 5 The application provides a kind of multi-modal data fusion's metal content quality inspection method in wet wipe, the method includes the following steps:
[0051] First, referring to Figure 3, S1: a hardware synchronization beacon module 1 is deployed at a detection station through which the wet wipe product to be tested passes, when a specific physical reference point of the wet wipe product to be tested enters a preset triggering area of the hardware synchronization beacon module 1, a globally unique synchronization identifier containing a nanosecond-level precision timestamp is generated by a high-frequency stable clock source and a unique identifier generation logic circuit embedded in the hardware synchronization beacon module 1, and a differential triggering signal with high anti-interference characteristics is synchronously sent to a visual data acquisition unit 2, a spectral data acquisition unit 3 and a production process log acquisition unit 4.
[0052] In particular, the hardware synchronization beacon module 1 is precisely installed and deployed at the key detection node on the product transmission path of the high-speed production line of the wet wipe product. The physical implementation of the hardware synchronization beacon module 1 contains a temperature-compensated crystal oscillator, which serves as the system master clock source and can provide an ultra-high stability clock signal with a frequency stability better than one part in a million. This high-stability clock signal is the physical basis for ensuring nanosecond-level timestamp accuracy, effectively avoiding time deviation caused by clock drift. The hardware synchronization beacon module 1 also contains a field programmable gate array chip, whose internal special logic circuit is responsible for performing the following core functions: first, real-time monitoring of the position of the wet wipe product on the conveyor belt. The monitoring mechanism is realized by a high-precision laser beam sensor array or a high-speed photoelectric switch. When the physical reference point of the wet wipe product to be tested, such as its front edge or a pre-set marker point, accurately enters the pre-set trigger area defined by the overlap of multiple sensors, the field programmable gate array chip immediately captures this event. Second, based on the system master clock signal, the synchronization identifier generation logic circuit inside the field programmable gate array chip generates a globally unique synchronization identifier that conforms to the Universal Unique Dictionary Sort Identifier specification in real time. This identifier not only contains a Unix timestamp with nanosecond-level accuracy, but also integrates a serial number or batch information to ensure its uniqueness and traceability in the entire system. The precise time point of generating this identifier is defined as the "zero point" time of the wet wipe product at the detection station. Third, the field programmable gate array chip integrates a low-voltage differential signal transmitter. At the same time of synchronization identifier generation, the transmitter immediately outputs a differential trigger signal with high anti-interference characteristics in the form of a differential pair to the pre-connected visual data acquisition unit 2, spectral data acquisition unit 3, and production process log acquisition unit 4. The differential signal transmission mechanism can effectively suppress the common electromagnetic noise interference in the industrial environment, ensuring the signal integrity and extremely low delay of the trigger signal during long-distance transmission, thereby ensuring that each acquisition unit can physically receive the trigger instruction with sub-microsecond-level or even nanosecond-level synchronization accuracy. The core of this step is to eliminate the inconsistency problem of multi-modal data caused by time sequence misalignment through a high-precision synchronization mechanism at the hardware level, laying a solid and reliable foundation for subsequent data fusion and intelligent analysis.
[0053] Secondly, S2: the visual data acquisition unit 2, the spectral data acquisition unit 3 and the production process log acquisition unit 4 perform their respective data acquisition operations in parallel at the same time when receiving the differential trigger signal. Specifically, the visual data acquisition unit 2 captures a frame of digital image data covering the surface of the wet tissue product under test; the spectral data acquisition unit 3 acquires the reflectance or transmittance spectral data of the predetermined detection points on the surface of the wet tissue product under test; the production process log acquisition unit 4 initiates an accurate query to the manufacturing execution system backend database 5 or real-time data interface to obtain the structured log data associated with the current production state. Each acquisition unit immediately binds the obtained synchronization identifier with the acquired data after completing the data acquisition, forming a data unit containing the synchronization identifier and the data payload.
[0054] Specifically, the core component of the visual data acquisition unit 2 is an industrial camera using a complementary metal-oxide-semiconductor image sensor with a global shutter function. The industrial camera has an ultra-high resolution of 2048 by 2048 pixels, equipped with a telecentric lens with a focal length of 75 mm. The telecentric lens can eliminate the perspective distortion caused by the thickness fluctuation of the wet tissue product under test in the optical design, ensuring that the imaging size and shape of the wet tissue remain consistent regardless of where it is within its depth of field, thereby ensuring measurement accuracy. The industrial camera is connected to the system through a gigabit Ethernet interface supporting the precision time protocol, and its internal clock is synchronized with the clock source of the hardware synchronization beacon module 1 to within sub-microseconds, further ensuring the accurate alignment of data acquisition time. When the visual data acquisition unit 2 receives the differential trigger signal from the hardware synchronization beacon module 1, the industrial camera immediately performs a single-frame image capture operation. The captured digital image data contains high-resolution visual information of the surface of the wet tissue product under test. After the acquisition is completed, the control logic of the visual data acquisition unit 2 binds the synchronization identifier received from the hardware synchronization beacon module 1 with the digital image data in the form of digital signature or metadata embedding.
[0055] The spectral data acquisition unit 3 employs a near-infrared spectrometer based on micro-electro-mechanical system Fabry-Perot interferometer technology. The working wavelength range of this spectrometer covers 900 nanometers to 1700 nanometers, with a spectral resolution of 10 nanometers, and it has the characteristics of high-speed scanning and high sensitivity. Its optical probe is precisely calibrated and fixed above the predetermined detection point on the surface of the wet wipe product, such as the central region. When the spectral data acquisition unit 3 receives the differential trigger signal, the near-infrared spectrometer immediately performs a reflection or transmission spectral scan on the predetermined detection point on the surface of the wet wipe product to be tested within the preset integration time, and obtains the spectral data of the point. The spectral data is represented as a sequence of 81 spectral intensity values sampled at 10 nanometer intervals within the wavelength range of 900 nanometers to 1700 nanometers, which reflects the absorption or reflection characteristics of the wet wipe material and possible foreign matter to specific near-infrared light. After completing data acquisition, the spectral data acquisition unit 3 also rigidly binds the received synchronization identifier with the obtained spectral data.
[0056] The production process log acquisition unit 4 is a special software agent program running on an industrial personal computer. This program establishes a persistent connection with the manufacturing execution system backend relational database through an open database connection interface, keeping the communication channel smooth. When the production process log acquisition unit 4 receives the differential trigger signal, the software agent program immediately executes a predefined structured query statement containing an accurate timestamp matching condition. The query statement obtains structured log data associated with the current production state from the manufacturing execution system backend database 5 in real time in an atomic operation. These log data include but are not limited to the batch number, real-time running speed, cutter number used, and cumulative working time of the cutter of the current production line, etc. The accurate timestamp matching condition ensures that the queried log data is highly consistent with the wet wipe product indicated by the synchronization identifier in the time dimension. After obtaining the log data, the production process log acquisition unit 4 rigidly binds the synchronization identifier with the obtained structured log data.
[0057] Through the above parallel and synchronous data acquisition and identifier binding mechanism, it is ensured that the heterogeneous data from vision, spectrum and production process log accurately correspond to the same wet wipe product to be tested in physical time. Each data unit generated by each acquisition unit contains the same synchronization identifier, which is used as the unique key value for subsequent data aggregation.
[0058] Again, referring to Figure 4S3: transmitting the data units generated by the three data acquisition units and bound with the same synchronization identifier to a data aggregation and preprocessing unit 6. The data aggregation and preprocessing unit 6 aggregates the image data, the spectral data and the structured log data into a structured multi-modal data package according to the synchronization identifier as a unique key value. The data package follows a predefined and strict data schema, ensuring the absolute one-to-one correspondence between different modal data.
[0059] In particular, the data aggregation and preprocessing unit 6 is deployed on a dedicated edge computing device. The edge computing device is configured with high-performance processors and large-capacity memories to meet the real-time processing requirements of high-speed data streams. The data aggregation and preprocessing unit 6 is connected to each data acquisition unit through a high-speed industrial Ethernet and receives the data unit stream sent by them and bound with the synchronization identifier. After receiving the data unit, the data scheduling module inside the data aggregation and preprocessing unit 6 first buffers and matches these data units according to their internal synchronization identifier. Once the visual data unit, the spectral data unit and the production process log data unit with the same synchronization identifier are received, the data scheduling module triggers the data aggregation process immediately.
[0060] The core of the data aggregation process is to accurately match and merge different modal data belonging to the same wet wipe product according to the synchronization identifier as a unique key value. The structured multi-modal data package is packaged in JavaScript Object Notation format, which is a lightweight data exchange format, easy to machine parse and generate, and has good scalability. The data package structure follows a predefined and strict data schema, ensuring the absolute one-to-one correspondence between different modal data and data integrity. The data package structure includes the following main fields: a string field storing the universally unique dictionary ordering identifier, which accurately records the synchronization identifier; a 64-bit integer field storing the nanosecond Unix timestamp, which indicates the generation time of the data package, highly consistent with the timestamp in the synchronization identifier; and a nested modal data object field.
[0061] Inside the modal data object field, there is a visual data object field. The visual data object field further includes image formats such as Portable Network Graphics or Joint Photographic Experts Group Standard; Base64 encoded image data, which converts binary image data into text transmissible strings, facilitating network transmission and storage; and resolution and exposure time metadata, which provides key parameter information of image acquisition.
[0062] The modal data object field further contains a spectrum data object field. The spectrum data object field contains an array recording all sampling wavelength points, which lists 81 wavelength points in the range of 900 nm to 1700 nm at intervals of 10 nm; an array recording the normalized reflection intensity values at each wavelength point, the normalization usually adopts standard normal variable transformation or multivariate scattering correction to eliminate baseline drift and particle size effect; and metadata such as integration time, which records the duration of spectrum acquisition.
[0063] The modal data object field further contains a process log data object field. The process log data object field directly stores structured log data obtained from the manufacturing execution system in the form of key-value pairs, including parameters such as batch number, production line speed, cutter number, and cutter cumulative working time. The key of each key-value pair is the parameter name, and the value is the corresponding parameter data.
[0064] After successfully generating the structured multi-modal data package, the data aggregation and preprocessing unit 6 transmits the data package to the central processing unit 7 in an asynchronous and reliable manner through a message-oriented middleware, specifically a zero message queue, in a publish-subscribe mode. The zero message queue provides high-performance and low-latency message transmission capabilities, and its publish-subscribe mode allows the central processing unit 7 to subscribe to the required data while ensuring reliable delivery of the data package. Even in the case of network transient interruption, the message persistence mechanism ensures that data is not lost. This step ensures complete aggregation and standardization of data at the logical level, providing a standardized and high-quality input for subsequent deep learning model inference.
[0065] Then, referring to Figure 2 , S4: feeding the structured multi-modal data package as input into a pre-trained bidirectional attention fusion and attribution model deployed on the central processing unit 7. The model performs a forward inference process in which the inter-modal collaborative attention mechanism within the model weights and fuses deep features from different modalities, calculates and outputs a quantitative risk score representing the likelihood of the presence of metal foreign matter in the wet wipe product under test.
[0066] Specifically, the central processing unit 7 is usually a high-performance server equipped with multiple graphics processing units to provide the powerful computing power required for deep learning model inference. The bidirectional attention fusion and attribution model is pre-trained on the central processing unit 7, and its training data set is composed of multi-modal data packages of a large number of normal wet wipes and wet wipe samples containing micro metal foreign matter, and is labeled with true values by expert annotation. The model starts the forward inference process as soon as it receives the structured multi-modal data package as input.
[0067] The specific network structure of the bidirectional attention fusion and attribution model is designed as a deep neural network architecture composed of three parallel feature extractor branches, an inter-modal collaborative attention fusion module, and a prediction and attribution head.
[0068] The three parallel feature extractor branches are respectively:
[0069] First, the visual feature extractor adopts a pre-trained residual network without top classification layers as the backbone network. The residual network is pre-trained on large-scale image datasets such as ImageNet, and has strong image feature extraction capability. The input is a normalized 2048x2048x1 single-channel grayscale image tensor. The normalization process usually includes operations such as scaling pixel values to the interval of zero to one, mean normalization, and standard deviation normalization to eliminate the effects of lighting and camera parameter differences. The visual feature extractor outputs a deep feature map with dimensions of 64x64x2048 through the processing of multiple convolutional layers, pooling layers, and residual blocks. This feature map encodes spatial and semantic information in the image.
[0070] Second, the spectral feature extractor adopts a one-dimensional transformer network composed of six encoder layers. The one-dimensional transformer network is specifically designed for processing sequential data and can capture long-range dependencies within the sequence. The input is a spectral intensity sequence with dimensions of 81x1, corresponding to 81 wavelength points. The sequence is first positionally encoded to inject relative position information of the wavelength points, and then fed into the transformer network. Each encoder layer of the transformer network contains multi-head self-attention mechanisms and feedforward networks. Through the processing of six encoder layers, the spectral feature extractor outputs a sequence feature embedding with dimensions of 81x512, which is rich in shape, peak, valley, and fine changes of the spectral curve in different wavelength regions.
[0071] Third, the log feature extractor adopts a multi-layer perceptron. For categorical log data such as cutter number, an embedding layer is used to convert it into a dense vector, thereby mapping discrete category information into a continuous vector space. For numerical log data such as production line speed and cutter cumulative working time, standardization processing is performed, usually using zero-mean unit variance scaling to eliminate dimensional differences. All processed log features, including embedding vectors and standardized numerical values, are concatenated into a one-dimensional vector, which is then nonlinearly transformed by the multi-layer perceptron network to output a log feature embedding with dimensions of 1x512, which represents the current production process state.
[0072] The inter-modal collaborative attention fusion module is the core of the model, which includes a collaborative attention submodule stacked by four layers. In each layer submodule, bidirectional attention calculation is performed to realize the deep and iterative interaction and fusion between multi-modal features. Specifically:
[0073] On the one hand, the visual feature map (dimension 64 by 64 by 2048) is first reshaped and flattened into a sequence of 4096 by 2048 as a query vector. At the same time, the spectral feature embedding (81 by 512) and the log feature embedding (1 by 512) are spliced to form a sequence of 82 by 512 as a key vector and a value vector. Then, the attention weight of the visual feature to the spectral and log features is calculated. This process can be represented as:
[0074] where Q is the visual query, which is the sequence obtained by reshaping and flattening the visual feature map, representing the visual feature vector used for querying. K is the spliced key of the spectrum and the log, which is the sequence formed by splicing the spectral feature embedding and the log feature embedding, as the "key" matched with the query vector. V is the spliced value of the spectrum and the log, which is the sequence formed by splicing the spectral feature embedding and the log feature embedding, which is the "value" corresponding to the "key" and will be weighted by the attention weight. is the dimension of the key K, used to scale the dot product result of Q and K to avoid the dot product value being too large due to high dimension.
[0075] Through this attention mechanism, a visual context vector modulated by spectral and log information is generated, which enhances the parts of the visual features related to spectral and log information.
[0076] On the other hand, the spectral and log feature embedding sequence (82 by 512) is taken as the query vector. At the same time, the reshaped sequence of the visual feature map (4096 by 2048) is taken as the key vector and the value vector. Similarly, using attention calculation, a vector containing spectral and log context information modulated by visual information is generated.
[0077] Subsequently, the original modal features, i.e. the original visual feature, the original spectral feature and the original log feature, are spliced with their corresponding context vectors. The spliced features are fused by a feedforward neural network, which contains an activation function and can introduce nonlinearity. This bidirectional attention calculation and feature fusion process is performed layer by layer in four submodules, with the fused features of the previous layer as input to realize the deep and iterative interaction and fusion between multi-modal features. Finally, the module outputs a fusion feature vector with a dimension of 1 by 1024, which highly condenses all modal information.
[0078] The prediction and attribution head includes two functional paths:
[0079] The forward prediction path takes the 1-by-1024 fused feature vector as input and feeds it into a classifier consisting of two fully connected layers and a sigmoid activation function. The fully connected layers are responsible for linear transformation of the fused features, and the sigmoid activation function compresses the output values into the interval of zero to one. The final output is a scalar value between zero and one, which is the risk score. The higher the risk score, the more likely it is that there is a metal foreign object in the wet wipe product under test.
[0080] Finally, referring to Figure 6 , S5: thresholding the risk score. When the risk score exceeds a pre-set risk threshold, the backward attribution function of the model is activated. This function uses the attention weight matrix generated during the forward inference process and stored inside the model to trace back and calculate the contribution of each input feature to the final high risk score. Accordingly, a set of key input features that play a decisive role in the high risk determination are accurately identified and output. The key input features are further associated with their original modalities to present in specific forms, including one or more specific pixel coordinate regions in the digital image, one or more specific absorption or reflection peak wavebands in the spectral data, and one or more specific parameter entries in the structured log data. Moreover, the risk score and the key input features identified and enhanced in the form of highlighting or marking are presented together on the human-computer interaction interface 8, providing direct and explicit decision-making basis for the operator.
[0081] Specifically, the pre-set risk threshold is set and calibrated through experiments on a large number of samples and expert experience before system deployment, for example, set to 0.85. When the risk score output by the forward prediction path of the model is greater than the threshold, the system triggers an alarm and activates the backward attribution function of the model.
[0082] When the attribution is triggered, the backward attribution path adopts a composite algorithm based on gradient and attention flow. First, the gradient of the risk score output value with respect to the last layer of attention weight matrix in the inter-modal collaborative attention fusion module is calculated. This gradient represents the influence strength and direction of each element in the attention weight matrix on the final risk score. Then, the gradient is element-wise Hadamard product with the attention weight matrix to obtain a significance score matrix of gradient attention fusion. Each element of this score matrix represents the contribution degree of a specific interaction between modal features to the final decision. This process can be represented as:
[0083]
[0084] wherein, This represents the Hadamard product. It's an element-wise multiplication operation used to combine the gradient with the attention weight matrix. The elements are multiplied one by one to obtain the fused significance score matrix. Significance (significance score matrix): The final matrix, where each element represents the contribution of a specific link in the interaction between modal features to the final decision (e.g., risk score). Gradient: The gradient of the risk score output value relative to the last layer of attention weight matrix in the inter-modal collaborative attention fusion module. It reflects the strength and direction of the influence of each element in the attention weight matrix on the final risk score. The attention weight matrix, derived from the last layer of the "intermodal collaborative attention fusion module," records the attention weight relationships between different features.
[0085] Next, this saliency score is backpropagated along the network up to the input layer. During backpropagation, the saliency score is inversely mapped back to the original input space through a feature extractor, thereby quantifying the contribution of the original input features. For visual data, this process generates a saliency map of the same size as the original image, where bright areas are the key pixel regions that play a decisive role in high-risk judgments, such as a dark spot with a diameter of three millimeters. For spectral data, the backpropagation results generate importance scores for each feature dimension, with the bands with the highest scores identified as the bands with key absorption or reflection peaks, such as the anomalous absorption peaks in the 1030 nm to 1050 nm band. For log data, importance scores are also generated for each parameter item, with the parameter item with the highest score being the key feature, such as the item that the cutter numbered three has accumulated more than eight hours of working time.
[0086] Finally, the central processing unit 7 transmits the calculated risk score along with the identified key input features, enhanced with highlights or markers, to the human-machine interface 8. The human-machine interface 8 presents this information in an intuitive and graphical manner. For example, in digital images, key pixel areas are highlighted with red borders or semi-transparent overlays; in spectral curves, key bands are marked with special colors or shaded areas; in production process logs, key parameter entries are bolded or highlighted. This integrated and transparent information presentation provides operators with direct and clear decision-making support, enabling them to quickly understand the source of risk and take appropriate corrective measures, such as immediately stopping the machine for inspection, replacing the cutter, or adjusting production process parameters. This step greatly improves decision-making efficiency and system reliability, achieving full transparency and traceability of the intelligent quality inspection process.
[0087] See Figure 1The application also provides a multi-modal data fusion wet wipe metal content quality inspection system, which comprises:
[0088] A hardware synchronization beacon module 1, which is used to generate a globally unique synchronization identifier containing a nanosecond timestamp when the wet wipe product enters the preset detection area, and synchronously send a high anti-interference differential trigger signal to each data acquisition unit. The physical implementation of the module includes a temperature-compensated crystal oscillator, which provides a system master clock signal with a frequency stability better than one part in a million. The internal field programmable gate array chip is responsible for generating a synchronization identifier that meets the general unique dictionary sorting identifier specification according to the master clock signal, and integrates a low-voltage differential signal transmitter for outputting the trigger signal in differential pair form, ensuring high signal-to-noise ratio and transmission integrity of the signal in the industrial electromagnetic environment.
[0089] A visual data acquisition unit 2 connected to the hardware synchronization beacon module 1. After receiving the differential trigger signal, the unit accurately captures a frame of high-resolution digital image data covering the surface of the wet wipe product to be tested. After the acquisition is completed, the synchronization identifier is rigidly bound with the captured image data. The core component of the visual data acquisition unit 2 is an industrial camera using a complementary metal oxide semiconductor image sensor with a global shutter function, with a pixel resolution of 2048 by 2048 pixels, equipped with a telecentric lens with a focal length of 75 mm, which is used to eliminate perspective distortion caused by the thickness fluctuation of the wet wipe product to be tested. The industrial camera is connected to the system through a gigabit Ethernet interface supporting the precise time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module 1 to within sub-microseconds.
[0090] A spectral data acquisition unit 3 connected to the hardware synchronization beacon module 1. After receiving the differential trigger signal, the unit accurately acquires the reflection or transmission spectral data of the predetermined detection point on the surface of the wet wipe product to be tested. After the acquisition is completed, the synchronization identifier is rigidly bound with the acquired spectral data. The spectral data acquisition unit 3 uses a near-infrared spectrometer based on a micro-electro-mechanical system Fabry-Perot interferometer technology, with a working wavelength range of 900 nanometers to 1700 nanometers and a spectral resolution of 10 nanometers, which can capture spectral information at high speed and high sensitivity.
[0091] A production process log collection unit 4 is connected to the hardware synchronization beacon module 1. Upon receiving the differential trigger signal, the unit initiates an exact query to the manufacturing execution system backend database 5 or real-time data interface to obtain the structured log data associated with the current production state. After obtaining the data, the synchronization identifier is rigidly bound to the obtained structured log data. The production process log collection unit 4 is a special software agent running on an industrial personal computer, which establishes a persistent connection with the manufacturing execution system backend relational database through an open database connection interface, and upon receiving the trigger signal, executes a structured query statement containing an exact timestamp matching condition to obtain key parameters such as batch number, running speed, cutting tool number used, and cumulative working time of the cutting tool of the current production line in an atomic operation.
[0092] A data aggregation and preprocessing unit 6 is connected to the visual data collection unit 2, the spectral data collection unit 3, and the production process log collection unit 4. The unit aggregates the received image data, spectral data, and structured log data into a structured multi-modal data package according to the synchronization identifier as the unique key value. The data aggregation and preprocessing unit 6 is deployed on an edge computing device, and the structured multi-modal data package is packaged in JavaScript Object Notation format. The data package structure includes a string field storing the universally unique dictionary ordering identifier, a 64-bit integer field storing the nanosecond Unix timestamp, and a nested modal data object field. The modal data object field internally includes a visual data object field containing the image format, Base64 encoded image data, and resolution and exposure time metadata; a spectral data object field containing an array recording all sampled wavelength points, an array recording the normalized reflectance intensity values corresponding to each wavelength point, and integration time metadata; and a process log data object field directly storing the structured log data obtained from the manufacturing execution system in the form of key-value pairs. After generating the data package, the data aggregation and preprocessing unit 6 transmits the data package asynchronously and reliably to the central processing unit 7 in a publish-subscribe mode through a message-oriented middleware, specifically a zero message queue.
[0093] A central processing unit 7 is connected to the data aggregation and pre-processing unit 6. The bidirectional attention fusion and attribution model is deployed in the central processing unit 7. The unit receives the multi-modal data package and performs a forward inference process to calculate a quantitative risk score. When the risk score exceeds a preset risk threshold, the reverse attribution function of the model is activated to trace back and identify the key input features that play a decisive role in the high-risk judgment, and finally output the risk score and the identified key input features. The network architecture of the bidirectional attention fusion and attribution model deployed in the central processing unit 7 includes three parallel feature extractor branches, an inter-modal collaborative attention fusion module, and a prediction and attribution head. The visual feature extractor adopts a pre-trained residual network. The spectral feature extractor adopts a one-dimensional transformer network. The log feature extractor adopts a multi-layer perceptron. The inter-modal collaborative attention fusion module includes four stacked collaborative attention sub-modules, each of which performs bidirectional attention calculation to realize deep interactive fusion of multi-modal features. The prediction and attribution head includes a forward prediction path for calculating the risk score and a reverse attribution path that adopts a composite algorithm based on gradient and attention flow to identify key input features.
[0094] A human-computer interaction interface 8 is connected to the central processing unit 7. The human-computer interaction interface 8 is used to receive and visually present the risk score and the identified key input features in an intuitive and graphical manner, for example, to enhance the display of key parts in image, spectral and log data in the form of highlighting or marking, thereby providing direct and explicit decision basis for the operator.
[0095] The above is only a specific embodiment of the present application, but the technical features of the present application are not limited thereto. Any simple change, equivalent replacement or modification made on the basis of the present application to solve the basically same technical problem and achieve the basically same technical effect is covered by the protection scope of the present application.
Claims
1. A method for detecting the content of metal in a multi-modal data fusion wet wipe, characterized in that, The method comprises the following steps: S1: deploying a hardware synchronization beacon module at a detection station through which a wet wipe product to be tested passes, when a specific physical reference point of the wet wipe product to be tested enters a preset triggering area of the hardware synchronization beacon module, the hardware synchronization beacon module generates a globally unique synchronization identifier and synchronously sends a differential triggering signal to a visual data acquisition unit, a spectral data acquisition unit, and a production process log acquisition unit; S2: the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit perform respective data acquisition operations in parallel at the same time when the differential triggering signal is received, the visual data acquisition unit captures a frame of digital image data, the spectral data acquisition unit acquires spectral data of a predetermined detection point on the surface of the wet wipe product to be tested, and the production process log acquisition unit acquires structured log data associated with the current production state from a manufacturing execution system, and each acquisition unit binds the obtained synchronization identifier with the acquired data to form a data unit containing the synchronization identifier and data payload; S3: transmitting the data units generated by the three data acquisition units and bound with the same synchronization identifier to a data aggregation and preprocessing unit, and the data aggregation and preprocessing unit aggregates the digital image data, the spectral data, and the structured log data into a structured multimodal data package according to the synchronization identifier as a unique key value; S4: feeding the structured multimodal data package as input into a pre-trained bidirectional attention fusion and attribution model deployed on a central processing unit, the model performs a forward inference process, calculates, and outputs a quantitative risk score; S5: performing threshold judgment on the risk score, when the risk score exceeds a preset risk threshold, activating the reverse attribution function of the model, identifying and outputting a set of key input features that play a decisive role in high-risk judgment, the key input features are further associated with their original modalities to be presented in specific forms, and the risk score and the key input features with enhanced visual processing are presented on a human-computer interaction interface; In steps S4 and S5, the network architecture of the bidirectional attention fusion and attribution model comprises three parallel feature extractor branches, an inter-modal collaborative attention fusion module, and a prediction and attribution head; the prediction and attribution head comprises a forward prediction path and a reverse attribution path, the forward prediction path is used to calculate the risk score, and the reverse attribution path adopts a composite algorithm based on gradient and attention flow, which is used to calculate the gradient of the risk score output value with respect to the last layer attention weight matrix in the inter-modal collaborative attention fusion module, perform Hadamard product on the gradient and the attention weight matrix to obtain a saliency score, and perform backward propagation of the saliency score along the network until the input layer to identify the key input features.
2. The method of claim 1, wherein, In the step S1, the physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator, which provides a system master clock signal with a frequency stability better than one part in a million; the hardware synchronization beacon module also includes a field programmable gate array chip, the internal logic circuit of which is responsible for generating the synchronization identifier conforming to the specification of the universally unique dictionary identifier based on the system master clock signal, and integrating a low-voltage differential signal transmitter for outputting the differential trigger signal in the form of a differential pair.
3. The method of claim 1, wherein, In the step S2: The core component of the visual data acquisition unit is an industrial camera using a CMOS image sensor with a global shutter function; the industrial camera is connected to the system through a gigabit Ethernet interface supporting the precision time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module to a sub-microsecond level; The spectral data acquisition unit uses a near-infrared spectrometer constructed based on the micro-electro-mechanical system Fabry-Perot interferometer technology; The production process log acquisition unit is a dedicated software agent program running on an industrial personal computer, which establishes a persistent connection with the manufacturing execution system back-end relational database through an open database connectivity interface, and after receiving a trigger signal, executes a structured query statement containing a precision timestamp matching condition to obtain the batch number, running speed, used cutter number and cumulative working time of the cutter parameters of the current production line in an atomic operation.
4. The method of claim 1, wherein, In the step S3, the data aggregation and preprocessing unit is deployed on an edge computing device; the structured multi-modal data packet is packaged in JavaScript Object Notation format, and the data packet structure includes a string field storing the universally unique dictionary identifier, a 64-bit integer field storing the nanosecond-level Unix timestamp, and a nested modal data object field; after generating the data packet, the data aggregation and preprocessing unit transmits the data packet to the central processing unit in a publish-subscribe mode through a message-oriented middleware, specifically a zero message queue, in an asynchronous and reliable manner.
5. A multimodal data fusion-based quality inspection system for metal content in wet wipes, characterized in that, It comprises: a hardware synchronization beacon module for generating a globally unique synchronization identifier and sending a differential trigger signal when a wet wipe product to be tested enters a preset trigger area; a visual data acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and capturing digital image data covering the surface of the wet wipe product to be tested, and binding the synchronization identifier with the digital image data; a spectral data acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and obtaining spectral data of predetermined detection points on the surface of the wet wipe product to be tested, and binding the synchronization identifier with the spectral data; a production process log acquisition unit connected to the hardware synchronization beacon module for receiving the differential trigger signal and obtaining structured log data associated with the current production state from the manufacturing execution system, and binding the synchronization identifier with the structured log data. a data aggregation and preprocessing unit connected to the visual data acquisition unit, the spectral data acquisition unit, and the production process log acquisition unit, configured to aggregate the received digital image data, spectral data, and structured log data into a structured multi-modal data package according to the synchronization identifier; a central processing unit connected to the data aggregation and preprocessing unit and deployed with a bidirectional attention fusion and attribution model, configured to receive the multi-modal data package and perform forward inference to calculate a risk score, and activate the reverse attribution function of the model to identify key input features when the risk score exceeds a preset threshold, and output the risk score and the key input features; a human-computer interaction interface connected to the central processing unit, configured to receive and visually present the risk score and the key input features; the network architecture of the bidirectional attention fusion and attribution model deployed in the central processing unit includes three parallel feature extractor branches, an inter-modal collaborative attention fusion module, and a prediction and attribution head; the three parallel feature extractor branches consist of a visual feature extractor, a spectral feature extractor, and a log feature extractor, the visual feature extractor adopts a pre-trained residual network, the spectral feature extractor adopts a one-dimensional transformer network, and the log feature extractor adopts a multi-layer perceptron; the inter-modal collaborative attention fusion module includes four stacked collaborative attention submodules, each submodule performs bidirectional attention calculation to realize deep interactive fusion of multi-modal features; the prediction and attribution head includes a forward prediction path and a reverse attribution path, the forward prediction path is used to calculate a risk score, and the reverse attribution path adopts a composite algorithm based on gradient and attention flow to identify key input features.
6. The system of claim 5, wherein, The physical implementation of the hardware synchronization beacon module includes a temperature-compensated crystal oscillator that provides a system master clock signal with a frequency stability better than one part in a million; the hardware synchronization beacon module also includes a field programmable gate array chip, the internal logic circuit of which is responsible for generating the synchronization identifier conforming to the universal unique dictionary ordering identifier specification based on the system master clock signal, and integrating a low-voltage differential signal transmitter for outputting the differential trigger signal in the form of a differential pair.
7. The system of claim 5, wherein: the core component of the visual data acquisition unit is an industrial camera using a CMOS image sensor with a global shutter function, with a pixel resolution of 2048 by 2048 pixels and equipped with a telecentric lens with a focal length of 75 mm; the industrial camera is connected to the system through a gigabit Ethernet interface supporting the precision time protocol, ensuring that its internal clock is synchronized with the clock source of the hardware synchronization beacon module to within sub-microsecond level; the spectral data acquisition unit uses a near-infrared spectrometer based on micro-electro-mechanical system Fabry-Perot interferometer technology, with a working wavelength range of 900 nanometers to 1700 nanometers and a spectral resolution of 10 nanometers; The production process log collection unit is a special software agent program running on an industrial personal computer, which establishes a persistent connection with a manufacturing execution system back-end relational database through an open database connection interface, and after receiving a trigger signal, executes a structured query statement containing a precise timestamp matching condition to obtain the batch number, running speed, used cutter number and cumulative working time of the cutter of the current production line in an atomic operation.
8. The system of claim 5, wherein, The data aggregation and preprocessing unit is deployed on an edge computing device; the structured multi-modal data packet is packaged in JavaScript Object Notation format, the data packet structure includes a string field storing a universally unique dictionary sorted identifier, a 64-bit integer field storing a nanosecond Unix timestamp, and a nested modal data object field; after generating the data packet, the data aggregation and preprocessing unit transmits the data packet to the central processing unit in an asynchronous and reliable manner through a message-oriented middleware, specifically a zero message queue, in a publish-subscribe mode.
Citation Information
Patent Citations
Online detection method and system for foreign matters in meat products
CN119619427A
Industrial process supervision and management system based on multi-modal data processing
CN120523091A