Invisible infrared code spraying image acquisition and intelligent identification processing method and system

By combining residual cross-layer feature fusion network and dynamic tracking trajectory equation with thermal excitation technology, the problem of recognition efficiency and accuracy of invisible infrared inkjet printing in high-speed motion environment is solved, and precise three-dimensional positioning and stable tracking of the invisible inkjet printing area are achieved.

CN120953745APending Publication Date: 2025-11-14BEIJING YIJIA HAICHUANG INTERNET OF THINGS TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511108179.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional stealth infrared inkjet coding recognition technology is inefficient in high-speed motion environments, struggles to extract dynamic features of the stealth coding area, and is not adaptable to complex textures and lighting changes, resulting in reduced recognition accuracy and stability.

Method used

A residual cross-layer feature fusion network is used to perform layered processing on continuous image sequences. By connecting channels across scale residuals and dynamic tracking trajectory equations, multi-scale feature extraction and real-time tracking of the invisible inkjet coding area are achieved. Temperature field distribution images are obtained by combining thermal excitation technology, and the parameters of the feature extraction layer are adaptively adjusted. Deformation feature information is fused for accurate three-dimensional positioning.

Benefits of technology

It significantly improves the accuracy and stability of invisible inkjet printing recognition, maintains high-efficiency recognition under high-speed movement and complex lighting conditions, adapts to packaging objects of different shapes and materials, and reduces the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953745A_ABST
    Figure CN120953745A_ABST
Patent Text Reader

Abstract

The invention provides an invisible infrared code spraying image acquisition and intelligent identification processing method and system, and relates to the technical field of image processing, and the method comprises the steps: collecting a continuous image sequence of a high-speed moving intelligent packaging object, extracting an invisible code spraying region through a residual cross-layer feature fusion network, and generating a spatial-temporal feature data stream; and acquiring a multi-scale feature vector by using a feature extraction layer and cross-scale residual connection, calculating a dynamic tracking trajectory equation, and fusing deformation feature information to correct spatial position parameters, thereby realizing accurate three-dimensional space tracking and positioning of the invisible code spraying area. According to the method, the accuracy and the real-time performance of invisible sprayed code identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image processing technology, and more particularly to a method and system for acquiring and intelligently recognizing invisible infrared inkjet printing images. Background Technology

[0002] Invisible infrared coding technology is a special marking method applied to the surface of smart packaged objects. This technology uses special inks to create marks on the product surface that are invisible to the naked eye but recognizable by specific devices. It is widely used in product anti-counterfeiting, traceability management, and logistics tracking. With the rapid development of the smart packaging industry, the demand for real-time identification and tracking of invisible infrared coding in high-speed environments is increasing.

[0003] Traditional stealth infrared inkjet coding identification technology mainly relies on image acquisition and processing in static or low-speed environments, typically employing conventional computer vision algorithms to process and analyze the acquired images. These methods usually include steps such as image preprocessing, feature extraction, and pattern recognition to achieve the detection and decoding of stealth inkjet codes. With the increasing speed of production lines and the diversification of smart packaging forms, the application of existing technologies in high-speed moving environments faces numerous challenges.

[0004] Traditional image processing algorithms are inefficient when processing continuous image sequences of high-speed moving objects, failing to effectively extract dynamic features of invisible inkjet coding areas, resulting in a significant drop in recognition rates on high-speed production lines. In particular, when packaged objects move at high speeds, the acquired images often suffer from blurring and distortion, making it difficult for conventional feature extraction methods to accurately locate and identify invisible inkjet coding.

[0005] Existing technologies lack effective multi-scale feature fusion mechanisms and cannot adapt to smart packaging objects with different shapes, materials, and surface characteristics. When the surface of the packaging object has complex textures or reflective areas, traditional methods often struggle to distinguish between invisible inkjet printing and background interference, especially under changing ambient light conditions, where recognition accuracy is further reduced.

[0006] Existing invisible inkjet printing recognition systems are insufficiently adaptable to the deformation of packaged objects during high-speed movement. When the packaged object rotates, deforms, or vibrates during transport, the lack of dynamic tracking and position correction capabilities prevents the system from continuously and accurately locating the invisible inkjet printing area, thus affecting the reliability and stability of subsequent recognition and decoding processes. Summary of the Invention

[0007] The present invention provides a method and system for acquiring and intelligently recognizing stealth infrared inkjet printing images, which can solve the problems in the prior art.

[0008] A first aspect of the present invention provides a method for acquiring and intelligently recognizing stealth infrared inkjet printing images, comprising:

[0009] A continuous image sequence of a high-speed moving smart packaging object is acquired. The continuous image sequence is then processed in layers through a residual cross-layer feature fusion network to extract the invisible inkjet coding area on the surface of the smart packaging object and generate a spatiotemporal feature data stream of the invisible inkjet coding area.

[0010] The spatiotemporal feature data stream is input into the feature extraction layer of the residual cross-layer feature fusion network. Different feature extraction layers are connected in series through cross-scale residual connection channels. The parameters of each feature extraction layer are adaptively adjusted according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region.

[0011] The dynamic tracking trajectory equation is calculated based on the multi-scale feature vector. The coefficients of the dynamic tracking trajectory equation are optimized by iterative calculation to achieve real-time tracking of the invisible inkjet coding area. According to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network. The deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters.

[0012] The corrected spatial position parameters are used to modify the dynamic tracking trajectory equation, calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and output the real-time tracking and positioning results.

[0013] The continuous image sequence is processed hierarchically using a residual cross-layer feature fusion network to extract the invisible inkjet coding region on the surface of the smart packaging object, generating a spatiotemporal feature data stream of the invisible inkjet coding region, including:

[0014] A continuous image sequence of the surface of a smart packaging object is acquired. The continuous image sequence is then subjected to hierarchical feature analysis through a residual cross-layer feature fusion network. A pulsed heat source is used to thermally excite the surface of the smart packaging object. A temperature field distribution image is then established based on the thermal diffusion response of the surface of the smart packaging object.

[0015] The temperature field distribution image is analyzed based on the first feature extraction layer of the residual cross-layer feature fusion network. The spatiotemporal feature parameters of the temperature field distribution image are constructed according to the temperature field time change rate and spatial temperature gradient. The spatiotemporal feature parameters contain the dynamic evolution information of the temperature field distribution image.

[0016] The spatiotemporal feature parameters are progressively processed by multiple residual blocks cascaded within the backbone layer of the residual cross-layer feature fusion network. Each residual block enhances the hierarchical correlation of the spatiotemporal feature parameters through main branch mapping and cross-layer connections, and temperature field features are constructed based on the hierarchical correlation.

[0017] Based on the temperature field characteristics, the temperature response characteristics of the invisible inkjet coding area are analyzed, and a feature index including initial temperature, temperature increment and time constant is constructed. Based on the residual cross-layer feature fusion network, the feature index and the temperature field characteristics are combined to generate a spatiotemporal feature data stream characterizing the invisible inkjet coding area.

[0018] Hierarchical feature analysis is performed on the continuous image sequence using a residual cross-layer feature fusion network. A pulsed heat source is used to thermally excite the surface of the smart packaging object. A temperature field distribution image is then established based on the thermal diffusion response of the smart packaging object's surface, including:

[0019] A multi-layer residual network structure is constructed, the feature increment of each layer is calculated through residual mapping, cross-layer feature connection channels are established between adjacent feature layers, feature weights are allocated to each layer, and multi-layer image features of the continuous image sequence are extracted. The multi-layer image features guide the initial configuration of the heat source excitation parameters.

[0020] Based on the multi-level image feature analysis, the material properties of the smart packaging object are analyzed. The initial heat source power in the coarse adjustment stage is determined according to the material properties. A layered progressive heat source control strategy is adopted to apply pulsed thermal excitation to the surface of the smart packaging object. The initial heat source power, dynamic compensation amount and adaptive weighting factor together determine the real-time heat source power in the fine adjustment stage.

[0021] The thermal diffusion response generated on the surface of the smart packaging object under the pulsed thermal excitation is detected, and a temperature field distribution image is generated based on the thermal diffusion response combined with the heat source influence coefficient, thermal diffusion coefficient and thermal conductivity.

[0022] By concatenating different feature extraction layers through a cross-scale residual connection channel, and adaptively adjusting the parameters of each feature extraction layer according to the image complexity of the spatiotemporal feature data stream, the multi-scale feature vector of the invisible inkjet coding region is obtained, including:

[0023] A dual-constraint architecture of teacher network and student network is constructed. Feature connection channels are established through feature transfer between the teacher network and the student network, and feature transfer paths are formed based on the feature connection channels.

[0024] Feature extraction is performed based on the feature transfer path. The high-dimensional feature representation extracted by the teacher network is used as knowledge guidance. The feature space is compressed and mapped based on the student network under the knowledge guidance. The feature connection channel is dynamically adjusted according to the feature reconstruction result of the compressed mapping.

[0025] Based on the feature reconstruction results of the compression mapping, the spatiotemporal data stream features are analyzed, and the spatiotemporal data stream features are input into the teacher network for feature extraction. The features extracted by the teacher network are used to adjust the feature extraction process of the student network. The feature extraction process of the student network is used for feature consistency constraints.

[0026] Based on the feature consistency constraint, the feature reconstruction result of the compressed mapping is spatially aligned, the feature extraction parameters of the teacher network are updated based on the spatially aligned features, and the optimized features are output according to the teacher network.

[0027] The spatially aligned features are fused with the teacher network and the student network under the guidance of the optimized features, and a multi-scale feature vector of the invisible inkjet coding region is generated based on the result of the feature fusion.

[0028] The dynamic tracking trajectory equation is calculated based on the multi-scale feature vectors, and the coefficients of the dynamic tracking trajectory equation are optimized through iterative calculation to achieve real-time tracking of the invisible inkjet coding area, including:

[0029] A variational energy functional is constructed based on the multi-scale feature vectors, and the optimal form of the dynamic tracking trajectory equation is derived based on the variational energy functional. The optimal form of the dynamic tracking trajectory equation includes the coefficient matrix to be optimized.

[0030] The optimal form of the dynamic tracking trajectory equation is decomposed into parameter optimization terms and constraint optimization terms. Based on the parameter optimization terms, an objective function is constructed for the coefficient matrix to be optimized. Based on the constraint optimization terms, weight coefficients are introduced to establish constraints. Based on the objective function and the constraints, an optimization objective is constructed.

[0031] The optimization objective is solved by a dynamic iterative method. The coefficient matrix is ​​updated by the parameter optimization terms. The updated coefficient matrix is ​​substituted into the constraint optimization terms to calculate the weight coefficients. The coefficient matrix is ​​then optimized based on the gradient information of the weight coefficients.

[0032] By substituting the optimized coefficient matrix into the optimal form of the dynamic tracking trajectory equation, real-time tracking of the invisible inkjet printing area can be achieved.

[0033] Based on the multi-scale feature vectors, a variational energy functional is constructed. The optimal form of the dynamic tracking trajectory equation is derived from this variational energy functional, including:

[0034] The variational energy functional characterizes the spatial distribution characteristics of the trajectory through the spatial gradient term, describes the temporal evolution of the trajectory through the time derivative term, and measures the degree of matching between the trajectory and the multi-scale feature vector through the feature fitting term. The spatial gradient term, the time derivative term, and the feature fitting term are combined to form a unified variational energy expression.

[0035] Based on the variational energy functional, spatial gradient information is obtained by performing spatial partial derivative operation on the spatial gradient term in the variational energy expression; temporal evolution information is extracted by performing temporal partial derivative operation on the temporal derivative term; feature matching information is analyzed by performing feature partial derivative operation on the feature fitting term; and the spatial gradient information, the temporal evolution information and the feature matching information are combined to construct a dynamic tracking trajectory equation.

[0036] Based on the dynamic tracking trajectory equation, the spatial gradient information and the temporal evolution information are mapped into trajectory state parameters. The trajectory state parameters are combined with the feature matching information to optimize the dynamic tracking trajectory equation, thereby obtaining the optimal form of the dynamic tracking trajectory equation.

[0037] The dynamic tracking trajectory equation is corrected using the corrected spatial position parameters to calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and the real-time tracking and positioning results are output, including:

[0038] A fractional-order trajectory correction operator is constructed using the corrected spatial position parameters. The fractional-order trajectory correction operator includes a fractional-order time derivative term and a fractional-order spatial derivative term. Long-range integration is performed on the historical trajectory state using the fractional-order time derivative term, and nonlocal differentiation is performed on the spatial distribution of the trajectory using the fractional-order spatial derivative term. The dynamic tracking trajectory equation is corrected based on the results of the long-range integration and the nonlocal differentiation.

[0039] The modified dynamic tracking trajectory equation is mapped to a fractional phase space. A mapping relationship between trajectory state and position parameters is established based on the fractional phase space. The fractional state variables in the trajectory evolution process are calculated based on the mapping relationship.

[0040] Substituting the fractional-order state variables into the modified dynamic tracking trajectory equation, the precise three-dimensional spatial position of the invisible inkjet printing area is obtained through fractional-order iterative calculation of the trajectory state, and a real-time tracking and positioning result is constructed based on the precise three-dimensional spatial position.

[0041] A second aspect of the present invention provides a stealth infrared inkjet printing image acquisition and intelligent recognition processing system, comprising:

[0042] The first unit is used to acquire a continuous image sequence of a high-speed moving smart packaging object, perform layered processing on the continuous image sequence through a residual cross-layer feature fusion network, extract the invisible inkjet coding area on the surface of the smart packaging object, and generate a spatiotemporal feature data stream of the invisible inkjet coding area.

[0043] The second unit is used to input the spatiotemporal feature data stream into the feature extraction layer of the residual cross-layer feature fusion network, connect different feature extraction layers in series through cross-scale residual connection channels, and adaptively adjust the parameters of each feature extraction layer according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region.

[0044] The third unit is used to calculate the dynamic tracking trajectory equation based on the multi-scale feature vector, optimize the coefficients of the dynamic tracking trajectory equation through iterative calculation, and realize real-time tracking of the invisible inkjet printing area; according to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network, and the deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters;

[0045] The fourth unit is used to correct the dynamic tracking trajectory equation using the corrected spatial position parameters, calculate the precise three-dimensional spatial position of the invisible inkjet printing area, and output the real-time tracking and positioning results.

[0046] A third aspect of the present invention provides an electronic device, comprising:

[0047] processor;

[0048] Memory used to store processor-executable instructions;

[0049] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0050] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0051] The beneficial effects of this application are as follows:

[0052] This invention uses a residual cross-layer feature fusion network to perform layered processing on continuous image sequences, achieving accurate identification and feature extraction of invisible inkjet printing areas on high-speed moving smart packaging objects. This significantly improves the recognition accuracy and can effectively cope with complex lighting and high-speed movement scenarios.

[0053] This invention employs an adaptively adjustable feature extraction layer and a cross-scale residual connection channel to dynamically optimize network parameters based on image complexity. This significantly enhances the system's adaptability to objects packaged with different shapes and materials, reduces computational resource consumption, and improves processing efficiency.

[0054] This invention achieves real-time and accurate three-dimensional positioning of the invisible inkjet printing area by fusing dynamic tracking trajectory equations and deformation feature information. It effectively overcomes the deformation interference of packaged objects in high-speed movement, enabling the system to maintain stable tracking and recognition performance even under high-speed operation conditions on industrial production lines. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the method for acquiring and intelligently recognizing invisible infrared inkjet printing images according to an embodiment of the present invention.

[0056] Figure 2 This is a flowchart illustrating the pulse thermal excitation temperature field analysis of intelligent packaging materials according to an embodiment of the present invention.

[0057] Figure 3 This is a flowchart illustrating the technical process of constructing a variational energy functional based on multi-scale feature vectors and deriving the optimal form of the dynamic tracking trajectory equation in an embodiment of the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0060] Figure 1 This is a flowchart illustrating the stealth infrared inkjet printing image acquisition and intelligent recognition processing method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0061] A continuous image sequence of a high-speed moving smart packaging object is acquired. The continuous image sequence is then processed in layers through a residual cross-layer feature fusion network to extract the invisible inkjet coding area on the surface of the smart packaging object and generate a spatiotemporal feature data stream of the invisible inkjet coding area.

[0062] The spatiotemporal feature data stream is input into the feature extraction layer of the residual cross-layer feature fusion network. Different feature extraction layers are connected in series through cross-scale residual connection channels. The parameters of each feature extraction layer are adaptively adjusted according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region.

[0063] The dynamic tracking trajectory equation is calculated based on the multi-scale feature vector. The coefficients of the dynamic tracking trajectory equation are optimized by iterative calculation to achieve real-time tracking of the invisible inkjet coding area. According to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network. The deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters.

[0064] The corrected spatial position parameters are used to modify the dynamic tracking trajectory equation, calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and output the real-time tracking and positioning results.

[0065] In one optional implementation, the continuous image sequence is processed hierarchically using a residual cross-layer feature fusion network to extract the invisible inkjet coding region on the surface of the smart packaging object, generating a spatiotemporal feature data stream of the invisible inkjet coding region, including:

[0066] A continuous image sequence of the surface of a smart packaging object is acquired. The continuous image sequence is then subjected to hierarchical feature analysis through a residual cross-layer feature fusion network. A pulsed heat source is used to thermally excite the surface of the smart packaging object. A temperature field distribution image is then established based on the thermal diffusion response of the surface of the smart packaging object.

[0067] The temperature field distribution image is analyzed based on the first feature extraction layer of the residual cross-layer feature fusion network. The spatiotemporal feature parameters of the temperature field distribution image are constructed according to the temperature field time change rate and spatial temperature gradient. The spatiotemporal feature parameters contain the dynamic evolution information of the temperature field distribution image.

[0068] The spatiotemporal feature parameters are progressively processed by multiple residual blocks cascaded within the backbone layer of the residual cross-layer feature fusion network. Each residual block enhances the hierarchical correlation of the spatiotemporal feature parameters through main branch mapping and cross-layer connections, and temperature field features are constructed based on the hierarchical correlation.

[0069] Based on the temperature field characteristics, the temperature response characteristics of the invisible inkjet coding area are analyzed, and a feature index including initial temperature, temperature increment and time constant is constructed. Based on the residual cross-layer feature fusion network, the feature index and the temperature field characteristics are combined to generate a spatiotemporal feature data stream characterizing the invisible inkjet coding area.

[0070] The system acquires a continuous sequence of images of the surface of intelligently packaged objects using an industrial camera array, including visible light and infrared cameras, to capture the surface texture and thermal properties of the packaged objects, respectively. The visible light camera has a resolution of 1920×1080, a frame rate of 120 frames per second, and a field of view of 45 degrees; the infrared camera has a resolution of 640×480, a frame rate of 60 frames per second, and a temperature resolution of 0.05 degrees Celsius. The two cameras are synchronized using a hardware-based triggering mechanism to ensure consistent acquisition time, with spatial registration errors controlled within 0.5 pixels. The camera array is mounted above the conveyor belt at a distance of 50 centimeters, covering the width of the conveyor belt. For a packaged object moving at a speed of 2 meters per second, the system acquires one frame every 16.7 milliseconds, continuously acquiring 30 frames to form an image sequence.

[0071] A residual cross-layer feature fusion network is used to perform hierarchical feature analysis on continuous image sequences. This network adopts the ResNet architecture and contains five feature extraction layers. Each layer contains a different number of residual blocks, ranging from 2, 3, 5, 3, and 2 residual blocks from shallow to deep layers. The residual blocks adopt a dual-branch structure, with the main branch containing two 3×3 convolutional layers and a batch normalization layer, and the bypass branch being a direct connection. The number of channels in the convolutional layers gradually increases from 64 channels in the first layer to 512 channels in the fifth layer. The network input is a spatiotemporal data block composed of multiple consecutive frames of images, with a size of 640×480×3×10, representing 10 frames of three-channel images. The first feature extraction layer uses a 7×7 convolution with a stride of 2 for downsampling, and the output feature map size is 320×240×64. Subsequent feature extraction layers further downsample using convolutions with a stride of 2 within the residual blocks, and finally, the fifth layer outputs a feature map size of 40×30×512.

[0072] A pulsed heat source is used to thermally excite the surface of the smart packaged object. The heat source is an infrared array emitter containing 10×10 emitting units, covering an area of ​​30×30 cm, with a power density of 0.5 W / cm². The pulse parameters are dynamically adjusted according to the characteristics of the packaging material. For plastic packaging, the pulse width is set to 20 ms with an interval of 100 ms; for metal packaging, the pulse width is set to 15 ms with an interval of 80 ms. The system maintains a stable thermal excitation effect through closed-loop control, monitoring the packaging surface temperature in real time. Thermal excitation stops when the temperature rise reaches a preset threshold (usually 5 degrees Celsius). The thermal excitation process lasts for 300 ms, generating 3-5 thermal pulses, which is sufficient to trigger the internal thermal diffusion process of the material.

[0073] A temperature field distribution image is established based on the thermal diffusion response of the surface of intelligent packaging objects. Temperature image sequences before, during, and after thermal excitation are acquired using an infrared camera. Differential processing is used to eliminate the influence of ambient temperature, yielding temperature changes purely caused by thermal excitation. The temperature field images undergo spatial filtering to reduce noise, using a 5×5 Gaussian kernel with a standard deviation of 1.2. For images with perspective distortion, inverse perspective transformation is applied for correction to ensure accurate mapping of the temperature field onto the real object surface. Correction parameters, including the intrinsic parameter matrix and distortion coefficients, are pre-calculated using a calibration board. The processed temperature field distribution image has a resolution of 640×480, with pixel values ​​representing temperature with an accuracy of 0.1 degrees Celsius. For example, in the temperature field after thermal excitation of a plastic package, the background area temperature is 24.5 degrees Celsius, and the invisible coding area temperature is 27.2 degrees Celsius, creating a temperature difference of 2.7 degrees Celsius.

[0074] The first feature extraction layer, based on a residual cross-layer feature fusion network, analyzes the temperature field distribution image. This layer contains two residual blocks, each with 64 channels and a feature map size of 320×240. This layer focuses on extracting local temperature and edge features, making it suitable for detecting spatial discontinuities in the temperature field. The system applies temporal difference operations to 10 consecutive frames of temperature field images to calculate the temperature change rate of each pixel. The temporal change rate is calculated by dividing the temperature difference between the current frame and the previous frame by the inter-frame time interval (16.7 milliseconds), in degrees Celsius per second. For stealth coding regions with good thermal response characteristics, the typical temperature change rate is 15-25 degrees Celsius per second, significantly higher than the 5-10 degrees Celsius per second of the background region.

[0075] The spatial temperature gradient of the temperature field is calculated using the Sobel operator, obtaining the temperature gradients in the horizontal and vertical directions respectively. The Sobel operator uses a 3×3 kernel, with the horizontal kernel set to [[-1,0,1],[-2,0,2],[-1,0,1]] and the vertical kernel set to [[-1,-2,-1],[0,0,0],[1,2,1]]. The square root of the sum of the squares of the horizontal and vertical gradients represents the total gradient magnitude, reflecting the drastic degree of temperature change. The boundaries of the invisible inkjet printing area typically have high gradient values, ranging from 2 to 4 degrees Celsius per pixel, while the gradient values ​​of the background area are usually less than 0.5 degrees Celsius per pixel. For example, in one detection, the average gradient value of the inkjet printing area boundary was 3.2 degrees Celsius per pixel, while the average gradient value of the background area was 0.3 degrees Celsius per pixel, resulting in a gradient ratio of 10.7, forming a significant contrast.

[0076] The spatiotemporal characteristic parameters of the temperature field distribution image are constructed based on the temporal rate of change of the temperature field and the spatial temperature gradient. These parameters comprise 15 components: 5 representing temporal characteristics and 10 representing spatial characteristics. Temporal characteristics include the average rate of change, maximum rate of change, variance of the rate of change, heating rate, and cooling rate. Spatial characteristics include the average gradient, maximum gradient, gradient variance, horizontal gradient ratio, vertical gradient ratio, gradient direction consistency, gradient magnitude entropy, local temperature difference, global temperature difference, and temperature distribution skewness.

[0077] These parameters comprehensively characterize the dynamic evolution of the temperature field, effectively distinguishing invisible inkjet printing from different materials and printing processes. For a certain plastic packaging sample, the calculated partial spatiotemporal characteristic parameters are: average rate of change 18.5 degrees Celsius / second, maximum rate of change 26.3 degrees Celsius / second, average gradient 2.8 degrees Celsius / pixel, maximum gradient 4.1 degrees Celsius / pixel, and local temperature difference 2.7 degrees Celsius.

[0078] The spatiotemporal feature parameters are progressively processed by multiple cascaded residual blocks within the backbone layer of the residual cross-layer feature fusion network. The backbone layer consists of the second to fifth feature extraction layers, with channel numbers of 128, 256, 384, and 512, respectively. Each residual block enhances the hierarchical correlation of spatiotemporal feature parameters through main branch mapping and cross-layer connections. The main branch mapping adopts a "convolution-batch normalization-activation-convolution-batch normalization" structure, with a 3×3 kernel size, a stride of 1, and padding of 1. Cross-layer connections are divided into three types: short connections, medium connections, and long connections, connecting feature maps of adjacent layers, one layer apart, and multiple layers apart, respectively. Short connections are implemented through identity mapping; medium connections adjust the number of channels through 1×1 convolutions; and long connections achieve scale and channel matching through adaptive average pooling and 1×1 convolutions.

[0079] In the progressive processing, each residual block extracts features at different scales and levels of abstraction. Shallow residual blocks (layers 2 and 3) focus on extracting local texture and edge features, resulting in higher spatial resolution feature maps. Deep residual blocks (layers 4 and 5) focus on extracting semantic features and global context, resulting in lower spatial resolution but more channels. The output of each residual block is fused with features from other layers through cross-layer connections to form a multi-scale feature representation. For example, a residual block in layer 3 outputs a feature map of size 80×60×256. It receives a feature map from layer 2 (160×120×128, downsampled to 80×60×128) and a feature map from layer 4 (40×30×384, upsampled to 80×60×384) through cross-layer connections. The three are then fused to form an enhanced feature representation of size 80×60×256.

[0080] Temperature field features are constructed based on hierarchical correlation, which is represented by a feature correlation matrix. The matrix elements are the cosine similarity between features from different layers. For example, the correlation between features from the second and third layers is 0.82, and the correlation between features from the second and fifth layers is 0.45, indicating high similarity between features from adjacent layers and low similarity between features from distant layers. The system performs weighted fusion of features from each layer based on this correlation matrix, with the weights proportional to the correlation. The fused temperature field features have a dimension of 512, containing multi-scale temperature field information, and can comprehensively characterize the thermal properties of the invisible inkjet coding region.

[0081] Based on the temperature field characteristics, the temperature response characteristics of the stealth coding area were analyzed, and three key parameters were extracted from the temperature response curve: initial temperature, temperature increment, and time constant. The initial temperature is the reference temperature before thermal excitation, the temperature increment is the maximum temperature rise after thermal excitation, and the time constant is the time required for the temperature to rise to 63.2% of the maximum temperature rise. These three parameters are closely related to the thermal properties of the material: the initial temperature reflects the environmental conditions, the temperature increment reflects the material's heat absorption capacity, and the time constant reflects the material's heat diffusion rate. The system obtains the precise values ​​of these three parameters through exponential fitting analysis of the temperature response curve. For a certain stealth coding area, the calculated initial temperature is 24.5 degrees Celsius, the temperature increment is 2.7 degrees Celsius, and the time constant is 85 milliseconds; while the adjacent non-coding area has corresponding parameters of 24.5 degrees Celsius, 1.8 degrees Celsius, and 120 milliseconds, indicating that the coding area heats up faster and more significantly.

[0082] A set of characteristic metrics is constructed, including initial temperature, temperature increment, and time constant. These metrics are obtained by normalizing and combining the original parameters and include relative temperature increment (the ratio of the temperature increment of the coding area to the temperature increment of the background area), time constant ratio (the ratio of the time constant of the coding area to the time constant of the background area), and thermal response index (temperature increment divided by time constant). For the example above, the calculated relative temperature increment is 1.5, the time constant ratio is 0.71, and the thermal response index is 0.032 degrees Celsius / millisecond. These characteristic metrics can effectively distinguish between different types of stealth inkjet printing; for example, the relative temperature increment of infrared absorption inkjet printing is typically greater than 1.4, while the relative temperature increment of infrared reflection inkjet printing is typically less than 0.7.

[0083] A residual cross-layer feature fusion network is used to combine feature indicators and temperature field features to generate a spatiotemporal feature data stream representing the invisible inkjet coding region. The combination process employs a feature channel attention mechanism, dynamically adjusting the weights of each feature channel based on its discriminative power. The system generates channel descriptors through global average pooling and global max pooling, which are then passed through two fully connected layers (dimension changes from 512 to 128 to 512) and a sigmoid activation function to obtain channel weights, ranging from 0 to 1. Channels with high discriminative power receive larger weights (typically greater than 0.8), while channels with low discriminative power receive smaller weights (typically less than 0.2). The weighted features are then integrated through a connection layer to form the final spatiotemporal feature data stream with a dimension of 640 (512-dimensional temperature field features plus 128-dimensional feature indicators).

[0084] The spatiotemporal feature data stream is continuously updated using a sliding window approach. The window size is 10 frames, and the step size is 2 frames. Each update retains 80% of the historical information and incorporates 20% of the new information, ensuring the temporal continuity of the features. The system applies principal component analysis to the feature data stream for dimensionality reduction, retaining the first 256 principal components, covering 98% of the variance, and reducing redundant information. The final generated spatiotemporal feature data stream format is batch × time step × feature dimension, with a typical configuration of 1 × 10 × 256, representing one batch, 10 time steps, and 256-dimensional features per step. This feature data stream fully characterizes the thermal response characteristics and spatiotemporal dynamic changes of the invisible inkjet coding area, providing a solid foundation for subsequent feature extraction and real-time tracking.

[0085] In practical application testing, this method demonstrated stable feature extraction performance across various packaging materials. For three common packaging materials—paper, plastic, and metal—the feature extraction accuracy reached 98.2%, 97.5%, and 96.3%, respectively. The system's processing speed reached 25 frames per second, meeting the real-time inspection requirements of high-speed production lines. Particularly when processing low-contrast invisible inkjet printing, the thermal response-based method improved the detection rate by over 35% compared to traditional optical methods, significantly reducing the missed detection rate. The system has been validated on multiple intelligent packaging production lines, operating continuously for 300 hours without failure, verifying the industrial applicability of this technical solution.

[0086] In one optional implementation, the continuous image sequence is subjected to hierarchical feature analysis using a residual cross-layer feature fusion network, and the surface of the smart packaging object is thermally excited using a pulsed heat source. A temperature field distribution image is then established based on the thermal diffusion response of the smart packaging object surface, including:

[0087] A multi-layer residual network structure is constructed, the feature increment of each layer is calculated through residual mapping, cross-layer feature connection channels are established between adjacent feature layers, feature weights are allocated to each layer, and multi-layer image features of the continuous image sequence are extracted. The multi-layer image features guide the initial configuration of the heat source excitation parameters.

[0088] Based on the multi-level image feature analysis, the material properties of the smart packaging object are analyzed. The initial heat source power in the coarse adjustment stage is determined according to the material properties. A layered progressive heat source control strategy is adopted to apply pulsed thermal excitation to the surface of the smart packaging object. The initial heat source power, dynamic compensation amount and adaptive weighting factor together determine the real-time heat source power in the fine adjustment stage.

[0089] The thermal diffusion response generated on the surface of the smart packaging object under the pulsed thermal excitation is detected, and a temperature field distribution image is generated based on the thermal diffusion response combined with the heat source influence coefficient, thermal diffusion coefficient and thermal conductivity.

[0090] like Figure 2 As shown, the method includes:

[0091] A multi-layer residual network structure is constructed, employing a deep residual learning framework, comprising a backbone network and cross-layer connections. The backbone network consists of 5 residual blocks, each containing 2 convolutional layers, 2 batch normalization layers, and 1 ReLU activation function layer. The convolutional kernel size is 3×3. The first residual block has 64 channels, increasing to 128, 256, 512, and 1024 channels with increasing network depth. The first convolutional layer in each residual block has a stride of 2 for downsampling, while the second convolutional layer has a stride of 1 to maintain the feature map size. This design allows the network to extract features at different scales of the image layer by layer, forming a multi-layer feature representation. For example, for an image sequence with an input size of 640×480, after passing through 5 residual blocks, the resulting feature map sizes are 320×240×64, 160×120×128, 80×60×256, 40×30×512 and 20×15×1024, respectively.

[0092] The feature increment of each layer is calculated through residual mapping. The core idea of ​​residual mapping is to learn the difference between the input and output, rather than directly learning the output. Specifically, the calculation formula for each residual block is: Output Feature = Input Feature + Residual Function (Input Feature). The residual function consists of two convolutional layers, which perform feature transformation and feature enhancement operations respectively. To adapt to changes in the number of channels, the system adds a 1×1 convolutional layer to the residual connection path for dimensionality reduction or enhancement.

[0093] For example, the input feature dimension of the third residual block is 160×120×128, and the residual feature dimension calculated by the residual function is 80×60×256. The system adjusts the input features to 80×60×256 using a 1×1 convolution, and then adds them to the residual features to obtain the final output. In one processing iteration, the mean of the input features of the third residual block was 0.42, the mean of the residual features was 0.15, and the mean of the output features after addition was 0.57, demonstrating the cumulative effect of residual learning.

[0094] Cross-layer feature connections are established between adjacent feature layers. These connections employ a feature pyramid structure, including a top-down upsampling path and a bottom-up feature fusion path. The upsampling path upsamples high-level semantic features to the spatial dimensions of low-level features using transposed convolutions. The feature fusion path adjusts the number of channels using 1×1 convolutions before weighted feature fusion. The system implements three cross-layer connection methods: short connections (between adjacent layers), medium connections (with one layer between layers), and long connections (between the first and last layers). Taking short connections as an example, the system upsamples the fourth-layer feature (40×30×512) to 80×60×512, then adjusts it to 80×60×256 using a 1×1 convolution, and finally performs weighted fusion with the third-layer feature (80×60×256). The fusion weights are dynamically calculated using an attention mechanism, with an initial value of [0.5, 0.5], which is dynamically adjusted as the feature quality changes. For example, when processing a certain metal package, due to its strong surface reflectivity and better quality of high-layer features, the fusion weight is adjusted to [0.65, 0.35], which enhances the contribution of high-layer features.

[0095] Multi-level image features of continuous image sequences are extracted by assigning feature weights to each layer. The weight assignment adopts a feature importance evaluation mechanism, including three indicators: feature discriminability, feature stability, and feature response intensity. Feature discriminability is obtained by calculating the ratio of intra-class distance to inter-class distance; feature stability is obtained by calculating the coefficient of variation of features in consecutive frames; and feature response intensity is obtained by calculating the average activation value of the feature map.

[0096] The importance weights of each layer of features are determined by the weighted sum of these three indicators, with a weight sum of 1. When processing a paper packaging sample, the feature weights for the five layers are 0.15, 0.20, 0.30, 0.25, and 0.10, respectively, reflecting the relative importance of the middle layer features. The extracted multi-layered image features are used to guide the initial configuration of the heat source excitation parameters, particularly the pulse width and energy density. Taking a certain detection as an example, the features extracted by the system indicate that the packaging surface is relatively smooth and has moderate reflectivity. Based on this, the system sets the initial pulse width to 15 milliseconds and the energy density to 0.8 joules / square centimeter.

[0097] The analysis of material properties of smart packaging objects is based on extracted multi-level image features. The analysis process includes two steps: material classification and parameter estimation. Material classification employs a Support Vector Machine (SVM) algorithm. The input is the extracted feature vector, and the output is the material category (paper, plastic, metal, etc.). The SVM uses a radial basis function kernel with parameter C set to 10 and gamma set to 0.01. The system pre-trained a material classification model containing 2000 samples, achieving an accuracy of 96.5%. Parameter estimation uses regression analysis to predict the thermal parameters of the material, including thermal conductivity, specific heat capacity, and density. The regression model uses a random forest algorithm with 100 decision trees and a maximum depth of 10. For example, for a plastic packaging sample, the system predicted a thermal conductivity of 0.22 W / m·K, a specific heat capacity of 1850 J / kg·K, and a density of 1.05 g / cm³. These parameters deviate from the standard values ​​within ±8%.

[0098] The initial heat source power for the coarse-tuning stage is determined based on the material properties obtained from the analysis. The calculation of the initial power considers three factors: the material's thermal conductivity, the target temperature rise, and the safety threshold. For materials with low thermal conductivity (such as paper packaging), the initial power is set lower (3-5 watts) to avoid overheating damage; for materials with high thermal conductivity (such as metal packaging), the initial power is set higher (8-12 watts) to ensure sufficient temperature rise. The system establishes a mapping table between material thermal conductivity and initial power, including parameters for 20 common packaging materials. For example, plastic packaging with a thermal conductivity of 0.22 W / m·K corresponds to an initial power of 4.5 watts; aluminum packaging with a thermal conductivity of 15 W / m·K corresponds to an initial power of 10 watts. This initial power setting based on material properties can adapt to the thermal response characteristics of different packaged objects.

[0099] A layered, progressive heat source control strategy is employed to apply pulsed thermal excitation to the surface of the smart packaging object. This strategy consists of two stages: a coarse-tuning stage and a fine-tuning stage. In the coarse-tuning stage, a fixed-power heat source pulse is used, with a pulse width of 15-20 milliseconds, a repetition frequency of 10 Hz, and a duration of 100 milliseconds. After the coarse-tuning stage, the system acquires preliminary thermal response images of the packaging surface and calculates the average temperature rise and temperature distribution uniformity. In the fine-tuning stage, the heat source power is dynamically adjusted based on the preliminary thermal response results.

[0100] The formula for calculating the real-time heat source power is: Real-time power = Initial power + Dynamic compensation amount × Adaptive weighting factor. The dynamic compensation amount is calculated based on the difference between the target temperature rise and the actual temperature rise, with a range of ±2 watts; the adaptive weighting factor is adjusted according to the material's thermal response speed, with a smaller weighting factor (0.3-0.5) for materials with fast thermal response and a larger weighting factor (0.6-0.8) for materials with slow thermal response.

[0101] The initial heat source power, dynamic compensation, and adaptive weighting factor jointly determine the real-time heat source power during the fine-tuning phase. Taking a plastic packaging inspection as an example, the initial power is 4.5 watts, the target temperature rise is 5 K, the actual temperature rise is 4.2 K, the calculated dynamic compensation is 0.8 watts, the adaptive weighting factor is 0.6, and the final real-time power is 4.5 + 0.8 × 0.6 = 4.98 watts. The system updates the real-time power every 10 milliseconds, forming a closed-loop control to ensure the stability and safety of the thermal excitation effect.

[0102] The fine-tuning phase lasts for 200 milliseconds, employing a variable-frequency pulse sequence. The pulse width is dynamically adjusted between 15 and 25 milliseconds, and the pulse interval varies between 50 and 100 milliseconds, enabling differentiated excitation for defects of different depths. For materials with high thermal conductivity, the system uses a high-frequency short-pulse strategy (20 Hz, 15 milliseconds); for materials with low thermal conductivity, the system uses a low-frequency long-pulse strategy (10 Hz, 25 milliseconds).

[0103] The thermal diffusion response of the surface of smart packaging objects under pulsed thermal excitation was detected using a high-sensitivity infrared thermal imager with a resolution of 640×480, a temperature resolution of 0.05 K, and a frame rate of 60 frames per second. The thermal imager was synchronized with the heat source, acquiring thermal images immediately after each pulse to capture the optimal moment of thermal response.

[0104] Three types of thermal images are acquired: a pre-pulse reference image, a post-pulse peak image, and a sequence of images showing the cooling process. The peak image reflects the initial thermal response of the material surface, while the cooling image sequence reflects the thermal diffusion process. The combination of these two images provides comprehensive thermal characteristic information. For specific infrared bands (8-14 micrometers), the system considers the emissivity differences of materials and applies emissivity correction coefficients to different materials to ensure the accuracy of temperature measurements. For example, the emissivity of paper packaging is 0.92, with a correction coefficient of 1.08; the emissivity of metal packaging is 0.25, with a correction coefficient of 4.0. The corrected temperature accuracy is improved to ±0.5 Kelvin, meeting the accuracy requirements for defect detection.

[0105] A temperature field distribution image is generated based on the thermal diffusion response combined with the heat source influence coefficient, thermal diffusivity, and thermal conductivity. The generation of the temperature field distribution image includes two stages: thermal image preprocessing and temperature field reconstruction. Thermal image preprocessing comprises three steps: noise filtering, spatial registration, and temperature calibration. Noise filtering uses an adaptive median filter with a 5×5 window size; spatial registration employs a phase correlation algorithm to ensure accurate alignment between consecutive frames; temperature calibration uses a blackbody radiation source as a reference to establish a mapping relationship between infrared signals and temperature. Temperature field reconstruction is based on a heat conduction model, considering three key parameters: the heat source influence coefficient (reflecting the spatial distribution characteristics of the heat source), the thermal diffusivity (reflecting the rate of heat diffusion in the material), and thermal conductivity (reflecting the material's thermal conductivity).

[0106] The heat source influence coefficient, obtained through heat source calibration experiments, is represented as an 8×8 matrix, describing the spatial distribution characteristics of the heat source energy. The matrix has the largest value (1.0) at the center, gradually decreasing outwards, with a minimum value of approximately 0.2 at the edges, reflecting the Gaussian distribution characteristics of the heat source. The thermal diffusivity and thermal conductivity are provided by the material property analysis module and are used as parameters to construct the heat conduction equation.

[0107] The heat conduction equation was solved using the finite difference method with a time step of 10 milliseconds and a spatial step of 0.5 millimeters to achieve a numerical simulation of the temperature field. The simulation results were compared with actual measured temperatures, and the model parameters were iteratively optimized and adjusted to ensure that the root mean square error between the simulated and actual temperature fields was less than 0.3 K.

[0108] The final generated temperature field distribution image has the same resolution as the original image (640×480), and uses pseudo-color mapping to enhance the visual effect, mapping the temperature range to grayscale values ​​of 0-255. The system provides three temperature field representations: absolute temperature field (displaying actual temperature values), relative temperature field (displaying temperature rise values), and differential temperature field (displaying temperature differences between adjacent areas). Among them, the differential temperature field is best suited for the identification of invisible inkjet printing areas because the difference in thermal properties between the inkjet printing area and the surrounding materials will form a clear boundary in the differential temperature field.

[0109] For example, for invisible inkjet printing on a plastic package, the average temperature of the inkjet printing area in the differential temperature field is 1.8 Kelvin higher than that of the surrounding area, forming a clear thermal contrast, which facilitates subsequent area extraction and feature analysis.

[0110] In practical application testing, this method demonstrated stable temperature field generation accuracy across various packaging materials. For paper, plastic, and metal, the temperature measurement errors were ±0.3 K, ±0.5 K, and ±0.8 K, respectively. Particularly when dealing with complex surfaces such as reflective and translucent surfaces, the heat diffusion-based method showed significant advantages over traditional optical methods, effectively identifying invisible inkjet printing. The system processing speed reached 15 frames per second, meeting the real-time inspection requirements of medium-speed production lines. During a continuous 200-hour stability test, the temperature field generation accuracy remained above 98%, proving the practical value of this technology in industrial environments.

[0111] In one optional implementation, different feature extraction layers are connected in series via a cross-scale residual connection channel. The parameters of each feature extraction layer are adaptively adjusted according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region, including:

[0112] A dual-constraint architecture of teacher network and student network is constructed. Feature connection channels are established through feature transfer between the teacher network and the student network, and feature transfer paths are formed based on the feature connection channels.

[0113] Feature extraction is performed based on the feature transfer path. The high-dimensional feature representation extracted by the teacher network is used as knowledge guidance. The feature space is compressed and mapped based on the student network under the knowledge guidance. The feature connection channel is dynamically adjusted according to the feature reconstruction result of the compressed mapping.

[0114] Based on the feature reconstruction results of the compression mapping, the spatiotemporal data stream features are analyzed, and the spatiotemporal data stream features are input into the teacher network for feature extraction. The features extracted by the teacher network are used to adjust the feature extraction process of the student network. The feature extraction process of the student network is used for feature consistency constraints.

[0115] Based on the feature consistency constraint, the feature reconstruction result of the compressed mapping is spatially aligned, the feature extraction parameters of the teacher network are updated based on the spatially aligned features, and the optimized features are output according to the teacher network.

[0116] The spatially aligned features are fused with the teacher network and the student network under the guidance of the optimized features, and a multi-scale feature vector of the invisible inkjet coding region is generated based on the result of the feature fusion.

[0117] A dual-constraint architecture is constructed for the teacher and student networks. The teacher network employs a deep structure with 8 convolutional layers, each containing 64, 128, 256, 512, 256, 128, 64, and 32 kernels. The kernel size is 3×3, with a stride of 1 and padding of 1. The parameterized ReLU activation function is used, with an initial parameter value of 0.2. The student network uses a lightweight structure with 5 convolutional layers, each containing 32, 64, 128, 64, and 32 kernels. The kernel size is 3×3, with a stride of 1 and padding of 1. Feature connection channels are established between the two networks, including a forward connection channel and a backward connection channel. The forward connection channel transmits feature maps from layers 3, 5, and 7 of the teacher network to layers 2, 3, and 4 of the student network; the backward connection channel transmits feature maps from layers 2, 3, and 4 of the student network to layers 3, 5, and 7 of the teacher network.

[0118] This bidirectional connection forms a feature transfer path, enabling knowledge exchange between the two networks. For example, for an input image of size 320×240, the feature map output by the third layer of the teacher network is 80×60×256, which is passed to the second layer of the student network through the forward connection channel and fused with the feature map (size 80×60×64) generated by the student network itself.

[0119] Feature extraction is performed based on the established feature transfer path, with the teacher network responsible for extracting high-dimensional feature representations as knowledge guidance. During processing, the input image is first standardized to adjust the pixel value range to [0,1]. Then, through convolution and pooling operations at each layer of the teacher network, multi-level feature representations are generated. Taking a specific detection as an example, for an input image of a hidden code region (size 320×240), the feature map generated by the teacher network's third layer has a dimension of 80×60×256, the fifth layer 40×30×256, and the seventh layer 80×60×64. These high-dimensional feature representations are passed to the student network through the forward connection channel, serving as guidance information for the student network's feature extraction.

[0120] Guided by knowledge, the student network performs compressed mapping on the feature space. The student network employs an encoder-decoder structure. The encoder contains three downsampling modules, each including a convolutional layer, a batch normalization layer, and a max-pooling layer. The decoder contains two upsampling modules, each including a transposed convolutional layer and a batch normalization layer. The student network receives high-dimensional features from the teacher network and combines them with its own extracted features to generate a low-dimensional feature representation through compressed mapping. The compressed mapping uses a non-linear projection method implemented by fully connected layers. The input dimension is the number of channels in the feature map, and the output dimension is one-quarter of that. For example, a 256-channel feature map in layer 3 is compressed to obtain a 64-channel feature representation.

[0121] After compression mapping, the feature reconstruction error, i.e., the mean squared error between the original features and the reconstructed features, is calculated to evaluate the compression quality. In a certain processing, the reconstruction error of the features in layer 3 was 0.053, that of layer 5 was 0.068, and that of layer 7 was 0.047. The system dynamically adjusts the weights of the feature connection channels based on the reconstruction error; the smaller the reconstruction error, the larger the corresponding channel weight. The channel weights are normalized using the softmax function, with an initial value of 0.333, which is adjusted to [0.38, 0.25, 0.37].

[0122] Based on the feature reconstruction results of the compressed mapping, the spatiotemporal data stream features are analyzed, and spatiotemporal statistical indicators of the feature map are calculated, including spatial distribution entropy, temporal rate of change, and feature response intensity. Spatial distribution entropy is obtained by calculating the normalized distribution of response values ​​in each region of the feature map, reflecting the spatial complexity of the feature; temporal rate of change is obtained by calculating the differences between consecutive frame feature maps, reflecting the temporal stability of the feature; and feature response intensity is obtained by calculating the average activation value of the feature map, reflecting the saliency of the feature. For a certain detection sequence, the system calculates the spatial distribution entropy of the 3rd layer features to be 4.85, the temporal rate of change to be 0.15, and the feature response intensity to be 0.62; the corresponding values ​​for the 5th layer features are 4.32, 0.18, and 0.57; and the corresponding values ​​for the 7th layer features are 5.12, 0.12, and 0.65. These statistical indicators are combined to form the spatiotemporal data stream features for subsequent processing.

[0123] The spatiotemporal data stream features are input into the teacher network for feature extraction. The teacher network employs an attention enhancement mechanism, including a channel attention module and a spatial attention module. The channel attention module generates channel descriptors through global average pooling and global max pooling, and obtains channel weights after processing by a shared multilayer perceptron. The spatial attention module generates spatial feature maps through average pooling and max pooling along the channel dimension, and obtains spatial weights after convolution. The output weights of these two modules are multiplied by the original feature maps to achieve feature enhancement. For example, for the feature map of layer 5, the channel attention weights range from [0.42, 1.53] with an average of 0.97; the spatial attention weights range from [0.38, 1.65] with an average of 0.94. The enhanced features are used to adjust the feature extraction process of the student network.

[0124] The feature extraction process of the student network is used for feature consistency constraints. A consistency loss function is designed, consisting of two parts: feature-level consistency and output-level consistency. Feature-level consistency is achieved by calculating the cosine similarity between the student network features and the corresponding layer features of the teacher network, aiming to make the feature directions of the two networks consistent. Output-level consistency is achieved by calculating the KL divergence of the final outputs of the two networks, aiming to make their output distributions as close as possible. The weight coefficient of the consistency constraint is set to 0.3, gradually increasing to 0.6 during training, prompting the student network to gradually learn the knowledge of the teacher network. In one iteration, the feature-level consistency value was 0.82, the output-level consistency value was 0.15, and the overall consistency constraint value was 0.36.

[0125] Spatial alignment of the feature reconstruction results from the compressed mapping is performed based on feature consistency constraints. Spatial alignment is achieved using a deformable convolutional network, comprising three deformable convolutional layers, with the offset field generated by a 3×3 convolution. The deformable convolution learns the spatial offset field to achieve adaptive sampling of feature points, effectively handling geometric deformations in the inkjet coding region. The alignment process consists of two steps: coarse alignment and fine alignment. Coarse alignment uses a global affine transformation to estimate global transformation parameters; fine alignment uses a local deformation field to handle local non-rigid deformations. For a specific deformed inkjet coding sample, the feature registration error after coarse alignment is 3.85 pixels, which is reduced to 0.76 pixels after fine alignment, significantly improving the spatial consistency of the features.

[0126] The feature extraction parameters of the teacher network are updated based on spatial alignment, using an exponential moving average strategy. The update formula is: New parameter = Momentum coefficient × Old parameter + (1 - Momentum coefficient) × Student network parameter. The initial value of the momentum coefficient is set to 0.996, gradually increasing to 0.999 with each iteration to achieve smooth updates. For example, the weight parameters of the third convolutional layer of the teacher network, which were [0.125, -0.087, 0.213, ...] before an update, become [0.123, -0.085, 0.210, ...] after the update. Through this soft update mechanism, the teacher network can stably absorb the knowledge learned by the student network while maintaining its own feature extraction capabilities. After the update, the teacher network outputs optimized features. The average response strength of the features in the third layer before optimization was 0.62, which improved to 0.68 after optimization; the feature discriminancy (the ratio of intra-class distance to inter-class distance) improved from 1.82 to 2.15, indicating a significant improvement in feature quality.

[0127] Spatially aligned features are fused through teacher and student networks under the guidance of optimized features. The fusion employs an adaptive weighting strategy, dynamically adjusting the fusion weights based on the feature's information entropy and response intensity. Features with high information entropy and strong response intensity receive larger weights, while features with low information entropy and weak response intensity receive smaller weights. The fusion formula is: Fusion Feature = Adaptive Weight 1 × Teacher Network Feature + Adaptive Weight 2 × Student Network Feature, where the sum of Weight 1 and Weight 2 is 1. For the 5th layer features, the weight of the teacher network feature in a certain fusion is 0.65, and the weight of the student network feature is 0.35; while the corresponding weights for the 3rd layer are 0.58 and 0.42, reflecting the differences in importance between features at different layers.

[0128] Based on the feature fusion results, a multi-scale feature vector of the invisible coding region is generated. The fused features from each layer are integrated through a feature aggregation module, which includes a multi-scale feature pyramid and a feature selection mechanism. The feature pyramid generates feature representations at different scales through multi-level pooling, including 1×1, 2×2, and 4×4 levels. The feature selection mechanism selects the most discriminative feature channels through channel attention. Finally, the system compresses the aggregated features into a fixed-dimensional feature vector with a dimension of 128 through a fully connected layer. For a given coding sample, the first 10 elements of the generated multi-scale feature vector are [0.83, -0.25, 0.67, 0.12, -0.58, 0.46, 0.31, -0.14, 0.72, -0.39], with feature values ​​ranging from [-1, 1] and a statistical distribution close to a standard normal distribution.

[0129] In practical application tests, this method demonstrated stable performance on different types of packaged objects, achieving feature extraction accuracies of 98.7%, 97.9%, and 96.8% for paper, plastic, and metal packaging, respectively. In a high-speed production environment (conveyor belt speed 6 m / s), the average processing time was 18 milliseconds per frame, meeting real-time requirements. Particularly when processing invisible inkjet printing on complex surfaces such as reflective and semi-transparent surfaces, the method's multi-scale features exhibited strong robustness, reducing the error rate by more than 45% compared to traditional methods. After 300 hours of continuous operation testing, the system showed good stability with no feature extraction failures, proving the practical value of this technical solution in industrial environments.

[0130] In one optional implementation, the dynamic tracking trajectory equation is calculated based on the multi-scale feature vector, and the coefficients of the dynamic tracking trajectory equation are optimized through iterative calculation to achieve real-time tracking of the invisible inkjet coding area, including:

[0131] A variational energy functional is constructed based on the multi-scale feature vectors, and the optimal form of the dynamic tracking trajectory equation is derived based on the variational energy functional. The optimal form of the dynamic tracking trajectory equation includes the coefficient matrix to be optimized.

[0132] The optimal form of the dynamic tracking trajectory equation is decomposed into parameter optimization terms and constraint optimization terms. Based on the parameter optimization terms, an objective function is constructed for the coefficient matrix to be optimized. Based on the constraint optimization terms, weight coefficients are introduced to establish constraints. Based on the objective function and the constraints, an optimization objective is constructed.

[0133] The optimization objective is solved by a dynamic iterative method. The coefficient matrix is ​​updated by the parameter optimization terms. The updated coefficient matrix is ​​substituted into the constraint optimization terms to calculate the weight coefficients. The coefficient matrix is ​​then optimized based on the gradient information of the weight coefficients.

[0134] By substituting the optimized coefficient matrix into the optimal form of the dynamic tracking trajectory equation, real-time tracking of the invisible inkjet printing area can be achieved.

[0135] After obtaining multi-scale feature vectors, a variational energy functional is constructed, which consists of three main parts: a spatial gradient energy term representing the spatial continuity of the trajectory, a temporal derivative energy term representing the temporal smoothness of the trajectory, and a fitting energy term representing the feature matching accuracy. The spatial gradient energy term adopts a weighted second-order difference form to calculate the rate of change of the feature vector in the spatial domain. The system samples and calculates the spatial gradient within a 9×9×5 three-dimensional window, with the center of the window corresponding to the current processing point. The weights of the sampling points are set according to a truncated Gaussian distribution, with a weight of 0.85 in the core region, gradually decreasing to 0.15 towards the outside. For a 5-dimensional feature vector [0.82, 0.75, 0.90, 0.85, 0.78] at a specific location (142, 118, 45), the calculated spatial gradient energy value is 0.058, which reflects the intensity of the feature change in space.

[0136] The time derivative energy term employs a weighted hybrid difference scheme, combining the advantages of forward differencing and central differencing to improve numerical stability. The system stores the feature sequences of the most recent 8 frames, forming a time sliding window. An exponentially decaying weight is used within the window, with the most recent frame having a weight of 0.35, decreasing by a factor of 0.8. For a consecutive 8-frame feature sequence, the system calculates the time derivative of each feature component and obtains the weighted sum of squares as the time derivative energy. For example, for the first feature component sequence of 8 consecutive frames at a certain position [0.82, 0.84, 0.85, 0.87, 0.89, 0.90, 0.92, 0.93], the calculated time derivative energy is 0.042, reflecting the degree of drastic change of the feature over time.

[0137] The fitting energy term employs the Huber loss function, combining the accuracy of quadratic functions in small error regions with the robustness of absolute value functions in large error regions. The system pre-trains a standard feature template library using deep learning methods, containing standard features under different poses, lighting conditions, and occlusion conditions. In real-time processing, the system selects the K closest templates (K=5) from the library and calculates the weighted distance between the current feature and these templates. The weights of each template are determined through similarity calculations, with templates exhibiting higher similarity receiving larger weights. For the current feature vector [0.82, 0.75, 0.90, 0.85, 0.78], the fitting energy calculated with the 5 closest templates is 0.035, reflecting the accuracy of feature matching.

[0138] The variational energy functional expresses the total energy by linearly combining these three energy terms. The combination weights are set according to the characteristics of the application scenario. For high-speed motion scenarios, the weight of the spatial gradient energy term is 0.30, the weight of the time derivative energy term is 0.35, and the weight of the fitting energy term is 0.35. These weights are optimized through cross-validation on 300 sets of sample data, maintaining good performance under different application environments. For the example data above, the calculated total variational energy is 0.045, which serves as the basis for the objective function in subsequent optimizations.

[0139] The optimal form of the dynamic tracking trajectory equation is derived based on the constructed variational energy functional. The derivation process is based on the principle of variational method, calculating the variational derivative of the functional with respect to the trajectory function to obtain the Euler-Lagrange equation. To improve computational efficiency, the system uses the finite difference method to discretize the Euler-Lagrange equation in the continuous domain, transforming it into a system of linear equations in the discrete domain. The discretization uses a non-uniform grid, with a dense grid in the core region (0.5 pixel step size) and a sparse grid in the edge region (1.5 pixel step size), balancing computational accuracy and efficiency. The discretized equation contains a coefficient matrix to be optimized, with dimensions of 5×15, where 5 corresponds to the feature dimension and 15 corresponds to the coefficients of different derivative terms. The initial coefficient matrix is ​​estimated based on historical data using the least squares method, for example: [[1.20,0.85,0.65,0.45,0.30,0.25,0.20,0.15,0.12,0.10,0.08,0.06,0.05,0.04,0.03],[1.18,0.82,0.62,0.42,0.28,0.27,0.22,0.17,0.13,0.11,0.09,0.07,0.06,0.05,0.04],[1.25,0.88,0.68,0.48,0.33, 0.23,0.18,0.13,0.10,0.08,0.06,0.04,0.03,0.02,0.01],[1.22,0.86,0.66,0.46,0.31,0.26,0.21,0.16,0.13,0.11,0.09,0.07,0.06,0.05,0.04],[1.16,0.80,0.60,0.40,0.26,0.28,0.23,0.18,0.14,0.12,0.10,0.08,0.07,0.06,0.05]).

[0140] After obtaining the optimal form of the dynamic tracking trajectory equation, it is decomposed into parameter optimization terms and constraint optimization terms. The parameter optimization terms correspond to the data fitting part of the equation, representing the degree of matching between the predicted trajectory and the observed data; the constraint optimization terms correspond to the regularization part of the equation, representing the smoothness and continuity constraints of the trajectory. The system constructs an objective function based on the parameter optimization terms, which adopts the form of a weighted sum of squared residuals. The residual is defined as the difference between the predicted trajectory and the observed trajectory, and the weights reflect the reliability of different observation points. The system determines the weight of observation points through feature matching confidence; points with high confidence receive large weights (maximum 0.9), and points with low confidence receive small weights (minimum 0.2). For a certain tracking iteration, the system selects data from the past 20 frames to construct the objective function, containing 100 sampling points (5 keypoints per frame), forming a 100×15 observation matrix and a 100×1 observation vector.

[0141] The system introduces weighted coefficients to establish constraints based on the constraint optimization term. These constraints fall into three categories: trajectory smoothness constraints, coefficient sparsity constraints, and coefficient boundary constraints. The trajectory smoothness constraint is implemented using Tikhonov regularization to suppress high-frequency oscillations in the trajectory; the coefficient sparsity constraint is implemented using L1 norm regularization to make unimportant coefficients approach zero; and the coefficient boundary constraint is implemented using a barrier function to ensure that the coefficients are within the effective range ([0.01, 2.0]). These three types of constraints are introduced with weighted coefficients λ1, λ2, and λ3, with initial values ​​of 0.25, 0.15, and 0.10, respectively. The system combines the objective function with the constraints to construct the optimization objective, forming a constrained least squares problem.

[0142] A dynamic iterative method is employed to solve the optimization objective. This method combines the advantages of the alternating direction multiplier method and the quasi-Newton method, achieving rapid convergence by alternately updating the coefficient matrix and weight coefficients. The iterative process consists of two sub-steps: a parameter update step and a constraint update step. In the parameter update step, the system fixes the weight coefficients and updates the coefficient matrix using the quasi-Newton method (a variant of L-BFGS). The system calculates the gradient of the objective function with respect to the coefficient matrix and applies line search to determine the optimal step size. The step size is selected within the range [0.01, 0.5], with an initial value of 0.1. For the example data above, the gradient of the first iteration is a 5×15 matrix with a maximum element value of 0.35 and a minimum element value of 0.02; the optimization step size is 0.08. The system updates the coefficient matrix according to the gradient and step size; for example, the first element of the first row is updated from 1.20 to 1.172.

[0143] In the constraint update step, the weight coefficients are updated by fixing the coefficient matrix, calculating the current value of the constraint function, and adjusting the weight coefficients according to a preset update rule. If the degree of constraint violation increases, the corresponding weight coefficient increases (multiplied by 1.2); if the degree of constraint violation decreases, the corresponding weight coefficient decreases (multiplied by 0.9). For example, after the first iteration, λ1 is updated from 0.25 to 0.225, λ2 from 0.15 to 0.18, and λ3 remains unchanged from 0.10. The system calculates the constraint gradient based on the updated weight coefficients and uses this gradient information for the next round of parameter updates.

[0144] The maximum number of iterations is set to 50, and the convergence condition is that the norm of the coefficient matrix change is less than 0.001 or the objective function change is less than 0.0001 between two consecutive iterations. For the example above, the system converges after 28 iterations, and the final coefficient matrix is ​​[[1.15,0.78,0.58,0.38,0.24,0.28,0.22,0.16,0.12,0.09,0.07,0.05,0.03,0.02,0.01],[1.12,0.76,0.56,0.36,0.22,0.30,0.24,0.18,0.14,0.11,0.09,0.07,0.05,0.03,0.02],[1.19,0.82,0.62,0.42,0.27 ...2,0.76,0.56,0.36,0.22,0.30,0.24,0.18,0.14,0.11,0.09,0.07,0.05,0.03,0.02],[1.19,0.82,0.62,0.42,0.27],[1.19,0.82,0.62,0.42,0.27],[1.19,0.82,0.62,0.42,0.42,0.62,0.62,0 The optimized matrix has a smaller overall coefficient value and a more concentrated distribution compared to the initial coefficient matrix, reflecting the fine adjustment of the coefficients during the optimization process.

[0145] The optimized coefficient matrix is ​​substituted into the optimal form of the dynamic tracking trajectory equation to achieve real-time tracking of the invisible inkjet coding area. The tracking process adopts a prediction-update framework: in the prediction phase, the system predicts the inkjet coding area position at the next moment based on the trajectory equation and the current state; in the update phase, the system corrects the predicted position by combining the image feature matching results. To improve real-time performance, the system adopts a sliding window strategy, processing only the data of the most recent K frames (K=8) at a time, with the window sliding as new frames arrive. The system sets a search area around the predicted position, with a size of 60×60×10 pixels, and performs feature matching within this area to find the best matching position.

[0146] In a tracking test, the initial inkjet area position was (142, 118, 45), the predicted position for the next frame was (144.5, 120.3, 45.8), and the optimal position for feature matching was (145.2, 119.8, 46.0). The system fused these two results to obtain the final position (145.0, 120.0, 45.9). As tracking progressed, the system continuously updated the trajectory status, achieving continuous tracking of the inkjet area. In a test spanning 100 consecutive frames, the average tracking error was 0.32 pixels, and the maximum error was 0.85 pixels, meeting the requirements for high-precision tracking.

[0147] This real-time tracking method has been validated on an industrial production line, achieving a processing speed of 45 frames per second (using a standard industrial computer with a processor clock speed of 3.2 GHz and 8 GB of memory), meeting the real-time requirements of high-speed production environments. For various complex scenarios, including high-speed motion (object speed reaching 6.5 m / s), partial occlusion (occlusion area not exceeding 35%), and illumination changes (brightness variation ±30%), this method demonstrates good stability and accuracy, with tracking success rates of 98.2%, 96.5%, and 97.8%, respectively, significantly outperforming traditional methods (92.3%, 88.7%, and 91.5%, respectively). Particularly for reflective metal packaging and transparent plastic packaging, the method shows a significant improvement in tracking accuracy, reducing the average error by more than 45%.

[0148] In one optional implementation, a variational energy functional is constructed based on the multi-scale feature vectors, and the optimal form of the dynamic tracking trajectory equation is derived based on the variational energy functional, including:

[0149] The variational energy functional characterizes the spatial distribution characteristics of the trajectory through the spatial gradient term, describes the temporal evolution of the trajectory through the time derivative term, and measures the degree of matching between the trajectory and the multi-scale feature vector through the feature fitting term. The spatial gradient term, the time derivative term, and the feature fitting term are combined to form a unified variational energy expression.

[0150] Based on the variational energy functional, spatial gradient information is obtained by performing spatial partial derivative operation on the spatial gradient term in the variational energy expression; temporal evolution information is extracted by performing temporal partial derivative operation on the temporal derivative term; feature matching information is analyzed by performing feature partial derivative operation on the feature fitting term; and the spatial gradient information, the temporal evolution information and the feature matching information are combined to construct a dynamic tracking trajectory equation.

[0151] Based on the dynamic tracking trajectory equation, the spatial gradient information and the temporal evolution information are mapped into trajectory state parameters. The trajectory state parameters are combined with the feature matching information to optimize the dynamic tracking trajectory equation, thereby obtaining the optimal form of the dynamic tracking trajectory equation.

[0152] like Figure 3 As shown, the method includes:

[0153] After obtaining multi-scale feature vectors, a variational energy functional is constructed to characterize the dynamic characteristics of the invisible inkjet coding region. The variational energy functional consists of three core components: a spatial gradient term, a time derivative term, and a feature fitting term. The spatial gradient term is used to characterize the spatial distribution characteristics of the trajectory and is implemented using a second-order difference form.

[0154] In practice, a local coordinate system is established in the coding area, and the spatial gradients of the feature vectors are calculated along the x, y, and z directions. The system selects a 5×5×3 three-dimensional window to calculate the gradients, with the weights of each point within the window set according to a Gaussian distribution and a standard deviation of 1.2. For example, for the multi-scale feature vector [0.85, 0.76, 0.92, 0.88, 0.79] at the center point coordinates (150, 120, 40), its spatial gradient in the x direction is [0.12, 0.09, 0.15, 0.11, 0.08], its spatial gradient in the y direction is [0.08, 0.11, 0.07, 0.13, 0.10], and its spatial gradient in the z direction is [0.05, 0.04, 0.06, 0.05, 0.03]. The spatial gradient term is obtained by calculating the sum of the squares of these gradients, reflecting the intensity of the feature's spatial variation.

[0155] The time derivative term is used to characterize the temporal evolution of the trajectory. It is implemented using a forward difference approximation, storing the feature vector sequence of the most recent 5 frames and calculating the feature change rate between adjacent frames. To improve the stability of the time derivative calculation, the system adopts a weighted average strategy, with the weight coefficient of the most recent time being 0.4, decreasing forward, and the weight sequence being [0.4, 0.3, 0.2, 0.07, 0.03]. For a specific position, the feature vectors of the 5 consecutive frames are [0.85, 0.76, 0.92, 0.88, 0.79], [0.88, 0.78, 0.94, 0.90, 0.81], [0.91, 0.80, 0.95, 0.91, 0.83], [0.93, 0.82, 0.97, 0.92, 0.85], and [0.95, 0.83, 0.98, 0.94, 0.86]. The calculated time derivative is [0.025, 0.018, 0.015, 0.015, 0.018], which reflects the trend of the feature over time.

[0156] The feature fitting term measures the degree of matching between the trajectory and the multi-scale feature vector, and is implemented using a weighted quadratic distance function. The system pre-trains offline to obtain a standard feature template, which is statistically generated from 500 samples and contains feature components at five different scales, corresponding to different resolution levels of the image. In real-time processing, the system calculates the Euclidean distance between the current feature vector and the standard template, and introduces scale weights to adjust the importance of features at different scales. The scale weights, from coarse to fine, are [0.15, 0.20, 0.30, 0.25, 0.10], reflecting the dominant role of intermediate-scale features in matching. For the current feature vector [0.85, 0.76, 0.92, 0.88, 0.79] and the standard template [0.82, 0.75, 0.90, 0.85, 0.78], the calculated weighted distance is 0.032, representing the degree of matching between the current feature and the template.

[0157] The variational energy functional combines these three terms to form a unified variational energy expression. The system assigns weight coefficients to each of the three components to reflect the importance of each part in the total energy. The spatial gradient term has a weight of 0.35, the temporal derivative term has a weight of 0.25, and the feature fitting term has a weight of 0.40. These weights were determined through cross-validation on 200 sets of test data, achieving an optimal balance between spatial continuity, temporal smoothness, and feature matching accuracy. For the example data above, the calculated total variational energy is 0.215; a smaller value indicates more accurate trajectory estimation.

[0158] After constructing the variational energy functional, partial derivatives are performed to extract key information about trajectory evolution. Spatial gradient information is obtained by performing spatial partial derivatives on the spatial gradient term. The system performs partial derivatives in the x, y, and z directions respectively to calculate the second derivative of the feature vector in each direction. In implementation, the system uses a central difference scheme and performs calculations within a 7×7×3 window to improve numerical stability. For the aforementioned example data, the second derivatives obtained from the spatial partial derivatives are: x-direction [0.04, 0.03, 0.05, 0.04, 0.02], y-direction [0.03, 0.04, 0.02, 0.05, 0.03], z-direction [0.01, 0.01, 0.02, 0.01, 0.01]. These data reflect the accelerated change characteristics of the features in space, which helps predict the curvature of the trajectory.

[0159] The system performs time partial derivative operations on the time derivative term to extract temporal evolution information and calculates the second-order time derivative of the feature vector, reflecting the acceleration of feature changes. In implementation, the system uses a second-order forward difference scheme, based on data from the most recent seven frames. To reduce noise, a moving average filter is applied with a window size of 3. For the example data above, the second-order time derivative obtained from the time partial derivative operation is [0.008, 0.006, 0.005, 0.004, 0.007], representing the trend of feature change rate. This information helps predict the acceleration or deceleration state of the trajectory.

[0160] The system performs partial derivative operations on the feature fitting terms, analyzes feature matching information, and calculates the gradient vector between the current feature vector and the standard template, reflecting the direction and magnitude of feature adjustment needed. In implementation, the system calculates the partial derivative for each feature component to form the feature gradient vector. For the example data above, the gradient vector obtained from the feature partial derivative operation is [0.018, 0.003, 0.012, 0.018, 0.006], indicating the degree to which the feature vector needs to be adjusted towards the template. This information helps optimize the feature matching process and improve the accuracy of trajectory recognition.

[0161] Based on the results of partial derivative operations, spatial gradient information, temporal evolution information, and feature matching information are combined to construct a dynamic tracking trajectory equation. The trajectory equation adopts the form of a second-order partial differential equation, including spatial second-order derivative terms, temporal second-order derivative terms, and feature matching terms. In specific implementation, the system uses the finite difference method to discretize the equation, transforming the continuous-domain partial differential equation into a system of discrete-domain algebraic equations. The system sets the grid size to 5×5×3 and the time step to 1 / 30 second (corresponding to an image acquisition rate of 30 frames / second). The discretized equation contains 15 coefficients, with initial values ​​determined based on historical data using the least squares method. For example, the initial coefficient vector is [1.20,0.85,0.65,0.45,0.30,0.25,0.20,0.15,0.12,0.10,0.08,0.06,0.05,0.04,0.03]. The first 5 coefficients correspond to the spatial derivative term, the middle 5 coefficients correspond to the temporal derivative term, and the last 5 coefficients correspond to the feature matching term.

[0162] Based on the constructed dynamic tracking trajectory equation, spatial gradient information and temporal evolution information are mapped to trajectory state parameters. The mapping process is implemented through a state observer, which includes a state prediction module and a state update module. The state prediction module predicts the state at the next moment based on the trajectory equation, and the state update module corrects the prediction result by combining the latest observation data. The system adopts an adaptive step size control strategy, dynamically adjusting the update step size according to the prediction error. A small step size (0.2) is used when the prediction error is large, and a large step size (0.8) is used when the prediction error is small. For a certain tracking instance, the initial trajectory state is [150,120,40,2.5,1.8,0.5] (the first three elements are position, and the last three elements are velocity), and the updated state after mapping is [152.5,121.8,40.5,2.6,1.9,0.5], reflecting the motion trend of the object in a short time.

[0163] By combining trajectory state parameters with feature matching information, the dynamic tracking trajectory equation is optimized. The optimization process employs an iterative reweighted least squares method, with each iteration including two steps: weight calculation and parameter update. The system calculates the weight of each sample point based on the current feature matching error, assigning larger weights to points with smaller matching errors and smaller weights to points with larger matching errors, thus achieving robust estimation. The weight function adopts a hyperbolic tangent form to ensure that the influence of outliers is effectively suppressed. The system sets the maximum number of iterations to 20 and the convergence threshold to 0.001. For the example above, the optimized equation coefficients become [1.18, 0.82, 0.62, 0.43, 0.28, 0.27, 0.22, 0.16, 0.13, 0.11, 0.07, 0.05, 0.04, 0.03, 0.02]. Compared with the initial coefficients, the spatial derivative coefficients decrease slightly, the temporal derivative coefficients increase slightly, and the feature matching coefficients remain basically unchanged, reflecting the system's enhanced emphasis on temporal continuity.

[0164] Through the above optimizations, the system obtains the optimal form of the dynamic tracking trajectory equation. This optimal trajectory equation possesses strong adaptive capabilities, automatically adjusting parameters according to the target's motion characteristics. For fast-moving targets, the influence of the time derivative term is enhanced; for targets against complex backgrounds, the influence of the feature matching term is enhanced; and for partially occluded targets, the influence of the spatial gradient term is enhanced. This adaptive mechanism enables the trajectory equation to cope with various complex scenarios. In actual testing, the optimal trajectory equation was used to track 100 sets of high-speed moving stealth code samples, achieving an average positioning error of 0.25 pixels and a success rate of 98.8%, significantly outperforming traditional methods (average error 0.68 pixels, success rate 92.5%).

[0165] This optimal trajectory equation can be directly used for real-time tracking of the invisible inkjet coding area, providing a reliable foundation for subsequent deformation feature extraction and spatial position correction. In practical applications, this method can achieve millisecond-level response time (average processing time 18 milliseconds / frame) and sub-pixel-level positioning accuracy (average error 0.22 pixels) even when objects are moving at high speeds (up to 6 meters per second), meeting the stringent requirements of the smart packaging industry for real-time detection of invisible infrared inkjet coding.

[0166] In one optional implementation, the dynamic tracking trajectory equation is corrected using the corrected spatial position parameters to calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and the real-time tracking and positioning results are output, including:

[0167] A fractional-order trajectory correction operator is constructed using the corrected spatial position parameters. The fractional-order trajectory correction operator includes a fractional-order time derivative term and a fractional-order spatial derivative term. Long-range integration is performed on the historical trajectory state using the fractional-order time derivative term, and nonlocal differentiation is performed on the spatial distribution of the trajectory using the fractional-order spatial derivative term. The dynamic tracking trajectory equation is corrected based on the results of the long-range integration and the nonlocal differentiation.

[0168] The modified dynamic tracking trajectory equation is mapped to a fractional phase space. A mapping relationship between trajectory state and position parameters is established based on the fractional phase space. The fractional state variables in the trajectory evolution process are calculated based on the mapping relationship.

[0169] Substituting the fractional-order state variables into the modified dynamic tracking trajectory equation, the precise three-dimensional spatial position of the invisible inkjet printing area is obtained through fractional-order iterative calculation of the trajectory state, and a real-time tracking and positioning result is constructed based on the precise three-dimensional spatial position.

[0170] A fractional-order trajectory correction operator was constructed using the corrected spatial position parameters. This operator consists of a fractional-order time derivative term and a fractional-order spatial derivative term, which are responsible for processing the temporal and spatial characteristics of the trajectory, respectively. The order of the fractional-order time derivative term was set to 0.82, a parameter obtained through experimental optimization, which can effectively capture the temporal characteristics of the invisible inkjet coding area during high-speed movement. The order of the fractional-order spatial derivative term was set to 1.18, which is suitable for processing the spatial distribution characteristics of the inkjet coding area.

[0171] The system performs long-range integration on the historical trajectory state using fractional-order time derivatives, with the integration interval selected from the past 32 frames of image data, forming a long-range memory mechanism in the time dimension. In implementation, the system assigns a weight coefficient to each frame of historical data, with the most recent frame having a weight of 0.85, decreasing at a decay rate of 0.92. For example, for an image sequence acquired at a frequency of 120Hz, the system performs weighted integration on the trajectory data within the past 266.7 milliseconds to obtain a comprehensive value (143.8, 115.2, 43.5) reflecting the influence of historical states.

[0172] Nonlocal differentiation of the trajectory spatial distribution is performed using fractional spatial derivative terms. A 45×45 pixel rectangular region is set in the current frame image, with the center of the invisible inkjet area as the origin, establishing a local coordinate system. Within this region, the system samples the image feature distribution along 12 directions (increasing by 30 degrees). Eight sampling points are selected in each direction, and the feature gradient is calculated. The nonlocal differentiation operation uses a weighted averaging method, with the weights decreasing by 0.88 with the distance of the sampling point from the center. This operation can effectively capture the spatial variation characteristics around the inkjet area. For example, for the sampling results of a certain test, the spatial gradient in the horizontal direction is (0.28, 0.21, 0.04), and the spatial gradient in the vertical direction is (0.15, 0.32, 0.06). After nonlocal differentiation, the combined spatial gradient value is (0.22, 0.28, 0.05).

[0173] The dynamic tracking trajectory equation is corrected based on the results of long-range integration and nonlocal differentiation. The correction process employs a two-layer weighting strategy: a time-domain correction weight of 0.65 and a spatial-domain correction weight of 0.35. The system constructs a correction coefficient matrix with dimensions 3×3, corresponding to the three coordinate axes in three-dimensional space. The diagonal elements of the matrix are 0.88, 0.92, and 0.94, representing the dominant correction factors for each coordinate axis; the off-diagonal elements range from 0.04 to 0.12, representing the coupling relationships between the coordinate axes. The system multiplies this correction matrix by the coefficient matrix of the original trajectory equation and applies the Gamma function for nonlinear adjustment to obtain the corrected trajectory equation coefficients. For example, the original trajectory equation coefficient matrix is ​​[[1.25,0.09,0.04],[0.06,1.38,0.08],[0.03,0.05,1.18]], and the corrected coefficient matrix becomes [[1.08,0.11,0.05],[0.07,1.25,0.10],[0.04,0.07,1.12]].

[0174] After correcting the trajectory equation, the corrected dynamic tracking trajectory equation is mapped to a fractional phase space. This is achieved through fractional phase space transformation, which converts the trajectory state from traditional Euclidean space to a fractional phase space that better expresses complex dynamic characteristics. The transformation uses a 3×9 transformation matrix, whose elements are determined by fractional parameters.

[0175] In the implementation, three sets of fractional-order parameters were selected: 0.82, 1.0, and 1.18, corresponding to the fractional-order characteristics of position, velocity, and acceleration, respectively. These parameters were determined through statistical analysis of 500 sets of inkjet printing area motion data, which can capture the nonlinear characteristics of high-speed moving packaged objects to the greatest extent. During the transformation process, the system expands the original 6-dimensional state vector (containing position and velocity) into a 9-dimensional fractional-order state vector (adding fractional-order state variables). For example, for the original state vector (152.4, 122.6, 46.1, 2.6, 1.9, 0.5), the state vector obtained after the fractional-order phase space transformation is (158.1, 126.5, 47.3, 2.3, 1.7, 0.4, 0.19, 0.24, 0.06).

[0176] A mapping relationship between trajectory states and position parameters is established in a fractional phase space. This mapping is implemented using a radial basis function network, comprising an input layer (9 nodes, corresponding to 9-dimensional fractional states), two hidden layers (containing 16 and 8 nodes respectively), and an output layer (3 nodes, corresponding to 3-dimensional position parameters). The hidden layers use the hyperbolic tangent activation function, and the output layer uses a linear activation function. The system is trained using 600 pre-collected trajectory data sets, with 480 sets used for training and 120 sets for validation. Training employs the Adam optimization algorithm with a learning rate of 0.001, a batch size of 32, and 300 training epochs. After training, the average accuracy of the mapping relationship reaches 99.2%. Based on this mapping relationship, the system calculates the fractional state variables during trajectory evolution. Specifically, for the state vector (158.1,126.5,47.3,2.3,1.7,0.4,0.19,0.24,0.06) in the fractional phase space, the fractional state variables obtained by mapping are (0.82,0.94,0.88).

[0177] The calculated fractional-order state variables are substituted into the modified dynamic tracking trajectory equation, and the precise three-dimensional spatial position of the invisible inkjet printing area is obtained through fractional-order iterative calculation of the trajectory state. The iterative calculation employs a fractional-order prediction-correction algorithm, which exhibits higher numerical stability and prediction accuracy compared to traditional integer-order algorithms. The iteration step size is set to 0.08 seconds, and the total number of iterations is 12.

[0178] Each iteration consists of two phases: prediction and correction. The prediction phase uses the current fractional-order state variables and the corrected trajectory equation to calculate the predicted position for the next moment. The correction phase combines image processing results and corrects the predicted value using the least squares method. The system sets the convergence condition for iterations to be that the Euclidean distance between two consecutive iterations is less than 0.03 pixels. For example, with an initial position of (158.1, 126.5, 47.3), after 12 iterations, it finally converges to the precise 3D spatial position (161.2, 130.4, 49.1), with the Euclidean distance between the last two iterations being 0.025 pixels.

[0179] Real-time tracking and positioning results are constructed based on precise 3D spatial location. The positioning results include the center coordinates, bounding box coordinates, attitude parameters, and detection confidence of the coding area. The center coordinates are directly obtained from iterative calculations, while the bounding box coordinates are calculated by combining the size information of the coding area.

[0180] In practical applications, different bounding box templates are preset according to product type, with a typical size of 85×45 pixels. Attitude parameters are determined by analyzing the principal axis direction and surface normal vector of the inkjet printing area, including pitch, roll, and yaw angles. Confidence is obtained by calculating the weighted average of the reciprocal of the iterative convergence error and the image feature matching degree, typically between 0.92 and 0.98. A complete positioning result example is as follows: center coordinates (161.2, 130.4, 49.1), bounding box coordinates [(118.7, 107.9, 49.1), (203.7, 107.9, 49.1), (203.7, 152.9, 49.1), (118.7, 152.9, 49.1)], attitude angles (5.8°, 4.2°, 11.3°), confidence 0.96.

[0181] In practical application testing, this method achieved an average positioning error of 0.28 pixels and an average processing time of 22 milliseconds under conditions of high-speed object movement (linear velocity reaching 5.5 m / s and angular velocity reaching 45 degrees / s). For different types of packaged objects (including cardboard boxes, plastic bottles, and metal cans), the method demonstrated good adaptability, with positioning success rates of 99.5%, 98.7%, and 97.8%, respectively. Partially, even with partial occlusion (maximum occlusion area not exceeding 40%), the method maintained a positioning success rate of over 95%, significantly outperforming traditional integer-order tracking methods (with a success rate of approximately 85% under the same conditions).

[0182] This method has completed a three-month verification test on a production line, processing over 10 million packaging samples. The average positioning accuracy is better than 0.3 pixels, meeting the industry's stringent requirements for real-time tracking of invisible infrared inkjet printing areas. Test data shows that using this method improved the accuracy of inkjet printing recognition by 8.5 percentage points, reduced the false alarm rate by 3.2 percentage points, and reduced the missed detection rate by 4.5 percentage points, providing effective technical support for intelligent packaging quality control.

[0183] A second aspect of the present invention provides a stealth infrared inkjet printing image acquisition and intelligent recognition processing system, comprising:

[0184] The first unit is used to acquire a continuous image sequence of a high-speed moving smart packaging object, perform layered processing on the continuous image sequence through a residual cross-layer feature fusion network, extract the invisible inkjet coding area on the surface of the smart packaging object, and generate a spatiotemporal feature data stream of the invisible inkjet coding area.

[0185] The second unit is used to input the spatiotemporal feature data stream into the feature extraction layer of the residual cross-layer feature fusion network, connect different feature extraction layers in series through cross-scale residual connection channels, and adaptively adjust the parameters of each feature extraction layer according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region.

[0186] The third unit is used to calculate the dynamic tracking trajectory equation based on the multi-scale feature vector, optimize the coefficients of the dynamic tracking trajectory equation through iterative calculation, and realize real-time tracking of the invisible inkjet printing area; according to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network, and the deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters;

[0187] The fourth unit is used to correct the dynamic tracking trajectory equation using the corrected spatial position parameters, calculate the precise three-dimensional spatial position of the invisible inkjet printing area, and output the real-time tracking and positioning results.

[0188] A third aspect of the present invention provides an electronic device, comprising:

[0189] processor;

[0190] Memory used to store processor-executable instructions;

[0191] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0192] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0193] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for acquiring and intelligently recognizing stealth infrared inkjet printing images, characterized in that, include: A continuous image sequence of a high-speed moving smart packaging object is acquired. The continuous image sequence is then processed in layers through a residual cross-layer feature fusion network to extract the invisible inkjet coding area on the surface of the smart packaging object and generate a spatiotemporal feature data stream of the invisible inkjet coding area. The spatiotemporal feature data stream is input into the feature extraction layer of the residual cross-layer feature fusion network. Different feature extraction layers are connected in series through cross-scale residual connection channels. The parameters of each feature extraction layer are adaptively adjusted according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region. The dynamic tracking trajectory equation is calculated based on the multi-scale feature vector. The coefficients of the dynamic tracking trajectory equation are optimized by iterative calculation to achieve real-time tracking of the invisible inkjet coding area. According to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network. The deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters. The corrected spatial position parameters are used to modify the dynamic tracking trajectory equation, calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and output the real-time tracking and positioning results.

2. The method according to claim 1, characterized in that, The continuous image sequence is processed hierarchically using a residual cross-layer feature fusion network to extract the invisible inkjet coding region on the surface of the smart packaging object, generating a spatiotemporal feature data stream of the invisible inkjet coding region, including: A continuous image sequence of the surface of a smart packaging object is acquired. The continuous image sequence is then subjected to hierarchical feature analysis through a residual cross-layer feature fusion network. A pulsed heat source is used to thermally excite the surface of the smart packaging object. A temperature field distribution image is then established based on the thermal diffusion response of the surface of the smart packaging object. The temperature field distribution image is analyzed based on the first feature extraction layer of the residual cross-layer feature fusion network. The spatiotemporal feature parameters of the temperature field distribution image are constructed according to the temperature field time change rate and spatial temperature gradient. The spatiotemporal feature parameters contain the dynamic evolution information of the temperature field distribution image. The spatiotemporal feature parameters are progressively processed by multiple residual blocks cascaded within the backbone layer of the residual cross-layer feature fusion network. Each residual block enhances the hierarchical correlation of the spatiotemporal feature parameters through main branch mapping and cross-layer connections, and temperature field features are constructed based on the hierarchical correlation. Based on the temperature field characteristics, the temperature response characteristics of the invisible inkjet coding area are analyzed, and a feature index including initial temperature, temperature increment and time constant is constructed. Based on the residual cross-layer feature fusion network, the feature index and the temperature field characteristics are combined to generate a spatiotemporal feature data stream characterizing the invisible inkjet coding area.

3. The method according to claim 2, characterized in that, Hierarchical feature analysis is performed on the continuous image sequence using a residual cross-layer feature fusion network. A pulsed heat source is used to thermally excite the surface of the smart packaging object. A temperature field distribution image is then established based on the thermal diffusion response of the smart packaging object's surface, including: A multi-layer residual network structure is constructed, the feature increment of each layer is calculated through residual mapping, cross-layer feature connection channels are established between adjacent feature layers, feature weights are allocated to each layer, and multi-layer image features of the continuous image sequence are extracted. The multi-layer image features guide the initial configuration of the heat source excitation parameters. Based on the multi-level image feature analysis, the material properties of the smart packaging object are analyzed. The initial heat source power in the coarse adjustment stage is determined according to the material properties. A layered progressive heat source control strategy is adopted to apply pulsed thermal excitation to the surface of the smart packaging object. The initial heat source power, dynamic compensation amount and adaptive weighting factor together determine the real-time heat source power in the fine adjustment stage. The thermal diffusion response generated on the surface of the smart packaging object under the pulsed thermal excitation is detected, and a temperature field distribution image is generated based on the thermal diffusion response combined with the heat source influence coefficient, thermal diffusion coefficient and thermal conductivity.

4. The method according to claim 1, characterized in that, By concatenating different feature extraction layers through a cross-scale residual connection channel, and adaptively adjusting the parameters of each feature extraction layer according to the image complexity of the spatiotemporal feature data stream, the multi-scale feature vector of the invisible inkjet coding region is obtained, including: A dual-constraint architecture of teacher network and student network is constructed. Feature connection channels are established through feature transfer between the teacher network and the student network, and feature transfer paths are formed based on the feature connection channels. Feature extraction is performed based on the feature transfer path. The high-dimensional feature representation extracted by the teacher network is used as knowledge guidance. The feature space is compressed and mapped based on the student network under the knowledge guidance. The feature connection channel is dynamically adjusted according to the feature reconstruction result of the compressed mapping. Based on the feature reconstruction results of the compression mapping, the spatiotemporal data stream features are analyzed, and the spatiotemporal data stream features are input into the teacher network for feature extraction. The features extracted by the teacher network are used to adjust the feature extraction process of the student network. The feature extraction process of the student network is used for feature consistency constraints. Based on the feature consistency constraint, the feature reconstruction result of the compressed mapping is spatially aligned, the feature extraction parameters of the teacher network are updated based on the spatially aligned features, and the optimized features are output according to the teacher network. The spatially aligned features are fused with the teacher network and the student network under the guidance of the optimized features, and a multi-scale feature vector of the invisible inkjet coding region is generated based on the result of the feature fusion.

5. The method according to claim 1, characterized in that, The dynamic tracking trajectory equation is calculated based on the multi-scale feature vectors, and the coefficients of the dynamic tracking trajectory equation are optimized through iterative calculation to achieve real-time tracking of the invisible inkjet coding area, including: A variational energy functional is constructed based on the multi-scale feature vectors, and the optimal form of the dynamic tracking trajectory equation is derived based on the variational energy functional. The optimal form of the dynamic tracking trajectory equation includes the coefficient matrix to be optimized. The optimal form of the dynamic tracking trajectory equation is decomposed into parameter optimization terms and constraint optimization terms. Based on the parameter optimization terms, an objective function is constructed for the coefficient matrix to be optimized. Based on the constraint optimization terms, weight coefficients are introduced to establish constraints. Based on the objective function and the constraints, an optimization objective is constructed. The optimization objective is solved by a dynamic iterative method. The coefficient matrix is ​​updated by the parameter optimization terms. The updated coefficient matrix is ​​substituted into the constraint optimization terms to calculate the weight coefficients. The coefficient matrix is ​​further optimized based on the gradient information of the weight coefficients. By substituting the optimized coefficient matrix into the optimal form of the dynamic tracking trajectory equation, real-time tracking of the invisible inkjet printing area can be achieved.

6. The method according to claim 5, characterized in that, Based on the multi-scale feature vectors, a variational energy functional is constructed. The optimal form of the dynamic tracking trajectory equation is derived from this variational energy functional, including: The variational energy functional characterizes the spatial distribution characteristics of the trajectory through the spatial gradient term, describes the temporal evolution of the trajectory through the time derivative term, and measures the degree of matching between the trajectory and the multi-scale feature vector through the feature fitting term. The spatial gradient term, the time derivative term, and the feature fitting term are combined to form a unified variational energy expression. Based on the variational energy functional, spatial gradient information is obtained by performing spatial partial derivative operation on the spatial gradient term in the variational energy expression; temporal evolution information is extracted by performing temporal partial derivative operation on the temporal derivative term; feature matching information is analyzed by performing feature partial derivative operation on the feature fitting term; and the spatial gradient information, the temporal evolution information and the feature matching information are combined to construct a dynamic tracking trajectory equation. Based on the dynamic tracking trajectory equation, the spatial gradient information and the temporal evolution information are mapped into trajectory state parameters. The trajectory state parameters are combined with the feature matching information to optimize the dynamic tracking trajectory equation, thereby obtaining the optimal form of the dynamic tracking trajectory equation.

7. The method according to claim 1, characterized in that, The dynamic tracking trajectory equation is corrected using the corrected spatial position parameters to calculate the precise three-dimensional spatial position of the invisible inkjet coding area, and the real-time tracking and positioning results are output, including: A fractional-order trajectory correction operator is constructed using the corrected spatial position parameters. The fractional-order trajectory correction operator includes a fractional-order time derivative term and a fractional-order spatial derivative term. Long-range integration is performed on the historical trajectory state using the fractional-order time derivative term, and nonlocal differentiation is performed on the spatial distribution of the trajectory using the fractional-order spatial derivative term. The dynamic tracking trajectory equation is corrected based on the results of the long-range integration and the nonlocal differentiation. The modified dynamic tracking trajectory equation is mapped to a fractional phase space. A mapping relationship between trajectory state and position parameters is established based on the fractional phase space. The fractional state variables in the trajectory evolution process are calculated based on the mapping relationship. Substituting the fractional-order state variables into the modified dynamic tracking trajectory equation, the precise three-dimensional spatial position of the invisible inkjet printing area is obtained through fractional-order iterative calculation of the trajectory state, and a real-time tracking and positioning result is constructed based on the precise three-dimensional spatial position.

8. A stealth infrared inkjet printing image acquisition and intelligent recognition processing system, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to acquire a continuous image sequence of a high-speed moving smart packaging object, perform layered processing on the continuous image sequence through a residual cross-layer feature fusion network, extract the invisible inkjet coding area on the surface of the smart packaging object, and generate a spatiotemporal feature data stream of the invisible inkjet coding area. The second unit is used to input the spatiotemporal feature data stream into the feature extraction layer of the residual cross-layer feature fusion network, connect different feature extraction layers in series through cross-scale residual connection channels, and adaptively adjust the parameters of each feature extraction layer according to the image complexity of the spatiotemporal feature data stream to obtain the multi-scale feature vector of the invisible inkjet coding region. The third unit is used to calculate the dynamic tracking trajectory equation based on the multi-scale feature vector, optimize the coefficients of the dynamic tracking trajectory equation through iterative calculation, and realize real-time tracking of the invisible inkjet printing area; according to the prediction result of the dynamic tracking trajectory equation, the deformation feature information of the surface of the smart packaging object is obtained from the feature extraction layer of the residual cross-layer feature fusion network, and the deformation feature information is fused with the dynamic tracking trajectory equation to obtain the corrected spatial position parameters; The fourth unit is used to correct the dynamic tracking trajectory equation using the corrected spatial position parameters, calculate the precise three-dimensional spatial position of the invisible inkjet printing area, and output the real-time tracking and positioning results.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Package change prevention identification method, device and equipment based on printed label, and storage medium

    CN121684964A