A food packaging image recognition method for low-light environments

By combining a temporal state modeling perceptual encoder, a material-aware spectral estimation network, and a physical illumination reconstruction module, the accuracy and stability issues of food packaging image recognition under low-light conditions are solved. High-quality semantic-physical dual closed-loop optimization is achieved, improving the recognition effect of food packaging images.

CN121837676BActive Publication Date: 2026-06-02SICHUAN FOOD INSPECTION INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN FOOD INSPECTION INST
Filing Date
2026-03-10
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for food packaging image recognition in low-light environments have shortcomings in photophysical modeling and semantic perception, resulting in insufficient recognition accuracy and stability. In particular, in low-light scenarios such as supermarket shelves, warehouse backlighting, and nighttime express delivery unloading, image quality deteriorates, making it difficult to effectively identify key information on food packaging.

Method used

By combining a temporal state modeling perceptual encoder, a material-aware spectral estimation network, a physical illumination reconstruction module, and a semantic graph prior enhancement module, a semantic-physical dual closed-loop optimization is formed through illumination state space modeling, material spectral matching, physical illumination compensation, and semantic feature extraction, thereby achieving the recognition of food packaging images.

Benefits of technology

It significantly improves image restoration quality and recognition reliability in low-light environments, can adaptively adjust compensation parameters, enhance visual consistency and physical realism, and ensure the stability and recognition accuracy of key semantic regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837676B_ABST
    Figure CN121837676B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of image recognition, and relates to a food packaging image recognition method in a low-light environment, comprising: modeling the light state space through a time sequence state modeling perception encoder to obtain a light state vector; matching the spectral prototype library of the food packaging material through a material perception spectrum estimation network to obtain a spectral reflectivity vector; performing physical light compensation through a physical light reconstruction module to obtain a reconstructed light distribution and a reconstructed enhanced image; performing a differentiable rendering optimization through a differentiable rendering module to obtain a rendered enhanced image; extracting semantic features through a semantic graph prior enhancement module to generate a semantic heat map, and performing semantic-physical double closed loop optimization based on the semantic heat map to obtain a final optimized enhanced image; performing image recognition on the final optimized enhanced image to obtain a food packaging recognition result; and improving the accuracy and stability of recognizing low-light food packaging images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition, and specifically discloses a method for recognizing food packaging images in low-light environments. Background Technology

[0002] Image recognition for food packaging is widely used in smart retail, warehouse management, and food traceability. These fields typically require automatic identification of product packaging to achieve functions such as inventory management and product traceability. However, in low-light environments, such as supermarket shelves, backlit warehouses, and nighttime express delivery unloading scenarios, image quality deteriorates significantly, affecting the accuracy and stability of image recognition systems. Under low light, images may exhibit blurriness, low contrast, and color distortion, causing the recognition system to be unable to effectively identify key information on food packaging (such as text, barcodes, and logos).

[0003] Most current image recognition methods rely on general image enhancement algorithms (such as histogram equalization and contrast enhancement), which do not model the photophysical characteristics of low-light environments. Since image quality degradation involves not only brightness but also physical effects such as light reflection and refraction, existing technologies fail to adequately consider these factors, making it difficult to achieve good image restoration results. Existing low-light image recognition methods typically do not model the higher-order semantic structures specific to certain domains (such as food packaging). These higher-order semantic structures include text, barcodes, and brand logos on the packaging; traditional methods are prone to errors or loss when recovering this information under low light. Furthermore, existing methods have significant shortcomings in the depth and differentiability of physical modeling. Most methods fail to embed optical physical models (such as BRDF and spectral reflectance) into neural networks in a differentiable form, resulting in enhancement results lacking physical realism. Simultaneously, the lack of cross-modal alignment mechanisms between semantic information and physical parameters makes it difficult for models to simultaneously guarantee semantic recognizability and physical plausibility under extreme low-light conditions. Therefore, existing low-light food packaging image recognition technologies have significant deficiencies in photophysical modeling and semantic perception.

[0004] In view of this, the present invention provides a food packaging image recognition method for low-light environments, which can recognize food packaging images in low-light environments by combining multiple aspects such as the physical lighting characteristics of low-light environments and the semantic information of food packaging, thereby improving the accuracy and stability of recognition. Summary of the Invention

[0005] The purpose of this invention is to provide a method for recognizing food packaging images in low-light environments, addressing the problem of improving the accuracy and stability of recognizing food packaging images in low-light conditions. The specific solution is as follows:

[0006] A method for food packaging image recognition under low-light conditions includes: performing illumination state space modeling on the original low-light image and semantic heatmap using a temporal state modeling perceptual encoder to obtain an illumination state vector; matching the original low-light image with a spectral prototype library of food packaging materials using a material-aware spectral estimation network to obtain a spectral reflectance vector; performing physical illumination compensation on the illumination state vector, spectral reflectance vector, and semantic heatmap using a physical illumination reconstruction module to obtain a reconstructed illumination distribution and a reconstructed enhanced image; performing differentiable rendering optimization on the reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap using a physically differentiable rendering module to obtain a rendered enhanced image; extracting semantic features from the rendered enhanced image using a semantic graph prior enhancement module to generate a semantic heatmap, and feeding the semantic heatmap back to the temporal state modeling perceptual encoder, material-aware spectral estimation network, physical illumination reconstruction module, and physically differentiable rendering module to achieve semantic-physical dual closed-loop optimization, obtaining the final optimized enhanced image; and performing image recognition on the final optimized enhanced image to obtain the food packaging recognition result.

[0007] Furthermore, a temporal state modeling perceptual encoder is used to perform illumination state space modeling on the original low-light image and semantic heatmap. This includes: employing a semantic modulation state space modeling method to establish an illumination state space model to capture temporal illumination changes and obtain the illumination state vector; the illumination state space model is as follows:

[0008] ;

[0009] in, Let be the illumination state vector at time t; For semantic modulation state update function; The original low-light image at time t; This is the illumination state vector at the previous time t-1; This is a semantic heatmap.

[0010] Furthermore, the material-aware spectral estimation network is used to match the original low-light image with a spectral prototype library of food packaging materials. This involves processing the illumination state vector and the input image using the material-aware spectral estimation network to achieve spectral prototype library matching for food packaging materials. This includes: constructing a material prototype library based on the spectral reflectance patterns of the materials; extracting pixel features from the original low-light image and performing feature matching between these pixel features and each material in the material prototype library to obtain the material attention weight for each material.

[0011] ;

[0012] in, Let x be the attention weight for pixel x belonging to material k; To represent along the material prototype dimension The normalized exponential function is used to map the matching scores of each material to an attention weight distribution that is non-negative and sums to 1; Let x be the local material encoding feature of pixel x; x is the pixel variable; T is the transpose of the matrix. A learnable feature matching matrix; Let k be the spectral reflectance mode of the k-th material. Then, process the multiple material attention weights and their corresponding spectral reflectance modes for this pixel to obtain a pixel-level spectral reflectance vector.

[0013] ;

[0014] in, is the spectral reflectance of pixel x; k is the material prototype variable; K is the total number of material prototypes.

[0015] Furthermore, it also includes updating the material prototype library based on semantic gradient feedback:

[0016] ;

[0017] in, and These are the spectral reflection modes of the k-th material after the t-th and t+1-th updates, respectively; The self-calibrating learning rate for the material prototype library; This represents the semantic consistency loss. For the spectral reflectance mode of the k-th material; Let be the partial derivative of the semantic consistency loss with respect to the k-th type of material prototype in the material prototype library.

[0018] Furthermore, the physical illumination reconstruction module performs physical illumination compensation on the illumination state vector, spectral reflectance vector, and semantic heatmap, including: constructing a physical illumination imaging model based on the original low-illumination image and surface reflectance.

[0019] ;

[0020] in, denoted as , where is the pixel value of the original low-light image; L(x) represents the incident light distribution; and R(x) represents the surface reflectivity. (x) represents photosensitive noise; based on the illumination state vector and spectral reflectance vector, the filter kernel parameters are adaptively adjusted to obtain the illumination compensation kernel function; the filter kernel parameters are:

[0021] ;

[0022] Where λ is the filter kernel parameter; It is an adaptive mapping network; Let be the illumination state vector at time t; Let x be the spectral reflectance of pixel x; based on the physical illumination imaging model and the illumination compensation kernel function, an illumination compensation framework is constructed to obtain the reconstructed illumination distribution:

[0023] ;

[0024] in, To reconstruct the light distribution; This represents the distribution of incident light. This is the illumination compensation kernel function; A semantic heatmap is generated; differential compensation intensities are assigned to different regions of the semantic heatmap to obtain semantic modulation weights:

[0025] ;

[0026] in, For semantic modulation weights; Use the Sigmoid activation function; These are the convolution weights; The semantic heatmap of pixel x is used; based on semantic modulation weights and illumination compensation kernel functions, the original low-light image is processed to obtain the reconstructed and enhanced image:

[0027] ;

[0028] in, To reconstruct and enhance the image; These are the pixel values ​​of the original low-light image; For semantic modulation weights; This is the illumination compensation kernel function; Let be the illumination state vector at time t; Let x be the spectral reflectance of pixel x; This is the semantic heatmap of pixel x.

[0029] Furthermore, the reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap are optimized using a physically differentiable rendering module. This includes processing the reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap using rendering operators to obtain an enhanced rendering image.

[0030] ;

[0031] in, Enhance the image for rendering; For rendering operators; To reconstruct the light distribution; Let x be the spectral reflectance of pixel x; This is the semantic heatmap of pixel x.

[0032] Furthermore, semantic features are extracted from the rendered enhanced image through a semantic graph prior enhancement module to generate a semantic heatmap. This includes: processing the rendered enhanced image through a semantic embedding network to obtain a semantic feature map; constructing a semantic graph based on the semantic feature map; the semantic graph includes node features and semantic similarity edges between nodes.

[0033] ;

[0034] ;

[0035] in, The semantic node features of pixel x; and These are the semantic feature maps for pixels x and y, respectively. Edges representing semantic similarity between nodes; and These are the node mapping function and the edge mapping function, respectively; the semantic graph is processed through a graph convolutional network to update node features and obtain a global semantic representation.

[0036] ;

[0037] in, The semantic global representation is used; GCN is a graph neural network; V is the set of semantic node features; E is the set of semantic similarity edges; the semantic global representation is processed through an attention mechanism to generate a semantic heatmap.

[0038] ;

[0039] in, For pixel x, a semantic heatmap; Use the Sigmoid activation function; These are learnable semantic global weights; This is a learnable semantic global bias.

[0040] Furthermore, the spectral regularization loss of the material-aware spectral estimation network is:

[0041] ;

[0042] in, This is the loss due to spectral regularization. The square of the L2 norm; β represents the gradient along the spectral wavelength direction; β is the energy constraint weight. Let x be the spectral reflectance of pixel x; is the reference spectrum for pixel x.

[0043] Furthermore, the illumination consistency loss of the physical illumination reconstruction module is:

[0044] ;

[0045] in, This is due to the loss of uniformity in illumination. The square of the L2 norm; To reconstruct the light distribution; The ideal illumination distribution is represented by γ, which is the weight of the temporal smoothing constraint. It is an L1 norm; Temporal difference of the enhancement results for consecutive frames; For time-series image enhancement.

[0046] Furthermore, the rendering loss of the physically differentiable rendering module includes rendering consistency loss, semantic consistency loss, and temporal consistency loss: the rendering consistency loss is:

[0047] ;

[0048] in, This results in a loss of rendering consistency. The square of the L2 norm; Reconstruct and enhance the image for pixel x; To reconstruct and enhance the image; To render an enhanced image; α and β are the first and second rendering consistency weight coefficients, respectively; To assess the structural similarity between reconstructed and rendered enhanced images; For semantic consistency constraints; the semantic consistency loss is:

[0049] ;

[0050] in, For semantic consistency loss; x is a pixel variable; For semantic weights; It is an L1 norm; The specular reference image is used; the temporal consistency loss is:

[0051] ;

[0052] in, This results in a loss of time-series consistency. and , respectively, are the rendered enhanced images at times t and t-1; γ is the temporal smoothing constraint weight; To reconstruct the difference in illumination distribution over time; To reconstruct the light distribution.

[0053] Furthermore, the semantic feature self-correction loss of the semantic graph prior enhancement module is:

[0054] ;

[0055] in, For semantic feature self-correction loss; For semantic recognition and reconstruction loss; , and These are the weight coefficients for the first, second, and third semantic features, respectively. This is due to the loss of uniformity in illumination. This results in a loss of rendering consistency. The temporal consistency loss is used; the feedback strength of the semantic graph prior enhancement module is:

[0056] ;

[0057] in, and , respectively, represent the feedback intensity at times t and t+1; η is the feedback intensity update step size coefficient, used to control the adjustment magnitude of the semantic feedback intensity between adjacent times; mean( () indicates the loss of illumination uniformity. Loss of rendering consistency The arithmetic mean is used to comprehensively characterize the overall error level of the physical module; δ is the error threshold of the physical module, used to determine whether the semantic feedback strength needs to be enhanced.

[0058] This invention proposes a food packaging image recognition method for low-light environments, which combines temporal state modeling, physical illumination compensation, semantic enhancement, cross-device joint learning, and multi-source data-driven dynamic training strategies, and has the following advantages and beneficial effects:

[0059] This invention is the first to introduce a dual closed-loop design of semantic feedback and physical compensation in low-light image recognition tasks. Information feedback is established between the semantic module (SRN) and the physical module (PCU, PDR), enabling semantically guided physical reconstruction and semantic enhancement fed back from physical modeling, significantly improving image restoration quality and recognition reliability.

[0060] This invention uses a temporal state modeling perceptual encoder (Mamba-Backbone) to model the temporal state space, capturing the dynamic characteristics of illumination changes over time. This breaks through the static limitations of traditional single-frame enhancement models, enabling the invention to understand illumination change trends and perform dynamic compensation.

[0061] This invention constructs a spectrum-driven physical illumination compensation mechanism through joint modeling using a Material Spectral Estimation Network (MSEN) and a Differentiable Compensation Kernel (PCU), achieving illumination reconstruction based on the material's reflectivity. The system can adaptively adjust compensation parameters for different packaging materials (plastic film, metal foil, paper labels, etc.), enhancing visual consistency and physical realism. Attached Figure Description

[0062] Figure 1 This is a schematic diagram of the process framework for a food packaging image recognition method for low-light environments provided by the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0064] This invention proposes a method for food packaging image recognition in low-light environments, including a low-light food packaging image recognition framework that combines physically differentiable rendering and semantic perception. For example... Figure 1As shown, the framework constructs an end-to-end closed-loop system from low-light physical imaging model to semantic structure reconstruction by introducing mechanisms such as material spectral estimation, temporal lighting modeling, differentiable physical compensation, and semantic graph prior enhancement. Within the complete technical system, this system revolves around five core modules: a temporal state modeling perceptual encoder (Mamba-Backbone), a material-aware spectral estimation network (MSEN), a physically-aware lighting compensation module (PCU), a physically-differentiable rendering module (PDR), and a semantic graph prior enhancement module (SRN). These modules, through deep coupling of semantic feedback and physical constraints, realize the workflow of lighting restoration, material reconstruction, semantic enhancement, and reflection optimization. The core innovation of the overall framework lies in proposing a semantic-physical dual-closed-loop collaborative perception mechanism, including a semantic module and a physical module. The semantic module generates a semantic heatmap based on structural importance, which is used to modulate lighting restoration, material estimation, and reflection rendering; the lighting consistency and physical errors generated by the physical module during enhancement and rendering are fed back to the semantic module, forming a continuously iterative semantic-physical bidirectional information flow, thereby solving the problem of information fragmentation in the traditional serial image enhancement-recognition process. Each module performs the following functions: The Temporal State Modeling Perceptual Encoder (Mamba-Backbone) captures illumination changes in low-light image sequences using an adjustable state-space model, improving illumination modeling and feature representation capabilities. Secondly, the Material Aware Spectral Estimation Network (MSEN) estimates the spectral reflectance of food packaging surfaces, achieving differentiable illumination modeling based on material differences. Subsequently, the Physical Illumination Compensation Module (PCU) generates a differentiable compensation kernel using illumination states and spectral parameters, performing locally accurate physical illumination restoration on the image. Finally, the Physically Differentiable Rendering Module (PDR) embeds illumination, reflectance, and semantic weights into the differentiable rendering equation, achieving more interpretable reflection reconstruction. The Semantic Graph Prior Enhancement Module (SRN) extracts semantic structures from the enhanced image to construct a semantic graph and generates a semantic heatmap to inversely modulate the physical module.

[0065] Through the collaborative design and closed-loop optimization mechanism of the above five modules, this invention constructs a method for recognizing low-light food packaging images with physical consistency and semantic sensitivity. This method significantly improves image quality and structural recognizability in complex low-light scenarios, providing a highly reliable and robust solution for applications such as food packaging inspection, smart retail, and warehouse management.

[0066] Step 1: The original low-light image and semantic heatmap are modeled in the illumination state space by a temporal state modeling perceptual encoder to obtain the illumination state vector.

[0067] The Mamba-Backbone temporal state modeling-aware encoder is a temporal-aware coding framework based on adjustable state space modeling, designed to dynamically capture illumination changes and semantic feature evolution in low-light image sequences. This module introduces a semantic heatmap from the Semantic Graph Prior Enhancement (SRN) module through a Semantic-modulated State Update (SMSU) mechanism to achieve semantically guided illumination state prediction and compensation, thereby enabling more accurate and interpretable low-light image restoration in the temporal dimension.

[0068] In some embodiments, a semantic modulation state space modeling method can be used to establish an illumination state space model to capture temporal illumination changes and obtain an illumination state vector.

[0069] In low-light environments, image illumination typically exhibits non-stationary temporal dynamics, meaning that brightness and reflection intensity change with time and viewing angle. Traditional enhancement methods are mostly based on static single-frame enhancement, making it difficult to characterize the temporal evolution of illumination. This invention proposes a state-space modeling method for semantic modulation, which captures temporal illumination changes while simultaneously using semantic heatmaps. The state update process is modulated so that the model prioritizes key semantic regions (such as text, barcodes, and logos) in the food packaging during illumination recovery. The formula for illumination state space modeling is defined as follows:

[0070] ;

[0071] in, Let be the illumination state vector at time t, representing the illumination distribution and brightness field of the image; The state update function for semantic modulation can be implemented using structures such as Mamba state space modules or LSTM. The original low-light image at time t; This is the illumination state vector at the previous time t-1; This is a semantic heatmap, generated by a SRN, used to adjust feature weights during the state update process. Unlike traditional RNNs / LSTMs, the semantic modulation state update function... The semantic modulation term was introduced. This leads to the formation of a learnable adaptive illumination vector adjustment (AIVA) mechanism, which can dynamically adjust the temporal illumination modeling weights according to scene semantics, thereby maintaining semantic consistency and visual stability in complex low-light environments.

[0072] After capturing the illumination changes in the image sequence, Mamba-Backbone performs illumination compensation based on the aforementioned state-space modeling results, obtaining temporally enhanced images. Unlike traditional static compensation, this process combines differentiable physical modeling with temporal semantic guidance to achieve adaptive restoration of low-light images. The illumination compensation process can be represented as:

[0073] ;

[0074] in, For temporal enhancement of images; This is the original low-light image; This is the illumination compensation kernel function, whose parameter λ is derived from the illumination state vector of the Mamba module. This is determined in conjunction with material properties. This formula makes illumination compensation interpretable and controllable. When the semantic module detects packaging text or logo areas, the semantic heatmap enhances the illumination recovery intensity of the corresponding areas. In non-critical areas, smoothing compensation is used to avoid distortion caused by over-enhancement. The final output image maintains physical consistency and visual naturalness while improving brightness and contrast.

[0075] The semantic modulation state-space model using Mamba-Backbone offers several significant advantages: It continuously tracks illumination changes over time, enabling adaptive updates of the illumination state vector. A semantic heatmap modulation mechanism ensures temporal stability of high-value regions such as packaging text and barcodes across consecutive frames. An adjustable illumination vector update mechanism dynamically adjusts the compensation magnitude based on temporal state and semantic feedback, avoiding over-brightening or artifacts common in traditional algorithms. The collaborative modeling of illumination state and semantic features effectively enhances the ability of subsequent recognition modules to discriminate key semantic regions.

[0076] The temporal state modeling perceptual encoder (Mamba-Backbone) receives semantic heatmaps from the semantic graph prior augmentation module (SRN) in real time during training. This ensures the priority of semantic region recovery, providing high-quality input for the subsequent Physical Illumination Reconstruction Unit (PCU).

[0077] Step 2: The original low-light image is matched with the spectral prototype library of food packaging materials through a material-aware spectral estimation network to obtain the spectral reflectance vector.

[0078] Material-Aware Spectral Estimation Network (MSEN) is an image material estimation network based on differentiable material spectral modeling and semantic feedback optimization. It aims to automatically extract material properties and spectral reflectance parameters of food packaging from low-light images. Unlike traditional statistical or texture-based material recognition methods, MSEN achieves a deep fusion of illumination modeling and semantic structure recognition by combining physical spectral parameter estimation with semantic heatmap alignment. The output of MSEN directly serves the subsequent Physically Reconstructed Illumination Unit (PCU) and Physically Differentiable Rendering Unit (PDR), providing them with pixel-level differentiable spectral reflectance mapping and interpretable material vectors.

[0079] In low-light environments, the absorption and reflection characteristics of complex materials on food packaging surfaces, such as reflective films, frosted plastics, and metallic coatings, vary significantly. Traditional enhancement methods operate only in the brightness domain, failing to consider the differences in the spectral response of materials. This can easily lead to overly bright reflective areas and undercompensated non-reflective areas, thus compromising the physical consistency of the image. Therefore, by constructing a learnable Material Prototype Bank (MPB) and a Semantic-feedback Adaptation (SFA) mechanism, adaptive estimation and dynamic optimization of material spectral features are achieved.

[0080] Step 2.1, the Material Prototype Encoder (MPE) submodule builds a learnable material prototype library based on the spectral reflectance patterns of materials:

[0081] ;

[0082] in, For the spectral reflection mode of the k-th type of material, the material prototype can include reflective film, metal foil and frosted plastic, etc.

[0083] Step 2.2: Extract pixel features from the original low-light image, and perform feature matching between the pixel features and each material in the material prototype library using attention weights to obtain the material attention weight for each material.

[0084] ;

[0085] in, Let x be the attention weight for pixel x belonging to material k; To represent along the material prototype dimension The normalized exponential function is used to map the matching scores of each material to an attention weight distribution that is non-negative and sums to 1; Let x be the local material encoding feature of pixel x; x is the pixel variable; T is the transpose of the matrix. is a learnable feature matching matrix.

[0086] Step 2.3: Process the multiple material attention weights and their corresponding spectral reflectance modes of the pixel to obtain the pixel-level spectral reflectance vector:

[0087] ;

[0088] in, is the spectral reflectance of pixel x; k is the material prototype variable; K is the total number of material prototypes.

[0089] To ensure that the spectral reflectance conforms to physical constraints, spectral regularization is added to the spectral reflectance of pixels through the SpectralEstimation Block (SEB), and the loss function of the Material Aware Spectral Estimation Network (MSEN) is constructed to obtain the spectral regularization loss:

[0090] ;

[0091] in, This is the loss due to spectral regularization. The square of the L2 norm; β represents the gradient along the spectral wavelength direction, constraining spectral continuity; β is the energy constraint weight. Let x be the spectral reflectance of pixel x; The reference spectrum for pixel x is obtained using a BRDF or Lambertian model. The first term of the spectral regularization loss is used to avoid spectral discontinuities and unnatural fluctuations, while the second term ensures that the spectrum conforms to the real material and that the estimated spectral distribution satisfies physical constraints (such as energy conservation and wavelength continuity).

[0092] In some embodiments, a semantic feedback self-correction mechanism (SFA) is also included to update the material prototype library based on semantic gradient feedback, using semantic gradient feedback obtained from the SRN. Update the material prototype library during training:

[0093] ;

[0094] in, and These are the spectral reflection modes of the k-th material after the t-th and t+1-th updates, respectively; The self-calibrating learning rate for the material prototype library; To mitigate semantic consistency loss, emphasis is placed on preserving the structural fidelity of text, barcodes, logos, and other elements. The spectral reflectance mode of the k-th material in the initial material prototype library; Let be the partial derivative of the semantic consistency loss with respect to the k-th type of material prototype in the material prototype library. By updating the material prototype library, material estimation can be made more accurate in semantically critical regions; material spectral estimation can be made semantically interpretable; and a semantically driven material modeling closed loop can be formed.

[0095] The Material Aware Spectral Estimation Network (MSEN) can achieve several effects through collaboration with other modules, including temporal collaboration with the Mamba-Backbone: MSEN obtains the illumination state vector output from Mamba. Extracting illumination information from the current frame to assist in intensity correction for spectral estimation; physical coordination with the PCU: the spectral reflectance vector output by MSEN. As the illumination compensation kernel function of PCU illumination reconstruction The input prior makes the compensation process more reasonable; and the semantic feedback of SRN: the semantic gradient feedback of SRN. It works inversely to MSEN's material prototyping library, enabling adaptive correction of spectral estimates in key semantic regions. It also works in conjunction with PDR's differentiable rendering. Including reflectivity as a parameter in the differentiable rendering equation improves the accuracy of reflection modeling.

[0096] This invention implements a low-light enhancement material estimation method through a Material Aware Spectral Estimation Network (MSEN). It can model differentiable material spectra and embed differentiable spectral reflectance estimation into a low-light visual network, making the material recovery process learnable and capable of backpropagation. A dynamically updatable material library is constructed through a learnable material prototype library (MPB), and feature self-correction is achieved through semantic feedback. Semantic-driven material adaptive optimization (SFA) allows semantic gradient signals to directly adjust material spectral parameters, forming a cross-modal optimization closed loop. Through a cross-module spectral sharing mechanism, the output spectral features can simultaneously serve both physical illumination compensation and differentiable rendering modules, improving overall consistency.

[0097] By combining spectral estimation with physical constraints, color drift and over-compensation of illumination are avoided; semantic feedback guides the adaptive correction of material spectra, making the recovery of key regions more accurate; all processes can be trained end-to-end and work seamlessly with Mamba, PCU, and SRN.

[0098] Step 3: Perform physical illumination compensation on the illumination state vector, spectral reflectance vector, and semantic heatmap using the physical illumination reconstruction module to obtain the reconstructed illumination distribution and the reconstructed enhanced image.

[0099] The Physics-aware Compensation Unit (PCU) is a reconstruction unit based on a semantically modulated differentiable physical illumination compensation mechanism, designed to restore the physical consistency of image brightness, color, and material reflection under low-light conditions. This is achieved by combining illumination state vectors from the Mamba-Backbone. Spectral reflectance of MSEN And introduce semantic heatmaps Together, they drive the adaptive update of the illumination compensation filter kernel function, enabling fine-grained enhancement and reflection modeling of low-light images.

[0100] Step 3.1: Based on the original low-light image and surface reflectivity, construct a physical illumination imaging model within a differentiable framework;

[0101] ;

[0102] in, L(x) represents the pixel values ​​of the original low-light image; L(x) represents the incident light distribution; and R(x) represents the surface reflectance, a spectral reflectance vector provided by the MSEN module. Mapped to obtain; (x) represents the photosensitive noise term.

[0103] Step 3.2: To achieve illumination compensation, the PCU defines a learnable illumination compensation kernel function. Based on the illumination state vector and spectral reflectance vector, the illumination compensation kernel function can be obtained by adaptively adjusting the filter kernel parameters.

[0104] Illumination compensation kernel function in PCU Employing a differentiable convolution kernel, its adaptively adjusted filter kernel parameter λ is determined by the illumination state output from the Mamba module. Together with the spectral reflectance vector of the MSEN module, it determines:

[0105] ;

[0106] in, This is an adaptive mapping network used to adjust the compensation kernel shape according to the lighting conditions and material characteristics of different frames. Through end-to-end backpropagation training, the filter kernel parameters are continuously learned, thus forming a temporally continuous and differentiable lighting compensation process. The adaptive update mechanism of this lighting compensation kernel function makes the lighting compensation kernel adaptive to material and spectral responses; the lighting changes between different time frames can be kept consistent through state vector smoothing constraints; and the use of semantic feedback to participate in kernel parameter adjustment can enhance the brightness recovery of key areas.

[0107] Step 3.3: Based on the physical illumination imaging model and the illumination compensation kernel function, the illumination distribution is adaptively reconstructed to form an illumination compensation framework, resulting in the reconstructed illumination distribution:

[0108]

[0109] in, To reconstruct the light distribution; This represents the distribution of incident light. This is a differentiable illumination compensation kernel function, controlled by the parameter λ, which is an adaptively updated filter kernel parameter during training. This filter kernel parameter is automatically adjusted during backpropagation based on the illumination reconstruction error, making the enhanced image more consistent with physical illumination characteristics.

[0110] Step 3.4: The PCU assigns differentiated compensation intensities to different regions using a semantic modulation illumination compensation mechanism combined with the semantic heatmap of the SRN, thus obtaining semantic modulation weights.

[0111] ;

[0112] in, For semantic modulation weights; Use the Sigmoid activation function; Learnable convolutional weights; This is a semantic heatmap representing the semantic importance of pixel x. The larger the value, the more critical the semantic meaning (e.g., text, logo, barcode, etc.).

[0113] Step 3.5: Based on semantic modulation weights and illumination compensation kernel functions, the original low-light image is processed to obtain the reconstructed and enhanced image.

[0114] ;

[0115] in, To reconstruct and enhance the image; These are the pixel values ​​of the original low-light image; For semantic modulation weights; This is the illumination compensation kernel function; Let be the illumination state vector at time t; Let x be the spectral reflectance of pixel x; This is a semantic heatmap of pixel x. A semantic modulation illumination compensation strategy ensures that the enhancement results are both visually natural and physically interpretable.

[0116] The PCU not only performs illumination compensation but also forms part of a semantic-physical dual closed-loop mechanism. It performs modulation compensation on the semantic heatmap through forward propagation and applies the illumination consistency error back to the SRN through back propagation to update the semantic map, thus obtaining the illumination consistency loss of the physical illumination reconstruction module.

[0117] ;

[0118] in, This is due to the loss of uniformity in illumination. The square of the L2 norm; To reconstruct the light distribution; The reference is the ideal illumination distribution generated by the illumination estimation network; γ is the weight of the temporal smoothing constraint. It is an L1 norm; Temporal difference for enhancing consecutive frames (to prevent video flicker); This mechanism enhances images temporally. It accurately restores the brightness of semantic regions, ensures brightness continuity across frames, and simultaneously optimizes physical and semantic consistency.

[0119] This module employs a differentiable illumination compensation mechanism based on semantic heatmap modulation and semantic modulation weights to achieve priority restoration of semantic regions. Its illumination compensation kernel function is jointly determined by illumination state and material features, ensuring illumination-material consistency. It combines Mamba temporal states to construct temporal consistency constraints, achieving continuous frame brightness smoothing and flicker-free enhancement. At the same time, through a dual closed-loop feedback path of semantic and physical modules, it improves semantic consistency and physical realism. In addition, it has independent deployment capabilities and can be used as a physical illumination restoration unit to adapt to other visual tasks such as detection and segmentation.

[0120] Step 4: The reconstructed illumination distribution, spectral reflectance vector, and semantic heat are optimized using the physically differentiable rendering module to obtain an enhanced rendering image.

[0121] Physically Differentiable Rendering (PDR) is a semantically driven rendering mechanism based on differentiable reflection modeling and spectral reconstruction. It is used to further restore the reflection details and visual consistency of an image after illumination compensation. Through a differentiable physically determined lighting rendering equation, it embeds illumination estimation results, spectral reflectance parameters, and semantic weights into the rendering process, thereby achieving interpretability-enhanced reconstruction under joint semantic and physical constraints. PDR maps the physical illumination propagation process of a real scene into a differentiable rendering operator, enabling the learning of optimal illumination and reflection function parameters during backpropagation, achieving semantically consistent high-fidelity visual restoration. The traditional rendering equation is:

[0122] ;

[0123] in, Let x be the observed intensity of pixel x; For the direction of origin Incident light illumination; is the reflection function; n is the direction of the surface normal.

[0124] To accommodate the learnability and gradient optimization capabilities of neural networks, the traditional rendering equation is made differentiable, allowing rendering operators to process the reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap to obtain an enhanced rendering image.

[0125] ;

[0126] in, Enhance the image for rendering; It is a differentiable rendering operator, and the parameter θ is learnable; The reconstructed illumination distribution obtained from PCU; Let x be the spectral reflectance of pixel x; This is a semantic heatmap of pixel x. The differentiable rendering function implements end-to-end differentiateable joint lighting-reflection modeling within a deep network framework, providing physical consistency and interpretability support for the system.

[0127] To achieve dynamic adaptability in physically based rendering, PDR introduces a Differentiable Reflectance Kernel Update (DRKU) mechanism, which continuously optimizes the rendering kernel parameters during training.

[0128] ;

[0129] in: Let be the rendering kernel parameters for the t-th iteration; For the updated rendering kernel parameters; The learning rate is used to render kernel parameters; This results in a loss of rendering consistency. The gradient is related to the rendering kernel parameters and is driven by both rendering error and semantic error.

[0130] The rendering loss of the physically differentiable rendering module includes rendering consistency loss, semantic consistency loss, and temporal consistency loss. Rendering consistency loss is defined as:

[0131] ;

[0132] in, This results in a loss of rendering consistency. The square of the L2 norm; The reconstructed and enhanced image of pixel x output by the PCU; Reconstructed and enhanced image output by PCU; The image is an enhanced rendering image after PDR rendering; α and β are the first and second rendering consistency weight coefficients, respectively; To assess the structural similarity between reconstructed and rendered enhanced images; This is a semantic consistency constraint that ensures the rendering result remains aligned with the semantic structure of the SRN. This mechanism guarantees that the rendering kernel dynamically converges under the combined constraints of lighting, materials, and semantics.

[0133] To enhance the rendering quality of semantically critical regions (such as brand text, barcodes, and patterns), this invention proposes Semantic Rendering Consistency (SRC) optimization. The semantic consistency loss is:

[0134] ;

[0135] in, For semantic consistency loss; x is a pixel variable; The semantic weights are generated from the SRN heatmap; It is an L1 norm; Use a specular reference image or pseudo-GT; semantic consistency loss reduces rendering errors in critical areas of PDR, achieving semantic-physical dual consistency.

[0136] To avoid common issues in low-light sequence enhancement such as reflection flicker and brightness jumps, PDR introduces a timing consistency loss:

[0137] ;

[0138] in, This results in a loss of time-series consistency. and , respectively, are the rendered enhanced images at times t and t-1; γ is the temporal smoothing constraint weight; To reconstruct the difference in illumination distribution over time, and to suppress illumination jumps; To reconstruct the lighting distribution, PDR maintains the continuity and stability of rendering in video scenes through temporal state vector constraints provided by Mamba-Backbone.

[0139] By embedding the traditional physical rendering process into a deep differentiable model through differentiable rendering operators, it supports self-learning of lighting and reflection functions during backpropagation. Through the Differentiable Reflection Kernel Update (DRKU) mechanism, the reflection kernel parameters are dynamically adjusted to achieve semantic-lighting-reflection co-optimization. Semantic Consistency Rendering Optimization (SRC) is used to enhance the rendering effect of key areas and ensure structural fidelity. At the same time, a temporal rendering stabilization mechanism is built using Mamba temporal states to maintain smooth consistency of lighting and reflection across multiple frames. Furthermore, through semantic-physical closed-loop control, the rendering output guides the semantic module update in reverse, forming an interpretable physical-semantic co-engineering system.

[0140] Step 5: Extract semantic features from the rendered enhanced image using the semantic graph prior enhancement module to generate a semantic heatmap. Feed the semantic heatmap back to the temporal state modeling perceptual encoder, the material-aware spectral estimation network, the physical illumination reconstruction module, and the physically differentiable rendering module to achieve semantic-physical dual closed-loop optimization and obtain the final optimized enhanced image.

[0141] The Semantic Reconstruction Network (SRN) is a core control unit with a semantic-physical dual closed loop. It extracts semantic information from the rendered and enhanced image and intermediate physical features, generates semantic graph priors and semantic heatmaps, and transmits these semantic feedback signals in real time to the Mamba, MSEN, PCU, and PDR modules to modulate their update and optimization processes. It not only achieves the recognition and enhancement of semantic information but also constructs a semantically-guided physical feedback (SPF) mechanism, enabling image reconstruction to possess self-learning and semantic consistency constraints.

[0142] In low-light scenarios, food packaging images often suffer from semantic feature blurring, edge breakage, and loss of text regions. SRN utilizes semantic graph modeling and attention heatmap mechanisms to perform semantic perception reconstruction of the augmentation results at the structural level. It extracts multi-layer semantic graph priors from the augmented image and generates dynamic semantic heatmaps based on differentiable semantic consistency constraints, enabling the reverse control of the physical layer by semantic information.

[0143] Step 5.1: The SRN processes the rendered and enhanced image using a learnable semantic embedding network to obtain a semantic feature map. .

[0144] Step 5.2: Based on the semantic feature map, construct the semantic graph prior G=(V,E); the semantic graph includes node features and semantic similarity edges between nodes. The node features and semantic similarity edges between nodes are defined as follows:

[0145] ;

[0146] ;

[0147] in, The semantic node features of pixel x; and These are the semantic feature maps for pixels x and y, respectively. Edges representing semantic similarity between nodes; and These are the learnable node mapping function and edge mapping function, respectively.

[0148] Step 5.3: Process the semantic graph prior G using a graph convolutional network to update node features and obtain the global semantic representation.

[0149] ;

[0150] in, For semantic global representation; GCN is a graph neural network; V is the set of semantic node features; E is the set of semantic similarity edges.

[0151] Step 5.4: SRN processes the global semantic representation through an attention mechanism to generate a semantic heatmap.

[0152] ;

[0153] in, For pixel x, a semantic heatmap; Use the Sigmoid activation function; These are learnable semantic global weights; This is a learnable semantic global bias. The heatmap reflects the importance of each pixel in the semantic structure and is used to modulate the computation process of other modules (Mamba, MSEN, PCU, PDR).

[0154] SRN not only generates semantic features but also participates in the backpropagation optimization of the temporal state modeling perceptual encoder, material-aware spectral estimation network, physically-oriented lighting reconstruction module, physically differentiable rendering module, and semantic graph prior enhancement module, constructing a semantic-physical co-optimization (SPC) mechanism. By collecting rendering consistency loss from the PDR module, lighting consistency loss from the PCU, and temporal stability feedback from Mamba, SRN performs semantic feature self-correction during backpropagation, obtaining the semantic feature self-correction loss used by the semantic graph prior enhancement module for joint optimization.

[0155] ;

[0156] in, For semantic feature self-correction loss; For semantic recognition and reconstruction loss; , and These are the weight coefficients for the first, second, and third semantic features, respectively. This is due to the loss of uniformity in illumination. This results in a loss of rendering consistency. The loss is for temporal consistency. This joint optimization achieves bidirectional information flow and consistency constraints from semantics to physical and back again.

[0157] The SRN module includes an Adaptive Semantic Feedback (ASF) regulator, which dynamically controls the semantic feedback strength based on the system's operating state. Let the feedback strength of the current module be... The update rule for the feedback strength of the semantic graph prior enhancement module is as follows:

[0158] ;

[0159] in, and , respectively, represent the feedback intensity at times t and t+1; η is the feedback intensity update step size coefficient, used to control the adjustment magnitude of the semantic feedback intensity between adjacent times; mean( () indicates the loss of illumination uniformity. Loss of rendering consistency The arithmetic mean is used to comprehensively characterize the overall error level of the physical module; δ is the error threshold of the physical module. When the error of the physical module is higher than the threshold δ, the SRN automatically enhances the semantic feedback strength to accelerate the alignment of illumination and semantics; conversely, it weakens semantic control to maintain visual naturalness. Adaptive semantic feedback adjustment enables adaptive dynamic scheduling of the semantic-physical feedback loop, giving the system online learning and self-stabilization capabilities.

[0160] The semantic graph prior enhancement module constructs a semantic-physical dual closed-loop optimization system. It achieves system-level adaptive enhancement by modulating illumination, reflection, and material estimation with semantic information. Combined with a differentiable semantic graph prior generation mechanism, the semantic graph structure is embedded into the physical enhancement network to construct an end-to-end learnable semantic graph. The adaptive semantic feedback modulator (ASF) dynamically adjusts the semantic feedback intensity to improve system stability and convergence speed. At the same time, it integrates multi-source joint optimization objective (SPC) to integrate semantic, illumination, rendering, and temporal multi-dimensional constraints to ensure semantic and physical consistency. It also has the advantage of interpretability, which can directly reflect the model's focus areas and enhancement logic through semantic heatmap visualization.

[0161] Step 6: Perform image recognition on the final optimized and enhanced image to obtain the food packaging recognition result.

[0162] Example 1

[0163] The low-light food packaging image recognition method proposed in this invention is based on a semantic-physical dual closed-loop structure, comprising five core modules: Mamba-Backbone, MSEN, PCU, PDR, and SRN. The system operation flow and functional implementation of this invention are described below with reference to specific embodiments.

[0164] (1) System input stage

[0165] The system first receives a sequence of low-light food packaging images from different devices and inputs it into a temporal state modeling perceptual encoder (Mamba-Backbone). This module captures the dynamic changes in illumination within the image sequence through temporal state modeling and outputs an illumination state vector. This provides timing information for subsequent lighting compensation and rendering.

[0166] (2) Material spectral estimation stage

[0167] The Material Aware Spectral Estimation Network (MSEN) receives the illumination state vector output from the temporal state modeling perceptual encoder (Mamba-Backbone) and the input image to estimate the spectral reflectance of food packaging materials. It generates pixel-level spectral reflectance vectors through a material prototype library and a semantic feedback self-correction mechanism. To depict the surface reaction characteristics of the packaging to light.

[0168] (3) Physical illumination reconstruction stage

[0169] The Physics-aware Compensation Unit (PCU) receives illumination during the illumination process. spectral vector With semantic heatmap Then, physical illumination compensation is performed. A scalable illumination filter kernel is then applied. Based on material and semantic information, the system adaptively updates and generates reconstructed and enhanced images. and reconstruction of light distribution This ensures the physical consistency of the image.

[0170] (4) Differentiable rendering stage

[0171] The Physically Differentiable Rendering (PDR) module optimizes the differentiable rendering of reconstructed and enhanced images by jointly driving the differentiable rendering operator using lighting, reflection, and semantic features. This achieves semantically driven reflection reconstruction and rendering consistency constraints, outputting enhanced rendering images. .

[0172] (5) Semantic graph prior enhancement stage

[0173] The Semantic Reconstruction Network (SRN) module, from the Semantic Graph Prior Enhancement Module, Extract multi-layer semantic features to generate semantic graph prior G(V,E) and semantic heatmap. The semantic feedback signal is then transmitted back to the Mamba, MSEN, PCU, and PDR modules to achieve semantic-physical dual closed-loop optimization.

[0174] (6) Image recognition stage

[0175] Image recognition technology is used to process the final optimized and enhanced image to obtain the identification results of the food packaging. This includes information such as the food name, type, and weight.

[0176] Through the above process, the system of the present invention realizes closed-loop optimization of the entire process from illumination modeling, material perception, physical compensation to semantic feedback, which significantly improves the recognizability and structural fidelity of low-light images and is suitable for various application scenarios such as automatic detection of food packaging, product recognition, and label text recovery.

[0177] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for food packaging image recognition in low-light environments, characterized in that, include: An illumination state space modeling method is used to model the original low-light image and semantic heatmap using a temporal state modeling perceptual encoder, resulting in an illumination state vector, including: A semantic modulation state-space modeling method is used to establish an illumination state-space model to capture temporal illumination changes and obtain the illumination state vector; the illumination state-space model is as follows: ; in, Let be the illumination state vector at time t; For semantic modulation state update function; The original low-light image at time t; The illumination state vector at the previous time t-1; A semantic heatmap; A material-aware spectral estimation network is used to match the original low-light image with a spectral prototype library of food packaging materials to obtain a spectral reflectance vector, including: A material prototype library is built based on the spectral reflectance patterns of the materials; Extract pixel features from the original low-light image, and perform feature matching between the pixel features and each material in the material prototype library to obtain the material attention weight of the pixel belonging to each material: ; in, Let x be the attention weight for pixel x belonging to material k; Let x be the local material encoding feature of pixel x; x is the pixel variable; T is the transpose of the matrix. A learnable feature matching matrix; For the spectral reflectance mode of the k-th material; This is used to normalize the model's output to obtain the weights for each material; By processing the multiple material attention weights of the pixel and their corresponding spectral reflectance modes, a pixel-level spectral reflectance vector is obtained: ; in, is the spectral reflectance of pixel x; k is the material prototype variable; K is the total number of material prototypes; It also includes updating the material prototype library based on semantic gradient feedback: ; in, and These are the spectral reflection modes of the k-th material after the t-th and t+1-th updates, respectively; The self-calibrating learning rate for the material prototype library; This represents the semantic consistency loss. For the spectral reflectance mode of the k-th material; Let be the partial derivative of the semantic consistency loss with respect to the k-th type of material prototype in the material prototype library; The physical illumination reconstruction module performs physical illumination compensation on the illumination state vector, spectral reflectance vector and semantic heatmap to obtain the reconstructed illumination distribution and the reconstructed enhanced image. The reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap are optimized by using a physically differentiable rendering module to obtain an enhanced rendering image. The semantic graph prior enhancement module extracts semantic features from the rendered enhanced image to generate a semantic heatmap. This heatmap is then fed back to the temporal state modeling perceptual encoder, material-aware spectral estimation network, physically-based illumination reconstruction module, and physically differentiable rendering module, achieving semantic-physical dual-loop optimization. Finally, an optimized enhanced image is generated, and this image is used for recognition to obtain food packaging recognition results, including: The semantic graph prior enhancement module processes the rendered enhanced image through a learnable semantic embedding network to obtain a semantic feature map; Based on the semantic feature graph, a semantic graph prior is constructed; the semantic graph includes node features and semantic similarity edges between nodes. The semantic graph prior is processed by a graph convolutional network to update node features and obtain a global semantic representation. The semantic graph prior enhancement module processes the global semantic representation through an attention mechanism to generate a semantic heatmap. By collecting rendering consistency loss from the physically differentiable rendering module, illumination consistency loss from the physically illuminated reconstruction module, and temporal stability feedback from the temporal state modeling perceptual encoder, the semantic graph prior enhancement module performs semantic feature self-correction during backpropagation, resulting in a semantic feature self-correction loss used for joint optimization. The semantic graph prior enhancement module includes an adaptive feedback regulator to dynamically control the semantic feedback intensity based on the system's operating state.

2. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The physical illumination reconstruction module performs physical illumination compensation on the illumination state vector, spectral reflectance vector, and semantic heatmap, including: Based on the original low-light image and surface reflectivity, a physical illumination imaging model is constructed: ; in, denoted as , where is the pixel value of the original low-light image; L(x) represents the incident light distribution; and R(x) represents the surface reflectivity. (x) represents photosensitive noise; Based on the illumination state vector and spectral reflectance vector, the filter kernel parameters are adaptively adjusted to obtain the illumination compensation kernel function; the filter kernel parameters are: ; Where λ is the filter kernel parameter; It is an adaptive mapping network; Let be the illumination state vector at time t; Let x be the spectral reflectance of pixel x; Based on the physical illumination imaging model and illumination compensation kernel function, an illumination compensation framework is constructed to obtain the reconstructed illumination distribution: ; in, To reconstruct the light distribution; This represents the distribution of incident light. This is the illumination compensation kernel function; A semantic heatmap; Differential compensation intensities are assigned to different regions of the semantic heatmap to obtain semantic modulation weights: ; in, For semantic modulation weights; Use the Sigmoid activation function; These are the convolution weights; For pixel x, a semantic heatmap; Based on semantic modulation weights and illumination compensation kernel functions, the original low-light image is processed to obtain a reconstructed and enhanced image: ; in, To reconstruct and enhance the image; These are the pixel values ​​of the original low-light image; For semantic modulation weights; This is the illumination compensation kernel function; Let be the illumination state vector at time t; Let x be the spectral reflectance of pixel x; This is the semantic heatmap of pixel x.

3. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap are optimized for differentiability rendering using a physically differentiable rendering module, including: The reconstructed illumination distribution, spectral reflectance vector, and semantic heatmap are processed using rendering operators to obtain an enhanced rendering image: ; in, Enhanced image rendering; For rendering operators; To reconstruct the light distribution; Let x be the spectral reflectance of pixel x; This is the semantic heatmap of pixel x.

4. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The semantic graph prior enhancement module extracts semantic features from the rendered enhanced image to generate a semantic heatmap, including: The node features and the semantic similarity edges between nodes are: ; ; in, The semantic node features of pixel x; and These are the semantic feature maps for pixels x and y, respectively. Edges representing semantic similarity between nodes; and These are the node mapping function and the edge mapping function, respectively; The semantic global representation is: ; in, For semantic global representation; GCN is a graph neural network; V is the set of semantic node features; E is the set of semantic similarity edges; The semantic heatmap is as follows: ; in, For pixel x, a semantic heatmap; Use the Sigmoid activation function; These are learnable semantic global weights; This is a learnable semantic global bias.

5. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The spectral regularization loss of the material-aware spectral estimation network is: ; in, This is the loss due to spectral regularization. The square of the L2 norm; β represents the gradient along the spectral wavelength direction; β is the energy constraint weight. Let x be the spectral reflectance of pixel x; is the reference spectrum for pixel x.

6. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The illumination consistency loss of the physical illumination reconstruction module is: ; in, This is due to the loss of uniformity in illumination. The square of the L2 norm; To reconstruct the light distribution; The ideal illumination distribution is represented by γ, which is the weight of the temporal smoothing constraint. It is an L1 norm; Temporal difference of the enhancement results for consecutive frames; For time-series image enhancement.

7. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The rendering loss of the physically differentiable rendering module includes rendering consistency loss, semantic consistency loss, and temporal consistency loss: The rendering consistency loss is: ; in, This results in a loss of rendering consistency. The square of the L2 norm; Reconstruct and enhance the image for pixel x; To reconstruct and enhance the image; To render an enhanced image; α and β are the first and second rendering consistency weight coefficients, respectively; To assess the structural similarity between reconstructed and rendered enhanced images; For semantic consistency constraints; The semantic consistency loss is: ; in, For semantic consistency loss; x is a pixel variable; For semantic weights; It is an L1 norm; For highlight reference image; The timing consistency loss is: ; in, This results in a loss of time-series consistency. and , respectively, are the rendered enhanced images at times t and t-1; γ is the temporal smoothing constraint weight; To reconstruct the difference in illumination distribution over time; To reconstruct the light distribution.

8. The food packaging image recognition method for low-light environments according to claim 1, characterized in that, The semantic feature self-correction loss of the semantic graph prior enhancement module is: ; in, For semantic feature self-correction loss; For semantic recognition and reconstruction loss; , and These are the weight coefficients for the first, second, and third semantic features, respectively. This is due to the loss of uniformity in illumination. This results in a loss of rendering consistency. This results in a loss of time-series consistency. The feedback strength of the semantic graph prior enhancement module is: ; in, and The feedback intensities at times t and t+1 are respectively; η is the feedback intensity update step size coefficient; mean( ) represents the arithmetic mean; δ represents the error threshold of the physical module.

Citation Information

Patent Citations

  • Low-light food package image recognition method, computer program product and terminal

    CN119785334A

  • Tree crown segmentation method based on CSAF framework

    CN121330303A