Intelligent food detection method, system, terminal and medium based on image recognition

By generating texture directional features through frequency domain phase analysis and spatial domain directional gradient, and combining frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree, the frequency domain aliasing problem of micro-texture and real defects in food surface defect detection is solved, and high-precision defect recognition is achieved.

CN120431571BActive Publication Date: 2025-09-26GUIZHOU SHIKEYUAN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510928671.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-26
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing food surface defect detection methods based on frequency domain analysis have difficulty in effectively distinguishing microtextures from real defects, resulting in decreased detection accuracy and the inability to effectively eliminate frequency domain aliasing effects.

Method used

Texture directional features are generated through frequency domain phase analysis and spatial domain directional gradient. Combined with the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree, mutually exclusive conflict areas are detected and spatial domain inverse transformation is performed to generate a de-aliased image, and finally defect recognition and judgment are performed.

Benefits of technology

It significantly improves the accuracy and reliability of food surface defect detection. Through multi-domain feature deep coupling and dynamic adaptive processing mechanism, it accurately captures the frequency domain aliasing characteristics of defects and background textures, solves the misjudgment problem in traditional methods, and ensures the integrity and robustness of defect feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431571B_ABST
    Figure CN120431571B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent food detection method and system, terminal and medium based on image recognition, which specifically relates to the field of food surface detection technology. The method is used to solve the problem in the prior art that it is difficult to distinguish between micro-textures on the food surface and real defects due to the frequency domain aliasing effect. By acquiring the food surface image and generating a frequency domain phase distribution map, the phase consistency feature and the spatial domain texture directional feature are extracted; the frequency domain response function is reconstructed based on the phase-texture dynamic correlation weight to generate an optimized frequency domain map; the frequency domain singular vector distribution entropy and phase coherence attenuation are combined to detect mutually exclusive conflicting areas, and a de-aliased image is obtained through spatial domain inverse transformation; defect recognition and judgment are realized based on adaptive morphological parameters and multi-feature classification; through cross-domain feature coupling analysis and dynamic parameter adjustment, the frequency domain aliasing interference under complex texture background is suppressed, and the sensitivity and anti-misjudgment ability of defect detection are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food surface detection, and more specifically, to an intelligent food detection method and system, terminal and medium based on image recognition. Background Art

[0002] In the field of food quality inspection, image recognition-based technologies have been widely used for surface defect detection. With the improvement of imaging device resolution, detailed information on food surface microtextures (such as the grooves on the surface of grains and the natural grain of nuts) can be captured with high precision. Existing methods often use frequency domain analysis to process images, separating noise from target features through frequency domain filtering to identify tiny defects (such as insect eggs and mold spots). However, food surface microtextures and tiny defects often exhibit similar energy distribution characteristics in the frequency domain, resulting in inherent technical bottlenecks in frequency domain analysis methods.

[0003] In the existing technology, defect detection methods based on frequency domain analysis have difficulty in effectively distinguishing the frequency domain characteristics of food surface microtextures and real defects. Since the energy distribution of microtextures and defects in the frequency domain highly overlaps, traditional filtering strategies that only rely on frequency domain amplitude information are prone to misjudge defect signals as background noise, or conversely, misjudge complex textures as defects, resulting in a significant decrease in detection accuracy and the inability to effectively eliminate the frequency domain aliasing effect of defects and background. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent food detection method and system based on image recognition, a terminal and a medium to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] The intelligent food detection method based on image recognition includes the following steps:

[0007] S1. Acquire a surface image of the food to be tested;

[0008] S2. performing frequency domain phase analysis on the surface image to generate a frequency domain phase distribution map of the surface image;

[0009] S3, extracting the phase consistency feature of the frequency domain phase distribution map, and generating the texture directional feature based on the spatial domain directional gradient of the surface image;

[0010] S4, reconstructing the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain map;

[0011] S5. Based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image, the mutually exclusive conflicting areas are detected and the spatial domain inverse transform is performed to obtain a de-aliased image;

[0012] S6. Perform defect recognition and determination based on the defect feature area in the de-aliased image.

[0013] In a preferred embodiment, obtaining a surface image of the food to be inspected includes:

[0014] Adjust the irradiation angle and polarization axis direction of the multi-angle light source so that the polarization axis direction of the light source is orthogonal to the normal direction of the food surface;

[0015] Under multi-angle light sources, the reflected light from the food surface is filtered through a polarizing filter to generate an initial image that suppresses reflection interference;

[0016] adjusting the exposure time of the image sensor based on the food surface transparency parameter, and capturing surface images under multi-angle polarized light according to the adjusted exposure time;

[0017] The surface image under multi-angle polarized light is checked for consistency in light intensity distribution. The target angle image is filtered according to the light intensity distribution consistency check result to generate the surface image of the food to be tested.

[0018] In a preferred embodiment, performing frequency domain phase analysis on the surface image to generate a frequency domain phase distribution map of the surface image includes:

[0019] Converting the surface image into a grayscale image, performing edge symmetry compensation processing on the grayscale image, and generating a preprocessed image;

[0020] Performing a two-dimensional discrete Fourier transform on the preprocessed image, extracting the phase component of the complex spectrum matrix, and generating an initial phase distribution map;

[0021] The phase threshold range is adjusted based on the local contrast parameter of the surface image, and the phase values ​​exceeding the threshold range in the initial phase distribution map are truncated and corrected;

[0022] According to the spatial frequency distribution characteristics of the surface image, the truncated and corrected phase distribution map is subjected to frequency band weighted fusion to generate a frequency domain phase distribution map.

[0023] In a preferred embodiment, extracting the phase consistency feature of the frequency domain phase distribution image and generating the texture directionality feature based on the spatial domain directional gradient of the surface image includes:

[0024] performing phase gradient direction analysis on the frequency domain phase distribution diagram to generate a phase gradient direction matrix;

[0025] Construct a phase gradient direction histogram based on the phase gradient direction matrix and extract phase consistency features;

[0026] The texture directional features are calculated based on the spatial directional gradient of the surface image. The spatial directional gradient is detected along the horizontal and vertical directions using the Sobel operator.

[0027] Normalizing the phase consistency feature and the texture directionality feature to generate a standardized phase consistency feature and a standardized texture directionality feature;

[0028] The phase-texture correlation matrix is ​​constructed based on the joint distribution of the normalized phase consistency features and the normalized texture directionality features.

[0029] In a preferred embodiment, reconstructing the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate the optimized frequency domain map includes:

[0030] A dynamic correlation weight matrix is ​​generated based on the joint distribution of phase consistency features and texture directionality features. The dynamic correlation weight matrix is ​​calculated by multiplying the joint distribution probability of the phase-texture correlation matrix and the normalized phase consistency feature.

[0031] Decomposing the initial frequency domain response function into a low-frequency band response function, a mid-frequency band response function and a high-frequency band response function;

[0032] The amplitude of each frequency band response function is adjusted according to the dynamic correlation weight matrix. The low frequency band response function is adjusted by weighted average method, the mid frequency band response function is adjusted by gradient direction constraint, and the high frequency band response function is adjusted by threshold truncation.

[0033] The adjusted low-frequency band response function, mid-frequency band response function and high-frequency band response function are superimposed to generate an optimized frequency domain graph.

[0034] In a preferred embodiment, based on optimizing the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the frequency domain image, mutually exclusive conflicting regions are detected and spatial inverse transformation is performed to obtain a de-aliased image, including:

[0035] Divide the optimized frequency domain image into overlapping sub-blocks, perform SVD decomposition on each sub-block, extract the left singular vector matrix and calculate the information entropy of the left singular vector matrix;

[0036] The optimized frequency domain image is divided into multiple frequency bands, the phase coherence between adjacent frequency bands is calculated by complex cross-correlation, and the coherence decay rate of each frequency band is statistically analyzed;

[0037] Marking a sub-block whose information entropy exceeds a preset information entropy threshold or a region whose coherence decay rate exceeds a preset decay rate threshold as a mutually exclusive conflict region and generating a mutually exclusive conflict region mask;

[0038] Perform spatial inverse transform processing on the frequency domain data covered by the mutually exclusive conflicting area mask to generate an anti-aliased image.

[0039] In a preferred embodiment, performing defect recognition and determination based on the defect feature area in the de-aliased image includes:

[0040] Construct a gray-level co-occurrence matrix for the de-aliased image and extract the texture contrast and energy characteristics of the defect area;

[0041] The geometric features of the defect area are extracted based on the morphological opening operation to generate a set of candidate defect areas. The structural element of the morphological opening operation is a circular kernel, and the radius of the circular kernel is dynamically adjusted according to the minimum size of the defect.

[0042] Input the texture contrast, energy features and geometric features in the candidate defect area set into the pre-trained support vector machine classifier, and output the defect category probability of each area in the candidate defect area set;

[0043] Generate preliminary defect determination results based on defect category probability and dynamic determination threshold. The dynamic determination threshold is dynamically adjusted based on the false detection rate distribution of historical defect samples.

[0044] The spatial continuity of the preliminary defect determination results is checked, and the final defect feature map is generated after the isolated abnormal areas in the candidate defect area set are eliminated.

[0045] In another aspect, the present invention provides an intelligent food detection system based on image recognition, comprising the following modules:

[0046] An image acquisition module, used to acquire a surface image of the food to be inspected;

[0047] A phase analysis module is used to perform frequency domain phase analysis on the surface image and generate a frequency domain phase distribution map of the surface image;

[0048] Texture analysis module, used to extract phase consistency features of the frequency domain phase distribution map and generate texture directional features based on the spatial domain directional gradient of the surface image;

[0049] A function reconstruction module is used to reconstruct the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain map;

[0050] The mixed image generation module is used to detect the mutually exclusive conflict area and perform spatial inverse transformation based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image to obtain the anti-aliased image;

[0051] The defect recognition module is used to perform defect recognition and judgment based on the defect feature area in the de-aliased image.

[0052] On the other hand, the present invention provides an intelligent food detection terminal based on image recognition, including: a processor, a memory, and a program or instruction stored in the memory and runnable on the processor. When the program or instruction is executed by the processor, an intelligent food detection method based on image recognition is implemented.

[0053] On the other hand, the present invention provides an intelligent food detection medium based on image recognition, on which a program or instruction is stored. When the program or instruction is executed by a processor, an intelligent food detection method based on image recognition is implemented.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. Through the deep coupling and dynamic adaptive processing mechanism of multi-domain features, the accuracy and reliability of food surface defect detection are significantly improved. First, based on the correlation weight design of frequency domain phase distribution and spatial domain texture directionality, the frequency domain aliasing characteristics of defects and background textures are accurately captured, and the defect signals under micro-texture interference are effectively separated through frequency domain response function reconstruction. The dynamic fusion of phase consistency features and texture directionality features enables frequency domain analysis to deeply decouple the coupled frequency domain responses of complex textures and real defects, solving the misjudgment problem caused by a single frequency domain feature in traditional methods. At the same time, combined with the multi-dimensional criteria of frequency domain singular value distribution entropy and phase coherence attenuation, an accurate identification rule for mutually exclusive conflict areas is constructed, which suppresses the aliasing effect from the dual perspectives of frequency domain energy distribution and phase continuity, ensuring the integrity and robustness of defect feature extraction.

[0056] 2. A closed-loop optimization of the detection process is achieved through the adaptive recognition and classification mechanism of defect feature areas. The morphological parameter setting based on dynamic adjustment of the minimum defect size enables the geometric feature extraction process to adaptively match the surface characteristics of different food categories, avoiding detection bias introduced by artificial experience parameters. The multi-level progressive processing of frequency domain images and de-aliased images is optimized to form a systematic mapping from the original image to the defect feature map, effectively filtering out high-frequency noise interference while retaining subtle defect features. Through the organic integration of multi-domain feature collaborative analysis, dynamic weight fusion and adaptive parameter adjustment, a highly sensitive defect recognition path is established in complex texture backgrounds, significantly improving the generalization ability and engineering applicability of the detection system. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a flow chart of the intelligent food detection method based on image recognition of the present invention;

[0058] Figure 2This is a structural diagram of the intelligent food detection system based on image recognition of the present invention. DETAILED DESCRIPTION

[0059] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0060] Example 1: Figure 1 The present invention provides an intelligent food detection method based on image recognition, which includes the following steps:

[0061] S1. Acquire a surface image of the food to be tested;

[0062] S2. performing frequency domain phase analysis on the surface image to generate a frequency domain phase distribution map of the surface image;

[0063] S3, extracting the phase consistency feature of the frequency domain phase distribution map, and generating the texture directional feature based on the spatial domain directional gradient of the surface image;

[0064] S4, reconstructing the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain map;

[0065] S5. Based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image, the mutually exclusive conflicting areas are detected and the spatial domain inverse transform is performed to obtain a de-aliased image;

[0066] S6. Perform defect recognition and determination based on the defect feature area in the de-aliased image.

[0067] S1. Obtaining a surface image of the food to be tested, specifically implemented as follows:

[0068] The illumination angles and polarization axes of the multi-angle light sources are adjusted so that the polarization axis is orthogonal to the food surface normal. The multi-angle light sources are arranged in a circular array, and the diameter of the circular array is dynamically adjusted based on the food surface curvature radius. The adjustment formula is: Circular array diameter = food surface curvature radius × 1.5. For example, when inspecting spherical food with a surface curvature radius of 10 cm, the circular array diameter is set to 15 cm. The spacing between the light sources is adjusted based on the maximum side length of the food's circumscribed rectangle. The adjustment rule is: Light source spacing = maximum side length of the food's circumscribed rectangle × 0.3. The polarization axes of adjacent light sources are arranged in an alternating orthogonal pattern: the first light source has a polarization axis of 0 degrees, the second light source has a polarization axis of 90 degrees, the third light source has a polarization axis of 0 degrees, and so on. The illumination angles of the light sources are measured in real time by a laser rangefinder to measure the food surface normal. A motor is used to rotate the light source bracket to ensure that the polarization axis is orthogonal to the food surface normal.

[0069] Under multi-angle light sources, polarizing filters filter the light reflected from the food surface to generate an initial image that suppresses reflection interference. The polarization axis of the polarizing filter is driven by a stepper motor, maintaining synchronization and parallelism with the polarization axis of the corresponding light source. For example, when the light source's polarization axis is 0 degrees, the filter's polarization axis is rotated to 0 degrees by the motor. For transparent or translucent foods, internal refracted light interference is suppressed by adjusting the combination of light source intensity and filter angle. The specific adjustment rule is: light source intensity = baseline intensity × (1 - transparency parameter), where the transparency parameter is preset in a range of 0.1 (low transparency) to 0.9 (high transparency). The initial image generation process involves sequentially illuminating each light source and synchronously triggering the image sensor to capture a single-angle polarization image, ultimately obtaining a sequence of original images under multi-angle polarized light.

[0070] The image sensor's exposure duration is adjusted based on the food's surface transparency parameter, capturing surface images under multi-angle polarized light using the adjusted exposure duration. The surface transparency parameter is obtained by matching against a pre-stored database of food types. For example, when the food category "apple" is entered, a transparency parameter of 0.5 is automatically used. The exposure duration adjustment formula is: Exposure duration = Baseline exposure duration × (1 + Transparency parameter × 2), where the base exposure duration is calibrated based on the ambient light intensity. For example, at an ambient light intensity of 1000 lux, the base exposure duration is set to 10 milliseconds. For packaged food with a transparency parameter of 0.8, the exposure duration is adjusted to 10 × (1 + 0.8 × 2) = 26 milliseconds. Based on the adjusted exposure duration, the image sensor simultaneously captures surface images under multi-angle polarized light and verifies whether the exposure results meet dynamic range requirements. If overexposure or underexposure occurs, the exposure parameters are readjusted.

[0071] The light intensity distribution consistency of surface images under multi-angle polarized light is checked. The target angle images are filtered based on the results of the light intensity distribution consistency check to generate the surface image of the food to be tested. The light intensity distribution consistency check includes the following steps: First, the light intensity distribution histogram of each single-angle image is calculated, and the histogram grayscale is divided into 256 levels; second, the histogram intersection algorithm is used to calculate the similarity score of the histograms of the images at each angle; finally, the similarity threshold is set to 0.85, and target angle images with scores above the threshold are filtered. For example, if the histogram intersection score of a certain angle image is lower than 0.85 due to local reflection, it is judged as an abnormal image and is eliminated. The filtered target angle images are fused at the pixel level to generate the final surface image. The fusion rule is: for the same pixel position, the median of the light intensity values ​​of all target angle images is taken as the fusion result.

[0072] The calculation formula for the similarity score is: ;in, represents the similarity score; represents the grayscale histogram data extracted from the first surface image, where Indicates that the grayscale value in the first image is The number of pixels; represents the grayscale histogram data extracted from the second surface image, where Indicates that the grayscale value in the second image is The number of pixels; Represents the interval index of the grayscale histogram.

[0073] The multi-angle light sources are arranged in a circular array, with the polarization axes of adjacent light sources alternating orthogonally. The radius of the circular array is adaptively adjusted based on the food size, and the number of light sources is set based on the food surface area. The setting rule is: number of light sources = 8 + (food surface area / 100 square centimeters) × 8, with the result rounded to an even value. For example, when inspecting food with a surface area of ​​50 square centimeters, the number of light sources is 8 + (50 / 100) × 8 = 12, rounded up to the nearest 12. The light source driver dynamically adjusts the light intensity based on the food surface reflectivity. The adjustment rule is: light intensity = reference light intensity × (1 + reflectivity), where reflectivity is obtained through matching a pre-stored database or real-time measurement. For example, when inspecting food with a reflectivity of 0.3, the reference light intensity is 100 lumens, and the actual light intensity is adjusted to 100 × (1 + 0.3) = 130 lumens.

[0074] S2. Perform frequency domain phase analysis on the surface image to generate a frequency domain phase distribution diagram of the surface image. The specific implementation is as follows:

[0075] The surface image is converted to a grayscale image, and edge symmetry compensation is performed on the grayscale image to generate a preprocessed image. Grayscale conversion uses the weighted averaging method defined in the international standard ITU-R BT.601. The red channel weight is 0.299, the green channel weight is 0.587, and the blue channel weight is 0.114. The calculation formula is: the grayscale value is equal to the red channel value multiplied by 0.299, the green channel value multiplied by 0.587, and the blue channel value multiplied by 0.114. Edge symmetry compensation includes the following steps: First, the Sobel edge detection operator is used to calculate the edge gradient magnitude and direction of the grayscale image. Regions with abrupt changes in edge gradient direction are defined as break regions. Specifically, the gradient direction of adjacent pixels changes by more than 30 degrees and the gradient magnitude difference is greater than 50. Second, symmetry compensation is performed on pixels within two pixels on either side of the break region. The compensation coefficient is dynamically adjusted based on the break length. For each additional pixel in the break length, the compensation coefficient increases by 0.1. The compensated pixel grayscale value is the original value multiplied by (1 plus the compensation coefficient). For example, if the break length is 3 pixels, the compensation coefficient is 0.3. The grayscale value of the pixel at the center of the break is adjusted to 1.3 times the original value, the first adjacent pixel on the left is adjusted to 1.2 times, the second adjacent pixel on the left is adjusted to 1.1 times, and the same applies to the right. The compensated grayscale image has significantly improved edge continuity and reduced phase jumps.

[0076] A two-dimensional discrete Fourier transform (DFT) is performed on the preprocessed image to extract the phase component of the complex spectrum matrix and generate an initial phase distribution map. This DFT is implemented using the fast Fourier transform (FFT) algorithm. The input is the grayscale matrix of the preprocessed image, and the output is a complex spectrum matrix. Each element of the complex spectrum matrix contains both real and imaginary components. The phase component of the complex spectrum matrix is ​​obtained by calculating the inverse tangent function of the complex number pixel by pixel. The calculation formula is: the phase value is equal to the imaginary component divided by the inverse tangent of the real component. The result is expressed in radians and is limited to the range from -π to +π. The initial phase distribution map has the same size as the preprocessed image, with one phase value for each pixel. For example, if the preprocessed image resolution is 1024 pixels by 1024 pixels, the initial phase distribution map is a 1024-by-1024-pixel floating-point matrix with phase values ​​ranging from -3.1416 to 3.1416 radians.

[0077] The phase threshold range is adjusted based on the local contrast parameter of the surface image, and phase values ​​in the initial phase distribution map that fall outside the threshold range are truncated. The local contrast parameter is obtained by dividing the grayscale image into 8-by-8 pixel blocks and calculating the average absolute grayscale difference between all pairs of adjacent pixels in each block. Adjacent pixel pairs include horizontal, vertical, and diagonal pixels. The phase threshold range is dynamically adjusted based on the local contrast parameter. The adjustment rule is: the upper phase threshold is equal to the base threshold multiplied by (1 plus the local contrast parameter multiplied by 0.5), and the lower phase threshold is equal to the base threshold multiplied by (1 minus the local contrast parameter multiplied by 0.5).

[0078] The baseline threshold is preset based on the food surface material: π / 2 radians for smooth surfaces and π / 3 radians for rough surfaces. For example, when testing smooth meat, the baseline threshold is 1.5708 radians. If the contrast parameter of a local block is 0.2, the upper limit of the phase threshold is adjusted to 1.5708 times 1.1, which equals 1.7279 radians, and the lower limit is adjusted to 1.5708 times 0.9, which equals 1.4137 radians. For pixels in the initial phase distribution map with a phase value greater than the upper threshold or less than the lower threshold, their phase values ​​are corrected to the corresponding threshold boundary values. This corrected phase distribution map eliminates abnormal phase changes caused by local high-contrast areas.

[0079] Based on the spatial frequency distribution characteristics of the surface image, the truncated and corrected phase distribution map is subjected to band-weighted fusion to generate a frequency-domain phase distribution map. The spatial frequency distribution characteristics are extracted using the spectral energy density function. The specific steps are: Gaussian filtering the amplitude component of the complex spectrum matrix to generate three frequency bands: low, mid, and high. The standard deviation of the Gaussian filtering is 10 pixels for the low band, 5 pixels for the mid band, and 2 pixels for the high band. The energy density of a frequency band is calculated as the sum of the squared amplitude values ​​of all pixels within the band divided by the total number of pixels in the band.

[0080] The weight distribution rule for band-weighted fusion is as follows: the low-band weight is equal to the ratio of the low-band energy density to the total energy multiplied by 0.6, the mid-band weight is equal to the ratio of the mid-band energy density to the total energy multiplied by 0.3, and the high-band weight is equal to the ratio of the high-band energy density to the total energy multiplied by 0.1. For example, if the low-band energy density is 1000, the mid-band is 800, the high-band is 200, and the total energy is 2000, then the low-band weight is (1000 / 2000) × 0.6 = 0.3, the mid-band weight is (800 / 2000) × 0.3 = 0.12, and the high-band weight is (200 / 2000) × 0.1 = 0.01. After normalization, the weights are adjusted to 0.698 for the low-band, 0.279 for the mid-band, and 0.023 for the high-band.

[0081] The truncated and corrected phase distribution maps are then weighted and superimposed according to the frequency band weights to generate the final frequency domain phase distribution map. The superposition formula is: the frequency domain phase distribution map is equal to the low-band phase map multiplied by the low-band weight, the mid-band phase map multiplied by the mid-band weight, and the high-band phase map multiplied by the high-band weight. This fused frequency domain phase distribution map preserves the key phase characteristics of the food surface texture while suppressing high-frequency noise interference.

[0082] The local contrast parameter is calculated by averaging the grayscale differences between adjacent pixels in the image. Adjacent pixel pairs include those adjacent horizontally, vertically, and diagonally. Pixel pairs with an absolute grayscale difference of less than 5 are excluded from the calculation to avoid noise interference. For example, within a local block of 8 pixels by 8 pixels, the absolute grayscale difference between adjacent pixels horizontally is 10, vertically is 8, and diagonally is 12. The local contrast parameter is (10 plus 8 plus 12) divided by 3, which equals 10. Spatial frequency distribution features are extracted using the spectral energy density function. The calculation of the spectral energy density function relies on the amplitude component of the complex spectrum matrix. The amplitude is equal to the square root of the square of the real component plus the square of the imaginary component. For example, if the real component of a complex element is 3 and the imaginary component is 4, the amplitude is 5.

[0083] The band-weighted fusion process requires an image processing library that supports floating-point operations, such as OpenCV 4.5 and above. The standard deviation parameter of the Gaussian filter is preset based on the frequency band characteristics. A standard deviation of 10 pixels in the low-frequency band corresponds to the spatial wavelength range covering the low-frequency areas of the image, 5 pixels in the mid-frequency band corresponds to moderate texture detail, and 2 pixels in the high-frequency band corresponds to minor defect features. The frequency-domain phase distribution map after weighted fusion is spatially smoothed using a bilinear interpolation algorithm to eliminate phase discontinuities at the band boundaries.

[0084] For extreme input conditions, such as completely black or white images, Phase Analysis automatically skips the truncation correction and band fusion steps, directly outputting the initial phase distribution map and recording the anomaly type in the log. If insufficient hardware resources are detected, such as GPU memory below 2GB, the system automatically switches to block processing mode, dividing the image into 512-pixel by 512-pixel sub-blocks and processing them sequentially to ensure the computation is feasible.

[0085] S3. Extract the phase consistency feature of the frequency domain phase distribution map, and generate the texture directional feature based on the spatial domain directional gradient of the surface image. The specific implementation is as follows:

[0086] Perform phase gradient direction analysis on the frequency domain phase distribution map to generate a phase gradient direction matrix. Phase gradient direction analysis is achieved by calculating the phase change direction of adjacent pixels in the frequency domain phase distribution map. The specific steps are: calculate the phase gradient value in the horizontal and vertical directions respectively. The phase gradient value is the difference between the current pixel phase value and the pixel phase value on the right (or bottom). The phase gradient direction is the inverse tangent angle of the gradient value. The calculation formula is: the phase gradient direction is equal to the vertical phase gradient value divided by the inverse tangent value of the horizontal phase gradient value. The result is expressed in angles, and the range is limited to 0 degrees to 180 degrees. Each element of the phase gradient direction matrix corresponds to the gradient direction value of the corresponding pixel in the frequency domain phase distribution map. For example, if the horizontal phase gradient value of a pixel in the frequency domain phase distribution map is 0.5 radians and the vertical phase gradient value is 1.0 radians, then the phase gradient direction is arctan(1.0 / 0.5)=63.43 degrees. The size of the phase gradient direction matrix is ​​related to the frequency domain phase distribution. Figure 1 Each element stores a floating-point angle value.

[0087] A phase gradient direction histogram is constructed based on the phase gradient direction matrix to extract phase consistency features. The interval division method of the phase gradient direction histogram is equal angle interval division, and 0 degrees to 180 degrees is evenly divided into 36 intervals, each interval spanning 5 degrees. The phase consistency feature is obtained by counting the proportion of pixels in each interval. The specific calculation formula is: the phase consistency feature is equal to the number of pixels in the interval divided by the total number of pixels. For example, if there are 1000 pixels in the interval from 0 degrees to 5 degrees and the total number of pixels is 100,000, then the phase consistency feature of this interval is 1000 / 100,000=0.01. The phase consistency feature reflects the concentration degree of the phase gradient direction. The higher the value, the stronger the directional consistency. During the histogram construction process, intervals with less than 10 pixels are ignored to avoid noise interference.

[0088] Texture directional features are calculated based on the spatial directional gradient of the surface image. Spatial directional gradients are detected along the horizontal and vertical directions using the Sobel operator. The steps for calculating spatial directional gradients are as follows: First, Sobel horizontal and vertical operators are applied to the surface image to extract the horizontal and vertical gradient components, respectively. The Sobel horizontal convolution kernel weights are -1 on the left, 0 in the center, and 1 on the right, while the vertical convolution kernel weights are -1 on the top, 0 in the center, and 1 on the bottom. Second, the gradient angle of each pixel is calculated using the formula: the gradient angle is the vertical gradient component divided by the inverse tangent of the horizontal gradient component. The result is expressed as an angle, limited to 0 to 180 degrees.

[0089] Texture directionality is generated by counting the number of pixels at each directional angle. Specifically, the gradient direction angle is divided into 5-degree intervals and the percentage of pixels within each interval is calculated. For example, the percentage of pixels within the gradient direction angle range of 45 to 50 degrees is 0.05, indicating a low proportion of texture in that direction. During the calculation, if the horizontal gradient component is zero, the gradient direction angle is set to 90 degrees to avoid division by zero errors.

[0090] The phase consistency feature and the texture directionality feature are normalized to generate standardized phase consistency features and standardized texture directionality features. The normalization process uses the maximum and minimum normalization method to map the original eigenvalues ​​to the range of 0 to 1. The calculation formula is: the standardized eigenvalue is equal to the difference between the original eigenvalue minus the minimum value and the maximum value minus the minimum value. For example, the original value range of the phase consistency feature is 0.01 to 0.15, and after normalization, it is (0.01-0.01) / (0.15-0.01)=0, (0.15-0.01) / (0.15-0.01)=1. The normalized features eliminate dimensional differences, which facilitates subsequent correlation analysis. The maximum and minimum values ​​are dynamically obtained by traversing all pixel feature values. For example, all interval proportions of the phase consistency feature are traversed to find the global maximum value of 0.15 and the minimum value of 0.01.

[0091] A phase-texture correlation matrix is ​​constructed based on the joint distribution of the standardized phase-consistency feature and the standardized texture directionality feature. The construction rule for the phase-texture correlation matrix is ​​to match the consistency strength of phase and texture direction on a pixel-by-pixel basis. Specifically, the standardized phase-consistency feature and the standardized texture directionality feature are aligned in the same direction interval. The joint distribution probability of each interval pair (phase interval, texture interval) is calculated. The joint distribution probability is equal to the ratio of the number of pixels belonging to both the phase interval and the texture interval to the total number of pixels.

[0092] For example, if the phase direction range is 30 to 35 degrees and the texture direction range is 30 to 35 degrees, the joint distribution probability is 0.12, indicating that the proportion of pixels with the two directions matching is 12%. The phase-texture correlation matrix has a dimension of 36 by 36, and the matrix elements are the joint distribution probabilities of each interval pair. During the construction process, if the joint distribution probability is less than 0.01, the interval pair is marked as invalid. For example, if the joint distribution probability of a phase interval and texture interval is 0.005, the interval pair is marked as invalid in the correlation matrix.

[0093] The phase gradient directional histogram is divided into intervals of equal angles, with each interval spanning 5 degrees, for a total of 36 intervals. The phase-texture correlation matrix is ​​constructed by matching the consistency strength of phase and texture directions pixel by pixel. During the matching process, interval pairs with a joint distribution probability less than 0.01 are ignored to avoid noise interference. The calculation of the spatial directional gradient relies on the horizontal and vertical convolution kernels of the Sobel operator. The convolution kernel size is 3 pixels by 3 pixels. The horizontal kernel weight is -1 on the left, 0 in the middle, and 1 on the right. The vertical kernel weight is -1 on the top, 0 in the middle, and 1 on the bottom. The calculation of the gradient direction angle must avoid division by zero errors. When the horizontal gradient component is zero, the gradient direction angle is directly set to 90 degrees.

[0094] The construction of the phase-texture correlation matrix requires support for high-dimensional matrix operations and relies on an image processing library with floating-point computation capabilities, such as OpenCV 4.5 and above. For extreme cases, such as a solid-color, untextured surface image, all elements of the phase-texture correlation matrix are automatically reset to zero, and the exception type is recorded in the log. When insufficient hardware resources are detected, such as less than 1GB of memory, the system automatically divides the image into 512-by-512-pixel blocks to ensure computational feasibility. During block processing, a 10-pixel overlap is maintained between adjacent blocks to prevent boundary effects.

[0095] S4. Reconstruct the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain graph, which is specifically implemented as follows:

[0096] A dynamic correlation weight matrix is ​​generated based on the joint distribution of phase congruency features and texture directionality features. The dynamic correlation weight matrix is ​​calculated by multiplying the joint distribution probability of the phase-texture correlation matrix by the normalized phase congruency feature. The joint distribution probability of the phase-texture correlation matrix is ​​obtained from the construction result of step S3, and the normalized phase congruency feature is obtained from the normalization processing result of step S3. The dynamic correlation weight matrix is ​​calculated as follows: the dynamic correlation weight matrix is ​​equal to the joint distribution probability matrix of the phase-texture correlation matrix multiplied by the normalized phase congruency feature matrix.

[0097] For example, if the joint distribution probability of a pixel is 0.12 and the normalized phase consistency feature is 0.8, then the weight value of the pixel in the dynamic correlation weight matrix is ​​0.12 multiplied by 0.8, which is equal to 0.096. The dimension of the dynamic correlation weight matrix is ​​similar to the frequency domain phase distribution. Figure 1 The value of each element ranges from 0 to 1. When generating the dynamic correlation weight matrix, if the joint distribution probability or the normalized phase consistency feature is an invalid value (such as NaN), the pixel weight is automatically reset to zero and an abnormality log is recorded.

[0098] The initial frequency domain response function is decomposed into a low-band response function, a mid-band response function, and a high-band response function. The initial frequency domain response function is obtained from the frequency domain phase distribution diagram, and the decomposition method is multi-scale band separation based on Gaussian filtering. The specific steps are: Gaussian low-pass filtering is performed on the initial frequency domain response function with a cutoff frequency of one-tenth of the image spatial frequency to generate the low-band response function; band-pass filtering is performed on the initial frequency domain response function with a cutoff frequency of one-tenth to one-fifth of the image spatial frequency to generate the mid-band response function; and Gaussian high-pass filtering is performed on the initial frequency domain response function with a cutoff frequency of one-fifth of the image spatial frequency to generate the high-band response function. For example, when the image spatial frequency is 10 line pairs per millimeter, the cutoff frequency of the low-band is 1 line pair / mm, the mid-band is 1 to 2 line pairs / mm, and the high-band is 2 line pairs / mm or higher. The band decomposition process depends on the image resolution parameter, which is measured in pixels / mm and is set automatically through metadata parsing or manually entered.

[0099] The amplitude of each frequency band response function is adjusted based on the dynamic relevance weight matrix. The low-band response function is adjusted using a weighted average method, the mid-band response function is adjusted using gradient direction constraints, and the high-band response function is adjusted using threshold truncation. The specific steps for adjusting the low-band response function using the weighted average method are: multiply the amplitude of the low-band response function by the weight value of the corresponding pixel in the dynamic relevance weight matrix, and then multiply it by the low-band weighting coefficient. The low-band weighting coefficient is dynamically set based on the concentration of the phase congruence feature. The concentration is calculated as: the concentration is the maximum minus the minimum value of the phase congruence feature.

[0100] For example, if the maximum value of the phase consistency feature is 0.15 and the minimum value is 0.01, the concentration is 0.14, and the low-frequency band weighting coefficient is set to 0.14 multiplied by 0.5, which equals 0.07. The specific steps for adjusting the mid-band response function with gradient direction constraints are: extracting texture directional features and directionally enhancing the amplitude of the mid-band response function along the texture direction. The enhancement coefficient is the square of the dynamic correlation weight matrix value.

[0101] The specific steps for adjusting the high-band response function using threshold truncation are: Set the high-band threshold equal to the dispersion of the texture directional characteristics multiplied by 0.1. The dispersion is calculated as follows: dispersion equals the standard deviation of the texture directional characteristics. For example, if the standard deviation of the texture directional characteristics is 0.2, the high-band threshold is 0.2 multiplied by 0.1, which equals 0.02. High-band response functions with amplitudes below 0.02 are set to zero. During the adjustment process, if the low-band weighting coefficient exceeds the preset upper limit of 1.0, it is automatically truncated to 1.0 and an alarm is recorded.

[0102] The adjusted low-band response function, mid-band response function, and high-band response function are superimposed to generate an optimized frequency domain map. The formula for band superposition is: the optimized frequency domain map equals the adjusted low-band response function plus the adjusted mid-band response function plus the adjusted high-band response function. The superimposed optimized frequency domain map preserves the overall structural information of the low-band, the directional texture characteristics of the mid-band, and the detailed information of the high-band, while suppressing noise interference. For example, if the adjusted amplitude of a pixel in the low-band is 0.5, the amplitude of the mid-band is 0.3, and the amplitude of the high-band is 0.1, then the amplitude of that pixel in the optimized frequency domain map is 0.5 plus 0.3 plus 0.1, which equals 0.9. The superimposed result is spatially smoothed using bilinear interpolation to eliminate sudden amplitude changes at band boundaries.

[0103] The weighted average adjustment coefficient for the low-band response function is dynamically set based on the concentration of the phase congruence feature. The threshold cutoff range for the high-band response function is determined based on the dispersion of the texture directional features. The concentration of the phase congruence feature is calculated as the difference between the maximum and minimum values ​​of the normalized phase congruence feature. For example, if the maximum value of the normalized phase congruence feature is 1.0 and the minimum value is 0.1, the concentration is 0.9, and the low-band weighting coefficient is set to 0.9 multiplied by 0.5, which equals 0.45. The dispersion of the texture directional feature is calculated based on the standard deviation of the normalized texture directional feature. For example, if the standard deviation is 0.25, the high-band threshold cutoff range is 0.25 multiplied by 0.1, which equals 0.025.

[0104] Calculation of the dynamic relevance weight matrix requires support for matrix multiplication and relies on an image processing library with floating-point capabilities, such as OpenCV 4.5 and above. During the band decomposition and adjustment process, the Gaussian filter cutoff frequency is dynamically adjusted based on the image resolution. The calculation formula is: the cutoff frequency equals the image resolution (in pixels / mm) multiplied by the scaling factor. For example, if the image resolution is 10 pixels / mm and the low-band scaling factor is 0.1, the cutoff frequency is 10 times 0.1, which equals 1 line pair / mm. For extreme input conditions, such as when the dynamic relevance weight matrix is ​​all zeros, the system automatically skips the band adjustment step and directly outputs the initial frequency domain response function as the optimized frequency domain plot. The exception type is recorded in the log. If insufficient hardware resources are detected, such as memory less than 2GB, the system automatically reduces the image resolution to one-quarter of the original before performing band decomposition and adjustment to ensure computational feasibility. During block processing, a 10-pixel overlap is maintained between adjacent blocks to prevent boundary effects.

[0105] Step S4 reconstructs the frequency domain response function to generate an optimized frequency domain graph by combining the dynamic correlation weights of the frequency domain phase consistency features and the spatial domain texture directional features. By using the phase-texture correlation matrix to quantify the coordinated change law of the frequency domain and spatial domain features, the response functions of the low-frequency, medium-frequency and high-frequency bands are adjusted accordingly: the low-frequency band enhances structural consistency through weighted averaging, the medium-frequency band strengthens detail continuity in combination with texture directional constraints, and the high-frequency band suppresses noise interference based on the discreteness threshold. Compared with the existing technology that only relies on single domain features or static filtering, the cross-domain dynamic weight fusion solves the problem of false detection or missed detection of defects caused by frequency domain aliasing. Breaking through the limitations of traditional frequency domain / spatial domain independent processing, the feature correlation is used to achieve precise control of the frequency band response, thereby improving the signal-to-noise ratio and accuracy of defect recognition under complex texture backgrounds.

[0106] S5. Based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image, the mutually exclusive conflict area is detected and the spatial domain inverse transform is performed to obtain a de-aliased image. The specific implementation is as follows:

[0107] The optimized frequency domain image is divided into overlapping sub-blocks. Singular value decomposition is performed on each sub-block to extract the left singular vector matrix and calculate its information entropy. The size of the overlapping sub-blocks is dynamically adjusted based on the resolution of the optimized frequency domain image. The adjustment rule is: the sub-block side length is equal to the resolution of the optimized frequency domain image (in pixels / mm) divided by 10, rounded up to the nearest even value. For example, if the optimized frequency domain image resolution is 95 pixels / mm, the sub-block side length is 95 / 10 = 9.5, rounded up to 10 pixels, resulting in a sub-block size of 10 pixels × 10 pixels.

[0108] The overlap area of ​​each sub-block is 25% of the sub-block side length. That is, the spacing between adjacent sub-blocks is the sub-block side length minus the overlap area. For example, if the sub-block side length is 10 pixels, the spacing is 10-2.5 = 7.5 pixels, rounded to 7 pixels. The singular value decomposition process is performed on the complex matrix of each sub-block to generate a left singular vector matrix. The information entropy of the left singular vector matrix is ​​calculated by calculating the probability distribution of the matrix elements. The specific formula is: information entropy is equal to the sum of the negative probability values ​​of all matrix elements multiplied by the base-2 logarithm of the probability value. The information entropy threshold is set by statistically analyzing the entropy distribution of normal food images. For example, if the entropy values ​​of 1000 images of defect-free food are collected, with a mean of 0.8 and a standard deviation of 0.05, the threshold is set to 0.8 + 3 × 0.05 = 0.95. If the entropy value of a sub-block exceeds 0.95, it is marked as abnormal.

[0109] The optimized frequency domain image is partitioned into multiple frequency bands. The phase coherence between adjacent frequency bands is calculated using complex cross-correlation, and the coherence decay rate for each frequency band is calculated. The multi-band partitioning uses Mel-scale bands, evenly dividing the spatial frequency range of the optimized frequency domain image into 24 frequency bands. The band boundaries are nonlinearly distributed according to the Mel scale. The complex cross-correlation between adjacent frequency bands is calculated as follows: the complex spectrum matrices of the two frequency bands are extracted, the complex spectrum of each band is expanded into a one-dimensional complex vector in row-major order, and the covariance matrix of the two vectors is calculated. The proportion of the maximum eigenvalue of the covariance matrix is ​​the phase coherence. For example, if the length of the complex vector of frequency band A is N and the length of the complex vector of frequency band B is M, the covariance matrix dimension is N×M. The maximum eigenvalue is solved using the power iteration method. If the maximum eigenvalue is 0.8, the phase coherence is 0.8. The coherence decay rate is calculated by comparing the phase coherence differences between adjacent frequency bands using the formula: decay rate equals 1 minus the phase coherence value between the current and next frequency bands. The decay rate threshold is set by statistically analyzing the decay rate distribution of 1,000 defect samples. For example, if the average decay rate of a cracked region is 0.6 with a standard deviation of 0.1, the threshold is set to 0.6-2×0.1=0.4. If the decay rate of a region exceeds 0.4, it is identified as a phase continuity break.

[0110] Among them, the Mel scale conversion formula is: ;in, Indicates Mel frequency, that is, the frequency value under the Mel scale, the unit is Mel; represents the linear frequency, that is, the original frequency value of the input signal, in Hertz (Hz); 2595 represents the scaling factor of the Mel scale; 700 represents the frequency normalization constant.

[0111] Sub-blocks with information entropy exceeding a preset entropy threshold or regions with a coherence decay rate exceeding a preset decay rate threshold are marked as mutually exclusive conflicting regions, and a mutually exclusive conflicting region mask is generated. The rule for generating the mutually exclusive conflicting region mask is as follows: if the information entropy of a sub-block exceeds the threshold or the decay rate of a region exceeds the threshold, the sub-block or region is marked as 1 in the mask; otherwise, it is marked as 0. For example, if the information entropy of a sub-block is 0.96 (threshold 0.95), it is marked regardless of its decay rate; and if the decay rate of a region is 0.45 (threshold 0.4), it is marked regardless of its entropy value. After the mask is generated, holes are filled and the boundaries are smoothed using a morphological closing operation. The structuring element of the morphological closing operation is a 3-pixel × 3-pixel rectangular kernel, and the number of iterations is 2 to ensure the continuity and integrity of the masked region.

[0112] An inverse spatial transform is performed on the frequency domain data masked by the mutually exclusive conflicting regions to generate a dealiased image. This inverse spatial transform utilizes an inverse fast Fourier transform (IFFT). The input is the complex spectrum data in the optimized frequency domain image marked as 1 by the mask, while the spectrum in the remaining regions is set to zero. This inverse transform generates a spatial image, which is then mapped to a range of 0 to 255 using grayscale stretching. The grayscale stretching formula is: stretched pixel value = (original pixel value - minimum value) × 255 / (maximum value - minimum value). For example, the complex spectrum of a conflicting region is inversely transformed to generate a local spatial image, which is then spliced ​​with the unprocessed region to form a complete image. If the mask is all zero (no conflicting region), the inverse transform of the original optimized frequency domain image is directly output.

[0113] The size of the overlapping sub-blocks is dynamically adjusted according to the resolution of the optimized frequency domain image. The resolution is obtained through image metadata or preset according to imaging device parameters. The complex cross-correlation is calculated by the eigenvalues ​​of the covariance matrix of the complex spectra between frequency bands. The construction rule of the covariance matrix is: the complex spectra of the two bands are expanded into a one-dimensional vector and then the covariance is calculated. The elements of the covariance matrix are Equal to Band A The conjugate of the element multiplied by the band B elements. For example ,in, represents the expectation operation; indicates conjugation; Indicates the covariance matrix of band A and band B. Rank The element value of the column; The complex spectrum of band A is expanded into a one-dimensional vector. elements; The complex spectrum of band B is expanded into a one-dimensional vector. elements.

[0114] Dynamic calibration of the information entropy threshold relies on a pre-stored database of normal samples. This database contains optimized frequency domain graph samples for different food categories (such as nuts, fruits and vegetables, and meat), with thresholds set independently for each category. For example, the information entropy threshold for nuts is 0.95, for fruits and vegetables it is 0.9, and for meat it is 0.93. The classification is based on the complexity of the food surface texture, which is quantified by the frequency domain energy variance. The phase coherence decay rate threshold is set by experimentally analyzing the decay rate distribution of defective samples. For example, the average decay rate in a cracked area is 0.6, and the threshold is set to 0.5 to cover 95% of defective samples, with a standard deviation of 0.1 and a confidence interval of ±2σ around the mean.

[0115] During processing, if insufficient hardware resources are detected (e.g., GPU memory is less than 4GB), the system automatically divides the optimized frequency domain image into 512×512 pixel blocks, with a 20-pixel overlap between blocks to eliminate boundary artifacts. The inverse-transformed spatial image blocks are fused using a weighted average method, with weights decreasing linearly based on the distance from the block center to the edge, with a center weight of 1 and an edge weight of 0.2. For abnormal inputs such as completely black or completely white, the system automatically skips the conflict detection step, directly performs the inverse transformation, and records the abnormality in a log format: "Anomaly Type: Completely Black Image; Processing Time: 2023-10-01 14:30:00."

[0116] Step S5 solves the frequency domain aliasing misdetection problem caused by single feature analysis in the existing technology by combining the frequency domain singular value decomposition (SVD) information entropy and the phase coherence attenuation degree of dual detection logic. Based on the SVD information entropy, the local structural complexity is quantified to identify abnormal sub-blocks, while the phase continuity break area is captured through multi-band phase coherence attenuation analysis. The two work together to detect mutually exclusive conflicting areas. Compared with traditional methods that rely only on independent judgments of frequency domain energy or spatial domain texture, this step integrates structural complexity and phase continuity features, introduces cross-domain feature correlation into frequency domain aliasing processing, and distinguishes between real defects and complex background textures through a dual verification mechanism, thereby ensuring detection accuracy while suppressing misjudgments. It is particularly suitable for complex scenarios where the micro-texture of the food surface overlaps with the frequency domain characteristics of the defect.

[0117] S6. Perform defect recognition and determination based on the defect feature area in the de-aliased image, specifically implemented as follows:

[0118] A gray-level co-occurrence matrix is ​​constructed for the de-aliased image to extract the texture contrast and energy features of the defect area. The gray-level co-occurrence matrix is ​​constructed in directions of 0, 45, 90, and 135 degrees, with a distance parameter of 1 pixel. The gray-level co-occurrence matrix is ​​generated by calculating the joint probability distribution of pixel pairs along the specified direction and distance for the gray-level value matrix of the de-aliased image. Texture contrast is calculated using the contrast formula, which is equal to the sum of the squared gray-level differences of all pixel pairs multiplied by the corresponding probability. The energy feature is equal to the sum of the squared probabilities of all pixel pairs. For example, if the gray-level difference of a pixel pair in the gray-level co-occurrence matrix is ​​3 and the probability is 0.01, the contrast contribution is 3 squared multiplied by 0.01, which equals 0.09, and the energy contribution is 0.01 squared, which equals 0.0001. Texture contrast and energy features reflect the edge sharpness and texture uniformity of the defect area, respectively. The direction and distance parameters of the gray-level co-occurrence matrix are dynamically adapted to the image resolution. For high-resolution images (greater than 200 pixels / mm), the distance parameter is set to 2 pixels to improve computational efficiency.

[0119] Based on the morphological opening operation, the geometric features of the defect area are extracted to generate a set of candidate defect areas. The structural element of the morphological opening operation is a circular kernel, and the radius of the circular kernel is dynamically adjusted according to the minimum size of the defect. The minimum size of the defect is obtained through statistics from the pre-stored defect sample library. For example, the minimum length of a crack on the surface of a nut is 5 pixels, so the radius of the circular kernel is set to 5 pixels multiplied by 0.6, which is equal to 3 pixels. The execution steps of the morphological opening operation are as follows: first, an erosion operation is performed on the de-aliased image to eliminate isolated noise points. The structural element of the erosion operation is a circular kernel with a radius of 3 pixels. Then, an expansion operation is performed to restore the original size of the retained area. The expanded structural element is consistent with the eroded structural element. The set of candidate defect areas consists of all connected areas with an area greater than 5 pixels. For example, if the retained area of ​​a region is 7 pixels after erosion and is restored to 8 pixels after expansion, it is determined to be a candidate defect area.

[0120] The texture contrast, energy, and geometric features of the candidate defect regions are input into a pre-trained support vector machine classifier, which outputs the defect category probability for each region in the candidate defect region set. The support vector machine classifier's training data includes normal texture samples and typical defect samples (such as mold spots, wormholes, and cracks). The feature vector dimension is 3, corresponding to the regional aspect ratios of the texture contrast, energy, and geometric features. The classifier outputs the probability value of the defect category. For example, a probability value of 0.92 for a candidate region indicates a 92% probability of being a defect. The pre-training process uses cross-validation to optimize the kernel function parameters. The kernel function type is the radial basis function, and the cross-validation fold is 5-fold to ensure model generalization.

[0121] A preliminary defect determination result is generated based on the defect category probability and the dynamic determination threshold. The dynamic determination threshold is dynamically adjusted based on the false positive rate distribution of historical defect samples. The dynamic determination threshold adjustment rule is as follows: if the false positive rate in historical data exceeds 5%, the threshold is increased by 0.05; if the false positive rate is less than 2%, the threshold is decreased by 0.03. For example, if the initial threshold is 0.9, when the false positive rate rises to 6%, the threshold is adjusted to 0.95. The preliminary defect determination result is all candidate areas with a probability value greater than or equal to the dynamic threshold. For example, when the threshold is 0.9, areas with a probability of 0.92 are determined to be defects, and areas with a probability of 0.88 are excluded. The historical data storage period is 30 days, and the threshold parameters are updated daily.

[0122] The preliminary defect determination results are checked for spatial continuity, and the final defect feature map is generated after the isolated abnormal areas in the set of candidate defect areas are eliminated. The spatial continuity check is achieved by filtering the connected area threshold. The connected area threshold matches the minimum defect size. The specific rule is that the area threshold is equal to the square of the minimum defect size multiplied by 3. For example, when the minimum defect size is 5 pixels, the area threshold is set to 5 squared multiplied by 3, which is equal to 75 pixels. Isolated areas with an area of ​​less than 75 pixels are eliminated. The steps for generating the final defect feature map are as follows: the retained defect area is smoothed by a morphological closing operation. The closing operation structure element is a 3-pixel by 3-pixel rectangular kernel, the number of iterations is 1, and the smoothed binary image is output as the defect feature map.

[0123] The training data for the support vector machine classifier must cover at least 1,000 samples from at least 10 food categories, with a balanced distribution of samples within each category to ensure classifier generalization. Training data preprocessing includes grayscale normalization and feature normalization. Grayscale normalization maps pixel values ​​to a range of 0 to 1, and feature normalization uses the Z-score method. The classifier model is saved as a binary file, and version compatibility is verified upon loading. Any version mismatch triggers a retraining process.

[0124] During processing, if the set of candidate defect regions is detected to be empty (i.e., there are no candidate regions), the system automatically determines the image as defect-free and outputs a null result. When hardware resources are insufficient, such as when the memory is less than 2GB, the system automatically reduces the image resolution to half of the original and re-executes the feature extraction and classification steps. Resolution adjustment uses a bilinear interpolation algorithm. For abnormal inputs that are completely white or completely black, the system skips the defect determination step and records an exception log. The log contains a timestamp and an identification of the exception type. The log is stored in the system's preset directory.

[0125] This method overcomes the limitations of traditional food defect detection, which relies on independent processing of frequency and spatial domain features, through collaborative analysis of multi-domain features and a dynamic weight fusion mechanism. Existing technologies typically rely on single-band filtering or spatial texture analysis, making it difficult to distinguish between frequency-domain aliasing of microtextures and true defects. This approach, however, introduces phase-texture correlation weights to reconstruct the frequency-domain response function (S4), combining the dual criteria of frequency-domain singular value distribution entropy and phase coherence decay (S5) to construct a defect detection framework that dynamically couples cross-domain features. Traditional methods fail to reveal the inherent connection between frequency-domain phase continuity and spatial-domain texture directionality. This method achieves collaborative quantification of frequency-domain structural complexity and spatial-domain texture discontinuity through a joint analysis of the phase gradient directional histogram (S3) and the frequency band coherence decay rate (S5). This allows for the design of dynamic marking rules for mutually exclusive conflicting regions, addressing the misdetection of minor defects in complex texture backgrounds. In addition, the morphological kernel parameters (S6) are dynamically adjusted based on the minimum defect size, making the geometric feature extraction adaptive to the surface characteristics of different foods, avoiding the subjective bias of manual experience parameter setting, and forming a closed-loop logic chain. For example, the optimized frequency domain image of S4 provides high-quality input for S5, and the de-aliased image of S5 provides accurate features for the defect classification of S6, achieving a balance between computational efficiency and detection accuracy. It is particularly suitable for real-time online detection scenarios of high-resolution food images.

[0126] Example 2: Figure 2 The following is a schematic diagram of the structure of the intelligent food detection system based on image recognition of the present invention. The intelligent food detection system based on image recognition includes the following modules:

[0127] An image acquisition module, used to acquire a surface image of the food to be inspected;

[0128] A phase analysis module is used to perform frequency domain phase analysis on the surface image and generate a frequency domain phase distribution map of the surface image;

[0129] Texture analysis module, used to extract phase consistency features of the frequency domain phase distribution map and generate texture directional features based on the spatial domain directional gradient of the surface image;

[0130] A function reconstruction module is used to reconstruct the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain map;

[0131] The mixed image generation module is used to detect the mutually exclusive conflict area and perform spatial inverse transformation based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image to obtain the anti-aliased image;

[0132] The defect recognition module is used to perform defect recognition and judgment based on the defect feature area in the de-aliased image.

[0133] Example 3: An intelligent food detection terminal based on image recognition includes: a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, an intelligent food detection method based on image recognition is implemented.

[0134] Example 4: An intelligent food detection medium based on image recognition, wherein a program or instruction is stored on the medium, and when the program or instruction is executed by a processor, an intelligent food detection method based on image recognition is implemented.

[0135] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.

[0136] It should be noted that the present invention can be deployed on the device itself to implement embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0137] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0138] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0139] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0140] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0141] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0142] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0143] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0144] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An intelligent food detection method based on image recognition, characterized in that: The steps include: S1. Acquire a surface image of the food to be tested; S2. performing frequency domain phase analysis on the surface image to generate a frequency domain phase distribution map of the surface image; S3, extracting the phase consistency feature of the frequency domain phase distribution map, and generating the texture directional feature based on the spatial domain directional gradient of the surface image; S4. Reconstructing the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain graph, including: A dynamic correlation weight matrix is ​​generated based on the joint distribution of phase consistency features and texture directionality features. The dynamic correlation weight matrix is ​​calculated by multiplying the joint distribution probability of the phase-texture correlation matrix and the normalized phase consistency feature. Decomposing the initial frequency domain response function into a low-frequency band response function, a mid-frequency band response function and a high-frequency band response function; The amplitude of each frequency band response function is adjusted according to the dynamic correlation weight matrix. The low frequency band response function is adjusted by weighted average method, the mid frequency band response function is adjusted by gradient direction constraint, and the high frequency band response function is adjusted by threshold truncation. Superimposing the adjusted low-frequency band response function, mid-frequency band response function and high-frequency band response function to generate an optimized frequency domain graph; S5. Based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image, the mutually exclusive conflicting areas are detected and the spatial domain inverse transform is performed to obtain a de-aliased image; S6. Perform defect recognition and determination based on the defect feature area in the de-aliased image.

2. The intelligent food detection method based on image recognition according to claim 1, characterized in that: Acquire surface images of the food to be inspected, including: Adjust the irradiation angle and polarization axis direction of the multi-angle light source so that the polarization axis direction of the light source is orthogonal to the normal direction of the food surface; Under multi-angle light sources, the reflected light from the food surface is filtered through a polarizing filter to generate an initial image that suppresses reflection interference; adjusting the exposure time of the image sensor based on the food surface transparency parameter, and capturing surface images under multi-angle polarized light according to the adjusted exposure time; The surface image under multi-angle polarized light is checked for consistency in light intensity distribution. The target angle image is filtered according to the light intensity distribution consistency check result to generate the surface image of the food to be tested.

3. The intelligent food detection method based on image recognition according to claim 1, characterized in that: Perform frequency domain phase analysis on the surface image to generate a frequency domain phase distribution map of the surface image, including: Converting the surface image into a grayscale image, performing edge symmetry compensation processing on the grayscale image, and generating a preprocessed image; Performing a two-dimensional discrete Fourier transform on the preprocessed image, extracting the phase component of the complex spectrum matrix, and generating an initial phase distribution map; The phase threshold range is adjusted based on the local contrast parameter of the surface image, and the phase values ​​exceeding the threshold range in the initial phase distribution map are truncated and corrected; According to the spatial frequency distribution characteristics of the surface image, the truncated and corrected phase distribution map is subjected to frequency band weighted fusion to generate a frequency domain phase distribution map.

4. The intelligent food detection method based on image recognition according to claim 1, characterized in that: Extract the phase consistency features of the frequency domain phase distribution map, and generate texture directional features based on the spatial directional gradient of the surface image, including: performing phase gradient direction analysis on the frequency domain phase distribution diagram to generate a phase gradient direction matrix; Construct a phase gradient direction histogram based on the phase gradient direction matrix and extract phase consistency features; The texture directional features are calculated based on the spatial directional gradient of the surface image. The spatial directional gradient is detected along the horizontal and vertical directions using the Sobel operator. Normalizing the phase consistency feature and the texture directionality feature to generate a standardized phase consistency feature and a standardized texture directionality feature; The phase-texture correlation matrix is ​​constructed based on the joint distribution of the normalized phase consistency features and the normalized texture directionality features.

5. The intelligent food detection method based on image recognition according to claim 1, characterized in that: Based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation of the optimized frequency domain image, the mutually exclusive conflict area is detected and the spatial domain inverse transform is performed to obtain the anti-aliased image, including: Divide the optimized frequency domain image into overlapping sub-blocks, perform SVD decomposition on each sub-block, extract the left singular vector matrix and calculate the information entropy of the left singular vector matrix; The optimized frequency domain image is divided into multiple frequency bands, the phase coherence between adjacent frequency bands is calculated by complex cross-correlation, and the coherence decay rate of each frequency band is statistically analyzed; Marking a sub-block whose information entropy exceeds a preset information entropy threshold or a region whose coherence decay rate exceeds a preset decay rate threshold as a mutually exclusive conflict region and generating a mutually exclusive conflict region mask; Perform spatial inverse transform processing on the frequency domain data covered by the mutually exclusive conflicting area mask to generate an anti-aliased image.

6. The intelligent food detection method based on image recognition according to claim 1, characterized in that: Defect recognition and judgment are performed based on the defect feature areas in the anti-aliased image, including: Construct a gray-level co-occurrence matrix for the de-aliased image and extract the texture contrast and energy characteristics of the defect area; The geometric features of the defect area are extracted based on the morphological opening operation to generate a set of candidate defect areas. The structural element of the morphological opening operation is a circular kernel, and the radius of the circular kernel is dynamically adjusted according to the minimum size of the defect. Input the texture contrast, energy features and geometric features in the candidate defect area set into the pre-trained support vector machine classifier, and output the defect category probability of each area in the candidate defect area set; Generate preliminary defect determination results based on defect category probability and dynamic determination threshold. The dynamic determination threshold is dynamically adjusted based on the false detection rate distribution of historical defect samples. The spatial continuity of the preliminary defect determination results is checked, and the final defect feature map is generated after the isolated abnormal areas in the candidate defect area set are eliminated.

7. An intelligent food detection system based on image recognition, used to implement the intelligent food detection method based on image recognition according to any one of claims 1 to 6, characterized in that: Includes the following modules: An image acquisition module, used to acquire a surface image of the food to be inspected; A phase analysis module is used to perform frequency domain phase analysis on the surface image and generate a frequency domain phase distribution map of the surface image; Texture analysis module, used to extract phase consistency features of the frequency domain phase distribution map and generate texture directional features based on the spatial domain directional gradient of the surface image; A function reconstruction module is used to reconstruct the frequency domain response function according to the dynamic correlation weight of the phase consistency feature and the texture directionality feature to generate an optimized frequency domain map; The mixed image generation module is used to detect the mutually exclusive conflict area and perform spatial inverse transformation based on the frequency domain SVD singular vector distribution entropy and phase coherence attenuation degree of the optimized frequency domain image to obtain the anti-aliased image; The defect recognition module is used to perform defect recognition and judgment based on the defect feature area in the de-aliased image.

8. Intelligent food detection terminal based on image recognition, characterized in that: include: A processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the intelligent food detection method based on image recognition as described in any one of claims 1 to 6 is implemented.

9. Intelligent food detection medium based on image recognition, characterized in that: The medium stores a program or instruction, and when the program or instruction is executed by the processor, the intelligent food detection method based on image recognition as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Rapid defect detection method and device based on machine vision, equipment and storage medium

    CN115984246A

  • Seamless steel tube defect detection method based on image processing

    CN119338809A