Naked eye 3D display grating film detection method based on AI large model
By applying AI large model and multimodal data fusion technology in the detection of grating film of naked-eye 3D displays, combining structured light distortion coefficient, ambient light interference coefficient and pixel consistency coefficient to generate a comprehensive detection index, the problems of low detection efficiency and high error detection rate in the existing technology are solved, and high-precision grating film quality detection and intelligent control throughout the process are achieved.
Patent Information
- Application Number
- CN202510671235.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve high-precision and high-efficiency bare-eye 3D display grating film quality detection, especially when multi-view data registration and fusion efficiency are low, ambient light interference is large, complex defect types are also needed, and traditional algorithm generalization capabilities are insufficient.
Using a detection method based on AI large model, the multi-view structured light image acquisition and multi-modal data are fused with multi-modal data, and the Transformer model is used to combine structured light distortion coefficients, ambient light interference coefficients and pixel consistency coefficients to generate a comprehensive detection index to achieve accurate quantification of grating film quality and defect classification. At the same time, three-level feedback calibration is implemented, including process parameter adjustment, ambient light control and model learning optimization.
It significantly improves the accuracy of surface defect recognition of grating films, reduces the error detection rate, realizes intelligent control of the quality of grating films throughout the process, and improves detection efficiency and production yield.
Smart Images

Figure CN120219901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI technology, and particularly to a method for detecting a grating film of a naked-eye 3D display based on an AI large model. Background Art
[0002] With the rapid popularization of naked-eye 3D display technology, its applications in fields such as consumer electronics, medical imaging, and virtual reality are becoming increasingly widespread. As the core optical component of a naked-eye 3D display, the accuracy of the surface structure of the grating film directly affects the 3D imaging effect and user experience. However, the manufacturing process of the grating film is complex, involving the uniform arrangement of micron-scale grating lines, the precise control of material thickness, and the requirement for environmental light stability, which makes its quality inspection a key challenge in the production process. Traditional inspection methods rely on manual visual inspection or machine vision systems with a single viewing angle, and it is difficult to meet the requirements of high-precision and high-efficiency industrial production.
[0003] Currently, the industry mainly uses inspection technologies based on structured light projection or multi-view image acquisition, combined with basic image processing algorithms for defect recognition. For example, some studies obtain the surface topography of the grating film through multi-angle scanning and use traditional machine learning models, such as support vector machines and random forests, to classify defects. However, the existing technologies have significant limitations: firstly, the registration and fusion efficiency of multi-view data is low, and it is difficult to process high-resolution images in real time; secondly, environmental light interference, such as fluctuations in illuminance and color temperature deviation, easily leads to unstable detection results; thirdly, when the defect types are complex, such as grating line offset and uneven material thickness, the generalization ability of traditional algorithms is insufficient, and it is impossible to accurately locate and classify defects. Furthermore, the current inspection systems generally lack dynamic feedback and adaptive optimization mechanisms. For example, the adjustment of production parameters depends on manual experience and cannot be automatically calibrated according to real-time detection results; environmental light control mostly uses fixed thresholds and it is difficult to cope with changes in complex working conditions; in addition, traditional models lack incremental learning ability and it is difficult to continuously optimize performance through historical data. These problems lead to low inspection efficiency and high misdetection rate, seriously restricting the mass production yield and cost control of naked-eye 3D displays. Therefore, there is an urgent need for an inspection method integrating an AI large model, multi-modal data fusion, and intelligent feedback to break through the bottleneck of existing technologies and achieve full-process intelligent control of the quality of grating films.
[0004] Therefore, it is very necessary to invent a method for detecting a grating film of a naked-eye 3D display based on an AI large model to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting a grating film of a naked-eye 3D display based on an AI large model to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: A method for detecting a lenticular film of a naked-eye 3D display based on an AI large model, including a lenticular film data acquisition terminal, a lenticular film intelligent analysis terminal, an AI analysis terminal, a calibration terminal, and an output terminal, specifically including the following steps: S1. The lenticular film data acquisition terminal performs line-by-line scanning and acquisition of the lenticular film of the naked-eye 3D display in a darkroom environment to obtain a five-view structured light image dataset, an ambient light dataset, and an image physical dataset containing spatial coordinate information; S2. The lenticular film intelligent analysis terminal preprocesses and analyzes the five-view structured light image dataset, the ambient light dataset, and the image physical dataset to obtain a structured light distortion coefficient, an ambient light interference coefficient, and a pixel consistency coefficient; S3. The AI analysis terminal inputs the structured light distortion coefficient, the ambient light interference coefficient, and the pixel consistency coefficient into the Transformer model, outputs a comprehensive detection index to be processed, and obtains a comprehensive detection index after normalization processing; S4. The AI analysis terminal grades the quality of the lenticular film based on the comprehensive detection index, and at the same time locates the defect type by combining the structured light distortion coefficient, the ambient light interference coefficient, and the pixel consistency coefficient, and sends a detection report generation instruction to the output terminal and a calibration instruction to the calibration terminal; S5. The output terminal receives the instruction sent by the AI analysis terminal and generates a visual detection report; S6. The calibration terminal receives the instruction sent by the AI analysis terminal and performs three-level feedback calibration.
[0007] Preferably, the five-view structured light image dataset includes: a front center view image pixel set, a horizontal left view image pixel set, a horizontal right view image pixel set, a vertical top view image pixel set, and a vertical bottom view image pixel set; the ambient light dataset includes: ambient light illuminance and color temperature; the image physical dataset includes: the width of the image, the height of the image, the coordinates of the image pixel points, and the actual surface height of the lenticular film.
[0008] Preferably, the structured light distortion coefficient is specifically: , where M is the total number of pixel points, specifically M = W×H, where W is the width of the image, H is the height of the image, x is the abscissa of the image pixel position, y is the ordinate of the image pixel position, max(Δh) is the maximum value among all height change amounts Δh(x, y), Δh(x, y) is the height change amount at the coordinate (x, y), where Δh(x, y)=Z actual (x, y)-Z ideal (x, y), where Z actualThe actual surface height of the grating film at the coordinates (x, y) is Z ideal The (x, y) is the ideal fitting plane.
[0009] Preferably, the ambient light interference coefficient is specifically: , where e is the Euler number, μ E is the ideal mean value of the ambient illumination, μ T is the ideal mean value of the color temperature, σ E is the standard deviation of the normalized value E′ of the illumination, σ T is the standard deviation of the normalized value T′ of the color temperature, is the average value of the normalized value E′ of the ambient illumination, is the average value of the normalized value T′ of the color temperature, where, , where E is the ambient illumination collected in real time, E min is the preset minimum value of the illumination, E max is the preset maximum value of the illumination, , where T is the ambient light color temperature collected in real time, T min is the preset minimum value of the color temperature, T max is the preset maximum value of the color temperature.
[0010] Preferably, the pixel consistency coefficient is specifically: , where γ j is the cosine similarity between a single registered image and the central view image, specifically: , where, is the pixel value of the registered image of the j-th view at (x, y), is the pixel value of the central view image at (x, y), is the pixel mean value of the registered image of the j-th view, is the pixel mean value of the central view image.
[0011] Preferably, the Transformer model is specifically: , where e is the Euler number, l is the layer index, L is the total number of layers of the Transformer model, α l is the attention weight of the l-th layer, F l is the representation of the input feature vector F at the l-th layer, FFN is the feed-forward network, MSA is the multi-head self-attention mechanism, and b is the bias term.
[0012] Preferably, the quality grading can be divided into the following levels: When the comprehensive detection index is within the preset threshold S1, it is classified as level one and determined to be qualified; When the comprehensive detection index is within the preset threshold S2, it is classified as level two and determined to require re-inspection; When the comprehensive detection index is within the preset threshold S3, it is classified as level three and determined to be unqualified.
[0013] Preferably, the defect types are divided into grating line offset, ambient light interference, and uneven material thickness distribution; the determination condition for grating line offset is that the structured light distortion coefficient exceeds the preset threshold and the pixel consistency coefficient is lower than the preset threshold; the determination condition for ambient light interference is that the ambient light interference exceeds the preset threshold; the determination condition for uneven material thickness distribution is that the structured light distortion coefficient is lower than the preset threshold and the pixel consistency coefficient is lower than the preset threshold.
[0014] Preferably, the three-level feedback calibration includes: process parameter adjustment, ambient light control, and model learning optimization. The process parameter adjustment is specifically that if a specific defect is detected continuously for multiple times, the production parameters of the grating film are automatically adjusted. The production parameters of the grating film include imprinting pressure, heating roller temperature, and the scanning angle of the multi-axis robotic arm; the ambient light control is specifically that when the ambient light interference coefficient exceeds the preset threshold, the darkroom light-shielding curtain and supplementary lighting fixtures are automatically adjusted; the model learning optimization is specifically that samples with the quality levels of level two and level three of the grating film are preferentially learned every week, and the model of the AI analysis terminal is incrementally trained using the contrastive learning algorithm.
[0015] The technical effects and advantages of the present invention: 1. Through multi-perspective structured light image acquisition and multi-modal data fusion, the present invention realizes full-coverage detection of the surface topography of the grating film. Using five-perspective structured light scanning technologies such as directly in front, horizontal left and right, and vertical up and down, combined with ambient light and physical height data, it effectively eliminates detection blind spots and significantly improves the recognition accuracy of surface defects such as grating line offset and uneven material thickness; 2. Through the AI large model Transformer and the dynamic feature weighting mechanism, the present invention breaks through the limitations of traditional algorithms. Based on the deep fusion of the structured light distortion coefficient, ambient light interference coefficient, and pixel consistency coefficient, a normalized comprehensive detection index is generated, which can accurately quantify the quality of the grating film and achieve highly robust classification of complex defects, reducing the false detection rate; 3. Through the calibration terminal to perform three-level feedback calibration, including process parameter adjustment, ambient light control, and model learning optimization, the present invention realizes automatic adjustment of production parameters, dynamic control of ambient light, and incremental training and optimization of the model, solves the problem that the traditional detection system lacks a dynamic feedback and adaptive optimization mechanism, improves the detection efficiency, reduces the false detection rate, and realizes the full-process intelligent management and control of the quality of the grating film. Brief Description of the Drawings
[0016] Figure 1 This is the system framework diagram of the present invention.
[0017] Figure 2 This is the flowchart of the method steps of the present invention. Detailed Embodiments
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] The present invention provides a method for detecting a lenticular film of a naked-eye 3D display based on an AI large model as shown in Figure 1 , including a lenticular film data acquisition terminal, a lenticular film intelligent analysis terminal, an AI analysis terminal, a calibration terminal, and an output terminal; It should be noted that the lenticular film data acquisition terminal is composed of the following devices: a high-resolution industrial camera, such as Sony IMX586, with 12 million pixels and a frame rate of 30fps, for collecting five-view structured light images; a multi-axis robotic arm, such as a UR10 six-axis robotic arm, preset with five-view coordinates, ±15° horizontally and ±10° vertically, to achieve precise angle positioning; a laser rangefinder, such as Keyence LK-H008, with an accuracy of ±0.01 mm, for measuring the actual surface height of the lenticular film; a spectroradiometer, such as KonicaMinolta CL-500A, for real-time collection of ambient illuminance in the range of 0 - 1000 lux and color temperature in the range of 2700 - 6500K; a constant temperature control system for maintaining the darkroom temperature at 20 - 25°C; a light-shielding curtain, such as a Somfy LT50 electric light-shielding curtain, for controlling the darkroom illuminance ≤10 lux; The lenticular film intelligent analysis terminal is equipped with an Intel Xeon processor and runs image preprocessing algorithms; The AI analysis terminal is composed of a GPU cluster server and a model inference module. The GPU cluster server is equipped with NVIDIA DGX A100, such as 8×A100 GPUs. The model inference module is based on the PyTorch framework and supports real-time calculation of the comprehensive detection index; The calibration terminal consists of a PLC controller, an LED fill light, a light-shielding curtain actuator, and an incremental training server. The PLC controller is a Siemens S7-1500, which adjusts the parameters of the grating film embossing machine according to the defect type, such as the pressure within ±5%. The LED fill light cooperates with a light sensor to dynamically adjust the brightness. The light-shielding curtain actuator is a Somfy LT50 electric motor, which automatically closes when the ambient light interference coefficient exceeds 0.5. The incremental training server is equipped with an AMD EPYC processor and runs the SimCLR contrastive learning algorithm, updating the model parameters weekly. The output terminal is equipped with an industrial-grade display, such as a Dell UltraSharp 32-inch, 4K resolution display, for real-time display of the detection report. The method of the present invention is as Figure 2 shown and includes the following steps: S1. The grating film data acquisition terminal performs a line-by-line scan of the grating film of the naked-eye 3D display in a darkroom environment to obtain a five-view structured light image dataset, an ambient light dataset, and an image physical dataset containing spatial coordinate information. Further, in the above technical solution, the five-view structured light image dataset includes: a front-center view image pixel set, a horizontal left view image pixel set, a horizontal right view image pixel set, a vertical top view image pixel set, and a vertical bottom view image pixel set; the ambient light dataset includes: ambient illuminance and color temperature; the image physical dataset includes: the width of the image, the height of the image, the coordinates of the image pixel points, and the actual surface height of the grating film. It should be noted that the front-center view is 0°, that is, the front view direction perpendicular to the center point of the grating film surface. Specifically, the axis of the camera lens is perpendicular to the grating film plane and aligned with the geometric center of the grating film. The purpose is to obtain a reference image for comparison with other view images to detect symmetry distortion or central region defects. The horizontal left view is -15°±2°, that is, a horizontal offset of 15° to the left from the front-center view, allowing a mechanical arm adjustment error of ±2°. Specifically, the camera rotates horizontally to the left, maintaining parallel to the grating film surface, covering the left edge area. The purpose is to detect the offset, stretching, or distortion of the grating lines in the horizontal left direction. The horizontal right view is +15°±2°, that is, a horizontal offset of 15° to the right from the front-center view, allowing a mechanical arm adjustment error of ±2°. Specifically, the camera rotates horizontally to the right, symmetric with the left view. The purpose is for the camera to rotate horizontally to the right, symmetric with the left view. The vertical upward viewing angle is +10° ± 2°, that is, a vertical offset of 10° upward from the center of the front, allowing a robotic arm adjustment error of ±2°. Specifically, the camera tilts upward along the vertical axis to cover the top area of the grating film, aiming to detect grating line breaks, uneven material thickness, or surface warping in the vertical direction; The vertical downward viewing angle is -10° ± 2°, that is, a vertical offset of 10° downward from the center of the front, allowing a robotic arm adjustment error of ±2°. Specifically, the camera tilts downward along the vertical axis to cover the bottom area of the grating film, aiming to detect the bottom grating line fit, uneven material deposition, or mechanical stress damage; The set of image pixels of the center front view is the set of all pixel points in the center front view image, the set of image pixels of the horizontal left view is the set of all pixel points in the horizontal left view image, the set of image pixels of the horizontal right view is the set of all pixel points in the horizontal right view image, the set of image pixels of the vertical upward view is the set of all pixel points in the vertical upward view image, and the set of image pixels of the vertical downward view is the set of all pixel points in the vertical downward view image; S2. The intelligent analysis terminal of the grating film preprocesses and analyzes the five-view structured light image dataset, the ambient light dataset, and the image physical dataset to obtain the structured light distortion coefficient, the ambient light interference coefficient, and the pixel consistency coefficient; Furthermore, in the above technical solution, the structured light distortion coefficient is specifically: , where M is the total number of pixel points, specifically M = W × H, where W is the width of the image and H is the height of the image, x is the abscissa of the pixel position in the image, y is the ordinate of the pixel position in the image, max(Δh) is the maximum value among all height change amounts Δh(x, y), Δh(x, y) is the height change amount at the coordinate (x, y), where Δh(x, y) = Z actual (x, y) - Z ideal (x, y), where Z actual (x, y) is the actual surface height of the grating film at the coordinate (x, y), and Z ideal (x, y) is the ideal fitting plane; It should be noted that the ideal fitting plane Z ideal (x, y) = a × x + b × y + c, where the parameters a, b, and c are calculated by least squares fitting of the actual height data.
[0020] In a specific embodiment, a laser rangefinder is used to scan the surface of the grating film to obtain the actual height Z of each pixel point (x, y). actual (x, y). For example, scanning a grid with W × H = 1000 × 1000 gives 106 For a pixel, through least squares fitting of the actual height data, it is calculated that a = 0.001, b = -0.0005, c = 2. If the actual height of a pixel at coordinates (500, 500) is 2.4 mm, then Δh(500, 500) = 0.25 mm. If max(Δh) = 0.5 mm, then through calculation, K = 0.08; Furthermore, in the above technical solution, the ambient light interference coefficient is specifically: , where e is the Euler number, μ E is the ideal mean value of the ambient illuminance, μ T is the ideal mean value of the color temperature, σ E is the standard deviation of the normalized illuminance value E′, σ T is the standard deviation of the normalized color temperature value T′, is the average value of the normalized ambient illuminance value E′, is the average value of the normalized color temperature value T′, where, , where E is the ambient illuminance collected in real time, E min is the preset minimum illuminance value, E max is the preset maximum illuminance value, , where T is the ambient light color temperature collected in real time, T min is the preset minimum color temperature value, T max is the preset maximum color temperature value; It should be noted that the ideal mean value μ of the ambient illuminance E = 0.5, and the ideal mean value μ of the color temperature T = 0.5; In a specific embodiment, the preset minimum illuminance value E min = 0 lux, the preset maximum illuminance value Emax = 10 lux, the preset minimum color temperature value T min = 2700 K, the preset maximum color temperature value illuminance T max = 6500 K, E = 8 lux. Through calculation of historical data, σ E = 0.2, σ T = 0.1. Finally, the calculated ambient light interference coefficient M ≈ 0.98, indicating that the ambient light interference is small.
[0021] Furthermore, in the above technical solution, the pixel consistency coefficient is specifically: , where γ j is the cosine similarity between a single-frame registered image and the central-view image, specifically: , Among them, is the pixel value at (x, y) of the registered image of the j-th perspective, is the pixel value at (x, y) of the central perspective image, is the pixel mean value of the registered image of the j-th perspective, is the pixel mean value of the central perspective image; It should be noted that j = 1 is the horizontal left perspective, j = 2 is the horizontal right perspective, j = 4 is the vertical upper perspective, and j = 5 is the vertical lower perspective; In a specific embodiment, the central perspective image S is set 3(x,y) and the registered perspective image is 2×2 pixels. Among them, the pixel values of the central perspective image are 10, 20, 30, 40; the pixel values of the horizontal left perspective image are 12, 18, 28, 32; the pixel values of the horizontal right perspective image are 11, 19, 29, 31; the pixel values of the vertical upper perspective image are 13, 17, 27, 33; the pixel values of the vertical lower perspective image are 14, 16, 26, 34; when j = 1, γ1≈0.986; when j = 2, γ2≈0.971, when j = 3, γ4≈0.986, when j = 4, γ4≈0.975; finally, it is calculated that P≈0.98, indicating that the registered image and the central perspective image have a high degree of consistency at the pixel level; S3. The AI analysis terminal inputs the structured light distortion coefficient, the ambient light interference coefficient, and the pixel consistency coefficient into the Transformer model, outputs the to-be-processed comprehensive detection index, and obtains the comprehensive detection index after normalization processing; Furthermore, in the above technical solution, the Transformer model is specifically: , where e is the Euler number, l is the layer index, L is the total number of layers of the Transformer model, α l is the attention weight of the l-th layer, F l is the representation of the input feature vector F at the l-th layer, FFN is the feed-forward network, MSA is the multi-head self-attention mechanism, and b is the bias term; It should be noted that the Transformer model architecture: 12-layer encoder, 8-head self-attention mechanism, and the feed-forward network dimension is 2048; In a specific embodiment, the three coefficients are combined into the input feature vector F = [K, M, P] = [0.08, 0.98, 0.98], assuming the bias term b = 0, and the attention weight of each layer is evenly distributed, that is, the attention weight , and the calculation process is specifically: Step 1, input the values F = [K, M, P] = [0.08, 0.98, 0.98]; Step 2, each attention head generates a query vector Q, a key vector K, and a value vector V for the input F, and calculates the attention scores: , where d K is the dimension of the key vector, set to 3. The output after concatenating the results of 8-head attention is MSA(F′), which represents the representation of the features weighted by the attention mechanism; Step 3, perform a non-linear transformation on MSA(F′): , where b1 and b2 are bias terms, W1 ∈ R 3×2048 , W2 ∈ R 2048×1 are learnable weights, where R represents the real number space, that is, the elements in the matrix are all real numbers. Finally, the output is obtained, and after passing through the Sigmoid function, the normalized is obtained, and the quality level is first level, judged as qualified.
[0022] S4. The AI analysis terminal performs quality grading on the grating film according to the comprehensive detection index, and at the same time combines the structural light distortion coefficient, the ambient light interference coefficient, and the pixel consistency coefficient to locate the defect type, and sends a detection report generation instruction to the output terminal and a calibration instruction to the calibration terminal; Further, in the above technical solution, the quality grading can be divided into the following levels: When the comprehensive detection index is within the preset threshold range S1, it is divided into the first level and judged as qualified; When the comprehensive detection index is within the preset threshold range S2, it is divided into the second level and judged as needing re-inspection; When the comprehensive detection index is within the preset threshold range S3, it is divided into the third level and judged as unqualified.
[0023] It should be noted that S1 ∈ [0.9, 1], S2 ∈ [0.7, 0.9), S3 ∈ [0, 0.7); Further, in the above technical solution, the defect types are divided into grating line offset, ambient light interference, and uneven material thickness distribution; the determination condition for the grating line offset is that the structural light distortion coefficient exceeds the preset threshold and the pixel consistency coefficient is lower than the preset threshold; the determination condition for the ambient light interference is that the ambient light interference exceeds the preset threshold; the determination condition for the uneven material thickness distribution is that the structural light distortion coefficient is lower than the preset threshold and the pixel consistency coefficient is lower than the preset threshold; It should be noted that the determination condition for the grating line offset is K > 0.1 and P < 0.9; the determination condition for the ambient light interference is M > 0.7; the determination condition for the uneven material thickness distribution is K < 0.02 and P < 0.85; S5. The output terminal generates a visual inspection report upon receiving the instruction sent by the AI analysis terminal; S6. The calibration terminal executes three - level feedback calibration upon receiving the instruction sent by the AI analysis terminal.
[0024] Further, in the above - mentioned technical solution, the three - level feedback calibration includes: process parameter adjustment, ambient light control, and model learning optimization. The process parameter adjustment specifically means that if a specific defect is detected continuously for multiple times, the production parameters of the grating film are automatically adjusted. The production parameters of the grating film include imprinting pressure, heating roller temperature, and the scanning angle of the multi - axis robotic arm. The ambient light control specifically means that when the ambient light interference coefficient exceeds a preset threshold, the light - proof curtain of the darkroom and the supplementary lighting fixtures are automatically adjusted. The model learning optimization specifically means that samples with a grating film quality level of two and three are preferentially learned every week, and a contrast learning algorithm is used to perform incremental training on the model of the AI analysis terminal.
[0025] It should be noted that the imprinting pressure is dynamically adjusted within the range of ±5% to improve the forming accuracy of the grating lines. The heating roller temperature is dynamically adjusted within the range of ±3°C to optimize the material melting uniformity. The scanning angle of the multi - axis robotic arm is dynamically adjusted within the range of ±1° to correct the detection perspective deviation caused by equipment wear. It should be noted that when the ambient light interference coefficient exceeds the preset threshold of 0.5, the calibration terminal automatically adjusts the light - proof curtain of the darkroom. The supplementary lighting fixture is an LED supplementary light combined with a light sensor. For example, Konica Minolta CL - 500A dynamically adjusts the brightness, and the color temperature range is adjustable from 2700 - 6500K. When the measured illuminance E > 10 lux, the light - proof curtain is automatically closed. If it still does not meet the standard, a low - power supplementary light, such as 50 lux, is turned on to balance the brightness. When the color temperature T deviates from the ideal average value μT = 0.5, corresponding to approximately 4600K, by more than ±200K, it is corrected through the spectral adjustment of the supplementary light. The contrast learning algorithm is the SimCLR contrast learning algorithm. When running the SimCLR contrast learning algorithm, every week, defect samples with a quality level of two and three are selected from the detection data. Geometric transformation, photometric perturbation, etc. are performed on their five - perspective structured light images to generate positive sample pairs, and other sample enhanced views are randomly selected as negative sample pairs. Image feature vectors are extracted through the ResNet encoder and the projection head, and the NT - Xent loss function is used to maximize the similarity of positive samples and minimize the similarity of negative samples for training. During training, only some parameters of the model are updated, and the learned image features are combined with traditional detection coefficients to perform incremental optimization on the Transformer model to improve the detection accuracy of grating film defects.
[0026] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting a lenticular film of a naked-eye 3D display based on an AI large model, characterized in that, It includes a grating film data acquisition terminal, a grating film intelligent analysis terminal, an AI analysis terminal, a calibration terminal and an output terminal. The specific steps are as follows: S1. The grating film data acquisition terminal performs line-by-line scanning and acquisition on the grating film of the naked-eye 3D display in a darkroom environment to obtain a five-view structured light image dataset containing spatial coordinate information, an ambient light dataset and an image physical dataset; S2. The grating film intelligent analysis terminal preprocesses and analyzes the five-view structured light image dataset, the ambient light dataset and the image physical dataset to obtain a structured light distortion coefficient, an ambient light interference coefficient and a pixel consistency coefficient; S3. The AI analysis terminal inputs the structured light distortion coefficient, the ambient light interference coefficient and the pixel consistency coefficient into the Transformer model, outputs a comprehensive detection index to be processed, and obtains a comprehensive detection index after normalization; S4. The AI analysis terminal grades the quality of the grating film according to the comprehensive detection index, and at the same time locates the defect type in combination with the combination of the structured light distortion coefficient, the ambient light interference coefficient and the pixel consistency coefficient, and sends a detection report generation instruction to the output terminal and a calibration instruction to the calibration terminal; S5. The output terminal receives the instruction sent by the AI analysis terminal and generates a visual detection report; S6. The calibration terminal receives the instruction sent by the AI analysis terminal and performs three-level feedback calibration.
2. The method for detecting the grating film of a naked-eye 3D display based on an AI large model according to claim 1, wherein, The five-view structured light image dataset includes: a front-center view image pixel set, a horizontal left view image pixel set, a horizontal right view image pixel set, a vertical top view image pixel set and a vertical bottom view image pixel set; the ambient light dataset includes: ambient light illuminance and color temperature; the image physical dataset includes: the width of the image, the height of the image, the coordinates of the image pixel points and the actual surface height of the grating film.
3. A method for detecting a lenticular film of a naked-eye 3D display based on an AI large model according to claim 1, wherein, The structured light distortion coefficient is specifically: , Among them, M is the total number of pixel points, specifically M = W × H, where W is the width of the image and H is the height of the image. x is the abscissa of the image pixel position, y is the ordinate of the image pixel position, max(Δh) is the maximum value among all height changes Δh(x, y), and Δh(x, y) is the height change at the coordinate (x, y), where Δh(x, y) = Z actual (x, y) - Z ideal (x, y), where Z actual (x, y) is the actual surface height of the grating film at the coordinate (x, y), and Z ideal (x, y) is the ideal fitting plane.
4. A method for detecting a lenticular film of a naked-eye 3D display based on an AI large model according to claim 1, characterized in that, The ambient light interference coefficient is specifically: , where e is the Euler number, μ E is the ideal mean value of the ambient illuminance, μ T is the ideal mean value of the color temperature, σ E is the standard deviation of the normalized illuminance value E′, σ T is the standard deviation of the normalized color temperature value T′, is the average value of the normalized ambient illuminance value E′, is the average value of the normalized color temperature value T′, where, , where E is the ambient illuminance collected in real time, E min is the preset minimum illuminance value, E max is the preset maximum illuminance value, , where T is the ambient light color temperature collected in real time, T min is the preset minimum color temperature value, T max is the preset maximum color temperature value.
5. A method for detecting a lenticular film of a naked-eye 3D display based on an AI large model according to claim 1, characterized in that, The pixel consistency coefficient is specifically: , where γ j is the cosine similarity between a single registered image and the central view image, specifically: , wherein, is the pixel value of the registered image of the j-th view at (x, y), is the pixel value of the central view image at (x, y), is the average pixel value of the registered image of the j-th view, is the average pixel value of the central view image.
6. The method for detecting a lenticular film of a naked-eye 3D display based on an AI large model according to claim 1, wherein The Transformer model is specifically: , Among them, e is the Euler number, l is the layer index, L is the total number of layers of the Transformer model, and α l is the attention weight of the l-th layer, F l is the representation of the input feature vector F at the l-th layer, FFN is the feed-forward network, MSA is the multi-head self-attention mechanism, and b is the bias term.
7. A method for detecting a grating film of a naked-eye 3D display based on an AI large model according to claim 1, characterized in that, The quality grading can be divided into the following levels: When the comprehensive detection index is within the preset threshold S1 range, it is classified as level one and judged to be qualified; When the comprehensive detection index is within the preset threshold S2 range, it is classified as level two and judged to require re-inspection; When the comprehensive detection index is within the preset threshold S3 range, it is classified as level three and judged to be unqualified.
8. A method for detecting a grating film of a naked-eye 3D display based on an AI large model according to claim 1, characterized in that The defect types are divided into grating line offset, ambient light interference and uneven material thickness distribution; the judgment condition for grating line offset is that the structured light distortion coefficient exceeds the preset threshold and the pixel consistency coefficient is lower than the preset threshold; the judgment condition for ambient light interference is that the ambient light interference exceeds the preset threshold; the judgment condition for uneven material thickness distribution is that the structured light distortion coefficient is lower than the preset threshold and the pixel consistency coefficient is lower than the preset threshold.
9. The method for detecting a grating film of a naked-eye 3D display based on an AI large model according to claim 1, wherein The three-level feedback calibration includes: process parameter adjustment, ambient light control, and model learning optimization. The process parameter adjustment specifically means that if a specific defect is detected continuously for multiple times, the production parameters of the grating film are automatically adjusted. The production parameters of the grating film include imprinting pressure, heating roller temperature, and scanning angle of the multi-axis robotic arm. The ambient light control specifically means that when the ambient light interference coefficient exceeds a preset threshold, the darkroom light-shielding curtain and supplementary lighting fixtures are automatically adjusted. The model learning optimization specifically means that samples with a grating film quality level of two and three are preferentially learned every week, and the model of the AI analysis terminal is incrementally trained using a contrastive learning algorithm.
Citation Information
Patent Citations
METHOD FOR SELECTING AN OPTIMIZED EVALUATION SUB-FEATURE FOR CONTROLLING A FREEFORM SURFACE AND METHOD FOR CONTROLLING A FREEFORM SURFACE
AT11770U1
Rapid defect detection method based on structured light projection
CN115184362A
Surface defect detection method and system based on computer vision
CN118396944A
Method and device for detecting appearance of thin-walled part
CN118761998A
PCB production line defect detection system and method based on image processing
CN119831966A