Artificial intelligence machine vision image acquisition system

By co-optimizing the multimodal perception layer, dynamic adaptive layer, and cognitive reasoning layer, and combining event camera and multispectral imaging technology, the problem of insufficient recognition accuracy of machine vision image acquisition systems in complex environments is solved, and efficient and stable image acquisition and decision optimization are achieved.

CN120912837APending Publication Date: 2025-11-07南昌理工学院
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510997478.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing machine vision image acquisition systems have insufficient recognition accuracy when faced with factors such as occlusion, deformation, and changes in lighting. In particular, metal reflections cause edge detection algorithms to fail in the inspection of industrial parts, affecting production efficiency and quality.

Method used

A multimodal perception layer is employed to suppress metallic reflection interference through adaptive optics modules and multispectral imaging units, combined with an event camera to capture edge changes of moving targets, and distortion correction and multi-source data alignment are performed during data preprocessing. A dynamic adaptive layer adjusts parameters in real time to enhance the algorithm's anti-interference capability and optimizes image acquisition using illumination and motion compensation submodules. A cognitive reasoning layer constructs an interpretable algorithm architecture, optimizes feature extraction through a dynamic routing network, and deploys a causal reasoning module to reduce data bias. A collaborative decision-making layer performs cloud-edge-device collaborative decision-making and constructs a human-machine collaborative interface to improve decision accuracy and efficiency.

Benefits of technology

It effectively improves the recognition accuracy and efficiency of machine vision image acquisition systems in complex environments, solves the problem of insufficient algorithm robustness, and achieves high-quality image acquisition and decision optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912837A_ABST
    Figure CN120912837A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence machine vision image acquisition system, and the system comprises a multi-mode perception layer which integrates a self-adaptive optical module, inhibits metal reflection, captures a visible light to short wave infrared image, and captures a motion edge; the dynamic adaptive layer adopts an illumination compensation and motion compensation module to dynamically adjust camera parameters and micro displacement compensation, feeds back an illumination trend, outputs a motion vector to the cognitive layer, generates a confrontation sample through a GAN, simulates virtual defects in combination with a physical engine, and expands training data; the cognitive reasoning layer is used for deploying a dynamic routing network, distributing computing resources according to image complexity and optimizing feature extraction efficiency; reducing data deviation through anti-fact analysis, and generating a thermodynamic diagram to explain a detection basis; and the collaborative decision-making layer is used for rapidly screening samples by edge nodes, training a global model by cloud aggregated data, automatically triggering manual rechecking when the confidence coefficient of the model is insufficient, synchronously optimizing a training set and a causal reasoning module by a rechecking result, and improving the labeling efficiency through AR assistance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of machine vision image acquisition, and in particular to an artificial intelligence machine vision image acquisition system. BACKGROUND

[0002] The machine vision image acquisition system can effectively improve detection efficiency, reduce human error, avoid missed detection and misjudgment caused by artificial fatigue, support data-driven decision-making, and has been deeply integrated into key fields such as industrial manufacturing, quality detection, medical health, and safety monitoring, and has become the core infrastructure for promoting the intelligent upgrading of various industries.

[0003] However, with the wide application of the machine vision image acquisition system, the problem of insufficient algorithm robustness has become increasingly prominent. Factors such as occlusion, deformation, and light changes significantly reduce the recognition accuracy. For example, in industrial part detection, metal reflection may cause the edge detection algorithm to fail, which seriously affects the efficiency and quality of industrial production. Therefore, an artificial intelligence machine vision image acquisition system is proposed. SUMMARY

[0004] The application aims to solve the problems in the prior art and provides an artificial intelligence machine vision image acquisition system.

[0005] In order to achieve the above-mentioned purpose, the application adopts the following technical scheme:

[0006] An artificial intelligence machine vision image acquisition system comprises:

[0007] A multi-modal perception layer: through the integration of multiple sensors and imaging technologies, multi-dimensional and high-precision data acquisition of the target scene is performed; an adaptive optical module is used to dynamically adjust the incident light by using a liquid lens and a polarizing filter array to suppress metal reflection interference; at the same time, a multi-spectral imaging unit is deployed to synchronously capture visible light, near-infrared, and short-wave infrared band images to enhance the separability of defect features; in addition, an event camera is introduced to capture the edge changes of moving targets with a time resolution of microseconds to solve the smear problem in high-speed scenes; and in the data preprocessing stage, real-time distortion correction and multi-source data alignment technology are used to adjust the accuracy and consistency of image data; these data are then transmitted to the dynamic adaptive layer for further environmental perception and parameter adjustment.

[0008] Dynamic self-adaptive layer: dynamically adjust parameters according to real-time environmental information, and improve the algorithm's resistance to complex interference through data enhancement technology; through the light adaptive sub-module and the motion compensation sub-module, real-time monitoring of environmental light changes and the motion state of the measured object, dynamically adjusting the camera gain, exposure time and micro displacement compensation, adjusting the stability and clarity of image acquisition, feeding back the light intensity change trend to the multi-modal perception layer, and generating a motion state vector input to the cognitive reasoning layer; at the same time, deploy a data enhancement engine, use a generative adversarial network to generate adversarial samples in real time, simulate complex interference such as oil stains and scratches, and generate virtual defect samples through a physical simulation engine, effectively solving the small sample problem; the processed multi-modal data is input to the cognitive reasoning layer as the input of the algorithm reasoning;

[0009] Cognitive reasoning layer: introduce the principles of cognitive science to build an algorithm architecture with explainability and anti-interference ability; adopt a dynamic routing network to automatically select the optimal feature extraction path according to the complexity of the input image, and real-time feedback the feature extraction efficiency index to the dynamic self-adaptive layer for rational allocation of computing resources and optimization of algorithm performance; at the same time, deploy a causal reasoning module to reduce data bias through counterfactual learning and generate a feature importance heat map to provide explainability output; the model prediction result and explainability output are then passed to the collaborative decision-making layer to trigger the collaborative decision-making process.

[0010] Collaborative decision-making layer: through the collaborative mechanism, break through the performance boundary of a single device, and improve the accuracy and efficiency of the overall decision; cloud, edge and end collaboration, edge computing nodes deploy lightweight models for preliminary screening, cloud aggregates multi-device data to train global models, and protects data security through differential privacy; at the same time, build a human-computer collaboration interface, automatically trigger manual review requests when the model confidence is insufficient, and feed back the review results to the training set and the cognitive reasoning layer to optimize the causal reasoning module; AR assisted labeling improves the efficiency of manual labeling; these designs can fully utilize the computing resources of the cloud and edge, and perform collaborative work of multiple devices and multiple models, while combining human intervention, significantly improving the decision-making accuracy and efficiency in complex scenarios; finally, the collaborative decision-making layer issues global optimization instructions to the multi-modal perception layer, forming an end-to-end closed-loop optimization.

[0011] The above technical solution further comprises:

[0012] Further, the adaptive optical module uses a liquid lens and a polarizing filter array to dynamically adjust the incident light and suppress metal reflection interference, including the following steps:

[0013] Liquid lens dynamic focusing:

[0014] Liquid lens consists of a container filled with optical liquid and an elastic polymer membrane; when detecting distance changes of the object, the shape of the liquid is changed by applying voltage or pressure, achieving fast zooming in milliseconds, and the image remains clear throughout the process, even when detecting reflective metal workpieces, the liquid lens can maintain the focus on the workpiece surface;

[0015] Polarization filter array reflection suppression:

[0016] The polarization filter array is composed of multiple polarization units, each unit only allows polarized light of a specific vibration direction to pass; when the non-polarized light reflected by the metal surface enters the array, part of the light of the polarization direction is blocked, effectively weakening the reflection intensity.

[0017] Adaptive optics cooperative work:

[0018] Liquid lens and polarization filter array work together, liquid lens is responsible for adjusting the focal length to adjust the clear imaging of the target; while the polarization filter array synchronously suppresses the metal reflection; for example, when detecting reflective metal workpieces, the liquid lens maintains the focus on the workpiece surface, while the polarization filter reduces the reflection intensity, both of which work together to obtain high-quality images;

[0019] Real-time environmental light monitoring compensation:

[0020] The light adaptive sub-module continuously monitors the intensity and color temperature of the ambient light through the photosensitive sensor; according to these monitoring data, dynamically adjust the optical parameters of the liquid lens, such as focal length and aperture, to obtain clear and stable images under different lighting conditions;

[0021] Wavefront distortion correction:

[0022] In high-end applications, the adaptive optics module can integrate a wavefront sensor. This sensor can measure the wavefront distortion of the incident light in real time and calculate the corresponding compensation amount. Then, drive the liquid lens to perform phase correction to further improve the resolution and quality of the image.

[0023] Closed-loop feedback optimization:

[0024] After image acquisition, analyze the image quality indicators, if the reflection residue or focus deviation is detected, automatically adjust the parameters of the liquid lens and the angle of the polarization filter, forming a closed-loop optimization mechanism to continuously improve the image quality;

[0025] Multi-spectrum fusion enhancement:

[0026] In combination with the data of the multispectral imaging unit, the image after suppressing the reflection is subjected to spectral fusion processing; for example, while suppressing the reflection in the visible light band, the short-wave infrared band is used to enhance the visibility of the metal surface defects; such multispectral fusion technology helps to improve the comprehensiveness and accuracy of detection.

[0027] Further, the introduction event camera captures the edge change of the moving target, and in the data preprocessing stage, the accuracy and consistency of the image data are adjusted through real-time distortion correction and multi-source data alignment technology, including the following steps:

[0028] Capture the edge change of the moving target:

[0029] Each pixel of the event camera independently monitors the brightness change, and when the change exceeds the preset threshold (such as ΔL≥15%), an event output is triggered, each event contains a timestamp (accuracy≤1μs), pixel coordinates (x, y) and polarity (brightness increase or decrease flag); the event camera only outputs the event stream of brightness change, reducing the data volume and retaining only the motion edge information; through the event clustering algorithm (such as Hough transform based on spatial and temporal proximity), the motion boundary is tracked in real time, and the target motion trajectory is generated;

[0030] Real-time distortion correction:

[0031] A polynomial model (such as Brown-Conrady model) is used to fit the radial distortion (k1, k2, k3) and tangential distortion (p1, p2), and the distortion coefficients are obtained through the chessboard calibration method; based on the atmospheric scattering model, the uneven light is compensated, the transmittance is estimated using the dark channel prior algorithm, and the image contrast is restored; the inconsistency of the sensor response is compensated, and the image brightness uniformity is adjusted;

[0032] Multi-source data alignment:

[0033] The precise time protocol (PTP) is used to adjust the timestamp synchronization accuracy of the event camera data and the multispectral image, the spatial alignment of the multi-source data is performed through the feature point matching algorithm (such as SIFT), and the registration error is less than 2 pixels; a cross-modal feature correlation matrix is constructed, the edge features of the event stream and the texture features of the multispectral image are fused, and the target recognition accuracy is improved; according to the scene dynamics, the fusion weight of the event data and the multispectral data is adjusted, and the target detection performance is optimized;

[0034] Adjust the accuracy and consistency of the image data:

[0035] Adopting decision-level fusion strategy, the event camera provides motion boundary localization (accuracy ±1 pixel), and the multispectral image supplements material classification information (accuracy ≥95%); geometric consistency is ensured through bidirectional projection verification (projecting the corrected data back to the original coordinate system, error ≤3%); histogram matching algorithm (correlation coefficient ≥0.9) is adopted to adjust the consistency of the radiation characteristics of the multi-source data; the continuity of the motion target trajectory is verified through time series analysis, and data jumps and loss are eliminated.

[0036] Further, the light adaptation sub-module and the motion compensation sub-module are used to monitor the ambient light change and the motion state of the measured object in real time, dynamically adjust the camera gain, exposure time and perform micro-displacement compensation, and adjust the stability and clarity of image acquisition, including the following steps:

[0037] The light adaptation sub-module comprises:

[0038] A high-precision photosensitive sensor array is used to collect the ambient light intensity (0.1 lux to 100,000 lux) and color temperature (2000K to 10000K) in real time, obtain spatial distribution data through multi-region sampling (such as a 3x3 grid) to identify uneven light areas, dynamically adjust the camera gain (0dB to 36dB) based on the PID algorithm, the response time is ≤5ms, and the signal intensity is adjusted in the best working interval (30% to 70% saturation); a bilinear interpolation algorithm is used to adaptively adjust the exposure time (1μs to 1s) according to the target motion speed (obtained through the motion compensation sub-module), the higher the motion speed, the shorter the exposure time; a tunable optical filter is used to dynamically adjust the spectral response curve to match the current ambient light characteristics and improve the signal-to-noise ratio (SNR≥40dB); the light intensity change trend (ΔL / Δt) is fed back to the liquid lens control module of the multi-modal perception layer to pre-adjust the optical parameters to adapt to the change of the ambient light;

[0039] The motion compensation sub-module comprises:

[0040] The event stream data is analyzed to extract the edge change rate (pixels / second) and direction vector of the moving target; combined with 6-axis IMU data (accelerometer ±16g range, gyroscope ±2000° / s range), the six-degree-of-freedom motion trajectory of the measured object is obtained through Kalman filter fusion; a motion prediction model is established based on the LSTM network to predict the displacement of the measured object in advance; a liquid lens is driven for active optical anti-shake, and a piezoelectric ceramic micro-displacement platform is used for mechanical compensation; a digital image stabilization algorithm based on feature point matching is used to compensate for residual jitter (≤0.1 pixel); a motion state vector (including speed, acceleration and direction information) is generated and input to the cognitive inference layer at a frequency of 1kHz;

[0041] Collaborative optimization mechanism

[0042] When high-speed motion (speed > 100 mm / s) is detected, the exposure time is shortened (to below 1 ms) and the gain is increased (to 24 dB) to maintain the signal-to-noise ratio; in low light environments (<100 lux), the exposure time is appropriately extended (to 100 ms), and motion blur is suppressed through motion compensation; a no-reference image quality assessment algorithm (such as BRISQUE) is used to calculate the image sharpness index (0-100 points) in real time; when the sharpness is <70 points, the parameter adjustment process is triggered, and the MTF (modulation transfer function) value is optimized to >0.3; the sub-module power consumption is dynamically adjusted according to the workload, and the idle state power consumption is <1W, and the full load state power consumption is less than 5W.

[0043] Further, the light intensity change trend is fed back to the multi-modal perception layer, and a motion state vector is input to the cognitive inference layer, comprising the following steps:

[0044] The light intensity change trend is fed back to the multi-modal perception layer:

[0045] The light intensity change trend is fed back to the multi-modal perception layer:

[0046] The motion state vector is input to the cognitive inference layer:

[0047] The motion compensation sub-module combines event camera data and IMU (inertial measurement unit) data to obtain a six-degree-of-freedom motion trajectory (x, y, z, roll, pitch, yaw) of a measured object through a Kalman filter fusion; the event camera provides high-precision edge change information (positioning accuracy ±1 pixel), and the IMU provides high-frequency motion data (sampling rate ≥1 kHz); based on the fused motion trajectory, motion parameters such as the speed (v), acceleration (a), and direction (θ) of the measured object are calculated, a motion state vector (v, a, θ) is generated, and three-dimensional space motion information and a timestamp (accuracy ≤1 μs) are included; the motion state vector is input to the cognitive inference layer at a frequency of 1 kHz through a high-speed bus (such as PCIe Gen4); in the cognitive inference layer, the motion state vector is spatiotemporally aligned with the multispectral image and event stream data, and the synchronization of feature extraction is adjusted; the dynamic routing network of the cognitive inference layer adjusts the feature extraction path in real time according to the motion state vector, for example, when high-speed motion (v>100 mm / s) is detected, a feature extraction algorithm with high computational efficiency and anti-motion blur is preferentially selected; the causal inference module analyzes the target motion mode using the motion state vector, optimizes the counterfactual learning strategy, and improves the decision robustness.

[0048] Further, the deployment data augmentation engine generates adversarial samples using a generative adversarial network (GAN), simulates complex interference, and generates virtual defect samples using a physics simulation engine, including the following steps:

[0049] The generative adversarial network (GAN) generates adversarial samples:

[0050] A conditional generative adversarial network (CGAN) is used, and a defect type label (such as oil stain, scratch, deformation) is introduced as a generation condition; the generator (Generator) is composed of a residual network (ResNet) and includes 9 residual blocks, and outputs adversarial samples with a resolution of 512×512; the discriminator (Discriminator) uses a PatchGAN structure and outputs a 70×70 probability map to locally determine the authenticity of the generated samples;

[0051] Features are extracted from a real defect dataset (such as DAGM 2007) and used as input conditions for the generator; the generator generates adversarial samples based on the input conditions, while the discriminator attempts to distinguish between real samples and generated samples; the training process is stabilized using a Wasserstein loss function (WGAN-GP) with a gradient penalty term coefficient λ=10;

[0052] The generator dynamically adjusts defect parameters (such as oil stain diffusion coefficient, scratch angle) to generate adversarial samples with randomness; combined with a multispectral simulation module, interference features in the visible, near-infrared, and short-wave infrared bands are added to the adversarial samples; the generated samples are mixed with the original data to form an augmented dataset, improving the robustness of the model to complex interference;

[0053] The physical simulation engine generates virtual defect samples:

[0054] Based on the Unreal Engine 5 engine, a virtual detection scene is built, a physical engine (such as PhysX) is integrated to simulate real mechanical behavior, a material library (such as metal, plastic, composite material) is defined, reflectivity, roughness, and refractive index optical parameters are set, the phase field method is used to simulate crack propagation, the initial crack length (1-10mm) and angle (0-180°) are randomly distributed, the Navier-Stokes equation is used to simulate oil diffusion, the viscosity coefficient (0.1-1000cP) and surface tension (20-72mN / m) are set, the finite element analysis (FEA) is used to simulate the deformation of the workpiece under heat and force, the Young's modulus (1-200GPa) and Poisson's ratio (0-0.5) are set, and the virtual camera module synchronously generates visible light (RGB), near-infrared (NIR), and short-wave infrared (SWIR) images, adds noise models (such as Gaussian noise and Poisson noise), and simulates the noise characteristics of real sensors;

[0055] The output includes complete data packages containing geometric information (.obj format), material information (.mtl format), and defect annotations (.json format);

[0056] Integrate data augmentation engine:

[0057] According to the training requirements, automatically allocate computing resources, the parallelism of GAN generated with physical simulation is adjustable (1:1 to 4:1), use priority queue mechanism, preferentially generate samples of rare defect types in the training set, evaluate the diversity of generated samples through FrechetInceptionDistance (FID), target FID≤50, physical simulation samples need to pass geometry verification (such as crack length error≤5%), and spectral verification (reflectivity error≤2%); periodically inject newly generated adversarial samples and virtual defect samples into the training set to keep the data distribution synchronized with the real scene, record sample generation logs (including timestamp, parameter configuration, evaluation results), support training process backtracking and optimization.

[0058] Further, the dynamic routing network is used to select the optimal feature extraction path according to the complexity of the input image, and the feature extraction efficiency index is fed back to the dynamic adaptive layer in real time, and the reasonable allocation of computing resources and the optimization of algorithm performance, including the following steps:

[0059] Build a dynamic routing network:

[0060] Three different complexity feature extraction branches are constructed: a lightweight branch (MobileNetV3) with 1.2M parameters, suitable for simple scenarios (such as uniform texture, single defect type); a standard branch (ResNet-18) with 11.7M parameters, suitable for medium complexity scenarios (such as multi-texture interweaving, mixed defects); an enhanced branch (ResNet-50+ attention module) with 25.6M parameters, targeting high complexity scenarios (such as tiny defects, dense interference); each branch output shares a feature pyramid, containing 4 scales (1 / 4, 1 / 8, 1 / 16, 1 / 32 original image size), with 256 channels per scale; the input image is first passed through a lightweight complexity evaluation network (3-layer CNN + global pooling), outputting a complexity score (0-1 range); preset thresholds: simple scene score less than 0.3, medium scene score between 0.3 and 0.7, complex scene score greater than 0.7; routing strategy: simple scene, only activate lightweight branch, close other branches; medium scene, run lightweight + standard branches in parallel, feature fusion uses weighted average (weight = 0.4:0.6); complex scene, full branch activation, feature fusion weight = 0.2:0.3:0.5;

[0061] Monitor feature extraction efficiency:

[0062] Embed performance probes within each feature extraction branch, measure computation time (ms / frame) using CUDA event API, query GPU memory usage using NVIDIA Management Library, calculate feature map histogram entropy, evaluate information richness (entropy value > 7.5 considered as valid features); statistics average time consumption (such as simple scene target <15ms / frame), memory peak (complex scene limit <4GB) by scene type; generate efficiency heat map, spatial dimension marks inefficient areas (such as feature entropy <6 redundant calculation block), time dimension tracks efficiency fluctuations;

[0063] Dynamic resource allocation:

[0064] After receiving efficiency indicators, dynamically adjust:

[0065] Simple scene: allocate 1 GPU stream processor (SM), close redundant calculation units;

[0066] Complex scene: activate 4 SMs, enable Tensor Core acceleration;

[0067] Adjust the momentum of the BN layer (set to 0.9 for simple scenarios to accelerate convergence; set to 0.1 for complex scenarios to enhance stability); dynamically modify the learning rate (use cosine annealing for complex scenarios, initial rate = 1e-3, period = 50 epochs); generate a "resource allocation performance report" every week, which includes resource utilization (GPU idle time ratio, target < 10%), algorithm acceleration ratio (the time ratio of the enhanced branch to the lightweight branch, target < 3.5); optimize the routing strategy based on the report, migrate high-frequency scenario features to the lightweight branch, and prune redundant channels.

[0068] Further, the human-machine collaborative interface is constructed, when the model confidence is insufficient, an artificial review request is automatically triggered, and the review result is fed back to the training set and the cognitive reasoning layer to optimize the causal reasoning module, including the following steps:

[0069] Model confidence monitoring:

[0070] The cognitive reasoning layer calculates a confidence score (0-1 range) for each detection result based on feature uncertainty (such as edge blur, spectral outliers) and historical error statistics (such as false positive rate of similar defects), when the confidence is < a preset threshold (such as 0.7), the result is automatically marked as "needs review", and an artificial review process is triggered; the human-machine collaborative interface pushes the request to the artificial reviewer through GUI / VUI / AR interface, displays the context information of the low-confidence sample, including multi-modal data fusion view (visible light + near-infrared + event stream overlay), historical detection record (spatiotemporal distribution heat map of similar defects), model's preliminary conclusion (such as "suspected scratch, confidence 0.65"), the reviewer can quickly locate the problem area through eye tracking or gesture interaction, and submit the correction result (such as adjusting the defect bounding box, modifying the defect type label);

[0071] Artificial review result feedback:

[0072] The review result (corrected label, bounding box coordinates) is written into the training set in real time through the message queue and marked as "artificially confirmed" data; a data version number (such as v2.1.3_AR001) is automatically generated, recording the reviewer ID, review timestamp, and modification content summary; the training set management module triggers incremental training every week, fine-tunes the last three layers of the model using the newly labeled data (learning rate = 1e-4), and uses the historical data playback mechanism to train the model in full every month, updating the BatchNorm layer statistics to prevent concept drift; extract causal relationship annotations (such as "oil stain causes infrared reflectance anomaly") from the review results, and update the causal graph model:

[0073] Add new edge: oil stain → infrared reflectance anomaly (weight 0.8);

[0074] Adjust edge weight: scratch → edge mutation (weight increased from 0.75 to 0.85);

[0075] Delete redundant associations: illumination change → color bias (verified as spurious correlation by human);

[0076] With the updated causal graph, simulate the detection results under the "no oil interference" scenario:

[0077] Input: original infrared image + counterfactual mask (remove oil area);

[0078] Output: the predicted defect type is corrected from "corrosion" to "normal wear", which is consistent with the manual review conclusion, verifying the effectiveness of the causal model;

[0079] The dynamic routing network selects the feature extraction path (such as SWIR band + morphological filtering) that is robust to oil pollution according to the updated causal relationship, and the causal reasoning module highlights the key causal evidence (such as "infrared reflectance anomaly is caused by oil pollution, not real defects") when explaining the detection results.

[0080] The present application has the following beneficial effects:

[0081] In the present application, the adaptive optical module is used to suppress metal reflection, the multispectral imaging unit is used to obtain multi-band images, the event camera is used to capture edge changes, and data preprocessing is used to ensure input data quality; The illumination and motion compensation sub-module is used to adapt to the state changes of the environment and the measured object, the data enhancement engine is used to expand the sample, and the algorithm adaptability is enhanced; The dynamic routing network is used to optimize feature extraction and resource allocation, and the causal reasoning module is used to reduce data bias; Through cloud edge end cooperation to improve the model performance, construct man-machine cooperation interface, use artificial review and feedback to optimize algorithm, and use AR auxiliary labeling to improve data quality, effectively solve the problem of insufficient algorithm robustness of machine vision image acquisition system. BRIEF DESCRIPTION OF DRAWINGS

[0082] Figure 1 A system block diagram of an artificial intelligence machine vision image acquisition system is provided. DETAILED DESCRIPTION

[0083] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0084] Please refer to Figure 1 The present application is an artificial intelligence machine vision image acquisition system, which comprises:

[0085] Multi-modal perception layer: Through the integration of multiple sensors and imaging technologies, multi-dimensional and high-precision data acquisition of the target scene is carried out; an adaptive optical module is used to dynamically adjust the incident light by using a liquid lens and a polarizing filter array to suppress metal reflection interference; at the same time, a multi-spectral imaging unit is deployed to simultaneously capture visible light, near-infrared and short-wave infrared band images, enhancing the separability of defect features; in addition, an event camera is introduced to capture the edge changes of moving targets with microsecond-level time resolution, solving the smear problem in high-speed scenes; and in the data preprocessing stage, real-time distortion correction and multi-source data alignment technology are used to adjust the accuracy and consistency of image data; these data are then transmitted to the dynamic adaptive layer for further environmental perception and parameter adjustment;

[0086] Dynamic adaptive layer: dynamically adjust parameters according to real-time environmental information, and improve the resistance of the algorithm to complex interference through data enhancement technology; through the light adaptive submodule and the motion compensation submodule, the environmental light change and the motion state of the measured object are monitored in real time, the camera gain, exposure time and micro-displacement compensation are dynamically adjusted, the stability and clarity of image acquisition are adjusted, the light intensity change trend is fed back to the multi-modal perception layer, and a motion state vector is generated to input the cognitive inference layer; at the same time, a data enhancement engine is deployed to generate adversarial samples in real time using a generative adversarial network, simulate complex interference such as oil stains and scratches, and generate virtual defect samples through a physical simulation engine to effectively solve the small sample problem; the processed multi-modal data is input to the cognitive inference layer as the input of the algorithm inference;

[0087] Cognitive inference layer: introduce the principles of cognitive science to build an algorithm architecture with explainability and anti-interference ability; adopt a dynamic routing network to automatically select the optimal feature extraction path according to the complexity of the input image, and real-time feedback the feature extraction efficiency index to the dynamic adaptive layer to reasonably allocate computing resources and optimize algorithm performance; at the same time, deploy a causal inference module to reduce data bias through counterfactual learning and generate a feature importance heat map to provide explainability output; the model prediction result and explainability output are then transmitted to the collaborative decision layer to trigger the collaborative decision-making process;

[0088] Synergistic decision-making layer: Break through the performance boundary of a single device through synergistic mechanisms to improve the accuracy and efficiency of overall decision-making; conduct cloud-edge-end synergy, deploy lightweight models on edge computing nodes for preliminary screening, aggregate multi-device data on the cloud to train global models, and protect data security through differential privacy; at the same time, build a human-machine collaborative interface that automatically triggers manual review requests when model confidence is insufficient, and feeds back the review results to the training set and the causal reasoning module of the cognitive reasoning layer for optimization; AR-assisted labeling improves the efficiency of manual labeling; these designs can fully utilize the computing resources of the cloud and edge, and perform collaborative work of multiple devices and multiple models, while combining human intervention to significantly improve the decision-making accuracy and efficiency in complex scenarios; finally, the synergistic decision-making layer issues global optimization instructions to the multi-modal perception layer, forming an end-to-end closed-loop optimization.

[0089] In one embodiment, the adaptive optics module uses a liquid lens and a polarizing filter array to dynamically adjust the incident light, suppress metal reflection interference, including the following steps:

[0090] Liquid lens dynamic focusing:

[0091] The liquid lens is composed of a container filled with optical liquid and an elastic polymer film; when the distance of the measured object is detected, the shape of the liquid is changed by applying voltage or pressure, so that the rapid zooming is achieved in milliseconds, and the image always remains clear, even when detecting reflective metal workpieces, the liquid lens can maintain the focusing of the workpiece surface;

[0092] Polarizing filter array reflection suppression:

[0093] The polarizing filter array is composed of multiple polarization units, each unit only allows polarized light of a specific vibration direction to pass; when the non-polarized light reflected by the metal surface enters the array, part of the light in the polarization direction is blocked, effectively reducing the reflection intensity.

[0094] Adaptive optics synergy:

[0095] The liquid lens and the polarizing filter array work together, the liquid lens is responsible for adjusting the focal length to make the target image clear; while the polarizing filter array synchronously suppresses the metal reflection; for example, when detecting reflective metal workpieces, the liquid lens maintains the focusing of the workpiece surface, while the polarizing filter reduces the reflection intensity, both of which work together to obtain high-quality images;

[0096] Real-time monitoring and compensation of ambient light:

[0097] The light adaptive sub-module continuously monitors the intensity and color temperature of the ambient light through the photosensitive sensor; according to these monitoring data, the optical parameters of the liquid lens such as focal length and aperture are dynamically adjusted, so that clear and stable images can be obtained under different lighting conditions;

[0098] Wavefront distortion correction:

[0099] In high-end applications, adaptive optics modules can integrate wavefront sensors. The sensor can measure the wavefront distortion of the incident light in real time and calculate the corresponding compensation amount. Then, the liquid lens is driven to perform phase correction to further improve the resolution and quality of imaging.

[0100] Closed-loop feedback optimization:

[0101] After image acquisition, the image quality indicators are analyzed, and if reflections or focusing deviations are detected, the parameters of the liquid lens and the angle of the polarizing filter are automatically adjusted to form a closed-loop optimization mechanism to continuously improve image quality.

[0102] Multi-spectral fusion enhancement:

[0103] In combination with the data of the multi-spectral imaging unit, the image after suppressing reflections is subjected to spectral fusion processing; for example, while suppressing reflections in the visible light band, the short-wave infrared band is used to enhance the visibility of metal surface defects; this multi-spectral fusion technology helps to improve the comprehensiveness and accuracy of detection.

[0104] In one embodiment, the introduction of the event camera captures the edge changes of moving targets, and in the data preprocessing stage, the accuracy and consistency of image data are adjusted through real-time distortion correction and multi-source data alignment technology, including the following steps:

[0105] Capture edge changes of moving targets:

[0106] Each pixel of the event camera independently monitors the brightness change, and when the change exceeds the preset threshold (such as ΔL≥15%), an event output is triggered, each event contains a timestamp (accuracy ≤1μs), pixel coordinates (x, y) and polarity (brightness increase or decrease flag); the event camera only outputs the event stream of brightness changes, reducing the data volume and retaining only the motion edge information; through the event clustering algorithm (such as Hough transform based on spatial and temporal proximity), the motion boundary is tracked in real time, and the target motion trajectory is generated;

[0107] Real-time distortion correction:

[0108] A polynomial model (such as the Brown-Conrady model) is used to fit the radial distortion (k1, k2, k3) and tangential distortion (p1, p2), and the distortion coefficients are obtained through the chessboard calibration method; based on the atmospheric scattering model, the uneven illumination is compensated, the transmittance is estimated using the dark channel prior algorithm, and the image contrast is restored; the inconsistency of sensor response is compensated to adjust the uniformity of image brightness;

[0109] Multi-source data alignment:

[0110] Adopting the Precision Time Protocol (PTP), the time stamp synchronization accuracy of event camera data and multispectral images is adjusted, and the spatial alignment of multi-source data is performed through a feature point matching algorithm (such as SIFT), with a registration error of less than 2 pixels; a cross-modal feature correlation matrix is constructed to fuse the edge features of event stream and the texture features of multispectral images, thereby improving the target recognition accuracy; the fusion weight of event data and multispectral data is adjusted according to the scene dynamics to optimize the target detection performance;

[0111] Adjusting the accuracy and consistency of image data:

[0112] Adopting a decision-level fusion strategy, the event camera provides motion boundary positioning (accuracy ±1 pixel), and the multispectral image supplements material classification information (accuracy ≥95%); through bidirectional projection verification (projecting the corrected data back to the original coordinate system, with an error ≤3%), the geometric consistency is ensured, and through a histogram matching algorithm (correlation coefficient ≥0.9), the consistency of the radiation characteristics of multi-source data is adjusted, and through time series analysis, the continuity of the motion target trajectory is verified, eliminating data jumps and loss.

[0113] In one embodiment, the real-time monitoring of environmental light changes and the motion state of the measured object is performed through the light adaptation sub-module and the motion compensation sub-module, and the camera gain, exposure time and micro-displacement compensation are dynamically adjusted to adjust the stability and clarity of image acquisition, including the following steps:

[0114] Light adaptation sub-module:

[0115] A high-precision photosensitive sensor array is used to real-time collect environmental light intensity (0.1 lux to 100,000 lux) and color temperature (2000K to 10000K), and through multi-region sampling (such as a 3x3 grid), spatial distribution data is obtained to identify uneven light areas; based on the PID algorithm, the camera gain is dynamically adjusted (0dB to 36dB), with a response time ≤5ms, and the signal intensity is adjusted within the optimal working interval (30% to 70% saturation); a bilinear interpolation algorithm is used to adaptively adjust the exposure time (1μs to 1s) according to the target motion speed (obtained through the motion compensation sub-module), and the higher the motion speed, the shorter the exposure time; a tunable optical filter is used to dynamically adjust the spectral response curve to match the current environmental light characteristics, thereby improving the signal-to-noise ratio (SNR≥40dB); the light intensity change trend (ΔL / Δt) is fed back to the liquid lens control module of the multi-modal perception layer to pre-adjust the optical parameters to adapt to environmental light changes;

[0116] Motion compensation sub-module:

[0117] Analyzing event stream data, extracting moving target edge change rate (pixels / s) and direction vector; combining 6-axis IMU data (accelerometer ±16g range, gyroscope ±2000° / s range), through Kalman filter fusion to obtain the measured object six degrees of freedom motion trajectory; based on LSTM network to establish motion prediction model, to predict the displacement of the measured object in advance; driving liquid lens for active optical anti-shake, combined with piezoelectric ceramic micro displacement platform for mechanical compensation; using a digital image stabilization algorithm based on feature point matching, compensating for residual jitter (≤0.1 pixels); generating motion state vector (including speed, acceleration, direction information) at a frequency of 1 kHz input to the cognitive inference layer;

[0118] Collaborative optimization mechanism

[0119] When high-speed motion (speed > 100 mm / s) is detected, the exposure time is shortened (to below 1 ms) and the gain is increased (to 24 dB) to maintain the signal-to-noise ratio; in low light environment (<100 lux), the exposure time is appropriately extended (to 100 ms), and motion blur is suppressed through motion compensation; using a no-reference image quality assessment algorithm (such as BRISQUE), the image sharpness index (0-100 points) is calculated in real time; when the sharpness is less than 70 points, the parameter readjustment process is triggered, and the MTF (modulation transfer function) value is preferentially optimized to above 0.3; dynamically adjusting the power consumption of the submodules according to the workload, the idle state power consumption is less than 1W, and the full load state power consumption is less than 5W.

[0120] In one embodiment, the light intensity change trend is fed back to the multi-modal perception layer, and a motion state vector is generated to input the cognitive inference layer, including the following steps:

[0121] The light intensity change trend is fed back to the multi-modal perception layer:

[0122] The light intensity change trend is fed back to the multi-modal perception layer:

[0123] The light intensity change trend is fed back to the multi-modal perception layer:

[0124] The motion compensation sub-module combines event camera data and IMU (inertial measurement unit) data to obtain the six-degree-of-freedom motion trajectory (x, y, z, roll, pitch, yaw) of the measured object through a Kalman filter; the event camera provides high-precision edge change information (positioning accuracy ±1 pixel), and the IMU provides high-frequency motion data (sampling rate ≥1 kHz); based on the fused motion trajectory, the motion parameters such as the speed (v), acceleration (a), and direction (θ) of the measured object are calculated, a motion state vector (v, a, θ) is generated, and the motion state vector contains three-dimensional space motion information and a timestamp (accuracy ≤1 μs); the motion state vector is input to the cognitive inference layer at a frequency of 1 kHz through a high-speed bus (such as PCIe Gen4); in the cognitive inference layer, the motion state vector is spatiotemporally aligned with the multispectral image and event stream data, and the synchronization of feature extraction is adjusted; the dynamic routing network of the cognitive inference layer adjusts the feature extraction path in real time according to the motion state vector, for example, when high-speed motion (v>100 mm / s) is detected, a feature extraction algorithm with high computational efficiency and anti-motion blur is preferentially selected; the causal inference module analyzes the target motion mode using the motion state vector, optimizes the counterfactual learning strategy, and improves the decision robustness.

[0125] In one embodiment, the deployment data augmentation engine generates adversarial samples using a generative adversarial network, simulates complex interference, and generates virtual defect samples through a physics simulation engine, including the following steps:

[0126] The generative adversarial network (GAN) generates adversarial samples:

[0127] A conditional generative adversarial network (CGAN) is used, and a defect type label (such as oil stain, scratch, and deformation) is introduced as a generation condition; a generator (Generator) is composed of a residual network (ResNet) and contains 9 residual blocks, and outputs adversarial samples with a resolution of 512×512; a discriminator (Discriminator) adopts a PatchGAN structure and outputs a 70×70 probability map to locally determine the authenticity of the generated samples;

[0128] Features are extracted from a real defect dataset (such as DAGM 2007) and used as input conditions for the generator; the generator generates adversarial samples according to the input conditions, and the discriminator attempts to distinguish between real samples and generated samples; a Wasserstein loss function (WGAN-GP) is used to stabilize the training process, and the gradient penalty coefficient λ is 10;

[0129] The generator dynamically adjusts defect parameters (such as oil diffusion coefficient, scratch angle) to generate adversarial samples with randomness; combined with a multi-spectral simulation module, the adversarial samples are added with interference features in the visible light, near-infrared, and short-wave infrared bands; the generated samples are mixed with the original data to form an enhanced data set, improving the robustness of the model to complex interference;

[0130] The physical simulation engine generates virtual defect samples:

[0131] Based on Unreal Engine 5 engine, a virtual detection scene is constructed, and a physical engine (such as PhysX) is integrated to simulate real mechanical behavior; a material library (such as metal, plastic, composite material) is defined, and reflectivity, roughness, and refractive index optical parameters are set; the phase field method is used to simulate crack propagation, and the initial crack length (1-10mm) and angle (0-180°) are randomly distributed; based on the Navier-Stokes equation, oil diffusion is simulated, and the viscosity coefficient (0.1-1000cP) and surface tension (20-72mN / m) are set; through finite element analysis (FEA), the deformation of the workpiece under heat and force is simulated, and the Young's modulus (1-200GPa) and Poisson's ratio (0-0.5) are set; the virtual camera module synchronously generates visible light (RGB), near-infrared (NIR), and short-wave infrared (SWIR) images, adds noise models (such as Gaussian noise, Poisson noise), and simulates the noise characteristics of real sensors;

[0132] The output includes complete data packages containing geometric information (.obj format), material information (.mtl format), and defect annotations (.json format);

[0133] Integration of data enhancement engine:

[0134] According to the training requirements, the computing resources are automatically allocated, the parallelism of GAN generation and physical simulation is adjustable (1:1 to 4:1), the priority queue mechanism is adopted, and the samples of rare defect types in the training set are generated preferentially; the FrechetInceptionDistance (FID) is used to evaluate the diversity of generated samples, the target FID is less than or equal to 50, and the physical simulation samples need to pass the geometric verification (such as crack length error less than or equal to 5%) and spectral verification (reflectivity error less than or equal to 2%); the newly generated adversarial samples and virtual defect samples are periodically injected into the training set to keep the data distribution synchronized with the real scene, and the sample generation log (including timestamp, parameter configuration, and evaluation results) is recorded to support the backtracking and optimization of the training process.

[0135] In one embodiment, the dynamic routing network is used to select the optimal feature extraction path according to the complexity of the input image, and the feature extraction efficiency index is fed back to the dynamic adaptive layer in real time, and the reasonable allocation of computing resources and the optimization of algorithm performance include the following steps:

[0136] Construct a dynamic routing network:

[0137] Construct a feature extraction branch containing three different complexities: a lightweight branch (MobileNetV3) with 1.2M parameters, suitable for simple scenarios (such as uniform texture and single defect type); a standard branch (ResNet-18) with 11.7M parameters, suitable for medium complexity scenarios (such as multi-texture interweaving and mixed defects); an enhanced branch (ResNet-50+ attention module) with 25.6M parameters, targeting high complexity scenarios (such as tiny defects and dense interference); the output shared feature pyramid of each branch contains four scales (1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original image size) with 256 channels each; the input image first passes through a lightweight complexity evaluation network (3-layer CNN + global pooling), outputting a complexity score (0-1 range); preset thresholds: simple scenario score less than 0.3, medium scenario score between 0.3 and 0.7, complex scenario score greater than 0.7; routing strategy: for simple scenarios, only activate the lightweight branch and close the other branches; for medium scenarios, run the lightweight and standard branches in parallel, and use weighted averaging (weight = 0.4:0.6) for feature fusion; for complex scenarios, activate all branches and use feature fusion weights = 0.2:0.3:0.5;

[0138] Monitor feature extraction efficiency:

[0139] Embed performance probes within each feature extraction branch, measure computation time (ms / frame) using CUDA event API, query GPU memory usage using NVIDIA Management Library, calculate feature map histogram entropy, and evaluate information richness (entropy value > 7.5 considered as valid features); statistically analyze average time consumption (e.g., simple scenario target < 15ms / frame) and memory peak (complex scenario limit < 4GB) by scene type; generate efficiency heat maps, spatial dimension marked inefficient areas (e.g., redundant calculation blocks with feature entropy < 6), and track efficiency fluctuations in time dimension;

[0140] Dynamic resource allocation:

[0141] After receiving efficiency indicators, dynamically adjust:

[0142] Simple scenario: allocate 1 GPU stream processor (SM) and close redundant calculation units;

[0143] Complex scenario: activate 4 SMs and enable Tensor Core acceleration;

[0144] Adjust the momentum of the BN layer (set to 0.9 for simple scenes to accelerate convergence; set to 0.1 for complex scenes to enhance stability); dynamically modify the learning rate (use cosine annealing for complex scenes, initial rate = 1e-3, period = 50 epochs); generate a "resource allocation performance report" every week, which includes resource utilization (GPU idle time ratio, target < 10%), algorithm acceleration ratio (enhanced branch relative to lightweight branch time-consuming ratio, target < 3.5); optimize routing strategies based on the report, migrate high-frequency scene features to lightweight branches, and prune redundant channels.

[0145] In one embodiment, the human-machine collaborative interface is constructed, and when the model confidence is insufficient, an automatic manual review request is triggered, and the review result is fed back to the training set and the cognitive reasoning layer optimization causal reasoning module, including the following steps:

[0146] Model confidence monitoring:

[0147] The cognitive reasoning layer calculates a confidence score (0-1 range) for each detection result based on feature uncertainty (such as edge blur, spectral outliers) and historical error statistics (such as false positive rate of similar defects), and when the confidence is < a preset threshold (such as 0.7), the result is automatically marked as "needs review" and a manual review process is triggered; the human-machine collaborative interface pushes the request to the manual reviewer through the GUI / VUI / AR interface, displays the context information of the low-confidence sample, including the multi-modal data fusion view (visible light + near-infrared + event stream overlay), historical detection record (spatiotemporal distribution heat map of similar defects), and model's preliminary conclusion (such as "suspected scratch, confidence 0.65"), the reviewer can quickly locate the problem area through eye tracking or gesture interaction, and submit the revised result (such as adjusting the defect bounding box, modifying the defect type label);

[0148] Manual review result feedback:

[0149] The review result (modified label, bounding box coordinates) is written into the training set in real time through the message queue and marked as "human-confirmed" data; a data version number (such as v2.1.3_AR001) is automatically generated, recording the reviewer ID, review timestamp, and modification content summary; the training set management module triggers incremental training every week, preferentially using new labeled data to fine-tune the last three layers of the model (learning rate = 1e-4), and the historical data playback mechanism trains the model in full every month, updating the BatchNorm layer statistics to prevent concept drift; extract causal relationship annotations (such as "oil stain causes infrared reflectance anomaly") from the review result and update the causal graph model:

[0150] Add new edge: oil stain → infrared reflectance anomaly (weight 0.8);

[0151] Adjust edge weight: scratch -> edge abruptness (weight increased from 0.75 to 0.85);

[0152] Remove redundant association: illumination change -> color bias (verified as spurious association by human);

[0153] With the updated causal graph, simulate the detection result under the "no oil interference" scenario:

[0154] Input: original infrared image + counterfactual mask (remove oil area);

[0155] Output: the predicted defect type is corrected from "corrosion" to "normal wear", consistent with the manual review conclusion, verifying the effectiveness of the causal model;

[0156] The dynamic routing network prioritizes feature extraction paths that are robust to oil (e.g. SWIR band + morphological filtering) based on the updated causal relationships. The causal reasoning module highlights key causal evidence (e.g. "infrared reflectance anomaly is caused by oil, not real defects") when explaining the detection result.

[0157] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. An artificial intelligence machine vision image acquisition system, comprising: Comprise: Multi-modal perception layer: adopt adaptive optical module, use liquid lens and polarizing filter array to dynamically adjust incident light, suppress metal reflection interference; At the same time, deploy multispectral imaging unit to capture visible light, near infrared and short wave infrared band image; Introduce event camera to capture the edge change of moving target, and in the data preprocessing stage, through real-time distortion correction and multi-source data alignment technology, adjust the accuracy and consistency of image data; The data collected by the multi-modal perception layer is transmitted to the dynamic adaptive layer; Dynamic adaptive layer: through light adaptive submodule and motion compensation submodule, real-time monitoring of ambient light changes and the motion state of the measured object, dynamically adjusting the camera gain, exposure time and micro displacement compensation, adjusting the stability and clarity of image acquisition, feeding back the light intensity change trend to the multi-modal perception layer, and generating a motion state vector input to the cognitive reasoning layer; Deploy data enhancement engine to generate adversarial network to generate adversarial samples, simulate complex interference, and generate virtual defect samples through physical simulation engine; Cognitive reasoning layer: adopt dynamic routing network, select the optimal feature extraction path according to the complexity of the input image, and feed back the feature extraction efficiency index to the dynamic adaptive layer to optimize the performance of the algorithm and the rational allocation of computing resources; At the same time, deploy causal reasoning module to reduce data bias through counterfactual learning, generate feature importance heat map to provide explainability output; Collaborative decision-making layer: cloud edge end collaboration, edge computing node deploys lightweight model for screening, cloud end aggregates device data to train global model, protects data security through differential privacy; Build human-computer collaborative interface, automatically trigger manual review request when model confidence is insufficient, feed back review results to training set and cognitive reasoning layer to optimize causal reasoning module; AR assisted labeling improves the efficiency of manual labeling.

2. The artificial intelligence machine vision image acquisition system according to claim 1, wherein, The adaptive optical module uses liquid lens and polarizing filter array to dynamically adjust incident light and suppress metal reflection interference, comprising the following steps: Liquid lens dynamic focusing: Liquid lens is composed of a container filled with optical liquid and an elastic polymer film; When the distance of the measured object changes, the shape of the liquid is changed by applying voltage or pressure, so that fast zooming is realized, and the image always remains clear, even when detecting reflective metal workpieces, the liquid lens can maintain the focus of the workpiece surface; Adaptive optical collaborative work: The polarizing filter array is composed of multiple polarization units, each unit only allows polarized light of a specific vibration direction to pass through; When the non-polarized light reflected by the metal surface enters the array, part of the polarized light in the direction is blocked, thereby weakening the reflection intensity; Liquid lens and polarizing filter array work together, liquid lens is responsible for adjusting the focal length to adjust the target imaging clarity, and polarizing filter array synchronously suppresses metal reflection; Real-time monitoring and compensation of ambient light: The light adaptive submodule continuously monitors the intensity and color temperature of the ambient light through photosensitive sensors; According to these monitoring data, the optical parameters of the liquid lens are dynamically adjusted; Wavefront distortion correction: Adaptive optics module integrates wavefront sensor, which measures wavefront distortion of incident light in real time and calculates corresponding compensation amount; then, it drives liquid lens to perform phase correction, further improving resolution and quality of imaging; Closed-loop feedback optimization: After image acquisition, image quality indicators are analyzed, and if reflections or focusing deviations are detected, the parameters of the liquid lens and the angle of the polarizing filter are automatically adjusted, forming a closed-loop optimization mechanism to continuously improve image quality; Multi-spectral fusion enhancement: Combining the data of the multi-spectral imaging unit, the image after suppressing reflections is subjected to spectral fusion processing; while suppressing reflections in the visible light band, the short-wave infrared band is used to enhance the visibility of metal surface defects.

3. The artificial intelligence machine vision image acquisition system according to claim 1, wherein, The introduction of event cameras captures the edge changes of moving targets and, in the data preprocessing stage, adjusts the accuracy and consistency of image data through real-time distortion correction and multi-source data alignment techniques, including the following steps: Capturing edge changes of moving targets: Each pixel of the event camera independently monitors brightness changes, and when the change exceeds a preset threshold, an event output is triggered, which includes a timestamp, pixel coordinates, and polarity; the event camera only outputs event streams of brightness changes, reducing data volume and retaining only motion edge information; real-time tracking of motion boundaries is achieved through event clustering algorithms, generating target motion trajectories; Real-time distortion correction: A polynomial model is used to fit radial distortion and tangential distortion, and distortion coefficients are obtained through a chessboard calibration method; based on an atmospheric scattering model, uneven illumination is compensated, dark channel prior algorithms are used to estimate transmittance, and image contrast is restored; inconsistencies in sensor response are compensated to adjust image brightness uniformity; Multi-source data alignment: The precise time protocol is used to adjust the timestamp synchronization accuracy of event camera data and multi-spectral images, and feature point matching algorithms are used for spatial alignment of multi-source data, with a registration error of less than 2 pixels; a cross-modal feature correlation matrix is constructed to fuse event stream edge features and multi-spectral image texture features, improving target recognition accuracy; the fusion weights of event data and multi-spectral data are adjusted according to the scene dynamics to optimize target detection performance; Adjusting the accuracy and consistency of image data: A decision-level fusion strategy is adopted, with event cameras providing motion boundary positioning and multi-spectral images providing material classification information; geometric consistency is verified through bidirectional projection, histogram matching algorithms are used to adjust the consistency of multi-source data radiation characteristics, and the continuity of motion target trajectories is verified through time series analysis.

4. The artificial intelligence machine vision image acquisition system of claim 1, wherein, The light adaptive sub-module and motion compensation sub-module monitor environmental light changes and the motion state of the measured object in real time, dynamically adjust the camera gain, exposure time, and perform micro-displacement compensation, and adjust the stability and clarity of image acquisition, including the following steps: Light adaptive sub-module: Adopt high-precision photosensitive sensor array, real-time collection of ambient light intensity and color temperature, through multi-region sampling to obtain spatial distribution data, identify uneven illumination area; Based on PID algorithm dynamic adjustment of camera gain, response time is not greater than 5ms, the adjustment signal intensity in the best working interval; Using bilinear interpolation algorithm, according to the target motion speed adaptive adjustment of exposure time, the higher the motion speed, the shorter the exposure time; Through the tunable optical filter dynamic adjustment of spectral response curve, match the current environment light characteristics, improve the signal-to-noise ratio; The light intensity trend feedback to the multi-modal perception layer of liquid lens control module, pre-adjustment of optical parameters to adapt to the change of environment light; Motion compensation sub-module: Analysis of event stream data, extraction of moving target edge change rate and direction vector; Combined with 6-axis IMU data, through Kalman filter fusion to obtain the six-degree-of-freedom motion trajectory of the measured object; Based on LSTM network to establish motion prediction model, to predict the displacement of the measured object; Drive liquid lens for active optical anti-shake, combined with piezoelectric ceramic micro displacement platform for mechanical compensation; Using digital image stabilization algorithm based on feature point matching, compensation of residual jitter; Generate motion state vector, input cognitive reasoning layer; Collaborative optimization mechanism When high-speed motion is detected, the exposure time is shortened and the gain is increased to maintain the signal-to-noise ratio; In low light environment, the exposure time is appropriately extended, and the motion blur is suppressed through motion compensation; Using no-reference image quality assessment algorithm, real-time calculation of image clarity index; When the clarity is less than 70 points, trigger parameter readjustment process, preferentially optimize MTF value to more than 0.3; According to the working load, dynamically adjust the power consumption of the sub-module, the idle state power consumption is less than 1W, and the full load state power consumption is less than 5W.

5. The artificial intelligence machine vision image acquisition system according to claim 1, wherein, The light intensity trend feedback to the multi-modal perception layer, and generate motion state vector input cognitive reasoning layer, including the following steps: Light intensity trend feedback to multi-modal perception layer: The light adaptive sub-module collects ambient light intensity and color temperature data in real time through high-precision photosensitive sensor array; Adopt multi-region sampling to obtain spatial distribution data, identify uneven illumination area; Through time series analysis algorithm to calculate the light intensity trend, predict the light change in the next 50ms; Combined with historical data to establish light change prediction model, improve the accuracy of trend prediction; The light intensity trend is fed back to the liquid lens control module of the multi-modal perception layer in the form of control signal; The liquid lens pre-adjusts the optical parameters according to the feedback signal, adapts to the change of environment light in advance, reduces the impact of light mutation on imaging quality; Generate motion state vector input cognitive reasoning layer: The motion compensation sub-module combines event camera data and IMU data to obtain a six-degree-of-freedom motion trajectory of the measured object through a Kalman filter; the event camera provides high-precision edge change information, and the IMU provides high-frequency motion data; based on the fused motion trajectory, the speed, acceleration, and directional motion parameters of the measured object are calculated to generate a motion state vector containing three-dimensional space motion information and a timestamp; the motion state vector is input into the cognitive inference layer through a high-speed bus; in the cognitive inference layer, the motion state vector is spatio-temporally aligned with the multispectral image and event stream data to adjust the synchronization of feature extraction; the dynamic routing network of the cognitive inference layer adjusts the feature extraction path in real time according to the motion state vector, and when high-speed motion is detected, a feature extraction algorithm with high computational efficiency and anti-motion blur is preferentially selected; the causal inference module analyzes the target motion mode using the motion state vector, optimizes the counterfactual learning strategy, and improves the decision robustness.

6. The artificial intelligence machine vision image acquisition system of claim 1, wherein, The deployment data enhancement engine generates adversarial network adversarial samples, simulates complex interference, and generates virtual defect samples through a physical simulation engine, including the following steps: The adversarial network generates adversarial samples: A conditional generative adversarial network is used, and a defect type label is introduced as a generation condition; the generator is composed of a residual network and contains 9 residual blocks, and outputs adversarial samples with a resolution of 512x512; the discriminator uses a PatchGAN structure and outputs a 70x70 probability map to locally determine the authenticity of the generated samples; Features are extracted from the real defect data set as input conditions for the generator; the generator generates adversarial samples based on the input conditions, while the discriminator attempts to distinguish between real samples and generated samples; the training process is stabilized by using a Wasserstein loss function, and the gradient penalty coefficient λ is set to 10; The generator dynamically adjusts the defect parameters to generate adversarial samples with randomness; combined with the multispectral simulation module, interference features in the visible light, near-infrared, and short-wave infrared bands are added to the adversarial samples; the generated samples are mixed with the original data to form an enhanced data set, improving the robustness of the model to complex interference; The physical simulation engine generates virtual defect samples: A virtual detection scene is built based on the Unreal Engine 5 engine, and a physical engine is integrated to simulate real mechanical behavior; a material library is defined, and reflectivity, roughness, and refractive index optical parameters are set; the phase field method is used to simulate crack propagation, and the initial crack length and angle are randomly distributed; the Navier-Stokes equation is used to simulate oil diffusion, and the viscosity coefficient and surface tension are set; finite element analysis is used to simulate workpiece deformation under heat and force, and Young's modulus and Poisson's ratio are set; a virtual camera module synchronously generates visible light, near-infrared, and short-wave infrared images, adds a noise model, and simulates the noise characteristics of real sensors; a complete data package containing geometric information, material information, and defect labels is output; Integrate the data enhancement engine: According to the training requirements, the computing resources are automatically allocated, the parallelism of the GAN generated with the physical simulation is adjustable, the priority queue mechanism is adopted, and the defect type samples in the training set are preferentially generated; the sample diversity is evaluated by Frechet Inception Distance, the target FID is less than or equal to 50, and the physical simulation samples need to pass the geometric verification and the spectrum verification; the newly generated adversarial samples and virtual defect samples are periodically injected into the training set, the data distribution is kept synchronized with the real scene, the sample generation log is recorded, and the training process is backtracked and optimized.

7. The artificial intelligence machine vision image acquisition system of claim 1, wherein, The dynamic routing network is constructed, the optimal feature extraction path is selected according to the complexity of the input image, the feature extraction efficiency index is fed back to the dynamic adaptive layer in real time, the computing resources are reasonably allocated, and the algorithm performance is optimized, including the following steps: Constructing a dynamic routing network: Three feature extraction branches with different complexities are constructed: a lightweight branch with a parameter amount of 1.2M, suitable for simple scenes; a standard branch with a parameter amount of 11.7M, suitable for medium complexity scenes; and an enhanced branch with a parameter amount of 25.6M, suitable for high complexity scenes; the outputs of the branches share a feature pyramid, which includes four scales and 256 channels per scale; the input image is first passed through a lightweight complexity evaluation network to output a complexity score; the preset threshold is: simple scene score less than 0.3, medium scene score between 0.3 and 0.7, and complex scene score greater than 0.7; the routing strategy is: for simple scenes, only the lightweight branch is activated, and the other branches are closed; for medium scenes, the lightweight and standard branches are run in parallel, and feature fusion uses weighted averaging; for complex scenes, all branches are activated, and the feature fusion weights are 0.2:0.3:0.5; Monitoring feature extraction efficiency: Performance probes are embedded in each feature extraction branch, computation time is measured using CUDA event API, memory usage is queried using NVIDIA Management Library, and feature map histogram entropy is calculated to evaluate information richness; average time consumption and memory peak are counted by scene type; efficiency heat map is generated, space dimension is marked for low efficiency area, and time dimension is tracked for efficiency fluctuation; Dynamic resource allocation: After receiving the efficiency index, the following adjustments are made: Simple scene: allocate 1 GPU stream processor, close redundant computing units; Complex scene: activate 4 SMs, enable Tensor Core acceleration; Adjust the momentum of the BN layer and dynamically modify the learning rate; generate a "resource allocation performance report" containing resource utilization and algorithm speedup ratio every week, optimize the routing strategy according to the report, migrate high-frequency scene features to the lightweight branch, and prune redundant channels.

8. The artificial intelligence machine vision image acquisition system of claim 1, wherein, The human-computer cooperation interface is constructed, the manual review request is automatically triggered when the model confidence is insufficient, the review result is fed back to the training set and the causal reasoning module of the cognitive reasoning layer, including the following steps: Model confidence monitoring: The cognitive inference layer calculates a confidence score for each detection result, based on feature uncertainty and historical error statistics. When the confidence is less than a preset threshold, the result is automatically labeled as "review required" and a human review process is triggered. The human-machine collaboration interface pushes the request to the human reviewer through a GUI / VUI / AR interface, displaying the context information of the low-confidence sample, including a multi-modal data fusion view, historical detection records, and the model's preliminary conclusion. The reviewer quickly locates the problem area through eye tracking or gesture interaction and submits the revised results. Human review result feedback: The review results are written into the training set in real time through a message queue and labeled as "human confirmed" data. A data version number is automatically generated, and the reviewer ID, review timestamp, and modification content summary are recorded. The training set management module triggers incremental training every week, fine-tuning the last three layers of the model using new labeled data. The historical data playback mechanism trains the model in full every month, updating the BatchNorm layer statistics to prevent concept drift. Causal relationship annotations are extracted from the review results to update the causal graph model. Add new edges: oil stain → infrared reflectance anomaly; Adjust edge weights: scratch → edge mutation; Delete redundant associations: lighting changes → color deviation; Using the updated causal graph, simulate the detection results under the "no oil stain interference" scenario: Input: original infrared image + counterfactual mask; Output: predicted defect type corrected from "corrosion" to "normal wear and tear", consistent with the human review conclusion, verifying the effectiveness of the causal model; The dynamic routing network selects feature extraction paths that are robust to oil stains based on the updated causal relationships. The causal reasoning module highlights key causal evidence when explaining the detection results.

Citation Information

Cited By

  • Defect visual classification detection method and system oriented to numerical control system

    CN121095686A

  • Concrete structure internal defect nondestructive testing method fusing big data feature extraction and deep learning

    CN121167572A

  • Intelligent imaging device based on multi-spectral linear array and adaptive light source

    CN121298610A

  • Intelligent imaging device based on multispectral linear array and adaptive light source

    CN121298610B

  • Workshop production quality online detection system based on machine vision and AI

    CN121481344A