Multi-modal visual imaging system and method based on intelligent sensor
By implementing structured acquisition, modeling standards, and multi-level collaborative mechanisms for multimodal visual imaging systems, a modular design and automatic focusing system was achieved. This solved the problems of low equipment integration and poor adaptability in existing technologies, and enabled efficient and accurate industrial inspection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU RUIXIN TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing multimodal vision imaging systems suffer from problems such as complex equipment installation and calibration, insufficient modular design, poor flexibility in functional configuration, poor lens compatibility, unstable imaging quality, poor multimodal coordination, and uncoordinated light source switching, making it difficult to meet the needs of high-precision and high-efficiency industrial inspection.
By structurally acquiring and standardizing the imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic, basic system data is generated. Multimodal imaging modeling specifications are designed, clarifying the imaging correlation thresholds and light source switching rules for 2D-3D-multispectral imaging. A liquid zoom lens is used to achieve automatic focusing without mechanical movement. Combined with centralized control of the control unit, a multi-level collaborative mechanism for multi-light source timing switching and modular detachable adaptation is set up to integrate and execute the multimodal imaging process.
It features a modular design that supports rapid assembly and disassembly, adapts to different testing scenarios, has strong depth-of-field adaptability and imaging stability, and offers efficient and accurate multimodal collaboration. The 3D reconstruction accuracy reaches ≤±2.5μm, ensuring accurate defect identification and improving testing efficiency and result reliability.
Smart Images

Figure CN121994801A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a multimodal visual imaging system and method based on intelligent sensors. Background Technology
[0002] In the fields of machine vision inspection and 3D measurement technology, multimodal vision imaging systems, which integrate 2D imaging, multispectral imaging, and 3D scanning functions, have been widely used in industrial inspection and other scenarios. However, existing technologies still have many shortcomings and cannot meet the demands for high-precision and high-efficiency inspection. Specific problems are as follows:
[0003] Some existing technologies employ a split design, with each functional module scattered throughout, leading to complex equipment installation and calibration processes. This requires operators to spend a significant amount of time on debugging, significantly reducing inspection efficiency. Furthermore, the lack of modular design results in poor flexibility in functional configuration, making it difficult for users to flexibly add or remove modules according to actual inspection needs, increasing initial investment costs and operational complexity. Traditional systems often use fixed-focal-length lenses, which cannot adapt to the inspection requirements of workpieces of varying heights. When facing different depth-of-field scenarios, manual lens changes are necessary, which is cumbersome and can easily disrupt inspection continuity. Even systems with focusing capabilities suffer from slow mechanical focusing speeds and low accuracy, making it difficult to quickly adapt to dynamic inspection scenarios.
[0004] Existing systems often use a single light source, making it difficult to adapt to the imaging needs of object surfaces with different colors and textures. In the detection of complex textured surfaces, problems such as reflection interference, loss of texture details, or blurred defect identification are prone to occur, resulting in unstable imaging quality and affecting the accuracy of detection results. Furthermore, the synergy between multispectral imaging and 2D and 3D imaging is poor, failing to fully leverage the complementary advantages of multimodal data.
[0005] Existing technologies lack clear multimodal collaborative constraint rules and modeling specifications. The correlation logic and threshold standards between 2D-3D-multispectral imaging are unclear. The timing coordination of light source switching and mode switching is poor, which can easily lead to illumination superposition interference or switching delay, resulting in an unsmooth imaging process. At the same time, there is a lack of dynamic optimization mechanisms for different scenarios, making it difficult to adapt to the diverse needs of complex industrial testing environments.
[0006] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0008] According to one aspect of this application, a multimodal visual imaging method based on intelligent sensors is provided, comprising: performing structured acquisition and standardized processing of imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic based on multimodal imaging requirements and the functional objectives of a composite system, generating system basic data including functional module types, optical path correlation strength, and imaging timing parameters; designing modeling rules for the system data based on multimodal collaborative imaging constraint characteristics, clarifying the imaging correlation thresholds of 2D-3D-multispectral imaging and the light source switching response coefficient, and generating multimodal imaging modeling specifications; and combining centralized management and control with the control unit and... To meet the multi-module collaboration requirements, a multi-level collaboration mechanism is established, including adaptive adjustment of the liquid zoom lens, multi-source timing switching, precise scanning of regions of interest, and modular detachable adaptation. The system integrates and executes basic data, multi-modal imaging modeling specifications, and multi-level collaboration mechanisms. The imaging effect is iteratively optimized through target image processing algorithms. The system follows a complete workflow of power-on, calibration, parameter setting, 2D appearance inspection, 3D scanning of suspected defect areas, data processing and analysis, generation of inspection reports, and power-off. This results in multi-modal imaging results that include module collaboration features, multispectral semantic information, dynamic focusing timing attributes, and 3D morphology reconstruction accuracy.
[0009] Another aspect of this application discloses a multimodal visual imaging system based on intelligent sensors, comprising: a system basic data acquisition and standardization module, used to acquire core data including imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic, and generate system basic data including functional module types, optical path correlation strength, and imaging timing parameters through structured acquisition and standardization processing; a multimodal imaging modeling specification design module, used to design system data modeling rules based on multimodal collaborative imaging constraint characteristics, clarify the imaging correlation thresholds of 2D-3D-multispectral imaging, and the light source switching response coefficients, and generate multimodal imaging modeling specifications; and a multi-level collaborative mechanism setting module, used to combine control... To meet the centralized control and multi-module collaboration requirements of the control unit, a multi-level collaboration mechanism is set up, including adaptive adjustment of the liquid zoom lens, multi-source timing switching, precise scanning of regions of interest, and modular detachable adaptation. The multi-modal imaging integration execution and result generation module is used to integrate the system's basic data, multi-modal imaging modeling specifications, and multi-level collaboration mechanism. It iteratively optimizes the imaging effect through target image processing algorithms, following a complete workflow of system power-on - calibration - parameter setting - 2D appearance inspection - 3D scanning of suspected defect areas - data processing and analysis - generation of inspection report - system power-off. It generates multi-modal imaging results that include module collaboration features, multispectral semantic information, dynamic focusing timing attributes, and 3D morphology reconstruction accuracy.
[0010] According to another aspect of this application, an electronic device includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-described multimodal visual imaging method based on a smart sensor by executing the executable instructions.
[0011] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described multimodal visual imaging method based on a smart sensor.
[0012] This application provides a multimodal visual imaging system and method based on intelligent sensors. Focusing on multimodal imaging requirements, it first collects and standardizes core data such as imaging modules and light sources to generate basic system data. Then, it designs modeling specifications, clarifying the 2D-3D-multispectral correlation thresholds and light source switching coefficients. It constructs multi-level collaborative mechanisms such as liquid zoom adaptive adjustment. Finally, it integrates data, specifications, and mechanisms, optimizing them through specialized algorithms such as Zhang's calibration algorithm and multispectral defect recognition algorithm. Following a complete "power-on-calibration-detection-analysis-reporting" process, it achieves multimodal imaging that combines module collaboration, multispectral information, dynamic focusing, and high-precision 3D reconstruction, solving problems such as low integration and poor adaptability in traditional systems.
[0013] This application boasts high integration and flexible adaptability. Its modular design supports rapid assembly and disassembly, allowing for on-demand configuration of functions, reducing investment costs and operational complexity, and adapting to various inspection scenarios. It exhibits strong depth-of-field adaptability and imaging stability. The liquid zoom lens enables autofocus without mechanical movement, and the multispectral independent lamp group adapts to complex textured surfaces, reducing reflections and detail loss. Multimodal collaboration ensures high efficiency and accuracy, with clearly defined collaborative constraint rules and timing logic, interference-free switching, and 3D reconstruction accuracy of ≤±2.5μm. Defect identification is accurate, improving inspection efficiency and result reliability.
[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0015] Figure 1 This document illustrates a flowchart of a multimodal visual imaging method based on a smart sensor, provided in an embodiment of this application.
[0016] Figure 2 A schematic diagram of the structure of a multimodal visual imaging system based on a smart sensor provided in an embodiment of this application is shown. Detailed Implementation
[0017] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0018] The following is combined with Figure 1 This application describes a multimodal visual imaging method based on smart sensors according to exemplary embodiments thereof. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application are applicable to any suitable scenario.
[0019] In one implementation, Figure 1 A schematic flowchart of a multimodal visual imaging method based on a smart sensor according to an embodiment of this application is shown, including:
[0020] S101, based on the requirements of multimodal imaging and the functional goals of composite systems, performs structured acquisition and standardized processing of imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic to generate basic system data including functional module type, optical path correlation intensity, and imaging timing parameters.
[0021] In one implementation, based on multimodal imaging requirements, the goal of composite system integration, and industrial inspection accuracy requirements, the core parameters of the imaging module, the characteristic parameters of the light source, the key parameters of structured light projection, and the timing control logic information are integrated to clarify the acquisition range and correlation requirements of each module's parameters. Specifically, based on multimodal imaging requirements, the goal of composite system integration, and industrial inspection accuracy requirements, the core information of each module is comprehensively integrated, and the acquisition boundaries and parameter correlation logic are clarified. Among the core parameters of the imaging module, the camera needs to acquire resolution and pixel size, while the liquid zoom lens needs to acquire working distance, high resolution, large depth of field, and low distortion characteristics. The correlation requirement between the two is that they are mechanically connected via a C-interface and the optical path is precisely aligned.
[0022] For light source characteristic parameters, the brightness adjustment range and fiber optic connection compatibility of the coaxial illumination source need to be collected. For the multispectral ring light source, the number of LED groups and the center wavelength of each group need to be collected. The correlation requirement is that the coaxial illumination source is embedded in the optical path of the imaging module, and the multispectral ring light source is coaxially surrounded on the outer periphery of the object side of the liquid zoom lens. For key parameters of structured light projection, the angle between the projection optical axis and the imaging module optical axis, the DLP chip resolution, and the fringe pattern accuracy need to be collected. The correlation requirement with the imaging module is that it can be detachably connected via a side mounting bracket, and the optical axis angle remains fixed. For timing control logic, the opening and closing sequence of multiple light sources and the triggering time of camera synchronous acquisition need to be collected. The correlation requirement is that the control unit uniformly manages and controls the system through an Ethernet interface to ensure that the optical path switching and image acquisition timing are coordinated.
[0023] Based on the need for standardized data processing, core parameters and auxiliary parameters are defined, and clear acquisition standards are established. Core parameters directly determine imaging accuracy and functional implementation. The camera resolution acquisition standard is a measured value of no less than 20 million pixels. For example, the total number of pixels in the camera output image is verified to meet the standard through professional testing equipment. The working distance acquisition standard for liquid zoom lenses is selected within the range of 100mm-500mm according to the actual application scenario. For example, 100mm-200mm is selected when inspecting small workpieces, and 300mm-500mm is selected when inspecting large workpieces. The light source wavelength range acquisition standard is that multispectral ring light sources must cover the visible light (400nm-760nm) and near-infrared bands (760nm-1400nm), and the center wavelength of each LED light group must be accurately measured and recorded, such as specific wavelengths like 475nm and 560nm. The auxiliary parameters ensure the system's installation, adaptation, and stable operation. The standard for acquiring the module installation angle is that the angle between the optical axes of the structured light projection module and the imaging module is strictly controlled at 45°, calibrated by the angle scale of the side mounting bracket. The standard for acquiring the connection interface type is that the camera and liquid zoom lens use a C interface, and the control unit and each module use an Ethernet interface. It is necessary to confirm that the interface specifications match and the connection is stable.
[0024] To address the stability requirements of imaging complex textured surfaces, parameter optimization rules were established for the independent control of the multispectral ring light source LED array and the precise adjustment of the liquid zoom lens voltage signal, ensuring data validity. Complex textured surfaces, due to their uneven color distribution and dense textures, are prone to blurring and loss of detail under single illumination and a fixed focal length. Therefore, targeted parameter optimization is necessary to adapt to different scenarios. For the multispectral ring light source, independent control rules for the LED array were established. This light source comprises multiple LED arrays emitting different center wavelengths, each of which can be individually turned on, off, and have its brightness adjusted to adapt to the imaging needs of objects with different colors and textures. When inspecting workpieces with complex dark textures, the dark surface absorbs visible light strongly, making it difficult to reveal the details of the texture. In this case, a separate near-infrared LED light group is turned on, utilizing the penetrability of near-infrared light to enhance the texture contrast and clearly reveal the subtle textures hidden on the dark surface. When inspecting workpieces with light textures, the light surface easily reflects visible light, causing glare and interfering with texture recognition. In this case, a specific light group adapted to the visible light band is turned on, while the brightness of the light group is reduced to avoid strong light reflection and ensure accurate capture of the edges and details of the light texture. For workpieces with complex textures of mixed colors, LED light groups of different wavelengths can be combined and turned on. Through multi-spectral illumination complementarity, the texture characteristics of each color area are fully restored, ensuring the integrity and clarity of the imaging data.
[0025] For liquid crystal zoom lenses, precise voltage signal adjustment rules are established. The focal length of a liquid crystal zoom lens is determined by the refractive index of its internal liquid, which can be adjusted by a voltage signal sent by the control unit to achieve rapid focusing without mechanical movement. When the detection distance is close, the control unit sends a higher voltage signal to the lens, increasing the refractive index of the liquid inside the lens and shortening the focal length accordingly. This precisely adapts to imaging close-range workpieces, allowing for clear imaging of fine textures and defects on the workpiece surface. When the detection distance is far, the control unit sends a lower voltage signal, decreasing the refractive index of the liquid and lengthening the focal length accordingly, ensuring complete capture of the overall texture and contour of distant workpieces. Through precise mapping between the voltage signal and the focal length, the lens can quickly adapt to different depth-of-field scenarios without manual replacement or adjustment, avoiding image blurring caused by unsuitable focal length. This ensures high-quality imaging data output at different detection distances, meeting the stable imaging requirements of complex textured surfaces.
[0026] The above parameter acquisition requirements, standardization processing needs, and optimization rules are integrated and processed to generate a complete and comprehensive system foundation data. Parameter types are clearly divided into four categories: imaging module parameters, light source characteristic parameters, structured light projection parameters, and timing control parameters. Acquisition specifications record the specific values or ranges of each parameter in detail, such as a camera resolution of 20 megapixels, a liquid zoom lens working distance of 100mm-500mm, a multispectral ring light source wavelength covering visible light + near-infrared, and a structured light projection angle of 45°. The correlation logic clarifies the adaptation relationships between each parameter, such as the optical path alignment requirements between the C-interface connection and the liquid zoom lens, and the 45° angle corresponding to the collaborative scanning relationship between structured light projection and the imaging module. The optimization strategy clarifies the independent control scheme for the multispectral LED light group and the voltage adjustment mapping table for the liquid zoom lens, such as the corresponding data of voltage values and focal length, and the light group activation combinations under different texture scenes. Ultimately, this forms the system foundation data containing all key information, which can be directly used for system initialization and operational control.
[0027] S102, based on the multimodal collaborative imaging constraint characteristics, designs the modeling rules for system data, clarifies the imaging association thresholds of 2D-3D-multispectral imaging and the light source switching response coefficients, and generates multimodal imaging modeling specifications.
[0028] In one implementation, based on the constraints of multimodal collaborative imaging, the requirements for optical path adaptation between modules, and the efficiency target of imaging mode switching, core information is classified and integrated to clarify the association type, switching logic, and adaptation range. Regarding association types, 2D imaging and multispectral imaging are "complementary associations." 2D imaging provides the location of macroscopic defects on the object's surface, while multispectral imaging supplements material spectral information to assist in defect characterization. Both share the camera and liquid zoom lens of the imaging module, and their optical paths must remain coaxially aligned. 2D imaging and 3D scanning are "progressive associations." 3D scanning must target the suspected defect area identified by 2D imaging to avoid wasting resources on full-area scanning; both rely on regional coordinate data transmission from the control unit. Multispectral imaging and 3D scanning are "time-divisional associations," requiring avoidance of light source interference and employing an alternating working mode.
[0029] In terms of switching logic, it follows the sequence of "multispectral imaging → 2D appearance inspection → 3D scanning of defect areas". For example, when inspecting the surface of electronic components, it first acquires spectral images by switching between wavelengths such as 475nm and 560nm using a multispectral ring light source, then turns on the coaxial illumination source for 2D defect localization, and finally turns off all light sources and starts the structured light projection module to scan the defect area. In terms of adaptability, 2D imaging is suitable for all workpiece surfaces within a working distance of 100mm-500mm, multispectral imaging is suitable for material analysis of workpieces with complex textures and different colors, and 3D scanning is suitable for the three-dimensional morphological reconstruction of suspected defect areas, and is only performed on the areas marked by 2D imaging.
[0030] The design logic of modeling rules is planned based on the requirements of imaging accuracy and process smoothness. The setting standards for imaging correlation threshold and light source switching response coefficient are clarified, and the core parameter information for rule design is generated. Imaging accuracy directly determines the reliability of industrial inspection results, while process smoothness affects inspection efficiency. Therefore, the design of modeling rules must take both into account. By clarifying the setting standards of key parameters, it is possible to ensure that multimodal imaging is coordinated and orderly and the results are accurate.
[0031] The setting of imaging association thresholds revolves around the synergistic matching degree between different imaging modes to ensure the accuracy of data association. For the regional association threshold between 2D imaging and 3D scanning, the standard is set at ≤±10 pixels. 2D imaging is mainly used to quickly locate suspected defect areas on the surface of an object, providing a precise target range for 3D scanning. If the regional deviation between the two is too large, it will cause the 3D scan to deviate from the core defect area or scan irrelevant areas, wasting inspection resources. For example, in the appearance inspection of electronic components, 2D imaging marks the boundary range of a defect area, and 3D scanning must strictly follow this range, controlling the pixel deviation within 10 pixels. This ensures complete coverage of the defect area while avoiding scanning unnecessary areas, guaranteeing the relevance and accuracy of 3D morphology reconstruction. For the material association threshold between multispectral imaging and 2D imaging, the standard is set at spectral similarity ≥90%. Multispectral imaging can acquire the spectral information of the object's material, while 2D imaging captures surface defect features. Combining the two allows for a more accurate determination of the defect type. For example, when inspecting plastic workpieces, if multispectral image analysis shows that the spectral information of a certain area of the material has a 92% similarity to the standard plastic spectrum, and the 2D image shows that there is a scratch defect in that area, it can be determined that the material is qualified but the surface is damaged; if the spectral similarity is only 85%, it is necessary to combine the defect characteristics to further determine whether the defect is caused by material abnormality. The high similarity threshold ensures the reliability of material judgment and assists in the accurate classification of defect types.
[0032] The setting of the light source switching response coefficient focuses on the timeliness and stability of imaging mode switching, avoiding the impact of light source switching on imaging quality and process efficiency. The standard setting for the light source on / off response time is ≤5ms. In multimodal imaging, different modes rely on different light sources. Excessive light source switching delay will increase the mode switching time, reduce detection efficiency, and may cause illumination superposition interference, affecting image clarity. For example, when performing 3D scanning, the multispectral ring light source and coaxial illumination source must be turned off before the structured light projection module is turned on. If the switching interval exceeds 5ms, it will prolong the detection time of a single workpiece, which will significantly reduce the overall efficiency, especially in batch detection scenarios. At the same time, residual illumination may interfere with the structured light pattern projection, affecting the accuracy of 3D coordinate calculation. The camera synchronous acquisition response coefficient is set to trigger acquisition ≤3ms after the light source state switch. After the light source switches, a short stabilization time is required to reach a stable illumination state. If acquisition is performed immediately, it will result in uneven image brightness and blurred details; if the waiting time is too long, it will waste detection time. For example, after the coaxial illumination source is turned on, the brightness of the light source reaches a stable state within 3ms. At this time, the camera is triggered to acquire 2D images, which can ensure uniform image brightness and clear defect details, providing high-quality data for subsequent defect identification. This ensures imaging accuracy without affecting the smoothness of the process.
[0033] By combining the imaging stability of complex textured surfaces with the adaptation requirements of different depths of field, optimization rules are set to dynamically adjust the correlation threshold and configure the response coefficient differently according to the imaging mode, ensuring the adaptability of multimodal collaboration. In actual inspection scenarios, there are significant differences in the complexity of workpiece texture and depth of field. Fixed correlation thresholds and response coefficients are difficult to adapt to all situations, which can easily lead to blurred imaging, correlation failure, or low efficiency. Therefore, accurate adaptation needs to be achieved through dynamic adjustment and differentiated configuration.
[0034] The dynamic adjustment rules for the association threshold are flexibly optimized based on scene characteristics to ensure the effectiveness of the association logic. For complex textured surfaces, the boundaries of defect areas marked by 2D imaging are easily blurred by dense textures. If the conventional region association threshold is used, it can easily cause 3D scanning to deviate from the actual defect range. In this case, the region association thresholds for 2D and 3D are reduced to ≤±5 pixels. By improving the region positioning accuracy, it is ensured that 3D scanning accurately focuses on the core defect area. For example, when inspecting plastic workpieces with dense textures, the complex texture can easily cause the defect boundary to be confused with the texture. After reducing the association threshold, the 3D scan can strictly fit the defect contour of the 2D imaging mark, avoiding scanning deviations caused by blurred boundaries and completely reconstructing the three-dimensional shape of the defect. For scenes with large depth of field, the increased working distance can easily lead to the attenuation of spectral information in multispectral images. If the original material association threshold is maintained, material association may fail, making it impossible to assist in defect characterization. In this case, the material association threshold is adjusted to ≥85%. By appropriately relaxing the standard, the material association logic between multispectral imaging and 2D imaging is ensured to be effective. For example, when inspecting large workpieces at a distance, the increased depth of field leads to a decrease in the intensity of the spectral signal. After adjusting the threshold, even if the spectral similarity decreases slightly, an effective association can still be established, and the defect type can be determined by combining 2D defect features.
[0035] The differentiated response coefficient configuration rules optimize switching efficiency and imaging quality based on the characteristics and requirements of each imaging mode. 2D imaging mode has extremely high requirements for illumination stability; sufficient time is needed for brightness to stabilize after light source switching, otherwise uneven image brightness and loss of defect details will occur. Therefore, the light source switching response coefficient is set to 3ms to ensure that acquisition is triggered only after the light source brightness stabilizes, guaranteeing the clarity of the 2D image and the accuracy of defect identification. Multispectral imaging mode requires switching multiple independent LED light groups. There are time differences in the start-up and wavelength stabilization of different wavelength light groups. If the response coefficient is too short, acquisition may occur before the wavelength stabilizes, affecting the accuracy of spectral information. Therefore, the response coefficient is set to 5ms to allow sufficient time for the light groups to complete start-up and wavelength calibration, ensuring the reliability of multispectral data. 3D scanning mode relies on a structured light projection module to project high-precision stripe patterns. This module has a fast start-up speed, and the stripe pattern can stabilize quickly after projection. Simultaneously, 3D scanning needs to efficiently capture deformed stripes to improve detection efficiency. Therefore, the response coefficient is set to 2ms to shorten the light source switching interval, improving the efficiency of the detection process while ensuring imaging quality, especially suitable for batch workpiece inspection scenarios.
[0036] The modeling basic information, core parameters of rule design, and optimization rules are comprehensively integrated to generate modeling specifications that include correlation constraint standards, timing control specifications, parameter adaptation requirements, and optimization strategies. The correlation constraint standards are clearly defined: the correlation between 2D, 3D, and multispectral systems must meet threshold requirements such as regional coordinate deviation and material similarity, and the optical paths must remain adapted. For example, the multispectral ring light source must coaxially surround the liquid zoom lens, and the optical axis of the structured light projection must form a 45° angle with the optical axis of the imaging module.
[0037] The timing control specifications are clearly defined: light source switching and camera acquisition must follow the timing logic of "multispectral off → 2D light source on → 2D acquisition → 2D light source off → 3D light source on → 3D acquisition," and the time intervals between each step must meet the response coefficient standard. Parameter adaptation requirements are clearly defined: imaging module parameters must match light source parameters. For example, when the working distance of the liquid zoom lens is 100mm, the short-wavelength LED group of the multispectral ring light source must be turned on; when the working distance is 500mm, the long-wavelength LED group must be turned on to ensure illumination intensity. The stripe pattern accuracy of the structured light projection module must be adapted to the camera resolution to ensure that the deformed stripe image can be accurately identified. The optimization strategy is clearly defined: the correlation threshold and response coefficient are dynamically adjusted according to the complexity of the workpiece texture and the depth of field range. Parameters are configured differently for different imaging modes, ultimately forming a multimodal imaging modeling specification that can adapt to various industrial inspection scenarios and ensure imaging accuracy and efficiency.
[0038] S103, in combination with the centralized control of the control unit and the need for multi-module collaboration, sets up a multi-level collaborative mechanism with adaptive adjustment of liquid zoom lens, timing switching of multiple light sources, precise scanning of region of interest, and modular detachable adaptation.
[0039] In one implementation, based on the centralized management and control requirements of the control unit and the goal of multi-module collaboration, the working parameters and interaction logic of each module are fully acquired, and key characteristics and adaptation requirements are extracted. In the imaging module, the liquid zoom lens features electronically controlled focusing. Responding to voltage signals from the control unit, it adjusts the focal length by changing the refractive index of the internal liquid. For example, when the working distance is adjusted from 200mm to 300mm, the control unit sends a corresponding voltage signal to achieve precise focal length adaptation. The light source switching sequence involves multiple light sources working in a time-sharing manner. During 3D scanning, the coaxial illumination source and multispectral ring light source must be turned off before the structured light projection module is turned on. During multispectral imaging, only the multispectral ring light source is turned on, and during 2D imaging, only the coaxial illumination source is turned on. The scanning area positioning accuracy depends on the control unit's analysis of the multispectral image, which can accurately identify the coordinates of the defect area, ensuring that the structured light projection module scans only the target area. For example, if the defect area coordinates are identified as (x3, y3) to (x4, y4), the scanning range is strictly limited to this range. The modular adaptation requirement is that the structured light projection module is detachably connected via a side-mounted bracket. The bracket includes a horizontal adjustment rail and a pitch adjustment mechanism. After reinstallation, the optical axis must be kept at a 45° angle to the imaging module using positioning pins and fine-tuning knobs.
[0040] The extracted module parameters and interaction logic were subjected to compliance verification to confirm their adaptability and generate verification results. For parameter adaptability verification, the matching degree between the camera resolution (≥20 million pixels) and the working distance of the liquid zoom lens (100mm-500mm) was checked to ensure that the camera could capture clear details across the entire working distance range. The verification result indicated that the parameter adaptability was qualified. For timing coordination verification, the light source switching process was simulated to verify whether the timing interval from turning off the multispectral ring light source to turning on the structured light projection module met the interference-free requirements. The measured interval was 3ms, with no light superposition interference, indicating that the timing coordination was qualified. For positioning accuracy verification, simulated defect areas were marked using a standard calibration board, and the deviation between the structured light scanning area and the marked area was tested. The deviation was ≤5 pixels, indicating that the positioning accuracy was qualified. For compatibility verification, the optical axis angle reset accuracy after disassembly and assembly of the structured light projection module was tested. The angle deviation after reset was ≤0.5°, indicating that the compatibility was qualified. Finally, a module collaborative verification result containing all verification conclusions was generated.
[0041] By combining different depth-of-field adaptation and complex texture imaging requirements, a multi-level collaborative mechanism is planned for implementation scenarios according to imaging modes. The activation conditions and execution standards of each mode are clearly defined to ensure the adaptability of the mechanism and the imaging effect. Different imaging modes target different detection targets and scene characteristics. By accurately defining the activation conditions and execution standards, efficient collaboration and accurate imaging of each mode can be achieved.
[0042] The 2D imaging mode focuses on rapidly capturing macroscopic defects on object surfaces. Its activation condition is set for scenarios requiring quick screening of object surface integrity, without the need for in-depth analysis of material or 3D morphology. In terms of execution standards, a coaxial illumination source is first activated. This source provides uniform illumination, effectively reducing interference from surface reflections and ensuring clear defect outlines. Exposure time and gain are simultaneously set to ensure moderate image brightness and discernible details. The liquid zoom lens needs to be flexibly adjusted to the corresponding working distance based on the workpiece height, ensuring the workpiece surface is within the clear imaging range. For example, when inspecting a workpiece of suitable height, the lens working distance is simultaneously adjusted to match the workpiece height, ensuring that macroscopic defects across the entire workpiece surface are clearly captured, providing basic defect location information for subsequent inspections.
[0043] Multispectral imaging mode focuses on material analysis and defect detection on complex textured surfaces. It is activated when it's necessary to distinguish whether a defect is caused by material abnormalities, or in scenarios where defects on complex textured surfaces are difficult to identify under conventional lighting. The standard procedure involves activating a multispectral ring light source, which comprises multiple independent LED groups emitting different center wavelengths. Illumination is switched sequentially according to a predetermined wavelength sequence, and the camera simultaneously acquires multispectral images at the corresponding wavelengths. Different wavelengths of light interact differently with the material, accurately reflecting differences in material composition. The ring-shaped arrangement of the LEDs ensures uniform illumination, suitable for complex textured surfaces. For example, when inspecting dark, complex-textured workpieces, the dark surface strongly absorbs visible light, easily obscuring texture details. In this case, the near-infrared LED group is activated, utilizing the penetrability of near-infrared light to enhance the contrast between texture and defects, clearly revealing subtle defects hidden within the complex texture. Simultaneously, multi-wavelength data comparison helps determine whether the defect is caused by material differences.
[0044] The 3D scanning mode aims to accurately reconstruct the three-dimensional morphology of suspected defect areas, providing a basis for quantitative defect analysis. It is activated when 2D or multispectral imaging has identified suspected defect areas, requiring further acquisition of 3D information such as the depth and size of the defects. The execution standard involves first turning off other light sources to avoid stray light interfering with the structured light pattern projection, ensuring 3D imaging accuracy. Then, the structured light projection module is activated, projecting Gray code and sinusoidal fringe patterns. These two patterns possess high-precision encoding characteristics, accurately reflecting changes in the object's surface morphology. The camera simultaneously captures the deformed fringe image, and combined with triangulation principles, calculates the 3D coordinates of the object's surface, accurately reconstructing the 3D morphology of the suspected defect area. Furthermore, the scanning process only targets the marked defect area, eliminating the need to scan the entire workpiece surface, significantly improving detection efficiency while avoiding interference from irrelevant data areas, ensuring the accuracy of the defect's 3D information.
[0045] Based on the actual requirements of industrial inspection efficiency and accuracy, collaborative fault-tolerance rules are set to clarify the tolerance range of various errors. The adjustment response delay tolerance range is ≤10ms after the liquid zoom lens receives the control signal; for example, after the control unit sends a voltage signal, the lens must complete the focal length adjustment from the adapted 200mm working distance to the 300mm working distance within 10ms to ensure a smooth inspection process. The switching interval deviation tolerance range is ≤2ms between the actual interval and the preset interval for light source switching; for example, after the preset multispectral ring light source is turned off, the structured light projection module is turned on 3ms later, and the actual interval is allowed to be between 1ms and 5ms. The scanning positioning error tolerance range is ≤8 pixels between the pixel deviation of the structured light scanning area and the marked defect area to avoid missing or mis-scanning areas due to positioning deviation. The adaptation reset error tolerance range is ≤1° after the structured light projection module is disassembled and reassembled, ensuring that the 3D scanning accuracy requirements are still met after reassembly.
[0046] By integrating scene planning parameters and fault tolerance rules, operational specifications were formulated, clarifying the triggering process, parameter adjustment method, and module linkage logic of each mechanism. The triggering process for the adaptive adjustment of the liquid zoom lens is as follows: after system initialization, the control unit triggers adjustment based on the preset working distance or initial imaging feedback. The parameter adjustment method involves the control unit sending a corresponding voltage signal. The module linkage logic involves the voltage signal being transmitted to the liquid zoom lens, where changes in the refractive index of the liquid inside the lens adjust the focal length. During the adjustment process, the camera remains in standby mode, and data acquisition is triggered once the focal length stabilizes.
[0047] The triggering process for multi-source timing switching is automatic based on the selected imaging mode. Parameter adjustment involves the control unit sending a switch command to control the light source's start and stop. The module linkage logic is that after the light source state is switched, the camera triggers image acquisition according to a preset response time; for example, the camera starts acquisition 3ms after the coaxial illumination source is turned on. The triggering process for precise region of interest scanning is triggered after multispectral image analysis identifies the defect area. Parameter adjustment involves the control unit transmitting the defect area coordinates to the structured light projection module. The module linkage logic is that the structured light projection module focuses on the target area and projects a stripe pattern, while the camera acquires the image simultaneously. The triggering process for modular, detachable, and adaptable systems is manually triggered by the user according to inspection needs. Parameter adjustment involves adjusting the position angle using the horizontal adjustment rail and pitch adjustment mechanism of the side mounting bracket. The module linkage logic is that after adjustment, the control unit confirms that the optical axis angle meets the standard before starting the imaging process.
[0048] The system integrates module collaboration verification results, scene planning parameters, fault tolerance rules, and operating specifications to generate complete multi-level collaborative mechanism implementation result information. The collaborative mechanism is defined by four core mechanisms: adaptive adjustment of the liquid zoom lens, multi-source timing switching, precise scanning of the region of interest, and modular detachable adaptation. All are designed around the goals of centralized control of the control unit and multi-module collaboration. Triggering conditions are clearly defined for 2D imaging, multispectral imaging, and 3D scanning scenarios, consistent with scene planning parameters. The execution standard details the operating steps, parameter ranges, and module linkage requirements for each mechanism, such as the working distance adjustment range of the liquid zoom lens and the wavelength switching sequence of the light source. The fault tolerance range clarifies the tolerable upper limits for various errors, such as adjustment response delay and switching interval deviation. The operating process is logically structured as "trigger-adjustment-linkage-acquisition," forming a closed-loop operating guide. Ultimately, this results in a set of multi-level collaborative mechanism implementation result information that can directly guide system operation.
[0049] S104 integrates and executes the system's basic data, multimodal imaging modeling specifications, and multi-level collaborative mechanisms. It iteratively optimizes the imaging effect through target image processing algorithms, following a complete workflow of system power-on, calibration, parameter setting, 2D appearance inspection, 3D scanning of suspected defect areas, data processing and analysis, generation of inspection reports, and system power-off. It generates multimodal imaging results that include module collaborative features, multispectral semantic information, dynamic focusing timing attributes, and 3D morphology reconstruction accuracy.
[0050] In one implementation, based on the requirements of multimodal imaging integration and the target of industrial detection accuracy, the system's basic data, multimodal imaging modeling specifications, and multi-level collaborative mechanisms are comprehensively integrated and utilized. Core information such as module collaboration parameters, imaging constraint standards, mechanism execution logic, and algorithm optimization requirements are deeply extracted, laying the foundation for efficient advancement and accuracy assurance of the subsequent imaging process. Module collaboration parameters are key data ensuring the orderly linkage of each module, covering liquid zoom lens voltage-focal length mapping parameters, multi-source switching timing intervals, and the angle between the structured light projection and the imaging module's optical axis. The liquid zoom lens voltage-focal length mapping parameters clarify the precise correspondence between different voltage signals and their corresponding focal lengths. Through these parameters, the control unit can send a specific voltage to the lens according to the detection distance, quickly adjusting the focal length to adapt to imaging requirements. The multi-source switching timing interval clarifies the time connection standard for the on and off of different light sources, ensuring interference-free light source switching and guaranteeing imaging quality. The angle between the structured light projection and the imaging module's optical axis is a fixed value; this angle setting maximizes the advantages of triangulation principles, providing a precise geometric basis for 3D topography reconstruction.
[0051] Imaging constraint standards are core criteria for defining the boundaries and quality baselines of the imaging process. They clarify the region correlation thresholds for 2D-3D-multispectral imaging, the light source switching response time, and the light source control requirements for 3D scanning. The region correlation threshold limits the deviation range for defect region localization under different imaging modes, ensuring accurate correspondence between the defect region marked in 2D imaging and the 3D scanning area, thus avoiding scanning deviations. The light source switching response time requirement reduces mode switching time, improves detection efficiency, and avoids interference from light superposition caused by slow switching. The requirement to shut down other light sources during 3D scanning eliminates the influence of stray light on the projection of structured light patterns, ensuring the clarity of the fringe image and the accuracy of deformation data.
[0052] The mechanism's execution logic standardizes the overall process and module adaptation method for multimodal imaging. Its core is a time-sharing collaborative process of "multispectral imaging → 2D detection → 3D scanning of defect areas," along with modular and detachable adaptation requirements. The time-sharing collaborative process proceeds according to the logic of "material analysis and defect screening first, followed by precise 3D reconstruction." First, multispectral imaging acquires material information of the object to help determine the nature of defects. Then, 2D detection quickly locates suspected defect areas. Finally, 3D scanning is performed on these areas, avoiding the inefficiency of full-area scanning. Modular and detachable adaptation allows for rapid assembly and disassembly of the structured light projection module via a side-mounted bracket. After reassembly, the optical axis angle can be precisely reset using positioning pins and fine-tuning knobs, meeting the functional configuration requirements of different detection scenarios.
[0053] Algorithm optimization is a core technological support for improving imaging accuracy and efficiency, encompassing 3D coordinate calculation algorithms based on triangulation principles, multispectral image defect recognition algorithms, and liquid zoom lens focal length adaptive calibration algorithms. The 3D coordinate calculation algorithm based on triangulation principles captures the deformed image of the stripe pattern projected by the structured light projection module onto the object's surface. Combining this with parameters such as the base distance between the camera and projector, lens focal length, and projection angle, it calculates the three-dimensional coordinates of each point on the object's surface, achieving accurate reconstruction of defect areas.
[0054] Multispectral image defect recognition algorithms accurately locate and distinguish defect types by comparing the differences in the spectral response of an object's surface at different wavelengths. The core logic is "spectral feature extraction + difference comparison + defect determination," adaptable to the detection of defects involving complex textures and material anomalies. For spectral feature extraction, images acquired from a multispectral ring light source at different wavelengths (such as 475nm, 560nm, 668nm) are processed to extract features such as spectral reflectance and absorbance for each pixel, forming a "pixel-spectrum" feature matrix. For example, the reflectance characteristics of dark plastic workpieces in the near-infrared band (842nm) differ significantly from those in the visible light band (560nm). The algorithm captures this difference and establishes a standard spectral library.
[0055] For difference comparison, the measured spectral features are compared with the spectral features of standard samples to calculate similarity (such as cosine similarity and Euclidean distance). At the same time, the grayscale and edge features of 2D images are combined to eliminate texture interference. For example, the normal area of a workpiece with complex texture exhibits regular fluctuations in its multi-wavelength spectral features, while defective areas (such as scratches and material impurities) will disrupt this regularity, resulting in a significant decrease in similarity.
[0056] For defect identification, a spectral similarity threshold is set (e.g., ≥90% for normal scenes and ≥85% for deep-field scenes). Areas with grayscale / edge abnormalities in the 2D image that are below the threshold are identified as defects and classified according to the type of spectral difference. For example, defects caused by material impurities show a sharp increase in spectral absorption at a specific wavelength, while surface scratches show a slight decrease in reflectivity across the entire wavelength.
[0057] The adaptive focal length calibration algorithm for liquid zoom lenses dynamically corrects the voltage-focal length mapping relationship by providing real-time feedback on image sharpness. This ensures accurate focal length adaptation under different depths of field and environmental conditions. The core logic is "sharpness evaluation + deviation correction + mapping update," addressing focal length shifts caused by mechanical wear and temperature changes. For sharpness evaluation, an image at the current focal length is acquired, and image sharpness indices (such as mean edge gradient and information entropy) are calculated using gradient operators (e.g., Sobel operator) and entropy methods. Higher sharpness indicates a better fit between the focal length and the current working distance. For example, for minute textures on a workpiece surface, the mean edge gradient of a sharp image is significantly higher than that of a blurry image.
[0058] For deviation correction, the measured sharpness is compared with a preset threshold (e.g., average edge gradient ≥ 80). If the threshold is not met, the deviation direction (over-focus / under-focus) is calculated, and the voltage signal output by the control unit is adjusted according to the degree of deviation. For example, when the working distance changes from 200mm to 300mm, the focal length corresponding to the initial voltage is too short, and the image sharpness is below the threshold. The algorithm determines "under-focus" and sends a command to reduce the voltage to the liquid zoom lens. For mapping updates, after multiple iterations of calibration, the "optimal voltage-focal length" correspondence at the current working distance is recorded, and the system's built-in voltage-focal length mapping table is updated. Subsequent working distances of the same type can directly call the calibrated parameters to achieve fast and accurate focusing.
[0059] Based on the complete workflow of "system power-on - calibration - parameter setting - 2D detection - 3D scanning - data processing - report generation - system power-off", a step-by-step execution plan is designed and the operation standards of each link are clarified. The imaging effect is iteratively optimized through dedicated image processing algorithms to generate clear and accurate 2D images, multispectral data and preliminary results of 3D morphology reconstruction, ensuring that the detection process is orderly and the imaging quality meets the standards.
[0060] After power-on, the system automatically initiates self-test programs for each component. The control unit verifies the circuit connections and hardware status of the camera, liquid zoom lens, light source, and structured light projection module one by one. Once confirmed to be fault-free, preset parameters and control programs are loaded. Preset parameters include camera resolution, default working distance of the liquid zoom lens, initial brightness of the light source, and initial angle of the structured light projection module. For example, the system loads a camera resolution of 20 megapixels, a default working distance of 200mm for the liquid zoom lens, and an initial brightness of medium for the multispectral ring light source. Simultaneously, it loads multimodal collaborative control programs and image acquisition trigger logic, laying the hardware and software foundation for subsequent stages. If the self-test detects a component abnormality (such as a non-responsive light source or a loose lens connection), the system will trigger an alarm and pause the process until the fault is resolved.
[0061] A standard calibration board is used to complete the camera's intrinsic and extrinsic parameters, as well as hand-eye calibration, ensuring imaging and measurement accuracy. The standard calibration board is placed in a preset position on the worktable. The control unit controls the camera to acquire images of the calibration board at different angles and distances. Zhang's calibration algorithm is used to calculate the camera's intrinsic parameters (such as focal length, principal point coordinates, and distortion coefficients) to correct imaging errors caused by lens distortion. The PnP algorithm is used to solve for the camera's extrinsic parameters (such as the relative position and attitude of the camera and calibration board), establishing a mapping relationship between the world coordinate system and the image coordinate system. Hand-eye calibration involves acquiring the stripe pattern projected onto the calibration board by the structured light projection module. Combined with the images acquired by the camera, the relative positional relationship between the structured light projection module and the camera is calibrated using an iterative nearest-point algorithm, ensuring that the angle between the projection optical axis and the imaging module's optical axis is precisely maintained at 45°, providing accurate geometric parameters for 3D coordinate calculations. After calibration, the control unit stores the calibration results for subsequent image correction and 3D reconstruction.
[0062] Based on the inspection requirements and workpiece characteristics, the core parameters of each imaging mode are configured through the control unit to ensure parameter adaptation to the scenario. In 2D imaging mode, exposure time and gain are set; for example, exposure time is set to 10ms and gain to 1, balancing image brightness and noise to avoid overexposure or underexposure. In multispectral imaging mode, wavelength channels are selected, such as 475nm, 560nm, 668nm, 717nm, and 842nm, covering the visible and near-infrared bands to adapt to the inspection needs of different materials and textures. In 3D mode, the structured light projection pattern type and projection frequency are set, using Gray code and sinusoidal fringe patterns. Gray code is used for rapid coarse positioning, while sinusoidal fringes are used for high-precision phase calculation. Simultaneously, the projection frequency is set to synchronize with the camera's frame rate to avoid pattern blurring. After parameter settings are completed, the system enters the ready-to-inspection state, waiting for the workpiece to be positioned.
[0063] The control unit activates the coaxial illumination source and adjusts the brightness to the appropriate level, ensuring uniform and glare-free illumination of the workpiece surface. The camera acquires 2D images of the workpiece surface according to preset parameters and processes the images using an adaptive threshold segmentation algorithm and a Canny edge detection algorithm: the adaptive threshold segmentation algorithm dynamically adjusts the segmentation threshold based on the local grayscale features of the image to distinguish the workpiece from the background and eliminate ambient light interference. The Canny edge detection algorithm extracts the edge contours of the workpiece surface and combines morphological processing (such as expansion and corrosion) to eliminate noise points and accurately identify macroscopic defects such as scratches, dents, and protrusions. If a suspected defect area is detected, the control unit records the image coordinate range of that area to provide target area information for subsequent 3D scanning; if no defect is detected, the system can directly proceed to the report generation stage or initiate multispectral imaging for further verification as needed.
[0064] The control unit shuts off the coaxial illumination source to prevent stray light from interfering with structured light imaging. It then activates the structured light projection module to project a preset Gray code and sinusoidal fringe pattern onto the previously marked suspected defect area. The camera simultaneously acquires the deformed fringe image. During acquisition, the liquid zoom lens dynamically adjusts its focus based on the workpiece distance to ensure a clear fringe image. A Gray code decoding algorithm is used to obtain coarse positioning information of the defect area and determine the initial phase of each point. A phase-shifting algorithm (such as the four-step phase-shifting method) is then used to calculate high-precision phase values, converting the phase information into pixel-level displacement data. Finally, combining the intrinsic and extrinsic parameters obtained during calibration with the structured light projection parameters, a triangulation algorithm is used to calculate the three-dimensional coordinates of each point in the defect area, initially reconstructing the three-dimensional shape of the defect area and forming point cloud data, thus completing the conversion from a two-dimensional image to a three-dimensional model.
[0065] Dedicated algorithms are used to denoise and enhance 2D images, multispectral data, and 3D point cloud data, improving data quality. For 2D images, median filtering removes salt-and-pepper noise, and histogram equalization enhances the contrast between defect and normal areas, making defect outlines clearer. For multispectral data, principal component analysis extracts spectral features at key wavelengths, eliminating redundant information from environmental interference and retaining effective spectral data relevant to material differences. For 3D point cloud data, pass-through filtering removes noise points beyond reasonable distances, radius filtering eliminates isolated points, and moving least squares smooths the sparse point cloud, filling in some missing points to generate a uniformly dense, low-noise 3D point cloud model, providing high-quality data support for subsequent analysis and report generation.
[0066] Based on imaging accuracy requirements and data validity standards, an optimization scheme was designed, and optimized multi-dimensional imaging data was generated through correction and verification. The threshold for 3D topography reconstruction accuracy was defined as ≤±2.5μm, and the defect area localization error range was ≤±8 pixels. The 3D coordinate calculation results were corrected using the triangulation principle; for example, by substituting parameters such as a camera pixel size of 5μm and a baseline length of 100mm into the formula ΔZ=(pd) / (ftanθ), the 3D coordinate deviation was corrected. Cross-validation of 2D image defect identification results was performed by comparing the material analysis results of the multispectral image with the 2D defect features. For example, if the multispectral image showed an abnormal material in a certain area, and the 2D image showed an appearance defect in that area, then the area was confirmed as a real defect, eliminating false detections. For blurred images of complex textured surfaces, an enhancement algorithm was used to optimize the contrast of the multispectral data, improving the accuracy of defect identification. Finally, multi-dimensional imaging data containing accurate 2D defect information, multispectral material data, and high-precision 3D topography data was generated.
[0067] Based on the requirements of multimodal information fusion and the specifications of testing reports, result integration and verification rules are set, and a result verification report is generated after verification. The verification includes: Module coordination feature consistency verification, checking whether the timing coordination of liquid zoom lens adjustment and light source switching conforms to the specifications, such as whether the light source switching is triggered in time after focus adjustment to ensure no coordination conflicts; Multispectral semantic information accuracy verification, verifying whether the material identification results of multispectral wavelength channels are consistent with standard samples, such as whether the result of identifying plastic materials using a 475nm wavelength is accurate; Dynamic focusing timing rationality verification, checking whether the focusing response time of the liquid zoom lens is ≤10ms to ensure efficiency in adapting to different depth-of-field scenes; 3D accuracy compliance verification, testing the 3D topography reconstruction accuracy using standard parts to confirm whether it meets the requirement of ≤±2.5μm; The verification results are compiled and summarized to generate a result verification report including verification items, standard requirements, measured results, and pass / fail judgments.
[0068] The system integrates step-by-step execution plans, result optimization plans, verification rules, and verification reports to generate a comprehensive multimodal imaging final result. Module collaboration features clearly define the collaborative effects of mechanisms such as adaptive adjustment of the liquid zoom lens and multi-source timing switching, for example, recording data such as a focus adjustment response time of 8ms and a light source switching interval of 3ms. Multispectral semantic information details the material analysis results for each wavelength channel, such as the reflectivity data of a certain region at 475nm wavelength and the corresponding material determination. Dynamic focusing timing attributes record the focusing time and focus changes at different working distances, such as the focusing process and timing when the working distance is adjusted from 150mm to 300mm. 3D morphology reconstruction accuracy clearly shows the measured accuracy value, such as ±2.2μm. Defect detection details include defect location coordinates, dimensions (e.g., 5mm long, 3mm wide), type (e.g., scratches, dents), and multimodal verification results, ultimately forming a complete and accurate multimodal imaging final result.
[0069] In one implementation, such as Figure 2 As shown, this application also provides a multimodal visual imaging system based on a smart sensor, comprising:
[0070] The system basic data acquisition and standardization module 201 is used to acquire core data including imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic. Through structured acquisition and standardization processing, it generates system basic data including functional module type, optical path correlation intensity, and imaging timing parameters.
[0071] The multimodal imaging modeling specification design module 202 is used to design system data modeling rules based on the multimodal collaborative imaging constraint characteristics, clarify the imaging association thresholds of 2D-3D-multispectral imaging and the light source switching response coefficients, and generate multimodal imaging modeling specifications.
[0072] The multi-level collaborative mechanism setting module 203 is used to set up a multi-level collaborative mechanism that combines the centralized control of the control unit and the collaborative needs of multiple modules, including adaptive adjustment of the liquid zoom lens, timing switching of multiple light sources, precise scanning of the region of interest, and modular detachable adaptation.
[0073] The multimodal imaging integration execution and result generation module 204 is used to integrate system basic data, multimodal imaging modeling specifications and multi-level collaborative mechanisms. It iteratively optimizes the imaging effect through target image processing algorithms, following a complete workflow of system power-on - calibration - parameter setting - 2D appearance inspection - 3D scanning of suspected defect areas - data processing and analysis - generation of inspection report - system power-off. It generates multimodal imaging results that include module collaborative features, multispectral semantic information, dynamic focusing timing attributes and 3D morphology reconstruction accuracy.
[0074] The computer-readable storage medium provided in the above embodiments of this application and the multimodal visual imaging method based on smart sensors provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
Claims
1. A multimodal visual imaging method based on intelligent sensors, characterized in that, include: Based on the requirements of multimodal imaging and the functional goals of composite systems, the imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic are collected and standardized in a structured manner to generate basic system data containing functional module types, optical path correlation strength, and imaging timing parameters. Based on the constraints of multimodal collaborative imaging, modeling rules for system data are designed, the imaging correlation thresholds of 2D-3D-multispectral imaging and the light source switching response coefficients are defined, and multimodal imaging modeling specifications are generated. Combining the centralized control and management of the control unit with the need for multi-module collaboration, a multi-level collaborative mechanism is set up, including adaptive adjustment of the liquid zoom lens, timing switching of multiple light sources, precise scanning of the region of interest, and modular detachable adaptation. The system integrates and executes basic data, multimodal imaging modeling specifications, and multi-level collaborative mechanisms. It iteratively optimizes imaging effects through target image processing algorithms, following a complete workflow of system power-on, calibration, parameter setting, 2D appearance inspection, 3D scanning of suspected defect areas, data processing and analysis, generation of inspection reports, and system power-off. This generates multimodal imaging results that include module collaborative features, multispectral semantic information, dynamic focusing timing attributes, and 3D morphology reconstruction accuracy.
2. The method as described in claim 1, characterized in that, Based on the requirements of multimodal imaging and the functional goals of the composite system, the imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic are structurally acquired and standardized to generate basic system data containing functional module types, optical path correlation strength, and imaging timing parameters, including: Based on the requirements of multimodal imaging, the goal of composite system integration, and the requirements of industrial detection accuracy, the core parameters of the imaging module, the characteristic parameters of the light source, the key parameters of structured light projection, and the timing control logic information are integrated to clarify the acquisition range and correlation requirements of the parameters of each module. The parameter acquisition logic is designed based on the data standardization processing requirements, and the acquisition standards for core parameters and auxiliary parameters are clarified. The core parameters include camera resolution, working distance of liquid zoom lens, and wavelength range of light source, while the auxiliary parameters include module installation angle and connection interface type information. In order to meet the requirements of imaging stability of complex textured surfaces, parameter optimization rules were set for independent control of multispectral ring light source LED lamp group and precise adjustment of liquid zoom lens voltage signal to ensure data validity; The system processes parameter acquisition requirements, data standardization requirements, and parameter optimization rules to generate basic system data that includes parameter types, acquisition specifications, correlation logic, and optimization strategies.
3. The method as described in claim 1, characterized in that, Modeling rules for system data are designed based on the constraints of multimodal collaborative imaging, clarifying the imaging correlation thresholds and light source switching response coefficients for 2D-3D-multispectral imaging, and generating multimodal imaging modeling specifications, including: Based on the constraints of multimodal collaborative imaging, the requirements for optical path adaptation between modules, and the efficiency target of imaging mode switching, the association logic, light source switching timing, and parameter adaptation standards of 2D-3D-multispectral imaging are classified and integrated to generate basic modeling information including association type, switching logic, and adaptation range. Based on the requirements of imaging accuracy and workflow smoothness, the design logic of modeling rules is planned, the setting standards of imaging association threshold and light source switching response coefficient are clarified, and the core parameter information of rule design is generated. Combining the imaging stability of complex textured surfaces and the adaptation requirements of different depths of field, optimization rules are set to dynamically adjust the associated threshold and configure the response coefficient differently according to the imaging mode, so as to ensure the adaptability of multimodal collaboration. The basic information for modeling, core parameters for rule design, and optimization rules are processed to generate a multimodal imaging modeling specification that includes correlation constraint standards, timing control specifications, parameter adaptation requirements, and optimization strategies.
4. The method as described in claim 1, characterized in that, Combining the centralized control and management requirements of the control unit with the need for multi-module collaboration, a multi-level collaborative mechanism is established, including adaptive adjustment of the liquid zoom lens, timing switching of multiple light sources, precise scanning of the region of interest, and modular, detachable, and adaptable features. Based on the centralized control requirements of the control unit and the goal of multi-module collaboration, the working parameters and interaction logic of the imaging module, light source, and structured light projection module are obtained, and the adjustment characteristics of the liquid zoom lens, the light source switching sequence, the scanning area positioning accuracy, and the modular adaptation requirements are extracted. Perform compliance verification on the extracted module parameters and interaction logic to confirm parameter adaptability, timing coordination, positioning accuracy and compatibility, and generate module collaboration verification results; Based on the different depth-of-field adaptation and complex texture imaging requirements, the implementation scenarios of the multi-level collaborative mechanism are planned. The activation conditions and execution standards of each mechanism are clarified according to the application scenarios of 2D imaging, multispectral imaging and 3D scanning, and scenario planning parameters are generated. Based on the actual requirements of industrial testing efficiency and accuracy, collaborative fault tolerance rules are set to clarify the tolerable adjustment response delay, switching interval deviation, scanning positioning error and adaptation reset error range. Integrate scenario planning parameters and fault tolerance rules, formulate multi-level collaborative mechanism operation specifications, and clarify the triggering process, parameter adjustment method and module linkage logic of each mechanism; The system integrates module collaboration verification results, scenario planning parameters, fault tolerance rules, and operation specifications to generate multi-level collaboration mechanism implementation result information, including collaboration mechanism definition, triggering conditions, execution standards, fault tolerance range, and operation procedures.
5. The method as described in claim 4, characterized in that, The system integrates and executes basic data, multimodal imaging modeling specifications, and multi-level collaborative mechanisms. Imaging effects are iteratively optimized through target image processing algorithms. Following a complete workflow—system power-on, calibration, parameter setting, 2D appearance inspection, 3D scanning of suspected defect areas, data processing and analysis, report generation, and system power-off—it generates multimodal imaging results containing module collaborative features, multispectral semantic information, dynamic focusing timing attributes, and 3D topography reconstruction accuracy. Based on the requirements of multimodal imaging integration and the industrial detection accuracy target, the system's basic data, multimodal imaging modeling specifications and multi-level collaborative mechanisms are integrated and invoked, and core information such as module collaborative parameters, imaging constraint standards, mechanism execution logic and algorithm optimization requirements are extracted. Based on a complete workflow, a step-by-step execution plan is designed, which clarifies the operational standards for each step, including system power-on initialization, calibration, parameter setting, 2D appearance inspection, 3D scanning of suspected defect areas, data processing and analysis, inspection report generation, and system power-off. The imaging effect is iteratively optimized through a dedicated image processing algorithm to generate 2D images, multispectral data, and preliminary results of 3D morphology reconstruction. Based on the imaging accuracy requirements and data validity standards, the design results optimization scheme is designed, the threshold for 3D shape reconstruction accuracy and the error range of defect area positioning are clarified, the 3D coordinate calculation results are corrected by combining the triangulation principle, the 2D image defect identification results are cross-validated, and optimized multi-dimensional imaging data is generated. Based on the requirements of multimodal information fusion and the standards of detection reports, result integration and verification rules are set to confirm the consistency of module collaborative features, the accuracy of multispectral semantic information, the rationality of dynamic focusing timing, and the compliance of 3D accuracy, and a result verification report is generated. By integrating the step-by-step execution plan, result optimization plan, verification rules and verification report, the final multimodal imaging result information is generated, which includes module collaborative features, multispectral semantic information, dynamic focusing timing attributes, 3D morphology reconstruction accuracy and defect detection details.
6. A multimodal visual imaging system based on intelligent sensors, characterized in that, The system includes: The system's basic data acquisition and standardization module is used to acquire core data including imaging module parameters, light source wavelength characteristics, structured light projection angle, and timing control logic. Through structured acquisition and standardization processing, it generates system basic data including functional module type, optical path correlation intensity, and imaging timing parameters. The multimodal imaging modeling specification design module is used to design system data modeling rules based on the constraint characteristics of multimodal collaborative imaging, clarify the imaging association thresholds of 2D-3D-multispectral imaging and the light source switching response coefficients, and generate multimodal imaging modeling specifications. The multi-level collaborative mechanism setting module is used to combine the centralized control of the control unit and the collaborative needs of multiple modules to set up a multi-level collaborative mechanism for adaptive adjustment of the liquid zoom lens, timing switching of multiple light sources, precise scanning of the region of interest, and modular detachable adaptation. The multimodal imaging integration execution and result generation module is used to integrate system basic data, multimodal imaging modeling specifications and multi-level collaborative mechanisms. It iteratively optimizes the imaging effect through target image processing algorithms, following a complete workflow of system power-on - calibration - parameter setting - 2D appearance inspection - 3D scanning of suspected defect areas - data processing and analysis - generation of inspection report - system power-off. It generates multimodal imaging results that include module collaborative features, multispectral semantic information, dynamic focusing timing attributes and 3D morphology reconstruction accuracy.
7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the multimodal visual imaging method based on a smart sensor as described in any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the multimodal visual imaging method based on smart sensors as described in any one of claims 1 to 5.