Panoramic stitching industrial camera device for complex scenes and stitching method
By combining a multi-camera array and an FPGA processor, sub-microsecond-level synchronous triggering and real-time image processing are achieved, solving the synchronization error and illumination vibration problems of traditional industrial panoramic imaging systems in complex environments, improving stitching success rate and detection accuracy, and reducing system power consumption.
Patent Information
- Application Number
- CN202510626514.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Traditional industrial panoramic imaging systems suffer from large synchronization errors, sudden changes in illumination, and image quality problems caused by mechanical vibrations in complex environments, making it difficult to meet the real-time and precision requirements for high-speed motion scenes and complex structure detection.
A combination of multi-camera array, FPGA processor and calibration module is adopted to achieve sub-microsecond synchronous triggering, real-time image preprocessing and fusion. Combined with multi-modal sensor monitoring of environmental conditions for dynamic adjustment, panoramic images are generated using feature point matching and multi-band fusion algorithms.
Reduce splicing misalignment error, improve splicing success rate, reduce system power consumption, enhance detection sensitivity and field of view coverage, and meet the high-precision detection needs of complex industrial scenarios.
Smart Images

Figure CN120263947B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of industrial cameras, in particular to a panoramic splicing industrial camera device and splicing method for complex scenes. BACKGROUND
[0002] Industrial cameras are image acquisition devices designed specifically for industrial environments, used in machine vision, automated inspection, monitoring, and other scenarios. It usually has high resolution, high frame rate, strong durability, anti-interference ability and other characteristics, and can work stably in complex environments. Common applications include production line quality inspection, robot navigation and logistics sorting, etc.
[0003] The traditional industrial panoramic imaging system has significant technical bottlenecks in multi-camera synchronization, environmental adaptability, and real-time processing: when using software triggering or discrete hardware triggering, the synchronization error between cameras is generally more than 100 microseconds, resulting in a misregistration artifact of ≥0.1mm in the spliced image under high-speed motion scenes (such as 1m / s conveyor belt), which cannot meet the precision detection requirements; sudden changes in light (such as arc welding instantaneous light intensity >10^5 lux) and mechanical vibration (5-500Hz frequency band vibration energy ratio >30%) in complex industrial environments can cause local overexposure, focus blur and geometric distortion accumulation. The existing automatic exposure and anti-shake technology has a response delay of >200ms and is not coordinated with the splicing algorithm, resulting in a dynamic scene splicing failure rate of more than 15%; when relying on CPU / GPU architecture in the image processing link, the SIFT feature matching and multi-band fusion algorithm takes more than 500ms to process 4K resolution images, and the power consumption is >50W, which is difficult to meet the real-time and mobile deployment requirements of industrial online detection. In addition, the traditional uniformly distributed camera array is prone to blind areas in the field of view when detecting curved surfaces or complex structures, and the insufficient resolution in the central region leads to an increase in the micron-level defect omission rate. SUMMARY
[0004] The main purpose of the present application is to provide a panoramic splicing industrial camera device and splicing method for complex scenes to solve the problems in the related art.
[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a panoramic splicing industrial camera device and splicing method for complex scenes are provided, which includes a multi-camera array containing at least two industrial cameras, and the viewing angle of each industrial camera forms an overlapping area to cover the target scene.
[0006] A synchronization control module is connected with the multi-camera array, used to generate a high-precision hardware trigger signal to trigger all industrial cameras in the multi-camera array to collect images at the same time with a synchronization error of less than 1 microsecond.
[0007] An image processing module connected with the multi-camera array, the image processing module comprising a FPGA processor, the FPGA processor being configured to:
[0008] receiving in real time a plurality of images synchronously captured by the multi-camera array;
[0009] performing in real time inside the device image pre-processing, image alignment based on feature point matching and image fusion algorithm to generate a panoramic image of the target scene;
[0010] a calibration module connected with the multi-camera array and / or the image processing module, configured to monitor at least one environmental condition in real time and dynamically adjust the imaging parameters of at least one industrial camera in the multi-camera array according to the monitored environmental condition to maintain image quality and stitching accuracy;
[0011] a communication module connected with the image processing module, configured to transmit the generated panoramic image to an external device.
[0012] Further, the synchronization control module is also configured to receive an external trigger signal input and trigger the multi-camera array synchronously based on the external trigger signal.
[0013] Further, the environmental condition comprises monitoring the change of light intensity and / or device vibration.
[0014] Further, the FPGA processor performs an image alignment algorithm based on feature point matching, which is a scale-invariant feature transform algorithm or an optimized variant thereof; the image fusion algorithm is a multi-band fusion algorithm, which is used to eliminate stitching marks.
[0015] Further, the industrial cameras in the multi-camera array are arranged in a matrix, the multi-camera array comprises a plurality of industrial cameras, the industrial cameras are installed on a support backboard of a support frame and arranged in a matrix or grid structure of at least two dimensions, the camera spacing in the center area of the matrix or grid structure is smaller than that in the edge area.
[0016] Further, the industrial cameras in the multi-camera array are arranged in a matrix, the multi-camera array comprises a plurality of industrial cameras, the industrial cameras are installed on a support backboard of a support frame and arranged in a matrix or grid structure of at least two dimensions, the camera spacing in the center area of the matrix or grid structure is smaller than that in the edge area.
[0017] A method for generating a panoramic image using the above device, comprising the following steps:
[0018] The synchronization control module is used to generate high-precision hardware trigger signals to trigger all the industrial cameras in the multi-camera array to capture images of their respective fields of view at the same time with a synchronization error of less than 1 microsecond;
[0019] The plurality of images captured synchronously are transmitted in real time to the image processing module;
[0020] Inside the device, the following operations are performed in real time by the FPGA processor in the image processing module:
[0021] Image preprocessing is performed on the received plurality of images;
[0022] Alignment of adjacent images is performed based on a feature point matching algorithm;
[0023] Image fusion algorithm is used to fuse the aligned images into a panoramic image;
[0024] During the image capturing and processing, the calibration module is used to monitor at least one environmental condition in real time, and the imaging parameters of at least one industrial camera in the multi-camera array are dynamically adjusted according to the monitored environmental condition;
[0025] The generated panoramic image is transmitted to an external device through the communication module.
[0026] Further, the step of synchronous triggering further includes receiving an external trigger signal and performing synchronous triggering based on the external trigger signal.
[0027] Further, the step of dynamically adjusting the imaging parameters includes monitoring the change of light intensity and / or device vibration in real time, and dynamically adjusting the exposure time and / or automatic focusing parameters of the industrial camera accordingly.
[0028] Further, the alignment step based on the feature point matching algorithm uses a scale-invariant feature transform algorithm or its optimized variants; the image fusion step uses a multi-band fusion algorithm to eliminate stitching marks.
[0029] Compared with the prior art, the present application has the following beneficial effects: the present application reduces the stitching misregistration error of high-speed motion scenes to below 0.01 mm by the sub-microsecond hardware synchronous trigger architecture (synchronization error < 0.8 mu s) and the optimized layout design of the ring / matrix camera array (horizontal overlap rate >= 30%, vertical overlap rate >= 15%), and eliminates 99.2% of ghosting artifacts; in combination with multi-modal sensing fusion (illumination / vibration / ranging) and deep reinforcement learning decision algorithm, the illumination mutation response time is <= 18 ms, the vibration compensation accuracy is <= 0.02 g RMS, and the stitching success rate is maintained at 98.7% in a strong interference environment; the system power consumption is reduced to 12 W while the defect detection sensitivity reaches 30 mu m level by using the FPGA hardening pipeline processing technology (feature matching delay <= 35 ms @ 4K) and the non-uniform matrix distribution (the resolution density in the central area is improved by 3.2 times); by using the adaptive noise Kalman filtering and the centripetal gaze optical layout, the field of view blind area is reduced by 82% in curved target detection, and the integrity and measurement accuracy of panoramic imaging in complex industrial scenes are ensured. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0031] The structures, proportions, sizes, etc. shown in the drawings of the present specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not define the limiting conditions for the implementation of the present application, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, which does not affect the effects and purposes that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application.
[0032] Figure 1 System block diagram of the preferred embodiment of the present application;
[0033] Figure 2 Calibration module schematic diagram of the preferred embodiment of the present application;
[0034] Figure 3 Multi-camera array arrangement schematic diagram of the preferred embodiment of the present application;
[0035] Figure 4 Local diagram of the multi-camera array of the preferred embodiment of the present application;
[0036] Figure 5A schematic diagram of a multi-camera array arrangement according to another preferred embodiment of the present application.
[0037] Reference signs: 1, multi-camera array; 2, image processing module; 3, synchronous control module; 4, calibration module; 401, multi-modal sensing unit; 402, data fusion processing unit; 403, decision unit; 404, dynamic execution unit; 5, communication module. DETAILED DESCRIPTION
[0038] In order to make the inventive purposes, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings of the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0039] In the description of the present application, it should be understood that the terms "upper", "lower", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. It should be noted that when a component is considered to be "connected" to another component, it can be directly connected to the other component or there can be a component disposed therebetween.
[0040] The technical solutions of the present application will be further described below in conjunction with the drawings and through specific embodiments.
[0041] A panoramic splicing industrial camera device for complex scenes, as shown in Figure 1 comprises:
[0042] A multi-camera array 1 comprising at least two industrial cameras, the viewing angle of each industrial camera forming an overlapping area to cover the target scene;
[0043] A synchronous control module 3 connected to the multi-camera array 1, for generating a high-precision hardware trigger signal to synchronously trigger all industrial cameras in the multi-camera array 1 to collect images at the same time with a synchronization error less than 1 microsecond;
[0044] The industrial cameras in the multi-camera array 1 are arranged in a ring shape, or the industrial cameras are arranged in a matrix.
[0045] When the industrial cameras are arranged in a ring shape:
[0046] At least two groups of industrial cameras, the industrial cameras being installed along the circumferential direction of the support frame;
[0047] wherein the optical centers of the first group of industrial cameras are located in the same plane, and their optical axes have a first preset tilt angle relative to the normal direction of the plane, the first preset tilt angle being other than zero, thereby defining a first vertical field of view bias;
[0048] the optical centers of the second group of industrial cameras are located in the same plane or another plane parallel to the same plane, and their optical axes have a second preset tilt angle relative to the normal direction of the plane, the second preset tilt angle being different from the first preset tilt angle and opposite in direction, thereby defining a second vertical field of view bias different from the first vertical field of view bias;
[0049] The arrangement of the industrial cameras ensures that the respective fields of view (FOV) of the cameras along the circumferential direction and between the first group of cameras and the second group of cameras in the vertical direction form overlapping regions, and the overlapping regions of the fields of view of adjacent cameras are not less than 20%, to support subsequent image stitching.
[0050] The first group of cameras and the second group of cameras are arranged in a staggered manner along the circumference of the support frame.
[0051] Specifically, the embodiment is used for the continuous and dead angle-free visual inspection of the inner wall of a large vertical cylindrical storage tank (10 meters in diameter and 15 meters in height) to find defects such as corrosion, cracks, and attachments, as shown in scene one (as shown in FIG. 1A) and scene two (as shown in FIG. 1B). Figure 3 、 Figure 4 The circular frame is made of high-strength aluminum alloy or carbon fiber material and is precisely machined to have a diameter of 1.5 meters. The frame is placed in the central region inside the storage tank through a central support rod.
[0052] Eight industrial cameras (N=8) are selected, and each camera is equipped with a sensor: Sony IMX264 (5 million pixels), and a lens: a focal length of 6 mm, providing a horizontal field of view angle (HFOV) of about 60° and a vertical field of view angle (VFOV) of 45°.
[0053] The eight cameras are evenly distributed along the circular frame, and one camera in the same group is installed every 45°. Cameras 1, 3, 5, and 7 have optical axes tilted upward by α=12° relative to the frame plane (horizontal plane). These cameras mainly cover the middle and upper regions of the inner wall of the storage tank and part of the top cover region. Cameras 2, 4, 6, and 8 have optical axes tilted downward by β=-12° relative to the frame plane (horizontal plane). These cameras mainly cover the middle and lower regions of the inner wall of the storage tank and part of the bottom plate region. According to the height-diameter ratio of the storage tank and the key detection regions, the tilt angles α and β can be adjusted in the range of 8° to 20°. A larger tilt angle can cover a wider vertical range, but may increase the difficulty of stitching and edge distortion.
[0054] Horizontal direction: 30%~35% FOV overlap between adjacent cameras (e.g. 1 and 2, 2 and 3). Calculation basis: overlap angle = HFOV - 360° / N = 60° - 45° = 15°. Overlap rate = (15° / 60°)*100% = 25%. To ensure robustness, the actual installation can slightly increase the overlap to more than 30%.
[0055] Vertical direction: The lower edge of the upwardly tilted group (1, 3, 5, 7) and the upper edge of the downwardly tilted group (2, 4, 6, 8) form a vertical FOV overlap zone of at least 10%~15% at a certain distance in front of the cameras (e.g. at the position of the tank wall), ensuring seamless stitching in the vertical direction. The overlap width is related to the tilt angle, camera VFOV and working distance.
[0056] On the outside of the circular frame, a set of multi-modal sensor units (TSL2561, AS7262, MPU6050) are installed at 0°, 120°, 240° positions to form a ring-shaped sensor network. The sensors are firmly installed to reduce their own vibration interference.
[0057] In another embodiment, the industrial cameras are arranged in a matrix:
[0058] The industrial cameras are installed on a support backboard and arranged in at least a two-dimensional matrix or grid structure;
[0059] The spatial distribution density of the industrial cameras on the support backboard is non-uniform, specifically, the camera spacing in the central region of the matrix or grid structure is smaller than that in the edge region;
[0060] The optical axes of the plurality of industrial cameras generally point to a common far field or target region, but are not strictly parallel; specifically, the optical axis of each or at least a portion of the industrial cameras has a small preset tilt angle inwardly pointing to the central normal direction, forming a centripetal pointing characteristic;
[0061] And the arrangement of the industrial cameras ensures that in the matrix or grid structure, the field of view (FOV) between each camera and its adjacent camera forms an overlap region to support subsequent image stitching.
[0062] Other embodiments are shown in scenario two (e.g. Figure 5 High-precision and high-efficiency surface defect detection (e.g. scratches, sand holes, machining residues) is performed on the top surface of the automobile engine cylinder block installed on the production line. The detection area is about 80cmx60cm, and the central key area (e.g. cylinder hole edge, sealing joint surface) needs to be checked in detail.
[0063] 9 industrial cameras (3x3 matrix) are chosen, each equipped with: Sensor: OnSemi Python12K (12.5 million pixels) for higher resolution. Lens: focal length 12mm for a horizontal field of view (HFOV) of about 30° and a vertical field of view (VFOV) of 22°. Longer focal length for higher detail.
[0064] A precision machined backplate of sufficient thickness (e.g. 15mm) is used, made of aluminum alloy or steel, with dimensions of about 50cm x 40cm, ensuring high rigidity and thermal stability. The backplate is mounted on the end of the mechanical arm of the detection station or on a fixed support.
[0065] Camera arrangement and tilt:
[0066] Center camera (C): located at the center of the backplate.
[0067] Peripheral cameras (M1-M4): surround the center camera at a horizontal / vertical distance d_center = 12cm from the center camera.
[0068] Corner cameras (E1-E4): located at the four corners of the matrix, at a horizontal / vertical distance d_edge = 18cm from the adjacent peripheral cameras. (i.e. denser center region cameras).
[0069] Concentric gaze setup (example with a detection distance of 1 meter):
[0070] Center camera (C): optical axis perpendicular to the backplate, γ = 0°.
[0071] Peripheral cameras (M1-M4): optical axis tilted inward, pointing towards the center of the array's field of view, with a tilt angle γ = 1.5°.
[0072] Corner cameras (E1-E4): optical axis tilted inward, pointing towards the center of the array's field of view, with a tilt angle γ = 2.5°.
[0073] Tilt angle range consideration: The tilt angle γ typically ranges between 0.5° and 5°. The specific value depends on the working distance, target depth / curvature, and the desired degree of edge perspective improvement. The closer the working distance and the greater the target curvature, the greater the tilt angle that may be required.
[0074] Overlap ratio:
[0075] Center region (between C and M1-M4, and between M and adjacent M): ensures a high overlap ratio of 40%~50%, providing data for high-precision stitching and possible super-resolution reconstruction.
[0076] Edge region (between M and E, and between E and adjacent E): ensures a standard overlap ratio of 25%~35%, ensuring complete coverage and stable stitching.
[0077] Sensor deployment (calibration module 401): 3-4 groups of multi-modal sensor units are installed near the four corner points of the backboard (outside the E1-E4 area) and the middle points of the upper and lower long edges, forming a sensing network around the array.
[0078] An image processing module 2 is connected to the multi-camera array 1, and the image processing module 2 includes an FPGA processor, which is used for:
[0079] receiving multiple images synchronously collected by the multi-camera array 1 in real time;
[0080] performing image preprocessing, image alignment based on feature point matching, and image fusion algorithm in real time inside the device to generate a panoramic image of the target scene;
[0081] Specifically, a separate physical receiving interface is configured on the FPGA for each industrial camera in the multi-camera array 1. According to the specific output of the industrial camera (such as MIPI CSI-2, LVDS, Camera Link, GigE Vision, etc.), the corresponding physical layer (PHY) and link layer / protocol layer controller is implemented.
[0082] The specific implementation steps are as follows:
[0083] Input: Raw image data stream from each camera input buffer.
[0084] Processing unit (in sequential pipeline):
[0085] a. Bad pixel correction: using the bad pixel map pre-stored in the FPGA internal memory (such as BRAM), the detected bad pixels (dead pixels, bright pixels) are replaced by interpolation.
[0086] b. Black level correction: subtract the black level reference value of the corresponding channel from the pixel value. This reference value can be fixed or updated to the FPGA register after dynamic fine-tuning by the calibration module 4.
[0087] c. Lens shading correction: using the LSC lookup table (LUT) or gain map stored in the FPGA BRAM, each pixel is compensated for radial and tangential brightness attenuation. This LUT should be generated during camera calibration. The dynamic adjustment of the calibration module 4 mainly affects the camera itself parameters, and the LSC remains the calibration value.
[0088] d. Demosaicing: if the input is in Bayer format, perform demosaicing algorithm to convert monochrome sensor data to full-color (such as RGB) image.
[0089] e. White balance gain: Although the calibration module 4 adjusts the camera's own white balance through the dynamic execution unit 404, the FPGA level can reserve a digital gain interface. The FPGA receives the digital gain calculated by the calibration module 4 at this stage, which needs to be fine-tuned after processing, mainly relying on the camera's own adjustment, and the FPGA applies a fixed or basic gain.
[0090] f. Color correction matrix: Apply a 3x3 color correction matrix to convert the image color space to a standard color space (such as sRGB), or perform color optimization according to specific industrial scene requirements. CCM parameters come from calibration.
[0091] g. Gamma correction: Apply a Gamma curve (implemented through a LUT) to adjust the image's non-linear brightness response to meet human eye perception or display device requirements.
[0092] h. Lens distortion correction: This is a crucial step before stitching. Use the distortion model parameters (such as radial and tangential distortion coefficients) stored in the FPGA memory. The FPGA calculates the sub-pixel coordinates on the input image corresponding to each output pixel, and obtains the pixel value through bilinear or bicubic interpolation to generate a distortion-free image.
[0093] Image alignment based on feature point matching:
[0094] Input: Multiple pre-processed images stored in the frame buffer.
[0095] Parallel processing (for adjacent overlapping image pairs):
[0096] a. Feature point detection: In the pre-defined overlapping area of each image, perform an efficient feature point detection algorithm in parallel. Considering FPGA resources and real-time performance, you can choose FAST, Harris, or an optimized version of the ORB (Oriented FAST and Rotated BRIEF) corner detector. FPGA implementation involves parallel operations such as window sliding, pixel comparison, and threshold judgment.
[0097] b. Feature point descriptor generation: Calculate the descriptor for each detected feature point. For FPGA, binary descriptors are particularly suitable due to their high computational and matching efficiency. FPGA achieves this by reading pixel blocks in parallel, performing comparisons, and generating bit strings.
[0098] c. Feature point matching: Match the feature point descriptor sets between the overlapping areas of adjacent images. FPGA uses a parallel brute-force matcher based on Hamming distance (very efficient for binary descriptors) or an approximate nearest neighbor search method.
[0099] d. RANSAC: Implement the RANSAC algorithm to reject mismatched pairs and estimate the geometric transformation (homography matrix H) between images. FPGA can accelerate the RANSAC process by parallelizing hypothesis generation, inlier counting, and model parameter computation (e.g., SVD decomposition or direct linear transformation optimization implementation).
[0100] Output: The accurate homography transformation matrix H computed for each pair of overlapping images. These matrices will be used in the next step of image fusion.
[0101] Output: Generate a multi-lane pre-processed, distortion-free, color-corrected image data stream. These data are stored in the FPGA-internal frame buffer (accessed through the FPGA's memory controller using external high-speed DRAM, such as DDR4 / LPDDR4) for subsequent stitching steps.
[0102] Image fusion to generate a panoramic image:
[0103] Input: Multi-lane pre-processed image data (stored in frame buffer) and computed homography matrices H.
[0104] a. Panorama canvas definition and coordinate mapping: Determine the coordinate system and dimensions of the final panoramic image based on the camera array's layout and calibration information. FPGA needs to compute the coordinates of each pixel on the panorama canvas corresponding to the source images (after H matrix transformation).
[0105] b. Image re-projection: Use the computed homography matrices H to "project" each source image onto the panorama canvas. FPGA implements real-time re-projection through efficient address generation logic and interpolation units (e.g., bilinear interpolation). Efficient access to source image data stored in external DRAM is required.
[0106] c. Exposure / gain compensation: Although the calibration module 4 strives to unify camera exposures, residual differences may exist. FPGA can analyze brightness differences in overlapping regions and compute pixel-level gain compensation factors to make transitions smoother.
[0107] d. Image stitching: In the overlapping regions, pixels from different source images are fused.
[0108] Output: The final panoramic image frame stored in the external DRAM controlled by the FPGA.
[0109] The calibration module 4 is connected with the multi-camera array 1 and / or the image processing module 2, for monitoring at least one environmental condition in real time, and dynamically adjusting the imaging parameters of at least one industrial camera in the multi-camera array 1 according to the monitored environmental condition, so as to maintain the image quality and stitching accuracy; although the synchronous control module 3 ensures the synchronization trigger of <1 µs, the data transmission path may introduce slight delay differences. The FPGA internal logic should be able to receive the accurate trigger timestamp (or synchronization frame start signal) from the synchronous control module 3, and align the input data streams based on this. And set up independent input FIFO (First In First Out) buffer for each camera data stream. This can absorb the slight jitter in data transmission, and decouple the speed difference of subsequent processing modules, to ensure data integrity. The buffer size needs to be determined according to the camera resolution, frame rate and the throughput capacity of the FPGA processing pipeline. The received raw data (e.g. RAW Bayer format) needs to be preliminarily arranged. The FPGA internal logic converts it into a unified internal processing format (e.g. pixel stream or specific block / Tile format), for pipeline processing.
[0110] Specifically, as shown in Figure 2 The calibration module 4 includes a multi-modal sensing unit 401 for real-time acquisition of light intensity, color temperature, and vibration data.
[0111] Among them, the light intensity: TSL2561 digital light sensor (range 0.1-40,000 lux, sampling rate 100Hz); color temperature: AMS AS7262 multispectral sensor (6-channel visible light detection); vibration: MPU6050 six-axis sensor (±16g acceleration, ±2000° / s angular velocity); deployment strategy: form a ring-shaped sensor network around the camera array, arrange a group of sensing nodes every 120 degrees.
[0112] And all sensor data are labeled by the global clock signal of the FPGA (accuracy ±10ns), a three-dimensional coordinate system conversion model is established, and each sensor data is mapped to the camera optical center coordinate system.
[0113] The data fusion processing unit 402 realizes multi-source information fusion based on an improved Kalman filter.
[0114] First, the state space modeling is performed:
[0115]
[0116] Among them, is the ambient light intensity at time k (unit: lux); is the ambient color temperature at time k (unit: Kelvin); is the three-axis acceleration at time k (unit: m / s²); Let k be the triaxial angular velocity at time k (unit: rad / s).
[0117] State transition equation:
[0118] ;
[0119] in, This is the state transition matrix (a diagonal matrix, with elements being decay factors). To control the input matrix; For external control input; This is process noise.
[0120] Establish the sensor observation equation: ;
[0121] in, The sensor observation vector; The observation matrix; To observe noise.
[0122] The initial value of the observed noise covariance is:
[0123] ;
[0124] , Light sensor noise variance (lux²);
[0125] , color temperature sensor noise variance (K²);
[0126] accelerometer noise variance ((m / s²)²);
[0127] , Gyroscope noise variance ((rad / s)²).
[0128] Furthermore, it employs adaptive noise covariance for updating;
[0129] Process noise adaptation:
[0130]
[0131] in, This indicates the degree of historical dependence of process noise;
[0132] ; Let be the state covariance matrix of the previous time step.
[0133] Observation noise adaptation:
[0134]
[0135] represents the temperature compensation coefficient; is the sensor temperature (℃).
[0136] Correlation between illumination and color temperature:
[0137]
[0138] wherein, is the correlation attenuation constant, (spectral distribution difference, is the 6-channel spectral data of AS7262); represents the spectral difference tolerance.
[0139] The correlation weight matrix is ;
[0140] Finally, enter the main loop of the improved Kalman filter;
[0141] Prediction:
[0142] ;
[0143]
[0144] Kalman gain:
[0145]
[0146] represents the Hadamard product;
[0147] State update:
[0148] ;
[0149]
[0150] Outlier suppression:
[0151] .
[0152] The data fusion processing unit 402 can make the absolute error of the illumination measurement ≤3%, the vibration parameter fusion error ≤0.02g (RMS) by dynamically adjusting the noise parameter and the data correlation weight through the above method, meeting the industrial level precision requirement.
[0153] The decision unit 403 generates a camera parameter adjustment strategy; the specific steps are as follows:
[0154] Taking the environmental state fused by the data fusion processing unit 402 as input, calculate the real-time change characteristics of illumination, vibration, and color temperature, and generate the parameter adjustment amount according to the preset rule:
[0155] Rate of change of illumination intensity: (unit: lux / s);
[0156] Acceleration vibration energy: (unit: m / s2);
[0157] Color temperature deviation from reference value: , ;
[0158] Exposure time adjustment rule:
[0159]
[0160] represents the illumination sensitivity coefficient, in the present embodiment ; represents the maximum range, , is the exposure time (ms) of the current camera, is the new exposure time after adjustment.
[0161] Constraint: exposure time change does not exceed ±30%.
[0162] Next, focus compensation calculation is performed:
[0163]
[0164] Fourier transform is performed on the speed data to calculate the power spectral density , and then frequency band integration is performed: low frequency band (5-200 Hz) compensation amplitude is large, and high frequency band (201-500 Hz) compensation amplitude is small;
[0165] The white balance adjustment strategy is:
[0166]
[0167] Constraint: red / blue channel gain ratio is limited to 0.7-2.0.
[0168] Each part of the adjustment strategy is controlled through dynamic weight distribution:
[0169]
[0170] According to the type of environmental change, the priority weight of parameter adjustment is dynamically allocated. When the illumination changes dramatically (ΔL k >1000): priority is given to exposure adjustment (weight 60%), followed by focus (30%), and finally white balance (10%); when strong vibration occurs (E_vib>2g): priority is given to focus compensation (weight 50%).
[0171] The dynamic execution unit 404 realizes the closed-loop control of exposure, focus, and white balance parameters. The parameter adjustment amount generated by the decision unit 403 is converted into an actual camera control signal, and stable adjustment is realized.
[0172] Exposure time control: PID controller is adopted:
[0173]
[0174] wherein, (Illumination error), according to the IMX sensor optimization coefficient: , , When the environment mutates, the Bang-Bang control mode is automatically switched to accelerate convergence.
[0175] The focus position compensation adopts a step adjustment algorithm:
[0176]
[0177] The maximum adjustment amount of each frame is 50 steps to avoid losing steps. The vibration compensation direction is opposite to the acceleration direction.
[0178] White balance closed-loop calibration:
[0179] 1. Obtain the current image white point coordinates (u, v);
[0180] 2. Calculate the color difference: ;
[0181] 3. Iterative adjustment:
[0182]
[0183] The control timing is as follows:
[0184] 0-2 ms: read fusion data ; 2-5 ms: execute strategy generation calculation; 5-10 ms: send control instructions to the camera; 10-16 ms: wait for execution to be completed and collect feedback images.
[0185] The complete closed loop from environment perception to parameter adjustment can be completed within 50 ms, ensuring that the sub-pixel level splicing accuracy is maintained under severe change conditions (such as illumination mutation 1000 lux / s, vibration 5g).
[0186] The communication module 5 is connected with the image processing module 2, and is configured to transmit the generated panoramic image to an external device.
[0187] The synchronization control module 3 is also used for receiving an external trigger signal input and synchronously triggering the multi-camera array 1 based on the external trigger signal.
[0188] The environmental conditions include monitoring the light intensity variation and / or the device vibration.
[0189] The feature point matching based image alignment algorithm executed by the FPGA processor is a scale-invariant feature transform algorithm or its optimized variant; the image fusion algorithm is a multi-band fusion algorithm for eliminating the stitching traces.
[0190] It should be noted that the optimized variant of the scale-invariant feature transform algorithm in the embodiment can use integral image and Hessian matrix approximation to accelerate feature detection, maintain scale invariance but significantly reduce the computational complexity;
[0191] A method for generating a panoramic image, comprising the following steps:
[0192] A high-precision hardware trigger signal is generated by using the synchronization control module 3 to trigger all the industrial cameras in the multi-camera array 1 to simultaneously collect images of their respective fields of view with a synchronization error of less than 1 microsecond;
[0193] The synchronously collected multiple images are transmitted to the image processing module 2 in real time;
[0194] Inside the device, the following operations are performed by the FPGA processor in the image processing module 2 in real time:
[0195] Image preprocessing is performed on the received multiple images;
[0196] The adjacent images are aligned based on a feature point matching algorithm;
[0197] The aligned images are fused into a panoramic image using an image fusion algorithm;
[0198] During the image collection and processing, at least one environmental condition is monitored in real time using the calibration module 4, and the imaging parameters of at least one industrial camera in the multi-camera array are dynamically adjusted according to the monitored environmental condition;
[0199] The generated panoramic image is transmitted to an external device through the communication module 5.
[0200] The step of synchronously triggering further includes receiving an external trigger signal and performing synchronous triggering based on the external trigger signal.
[0201] The step of dynamically adjusting the imaging parameters includes monitoring the light intensity variation and / or the device vibration in real time, and dynamically adjusting the exposure time and / or the automatic focusing parameters of the industrial camera accordingly.
[0202] The alignment step based on the feature point matching algorithm uses a scale-invariant feature transform algorithm or its optimized variant; the image fusion step uses a multi-band fusion algorithm to eliminate the stitching traces.
[0203] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, as the scope of the exemplary embodiments of this application is limited only by the appended claims. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0204] It is to be understood that the terms "first", "second", and the like, used herein do not necessarily connote any order, quantity, composition, or importance, but are used to distinguish one element from another, and do not necessarily indicate a requirement or order. It is to be understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0205] The preferred embodiments of the application are described above in detail. The application is not limited to the embodiments described above, but can vary and be modified within the scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall fall within the scope of the application.
Claims
1. A panoramic stitching industrial camera device for complex scenes, characterized in that, include: A multi-camera array (1) comprising at least two industrial cameras, wherein the viewpoints of each of the industrial cameras form an overlapping area to cover the target scene; A synchronization control module (3), connected to the multi-camera array (1), is used to generate a high-precision hardware trigger signal to synchronously trigger all industrial cameras in the multi-camera array (1) to acquire images simultaneously with a synchronization error of less than 1 microsecond; an image processing module (2), connected to the multi-camera array (1), includes an FPGA processor for receiving and processing multiple images synchronously acquired by the multi-camera array (1) in real time to generate a panoramic image of the target scene; A communication module (5) is connected to the image processing module (2) and configured to transmit the generated panoramic image to an external device; The system also includes a calibration system for maintaining image quality and stitching accuracy, comprising: a) a multimodal sensing unit (401) for real-time acquisition of various environmental data, including at least light intensity and device vibration; b) a data fusion processing unit (402) connected to the multimodal sensing unit (401), wherein the data fusion processing unit (402) is configured to: use an improved Kalman filter algorithm to fuse the various environmental data to generate a high-precision fused environmental state estimate; c) a decision unit (403) connected to the data fusion processing unit (402), wherein the decision unit (403) is configured to: generate an adjustment strategy for the imaging parameters of at least one industrial camera in the multi-camera array (1) based on the changing characteristics of the fused environmental state estimate and a preset dynamic priority weight allocation rule; and d) a dynamic execution unit (404) for converting the adjustment strategy into camera control signals and dynamically adjusting the imaging parameters of the at least one industrial camera through a closed-loop control method.
2. The apparatus according to claim 1, characterized in that, The synchronization control module (3) is also used to receive external trigger signal input and synchronously trigger the multi-camera array (1) based on the external trigger signal.
3. The apparatus according to claim 1, characterized in that, The image alignment algorithm based on feature point matching executed by the FPGA processor is a scale-invariant feature transformation algorithm or an optimized variant thereof; the image fusion algorithm is a multi-band fusion algorithm used to eliminate stitching marks.
4. The apparatus according to claim 1, characterized in that, The industrial cameras in the multi-camera array (1) are arranged in a ring. The multi-camera array (1) includes at least two sets of industrial cameras. The industrial cameras are installed in a circumferential distribution along the support frame. The optical centers of the first set of industrial cameras are located in the same plane, and their optical axes have a first preset tilt angle relative to the normal direction of the plane. The first preset tilt angle is not zero degrees. The optical centers of the second set of industrial cameras are located in the same plane or in another plane parallel to it, and their optical axes have a second preset tilt angle relative to the normal direction of the plane. The second preset tilt angle is different from the first preset tilt angle and is in the opposite direction.
5. The apparatus according to claim 1, characterized in that, The industrial cameras in the multi-camera array (1) are distributed in a matrix. The multi-camera array (1) includes a plurality of industrial cameras. The industrial cameras are mounted on the support back plate of the support frame and arranged in a matrix or grid structure of at least two dimensions. The camera spacing in the central region of the matrix or grid structure is smaller than the camera spacing in the edge region.
6. A method for generating panoramic images, characterized in that, Includes the following steps: The synchronization control module (3) generates a high-precision hardware trigger signal to synchronously trigger all industrial cameras in the multi-camera array (1) to simultaneously acquire images of their respective fields of view with a synchronization error of less than 1 microsecond; the synchronously acquired multiple images are transmitted in real time to the image processing module (2), which generates a panoramic image; the generated panoramic image is transmitted to an external device through the communication module (5); And during the image acquisition and processing, a dynamic calibration closed-loop method is executed to maintain image quality and stitching accuracy, the method comprising: a) using a multimodal sensing unit (401) to acquire in real time various environmental data including at least light intensity and device vibration; b) An improved Kalman filter algorithm is used to fuse the various environmental data to generate a high-precision fused environmental state estimate; c) Based on the changing characteristics of the fused environmental state estimate and a preset dynamic priority weight allocation rule, an adjustment strategy for the imaging parameters of at least one industrial camera in the multi-camera array (1) is generated; and d) The adjustment strategy is converted into a camera control signal, and the imaging parameters of the at least one industrial camera are dynamically adjusted through a closed-loop control method.
7. The method according to claim 6, characterized in that, The synchronization triggering step further includes: receiving an external triggering signal and performing synchronization triggering based on the external triggering signal.
8. The method according to claim 6, characterized in that, The step of dynamically adjusting imaging parameters includes: dynamically adjusting the exposure time and / or autofocus parameters of the industrial camera according to the adjustment strategy.
9. The method according to claim 6, characterized in that, The steps of generating a panoramic image by the image processing module (2) include: aligning adjacent images based on a feature point matching algorithm, and fusing the aligned images into a panoramic image using an image fusion algorithm; wherein the alignment step based on the feature point matching algorithm uses a scale-invariant feature transformation algorithm or an optimized variant thereof; and the image fusion step uses a multi-band fusion algorithm to eliminate stitching marks.
Citation Information
Patent Citations
Device and method for adjusting view field of spliced panoramic camera
CN101833231A
Panoramic data acquisition device based on industrial camera and industrial lens
CN109379524A
Industrial visual inspection system based on multi-modal fusion
CN119444686A
Increasing spatial resolution of panoramic video captured by a camera array
US20170019594A1