Multi-scale light source fusion OCR enhancement method
Through the OCR enhancement method of multi-scale light source fusion, U-Net neural network is used to fuse image features under multi-light conditions to solve the problem of recognition of a single light source in complex environments, achieving higher OCR recognition accuracy and clearer image output.
Patent Information
- Application Number
- CN202510073958.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
In complex shooting environments, it is difficult for a single light source to effectively recognize text, especially under extreme lighting conditions, traditional image preprocessing methods have limited effects, resulting in a decrease in recognition accuracy.
The OCR enhancement method of multi-scale light source fusion is adopted to take pictures through lighting conditions of multiple light intensity and colors, and preprocess the image and input it into the U-Net neural network to fuse the image features under multi-light conditions, and finally output a normalized fusion image.
It reduces the impact of uneven light, shadows, reflections, etc., generates clearer and more informative images, improves OCR recognition accuracy, and avoids the need for on-site light source debugging.
Smart Images

Figure CN119991464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of light source fusion technology, and in particular to an OCR enhancement method of multi-scale light source fusion based on AI technology. Background Art
[0002] In actual shooting environments, the lighting environment is quite complex. It is obviously very difficult to recognize text in a changing environment on site with a single light source of intensity and color. In addition, traditional image preprocessing methods have limited effects under extreme lighting conditions, and the accuracy of unclear image recognition results will drop significantly. Even if neural network recognition methods are used, large errors will easily occur under different lighting conditions, and new training images have to be added, which further increases costs.
[0003] In practice, people mostly measure by selecting a single color light source. However, if the instrument to be inspected is a dashboard with built-in backlight, the background color may be green, gray, black, etc. At this time, if you still select a single intensity of light or white light, the text on the dial will not be clear, affecting the OCR recognition effect. In addition, when identifying certain large equipment, it is often necessary to conduct on-site testing. The complex environment on site, which is different from that in the laboratory, will make the effect of a single light source limited. Due to the timeliness requirements of on-site work, it is also difficult to debug and replace the on-site light source to adapt to actual use. Summary of the invention
[0004] In order to solve the above technical problems, an object of the present invention is to provide an OCR enhancement method of multi-scale light source fusion.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A multi-scale light source fusion OCR enhancement method, comprising:
[0007] A takes pictures under various light intensities and colors;
[0008] B preprocesses the image and inputs it into the U-Net neural network;
[0009] C fuses image features under multiple lighting conditions;
[0010] D outputs the fused image and normalizes it.
[0011] Compared with the prior art, one or more embodiments of the present invention may have the following advantages:
[0012] The final image used for recognition can minimize the effects of uneven lighting, shadows, reflections, etc.; multi-scale image features are fused through the jump connection in the U-Net network to obtain a clearer and more informative output image. By fusing images under different lighting conditions, the need to adjust the light source to obtain the best shooting results during on-site testing is eliminated. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flow chart of the OCR enhancement method of multi-scale light source fusion. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with embodiments and drawings.
[0015] like Figure 1 As shown, the OCR enhancement method flow of multi-scale light source fusion includes:
[0016] Take pictures under various light intensities and colors;
[0017] Preprocess the image and input it into the U-Net neural network;
[0018] Fusion of image features under multiple lighting conditions;
[0019] Output the fused image and normalize it.
[0020] The light intensity is controlled by the PWM signal, and the relationship between the light intensity I and the proportion D is:
[0021] I=I max ·D (1)
[0022] Among them I max is the maximum light intensity, D∈[0,1].
[0023] Spectral control: Different wavelengths of light can highlight specific features in an image. For example, black text on a red background may have higher contrast under blue light, and can effectively reduce the highlight reflections produced by certain material surfaces (such as plastics and metals).
[0024] Through multi-channel LED, each channel emits light of different wavelengths to form different colors; let the output intensity of each channel be w i , then the final output spectrum S(λ) is:
[0025]
[0026] S i (λ) is the spectral distribution of channel i, w iis the intensity of channel i; its specific intensity can be adjusted by the relationship between light intensity I and proportion D.
[0027] In order to ensure the accuracy of subsequent image fusion, the same light with different light intensity levels and spectral colors are used for shooting.
[0028] Light condition settings:
[0029] Choose from 5 light intensity levels to cover common scenarios:
[0030] · Low light (weak light environment): 50lux
[0031] · Normal indoor light: 200lux
[0032] · Bright light (bright scene): 500lux
[0033] · Outdoor direct sunlight: 1000lux
[0034] · Extremely strong light (high contrast test): 2000 lux
[0035] By mixing the RGB spectrum, the following 7 typical colors of light can be generated:
[0036] 1. Red light (620-750nm)
[0037] 2. Green light (495-570nm)
[0038] 3. Blue light (450-495nm)
[0039] 4. Yellow light (red + green mixed)
[0040] 5. Cyan light (green + blue mixture)
[0041] 6. Magenta light (red + blue mixture)
[0042] 7. White light (balanced mixture of red + green + blue)
[0043] Then the total number of photos taken is 5*7=35. This is just an example. In actual applications, not too many colors may be needed for testing. Users can adjust the number of photos taken in the combination according to their own preferences to increase the detection speed.
[0044] Since there are many photos taken, full-automatic light source adjustment and automatic shooting are adopted to ensure efficiency. Therefore, in order to ensure the synchronization of light source and camera during shooting, the following time synchronization formula is given:
[0045] t capture =ttrigger +Δt (1)
[0046] where t captue is the image acquisition time, t trigger Adjust the signal trigger time for the light source, Δt is the delay time
[0047] The influencing factors of the specific delay time Δt are shown in the following formula:
[0048] Δt=T light +T exposure +T transfer (2)
[0049] T light is the light source response time, T exposure is the exposure time, T transfer The specific time selection can be adjusted according to the user's own camera and light source device usage.
[0050] The U-Net neural network is used for image fusion. Through its unique encoder-decoder structure and jump connection mechanism, images under multiple lighting conditions are fused to generate a visually clearer and more informative image, thereby improving OCR recognition accuracy.
[0051] train
[0052] The U-Net encoder-decoder symmetric structure is used for feature extraction and image reconstruction, where the encoder downsamples to extract multi-level features of the input image; the feature extraction formula is:
[0053] F i =σ(w i *F i_1 +b i ) (5)
[0054] w i is the convolution kernel weight, F i-1 is the feature map of the previous layer, b i is the bias and σ is the activation function.
[0055] The decoder upsamples to restore the spatial resolution of the fused feature image;
[0056] Upsampling to restore spatial resolution formula:
[0057]
[0058] Upsampling is a method of enlarging a reduced-resolution image to a higher resolution, where Upsample refers to the corresponding up-sampling method, such as nearest neighbor interpolation or bilinear interpolation.
[0059] The multi-scale image features are fused through the jump connection in the U-Net network, and the low-level features are fused:
[0060]
[0061] where F j It is the feature map of the corresponding layer of the encoder. Concat() refers to the operation of concatenating or merging multiple tensors or matrices along a specific dimension.
[0062] Output the fused image through the output layer:
[0063] Y pred = =σ(w our *F final +b out ) (8)
[0064] Construct an image loss function consisting of pixel-level mean square error (MSE) and structural similarity (SSIM) loss;
[0065]
[0066] L SSIM =1-SSIM(r,Y pred ) (10)
[0067] Where SSIM expansion:
[0068]
[0069] where μ x, μ y The mean of x and y entered respectively, is the variance of x and y, σ xy is the covariance of x and y, C1 and C2 are constants introduced to avoid the denominator being zero;
[0070] The total loss is:
[0071] L total =αL MSE +βL SSIM (12)
[0072] Among them, α and β are weight hyperparameters, which are adjusted during training.
[0073] According to the total loss function, the model parameter optimization objective function is obtained:
[0074] θ * =argminL total (X, r; θ) (13)
[0075] θ is the model parameter. After training is completed, the optimal θ is output with the help of the loss function and used for model fusion reasoning.
[0076] The image fusion reasoning includes:
[0077] The input multi-channel image collection contains multiple lighting conditions or color channels, represented as:
[0078] X test = {X1, X2, ..., X N} (14)
[0079] The generation formula of the fused image is:
[0080] Y fused =f(X test θ * ) (15)
[0081] f(·): represents the previously trained U-Net model.
[0082] In order to adjust the pixel value range of the output image to meet the needs of human visual perception or subsequent processing, Y fused Perform normalization:
[0083]
[0084] max(Y fused ) and min(Y fused ) are the maximum and minimum pixel values of the fused image, respectively.
[0085] The image fusion method provided in the above embodiment uses multiple images taken under different lighting conditions as input to obtain richer image information and improve the accuracy of image recognition and reconstruction. The lighting adjustment and color change of the input image are based on a specific experimental design to ensure that the input image covers all possible lighting change conditions.
[0086] Although the embodiments disclosed in the present invention are as above, the above contents are only embodiments adopted for facilitating the understanding of the present invention and are not intended to limit the present invention. Any technician in the technical field to which the present invention belongs can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present invention, but the patent protection scope of the present invention shall still be subject to the scope defined in the attached claims.
Claims
1. A multi-scale light source fusion OCR enhancement method, characterized in that: The following steps are involved: A takes pictures under various light intensities and colors; B preprocesses the image and inputs it into the U-Net neural network; C fuses image features under multiple lighting conditions; D outputs the fused image and normalizes it.
2. The OCR enhancement method of multi-scale light source fusion according to claim 1, characterized in that: The light intensity in A is controlled by a PWM signal, and the relationship between the light intensity I and the proportion D is: I=I max ·D Among them I max is the maximum light intensity, D∈[0,1].
3. The OCR enhancement method of multi-scale light source fusion according to claim 1, characterized in that: In B, through multi-channel LED, each channel emits light of different wavelengths to form different colors; let the output intensity of each channel be w i , then the final output spectrum S(λ) is: S i (λ) is the spectral distribution of channel i, w i is the intensity of channel i; its specific intensity can be adjusted by the relationship between light intensity I and proportion D.
4. The OCR enhancement method of multi-scale light source fusion according to claim 1, characterized in that: In C, a U-Net neural network is used for image fusion, and a U-Net encoder-decoder symmetric structure is used for feature extraction and image reconstruction, wherein the encoder downsamples to extract multi-level features of the input image; the feature extraction formula is: F i =σ(w i *F i-1 +b i ) w i is the convolution kernel weight, F i-1 is the feature map of the previous layer, b i is the bias and σ is the activation function.
5. The OCR enhancement method of multi-scale light source fusion according to claim 4, characterized in that: The decoder upsamples to restore the spatial resolution of the fused feature image; Upsampling restores spatial resolution formula: Among them, Upsample refers to the corresponding upsample method, namely, nearest neighbor interpolation and bilinear interpolation.
6. The OCR enhancement method of multi-scale light source fusion according to claim 5, characterized in that: The multi-scale image features are fused through the jump connection in the U-Net network, and the low-level features are fused: where F j It is the feature map of the corresponding layer of the encoder. Concat() refers to the operation of concatenating or merging multiple tensors or matrices along a specific dimension.
7. The OCR enhancement method of multi-scale light source fusion according to claim 6, characterized in that: Output the fused image through the output layer: Y pred ==σ(w out *F final +b out ) Construct an image loss function consisting of pixel-level mean square error (MSE) and structural similarity (SSIM) loss; Where SSIM expansion: where μ x , μ y The mean of x and y entered respectively, is the variance of x and y, σ xy is the covariance of x and y, C1 and C2 are constants introduced to avoid the denominator being zero; The total loss is: L total =αL MSE +βL SSIM Among them, α and β are weight hyperparameters, which are adjusted during training.
8. The multi-scale light source fusion OCR enhancement method according to claim 7, characterized in that: According to the total loss function, the model parameter optimization objective function is obtained: θ is the model parameter. After training, the optimal θ is output with the help of the loss function. * It is then used for model fusion reasoning.
9. The OCR enhancement method of multi-scale light source fusion according to claim 8, characterized in that: The image fusion reasoning includes: The input multi-channel image collection contains multiple lighting conditions or color channels, represented as: X test ={X1,X2,...,X N } The generation formula of the fused image is: Y fused =f(X test ;θ * ) f(·): represents the previously trained U-Net model.
10. The OCR enhancement method of multi-scale light source fusion according to claim 9, characterized in that: To adjust the pixel value range of the output image, fused Perform normalization: max(Y fused ) and min(Y fused ) are the maximum and minimum pixel values of the fused image, respectively.