Improved MSRCR and Random Forest Based Quantitative Identification Method for Test Strip Color Change
Through the improved MSRCR algorithm combined with random forests, the test strip color discoloration is automatically recognized, which solves the problems of inaccurate and low efficiency of naked eyes in the prior art, and realizes efficient and accurate quantitative analysis of test strip color discoloration.
Patent Information
- Application Number
- CN202411788918.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-06
AI Technical Summary
In the existing cable thermal stability test, the detection results of the naked eye identification test strips are inaccurate, the detection efficiency is low, and the data utilization is insufficient.
The improved multi-scale adaptive gain MSRCR algorithm combined with random forests is used to quantify the color discoloration of the test strip by collecting test strip photos through the camera, image enhancement and normalization processing is performed, and the test strip area is located and segmented, features are extracted and the color discoloration degree is tested using random forests.
The automation of test strip color change recognition is achieved, the inspection efficiency is improved, subjectivity is reduced, and the data is fully utilized, and a new and efficient quantitative analysis method for test strip color change is provided.
Smart Images

Figure CN119274001B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for quantitatively identifying the color change of test paper based on improved MSRCR and random forest, belonging to the technical field of cable thermal stability performance detection. Background Art
[0002] The chemical stability of cable raw materials is of crucial significance to the performance, reliability, and safety of cables. High chemical stability means that cable raw materials can resist the erosion of chemical substances in the environment, such as acids, alkalis, salts, etc., thereby ensuring stable performance of cables during long-term use and being not easily aged or damaged. Stable chemical properties help maintain the electrical properties of cables, such as resistivity, insulation strength, etc., ensuring the reliability and efficiency of signal transmission; Chemically unstable raw materials may release harmful substances under specific conditions, posing a threat to the environment and human health. Raw materials with high chemical stability can reduce this risk and ensure the safety of cables during use; In emergency situations such as fires, chemically stable cable materials can reduce the release of toxic smoke.
[0003] The prior art usually uses a thermal stability test to evaluate the chemical stability of cable raw materials. This test focuses on monitoring the chemical changes of cable materials in an oil bath environment, especially paying attention to the release rate of acidic substances. Currently, the cable thermal stability test mainly has the following disadvantages:
[0004] 1. The frequent observation process of inspectors limits their ability to conduct other tests simultaneously, significantly reducing the overall inspection efficiency.
[0005] 2. When judging whether the test paper has changed color, inspectors mainly rely on visual observation. This method has obvious subjectivity and it is difficult to ensure the consistency of judgment criteria among different inspectors, which may affect the reliability of test results.
[0006] 3. The test data only includes whether the color has changed and the color change time. Although the test can be completed according to the standard, the data is not fully utilized, and the test itself and the test materials are not further analyzed and evaluated. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for quantitatively identifying the color change of test paper based on improved MSRCR and random forest, so as to solve the technical problems of inaccurate detection results and low detection efficiency in the naked-eye identification of the color change of test paper in the cable thermal stability test of the prior art.
[0008] The present invention adopts the following technical solutions: A method for quantitatively identifying the color change of test paper based on improved MSRCR and random forest, which includes the following steps:
[0009] S1. Collect photos of test strips in different states by a camera as target images. The photos of test strips in different states include photos of test strips with different light intensities, different discoloration states, different discoloration areas, and photos of the same test strip taken at different positions and angles.
[0010] S2. Enhance and normalize each target image through the improved multi-scale adaptive gain MSRCR. The specific steps are as follows:
[0011] S2.1. Enhance the edge information in the image through the SSR algorithm, remove the low-frequency illumination part in the original image, and leave the high-frequency component corresponding to the original image:
[0012] Since the incident light irradiates on the reflective object, through the reflection of the reflective object, the reflected light enters the human eye, and the finally formed image is expressed by the following formula:
[0013] (1)
[0014] In formula (1), R(x, y) represents the reflection property of the object, that is, the intrinsic attribute of the image; L(x, y) represents the incident light image. The original image is S(x, y), the reflection image is R(x, y), and the brightness image is L(x, y). Formula (2) can be obtained:
[0015] (2)
[0016] In formula (2), F(x, y) is the center surround function, and the convolution operation with S(x, y) can be expressed as:
[0017] (3)
[0018] In formula (3), E is the Gaussian surround scale, is the scale, and its value needs to satisfy:
[0019] (4)
[0020] S2.2. On the basis of the SSR algorithm, add the Gaussian center surround function to develop into the multi-scale MSR algorithm. Its formula is:
[0021] (5)
[0022] In formula (5), K is the number of Gaussian center surround functions, represents the scale parameter corresponding to the Gaussian center surround function, 0 < < 1, and + + …… + = 1;
[0023] S2.3. On the basis of MSR, a color restoration factor C is added to adjust the defect of color distortion caused by the enhancement of local image contrast. The formula is as follows:
[0024] (6)
[0025] (7)
[0026] (8)
[0027] Among them, represents the color restoration factor of the th channel, which is used to adjust the color ratio of the three channels in the RGB color channels; represents the image of the th channel; f represents the mapping function of the color space; β is the gain constant; α is the controlled non-linear intensity;
[0028] S2.4. The algorithm of MSRCR is improved with an adaptive color restoration factor. Assume that represents the image of the th channel, =R, G, B. The following steps are designed for the adaptive color restoration factor C:
[0029] Step S2.4.1. Use the Sobel operator to calculate the gradients of the image in the x and y directions and , and calculate the gradient magnitude:
[0030] (9)
[0031] Step S2.4.2. Define a function f(G(x, y)) related to the gradient magnitude G(x, y), and re-define the adaptive color restoration factor C with this function. C can be expressed as:
[0032] (10)
[0033] Among them, the exp part of formula (10) is a Gaussian function, which is used to adjust the color restoration factor according to the gradient magnitude G(x,y); and are the mean and standard deviation of the gradient magnitude G(x,y) respectively; the latter part of formula (10) is the contrast enhancement part, which is used to enhance the color restoration factor according to the local contrast. L(x,y) is the local brightness component of the image; and are the mean and standard deviation of the local brightness L(x,y) respectively;
[0034] Step S2.4.3: Use the new adaptive color restoration factor C in formula (10) to replace the color factor in formula (7) as the new color factor of MSRCR, calculate the Laplacian operator Laplacian of the image to obtain the high-frequency component, enhance the high-frequency component, and then add the enhanced high-frequency component back to the original image for sharpening intensity adjustment to increase the clarity of the image;
[0035] S3: Based on the Euclidean distance and Otsu's dichotomy method, complete the test strip positioning and target area segmentation for the enhanced image. The specific steps are as follows:
[0036] Automatically locate the target test strip area in the large-scene image: Traverse each pixel point (R, G, B) in the image and calculate the Euclidean distance from it to the pure red point (255, 0, 0) and the pure yellow point (255, 255, 0) respectively. The Euclidean distance of the test strip pixel color is significantly different from that of the surrounding environment pixels. Therefore, use Otsu's maximum inter-class variance method to perform binary classification on the Euclidean distances corresponding to these pixel points, and take the smaller one as the target pixel point; then continue to use Otsu's method for binary classification among the target pixel points to carefully distinguish the discolored red area and yellow area; then, continue to traverse the Euclidean coordinate distances between these pixel points, and cluster the closest ones to each other on the same test strip into a group; finally, use a minimum bounding orthogonal rectangle to frame the target area, and the framed area is the test strip positioning area;
[0037] S4: Use the Euclidean distance in the RGB color space to analyze the degree of color change from yellow to red of the pixel points in the test strip positioning area. Define the Euclidean distance in the RGB color space as:
[0038] (15)
[0039] , are the RGB values of color A respectively; , are the RGB values of color B respectively;
[0040] Extract the following five features for each target image according to formula (15):
[0041] Feature 1: The average Euclidean distance Ly of the target image from pure yellow (255, 255, 0);
[0042] Feature 2: The average Euclidean distance Lr of the target image from pure red (255, 0, 0);
[0043] Feature 3: Parameter P = Ly / Lr;
[0044] Feature 4: The closest Euclidean distance Lymin between the target image and pure yellow (255, 255, 0);
[0045] Feature 5: The closest Euclidean distance Lrmin between the target image and pure red (255, 0, 0);
[0046] Feature 5: The closest Euclidean distance Lrmin between the target image and pure red (255, 0, 0);
[0047] S5. Have experienced experimenters rate the degree of color change of the test strip photos collected in step (1). If the experimenter believes there is no color change, it is rated 0 points. If the test strip changes color, it is rated on a scale of 1 - 10. The basis for the rating is: the degree of color change visually observed by the experimenter and the ratio of the visually observed color-changing area to the total area of the test strip. The higher the degree of color change and the larger the ratio of the color-changing area to the total area of the test strip, the higher the rating value;
[0048] Make the rating result target and all the features extracted in step (4) into a data set;
[0049] S6. Random forest training and testing: Import the data set into the random forest algorithm for training, and use the trained random forest algorithm to test and identify the degree of color change of the test strip photos to be recognized.
[0050] In step S2.2, the number of Gaussian center surround functions is three, namely Gaussian center surround functions with different scales of high, medium, and low. K = 3, and = 1 / 3; = 15; = 80; = 200.
[0051] In step S5, have experienced experimenters rate the degree of color change of the test strip photos, and finally take the average of the ratings of three experimenters as the final rating of the test strip photo.
[0052] The data set is in csv format.
[0053] In step S1, the number of test strip photos collected in different states is 1000. The data set obtained in step S5 contains 1000 test strip photo data, among which 700 test strip photo data are used for random forest training and 300 test strip photo data are used for random forest testing.
[0054] In step S6, when performing random forest training, optimize the random forest parameters as follows:
[0055] S6.1. Weighted Random Forest: Higher weights are assigned to important features so that they are more likely to be selected during the construction of the tree. When calculating the mean squared error (MSE) of the child nodes, the contribution of each feature is weighted. The formula is as follows:
[0056] (16)
[0057] where, is the weight of feature , is the actual value, is the test value;
[0058] S6.2. Weighted Information Splitting Criterion: When using the random forest algorithm in a regression task, if it is known that some features have a strong relationship with the target parameter while others do not, the following steps are used for optimization:
[0059] In weighted information gain, the contribution of each feature is multiplied by its weight; for each candidate split, the reduction in impurity before and after the split is calculated and multiplied by the weight of the feature, expressed as:
[0060] (17)
[0061] where, is the weight of feature A, Impurity(D) is the impurity of the dataset, is the dataset of child node v after the split;
[0062] Optimize the regression forest parameters according to the above steps and import the dataset for training.
[0063] In the dataset, the importance ranking of the five features is: P ≥ Lrmin ≥ Ly ≥ Lymin ≥ Lr.
[0064] In the dataset, the importance proportion of the five features is: P is 0.35, Lrmin is 0.25, Ly is 0.2, Lymin is 0.1, and Lr is 0.1.
[0065] Advantages of the present invention: The present invention combines the improved MSRCR algorithm with the random forest. The improved MSRCR algorithm can adaptively enhance the image and improve the recognition stability, while the random forest algorithm accurately measures the degree of test paper discoloration by integrating multiple decision trees, showing good performance. By introducing machine vision technology, the present invention realizes the automation of test paper discoloration recognition, significantly improves the efficiency of inspection work, and effectively solves the problems of low inspection efficiency, strong subjectivity and insufficient data utilization in traditional methods. The present invention provides a new and efficient quantitative analysis method for test paper discoloration in cable thermal stability tests, which is not only applicable to the recognition of test paper discoloration in cable thermal stability tests, but also can be applied to the recognition of test paper discoloration in other fields, providing useful reference for similar problems in other fields.
[0066] The advantages of the present invention are as follows:
[0067] 1. By introducing machine vision technology, the test identification process is automated, getting rid of the shackles of manual visual judgment, thus significantly improving the efficiency of inspection work;
[0068] 2. The improved MSRCR algorithm can adaptively enhance the image under different lighting conditions, and the device has strong robustness to ambient light, so the recognition stability is strong;
[0069] 3. The random forest algorithm is an ensemble method of decision trees, which usually has better generalization ability and test accuracy than a single decision tree. The test paper discoloration problem is transformed into a regression problem, and its discoloration degree is regarded as a continuous value in the range of 0-10, and this value can be directly output as a quantitative representation of the discoloration degree;
[0070] 4. The quantitative test paper discoloration degree parameter is beneficial to further characterize the trend of the speed of test paper discoloration, and extract more data information of the thermal stability test, such as predicting the discoloration time of the same batch of the same sample based on hypothesis testing and quantitatively evaluating the thermal stability of different samples. Description of the Drawings
[0071] Figure 1 It is a comparison diagram of the original image, MSRCR, and improved MSRCR effects of the large-scale image A;
[0072] Figure 2 It is a comparison diagram of the original image, MSRCR, and improved MSRCR effects of test papers B, C, and D;
[0073] Figure 3 It is the original image, test paper red-yellow detection, test paper pixel clustering, and test paper position positioning process images;
[0074] Figure 4 It is the feature parameter and target evaluation table corresponding to each target image;
[0075] Figure 5 It is the scatter plot distribution of the actual value target and the test value;
[0076] Figure 6 It is the ranking diagram of the importance of each parameter. Specific implementation manner
[0077] The present invention will be described in detail below in conjunction with the accompanying drawings and specific principles.
[0078] An embodiment of the present invention is a quantitative recognition method for test strip color change based on improved MSRCR and random forest, which includes the following steps:
[0079] S1. The camera captures test strip photos in different states as target images. The test strip photos in different states include test strip photos with different light intensities, different color change states, different color change areas, and photos of the same test strip taken at different positions and angles;
[0080] Theoretically, the more target images, the better. However, due to too many images, the workload of subsequent dataset creation is too large. Therefore, in this embodiment, 1000 test strip photos are used as target images.
[0081] S2. Each target image is enhanced and normalized through the improved multi-scale adaptive gain MSRCR. The purpose of this step is to enhance and normalize the target image so that the test strip colors under different light conditions can present the original colors of the test strip, without being affected by external light, and form a unified standard for subsequent processing.
[0082] The improved multi-scale adaptive gain MSRCR (Multi-Scale Retinex with Color Restore) algorithm of the present invention is an image enhancement algorithm, which is an improvement based on the traditional Retinex algorithm, especially in color restoration. The MSRCR algorithm is based on the Retinex theory, which holds that the color of an object is determined by the object's reflection ability to long-wave (red), medium-wave (green), and short-wave (blue) light, rather than by the absolute value of the reflected light intensity. The MSRCR algorithm enhances the dynamic range of the image through multi-scale Retinex processing and adjusts the color distortion that may be caused by the enhancement process through a color restoration factor. This algorithm is widely used in the field of image enhancement normalization. In the research on haze image enhancement, by performing HE and improved MSRCR enhancement on the haze image respectively and then performing weighted fusion according to certain image fusion rules, the defogging effect of the haze image is effectively improved, but there may be problems with algorithm stability and robustness. In the research on the image enhancement algorithm for power equipment, aiming at the problems such as halation phenomenon and color distortion existing in the traditional MSRCR algorithm when processing power equipment images, an improved Retinex image enhancement method is used. After converting the original image to the frequency domain through Fourier transform, convolution calculation is performed to obtain the reflection component of the original image, and combined with histogram equalization processing, the halation phenomenon is effectively eliminated, but the applicability of the improved algorithm when processing other types of images (such as natural scene images) and the computational complexity problem of the frequency domain transformation that may be introduced are not mentioned.
[0083] The thermal stability test usually takes a long time, the ambient light changes significantly, and fog often appears on the test tube wall due to temperature difference, affecting the observation of the test strip. Therefore, a pre-image processing algorithm is needed that can adjust the brightness without affecting the color, so that images under various lighting conditions can maintain a good visual effect; can effectively maintain and restore the original color of the image, so that the color of the image remains bright and natural while the brightness is enhanced; can significantly enhance the contrast of the image, especially in areas where there are shadows or highlights in the image, and can better display details, improving the clarity and visual effect of the image. Based on this, this step is optimized on the basis of the MSRCR algorithm. The specific steps are as follows:
[0084] S2.1. In this step, the edge information in the image is enhanced through the SSR algorithm, the low-frequency illumination part in the original image is removed, and the high-frequency component corresponding to the original image is left:
[0085] First, the basic algorithm of the MSRCR algorithm, the single-scale SSR (Single Scale Retinex) algorithm, is derived. An image can be regarded as composed of an incident image and a reflection image. Since the incident light irradiates on the reflecting object, through the reflection of the reflecting object, the reflected light enters the human eye, and the finally formed image is expressed by the following formula:
[0086] (1)
[0087] In formula (1), R(x, y) represents the reflection property of the object, that is, the intrinsic attribute of the image, which should be retained to the greatest extent; L(x, y) represents the incident light image, which determines the dynamic range that the image pixels can reach and should be removed as much as possible; assuming that the illumination image is estimated as a spatially smooth image, the original image is S(x, y), the reflection image is R(x, y), and the brightness image is L(x, y), formula (2) can be obtained:
[0088] (2)
[0089] In formula (2), F(x, y) is the center surround function, which is convolved with S(x, y) and can be expressed as:
[0090] (3)
[0091] In formula (3), E is the Gaussian surround scale, is the scale, and its value needs to satisfy:
[0092] (4)
[0093] It can be seen from formulas (1) to (4) that the convolution in the SSR algorithm is the calculation of the incident image. Its physical meaning is to calculate the change of illumination in the image by calculating the weighted average of the pixel points and the surrounding area, and remove L(x, y), only retaining the attribute of S(x, y). The center surround function F(x, y) uses a low-pass function, which can calculate the low-frequency part of the incident image corresponding to the original image in the algorithm. Removing the low-frequency illumination part from the original image and leaving the high-frequency component corresponding to the original image, so the SSR algorithm can better enhance the edge information in the image.
[0094] S2.2. On the basis of the SSR algorithm, a Gaussian center surround function is added to develop into a multi-scale MSR algorithm, and its formula is:
[0095] (5)
[0096] In formula (5), K is the number of Gaussian center surround functions. When K = 1, MSR degenerates to SSR; represents the scale parameter corresponding to the Gaussian center surround function, 0 < < 1, and + + …… + = 1;
[0097] In this embodiment, in order to ensure that the advantages of the high, medium, and low scales of SSR are combined, K = 3. Formula (5) is a three-scale split of formula (2). The number of Gaussian center surround functions is three, namely Gaussian center surround functions with different high, medium, and low scales. The specific meaning of the scale here is: The high scale usually corresponds to a smaller scale parameter of the Gaussian surround function, which means that the Gaussian surround function has a smaller spatial expansion range. Therefore, it mainly captures high-frequency detail information in the image, such as edges and textures; the medium scale corresponds to a moderate scale parameter of the Gaussian surround function. This scale can balance the high-frequency details and low-frequency information in the image, maintaining certain edge details while reflecting the overall illumination change of the image; the low scale corresponds to a larger scale parameter of the Gaussian surround function. This scale mainly captures the low-frequency information in the image, such as large-scale illumination changes, which helps to compress the dynamic range of the image.
[0098] And for the balanced processing of the high, medium, and low scales, = 1 / 3; According to the experimental empirical value = 15, = 80, = 200, corresponding to their respective Gaussian surround scales respectively.
[0099] S2.3. Although the performance of MSR is improved compared with SSR, there is still a color cast effect. Therefore, the MSR method with color restoration, MSRCR (Multi-Scale Retinex with Color Restoration), is proposed. On the basis of MSR, MSRCR adds a color restoration factor C to adjust the defect of color distortion caused by the enhancement of the local contrast of the image. Its formula is:
[0100] (6)
[0101] (7)
[0102] (8)
[0103] Among them, the meaning of is the image of the th channel; j takes three values 1, 2, 3, indicating the summation of the of the three channels of the image; N = 3; represents the The color restoration factor for each channel is used to adjust the color ratio of the three channels in the RGB color channels; Indicates the image of the channel; f represents the mapping function of the color space; β is the gain constant; α is the controlled non-linear intensity; according to the empirical data of the experiment, α = β = 1.5.
[0104] S2.4. The MSRCR algorithm uses the color restoration factor C to adjust the proportional relationship (formula (8)) between the three color channels (RGB color channels, RED, GREEN, BLUE) in the original image, so as to highlight the information in the relatively dark areas and achieve the defect of eliminating image color distortion.
[0105] The local contrast of the processed image is improved, and the brightness is similar to the real scene. Under people's visual perception, the image appears more realistic. However, the MSRCR algorithm has many empirical parameters, such as α and β in formula (8), which rely on parameter tuning and affect the software implementation and generalization to a certain extent. Moreover, when processing the strong light shadow transition area, the MSRCR algorithm may not be able to perfectly estimate the illumination, resulting in the appearance of the halo phenomenon. Based on the above problems, the present invention proposes an algorithm for improving MSRCR with an adaptive color restoration factor. Assume Indicates the image of the channel ( = R, G, B), and design the following adaptive color restoration factor C, and its steps are as follows:
[0106] Step S2.4.1: Use the Sobel operator to calculate the gradients of the image in the x and y directions and , and calculate the gradient magnitude:
[0107] (9)
[0108] Step S2.4.2: Define a function f(G(x, y)) related to the gradient magnitude G(x, y), and re-define the adaptive color restoration factor C with this function. C can be expressed as:
[0109] (10)
[0110] Among them, the exp part of formula (10) is a Gaussian function, which is used to adjust the color restoration factor according to the gradient magnitude G(x, y); and are the mean and standard deviation of the gradient magnitude G(x, y) respectively, and they can be calculated in the local area of the image (such as within a small window); the characteristic of this function is that when G(x, y) is close to , the output is close to 1, and when G(x, y) is far from The output decreases rapidly at this time. This helps to reduce color restoration in edge and detail areas (where the gradient magnitude is large), while increasing color restoration in smooth areas (where the gradient magnitude is small).
[0111] The second half of formula (10) is the contrast enhancement part, which is a simple linear transformation used to enhance the color restoration factor according to local contrast. L(x, y) is the local brightness component of the image and can be obtained through Gaussian blur; and are the mean and standard deviation of the local brightness L(x, y) respectively, also calculated within the local area. When the local contrast is high, the output of this part will increase, thus enhancing color restoration. This helps to restore more color information in areas with low contrast (such as shadows or dark parts).
[0112] Step S2.4.3: Use the new adaptive color restoration factor C of formula (10) to replace formula (7) as the new color factor of MSRCR, calculate the Laplacian of the image to obtain the high-frequency component, enhance the high-frequency component, and then add the enhanced high-frequency component back to the original image for sharpening intensity adjustment to increase the clarity of the image.
[0113] The experimental verification of step S2 is as follows:
[0114] Use the original image, MSRCR and the improved MSRCR algorithm for comparison, and evaluate the results using three evaluation indicators: image contrast, information entropy, and peak signal-to-noise ratio.
[0115] 1) The calculation formula of contrast Contrast is as follows:
[0116] (11)
[0117] In the formula , is the gray-level difference between adjacent pixels, is the pixel distribution probability that the gray-level difference between adjacent pixels is . Contrast can directly characterize the vividness of the image. The greater the contrast, the more vivid the color.
[0118] 2) The calculation formula of information entropy is as formula (12):
[0119] (12)
[0120] In formula (12), represents the probability that the gray value appears in the image. Information entropy reflects the amount of information in the image. The larger the information entropy value, the richer the detailed information contained in the image.
[0121] 3) The peak signal-to-noise ratio (PSNR) is used to evaluate the fidelity after image processing. The larger the PSNR value, the higher the image fidelity. Its calculation formula is:
[0122] (13)
[0123] where MSE is the mean square error of the pixel values after grayscale conversion of the original image I and the processed image K (both with dimensions of m×n), and the expression is:
[0124] (14)
[0125] To characterize the versatility, large-scale scene image A and test strip target images B, C, and D are selected for comparative analysis of the original image, MSRCR, and improved MSRCR algorithms. The comparison images are as Figure 1 、 Figure 2 shown. As can be seen from Figure 1 、 Figure 2 , while significantly enhancing the color, the improved MSRCR algorithm does not change the main color tone of the original image on a large scale like the initial MSRCR algorithm. In particular, for scene A with relatively dim ambient light and scene B with relatively strong ambient light, the improved MSRCR algorithm can restore and enhance the color of the test strip, with remarkable effects.
[0126] The evaluation results of each scene are shown in Table 1: From the data in Table 1, it can be observed that the improved MSRCR algorithm shows significant improvement in three evaluation indicators: contrast, entropy, and peak signal-to-noise ratio (PSNR). Specifically, for all test images A, B, C, and D, the improved MSRCR algorithm has greatly increased the contrast value compared to the original image and the traditional MSRCR algorithm, which means that the brightness difference in the image is more obvious and the image details are more prominent. At the same time, the increase in entropy indicates that the improved algorithm can retain more information when processing the image, making the image content richer and more delicate. In addition, in terms of the peak signal-to-noise ratio, the performance of the improved MSRCR algorithm has also been improved, indicating that the algorithm effectively reduces noise interference during the image enhancement process and improves the overall quality of the image. In summary, the improved MSRCR algorithm performs excellently in enhancing the visual effect of the image and retaining image information.
[0127] Table 1 Evaluation results of test images A, B, C, and D under different algorithms
[0128] Image Algorithm Contrast Entropy Peak Signal-to-Noise Ratio (PSNR) A Original Image 46.51 6.59 —— A MSRCR 48.81 6.67 27.28 A Improved MSRCR 55.65 6.93 27.27 B Original Image 20.95 5.91 —— B MSRCR 38.10 6.75 27.19 B Improved MSRCR 47.82 7.21 27.64 C Original Image 26.01 6.22 —— C MSRCR 39.12 6.84 27.61 C Improved MSRCR 43.18 7.14 28.09 D Original Image 16.21 5.34 —— D MSRCR 25.62 6.05 27.87 D Improved MSRCR 28.75 6.59 27.95
[0129] S3. During the actual experiment and original data collection, the sample is photographed. The original images collected by the camera contain useless information such as the background. That is to say, the initial images are necessarily large-scene images similar to Scene A that contain the external environment. However, the direct object of this study is the test strip area as shown in Pictures B, C, and D. Therefore, it is necessary to first complete the automatic positioning of the target test strip area in the large-scene image, that is, to remove the background information of the original photo and only retain the test strip area.
[0130] Therefore, in this step, the enhanced image is used to complete the test strip positioning and target area segmentation based on the Euclidean distance and Otsu's binarization method. The specific steps are as follows:
[0131] Automatic positioning of the target test strip area in the large-scene image: Traverse each pixel point (R, G, B) in the image and calculate the Euclidean distances from it to the pure red point (255, 0, 0) and the pure yellow point (255, 255, 0) respectively. The Euclidean distance of the test strip pixel color is significantly different from that of the surrounding environment pixels. Therefore, the Otsu's method of maximum inter-class variance is used to perform binary classification on the Euclidean distances corresponding to these pixel points, and the smaller one is taken as the target pixel point; and the Otsu's method is continued to be used for binary classification among the target pixel points to carefully distinguish the discolored red area and the yellow area; then, continue to traverse the Euclidean coordinate distances between these pixel points, and cluster the ones with the closest distance to each other on the same test strip into a group; finally, a minimum bounding orthogonal rectangle is used to frame the target area, and the framed area is the test strip positioning area. This process is as Figure 3 shown Figure 3 in. The leftmost one is the original image, the second from the left is the positioning of the reddest area and the yellowest area, the second from the right is the positioning of the same test strip area, and the rightmost one is the minimum bounding rectangle of the same test strip.
[0132] The basic principle of test strip positioning and target area segmentation is: The Euclidean distance of the test strip pixel color (whether it changes color or not) is significantly different from that of the surrounding environment pixels. Therefore, taking pure red (255, 0, 0)) and pure yellow (255, 255, 0) as the target pixels, calculate the RGB channel values of all pixels in the image and the Euclidean distances from these two points. The part closest to red can be regarded as the area where the test strip turns red, and the area closest to yellow can be regarded as the area where the test strip does not change color, that is, the yellow area, which is Figure 3 the effect shown in the second picture from the left. The purpose of this step is to find all the pixel point areas belonging to the test strip; then, cluster all the pixel points belonging to the test strip in the picture according to the information of their specific position coordinates in the photo. The purpose of this step is to distinguish how many test strips there are in the original picture, that is, Figure 3 the effect shown in the second picture from the right. Finally, a minimum bounding orthogonal rectangle is used to frame the target area, as Figure 3 shown in the rightmost picture.
[0133] S4. Use the Euclidean distance in the RGB color space to analyze the degree of color gradient from yellow to red in the test strip positioning area. The principle is to calculate the distance between the RGB values of each pixel point and the RGB values of pure yellow and red, and then judge the degree of color change according to the average distance of each pixel in the image from the standard pixel. Define the Euclidean distance in the RGB color space as:
[0134] (15)
[0135] , are the RGB values of color A respectively; , are the RGB values of color B respectively;
[0136] Extract the following five features from each target image according to formula (15):
[0137] Feature 1: The average Euclidean distance Ly between the target image and pure yellow (255, 255, 0);
[0138] Feature 2: The average Euclidean distance Lr between the target image and pure red (255, 0, 0);
[0139] Feature 3: Parameter P = Ly / Lr;
[0140] Feature 4: The closest Euclidean distance Lymin between the target image and pure yellow (255, 255, 0);
[0141] Feature 5: The closest Euclidean distance Lrmin between the target image and pure red (255, 0, 0).
[0142] S5. Have experienced experimenters rate the degree of color change of the test strip photos collected in step (1). If the experimenter believes there is no color change, it is rated 0 points. If the test strip changes color, it is rated on a scale of 1 - 10. The basis for the rating is: the degree of color change visually observed by the experimenter (usually the test strip changes color in the process of orange - red - magenta) and the ratio of the visually observed color - changing area to the overall area of the test strip. The higher the degree of color change and the larger the ratio of the color - changing area to the overall area of the test strip, the higher the rating value. In this step, have experienced experimenters rate the degree of color change of the test strip photos, and finally take the average of the ratings of three experimenters as the final rating of the test strip photo. Make the rating result target and all the features extracted in step (4) into a data set; the data set is in csv format. The data set obtained in step S5 contains 1000 test strip photo data, among which 700 test strip photo data are used for the next - step random forest training, and 300 test strip photo data are used for the next - step random forest testing.
[0143] Preliminary statistics on the scoring results and five features of the target images of 8 test strips in different states were carried out through OpenCV programming. The results are as Figure 4 shown in the table below. It can be seen from Figure 4 it that there is no perfect linear relationship between a single parameter and the target evaluation. However, the Ly and P values have a certain correlation with the target evaluation. Especially when the Ly value is small and the P value is also small, the target evaluation is low or zero. When the Ly value increases and the P value deviates from 1, the target evaluation tends to be high. Logically speaking, the above parameters should all have a relationship with the color change degree of the test strip. Therefore, the random forest can be used to deeply explore the internal relationship of the above parameters.
[0144] S6. Random forest training and testing:
[0145] The dataset was imported into the random forest algorithm for training, and the trained random forest algorithm was used to test and identify the color change degree of the test strip photos to be recognized. 700 groups of the obtained data were used for training, and 300 groups of data were used for verification, and finally a random forest decision system was formed.
[0146] Random Forest, as an ensemble learning method, obtains the final test value by constructing multiple decision trees and voting or averaging their results. The main advantage of this method is that by integrating multiple models, it can effectively handle the overfitting problem and improve the test accuracy and generalization ability of the model. In the random forest model of this step, the target score in the dataset is obtained by averaging the manual scores of three experienced experimenters, aiming to simulate the decision-making process of human experts and can be regarded as an expert system in the field of artificial intelligence.
[0147] When using the random forest algorithm in a regression task, the existing RandomForestRegression (RFR) improves the test accuracy by constructing multiple decision trees and partitioning the feature space in each tree, and finally integrating the test results of each tree. If it is known that some features (such as 1 or 2) have a great relationship with the target parameter, while other features have a small relationship, the random forest parameters can be optimized through the following steps, specifically as follows:
[0148] S6.1. Weighted random forest: Assign higher weights to important features so that they are more likely to be selected in the process of tree construction. When calculating the mean squared error MSE of the child nodes, the contribution of each feature is weighted. The formula is:
[0149] (16)
[0150] Among them, is the weight of the feature , is the actual value is the test value;
[0151] S6.2. Weighted Information Splitting Criterion:
[0152] Traditional information gain measures the reduction in the uncertainty of the data set after splitting using a certain feature. In weighted information gain, the contribution of each feature can be multiplied by its weight to pay more attention to important features. For each candidate split, calculate the reduction in impurity before and after the split and multiply it by the weight of the feature, which is expressed as:
[0153] (17)
[0154] where is the weight of feature A, Impurity(D) is the impurity of the data set is the data set of child node v after splitting;
[0155] Optimize the regression forest parameters according to the above steps, and import the data set for training. The final formed data is shown in Table 2 and Figures 5 to 6 as shown
[0156] In the data set, the importance ranking of the five features is: P≥Lrmin≥Ly≥Lymin≥Lr. In the data set Figure 6 as shown, the proportion of the importance of the five features is: P is 0.35, Lrmin is 0.25, Ly is 0.2, Lymin is 0.1, and Lr is 0.1.
[0157] Table 2 Evaluation Results of Random Forest Model Performance Parameters
[0158] (MSE) - Mean Squared Error R^2 Score - Coefficient of Determination (MAE) - Mean Absolute Error Median Absolute Error 0.15 0.85 0.35 0.25
[0159] As can be seen from Table 2, the test values are close to the actual values (target); the data in Table 2 show that MSE = 0.15: this value is relatively small, indicating that the test error of the model as a whole is not large. However, since MSE is more sensitive to larger errors, this value may still be affected by some larger errors; R^2 = 0.85: this value is relatively high, indicating that the model can explain 85% of the variability in the data. This is a very good result, showing that the model performs quite well in fitting the data; MAE = 0.35: this value is larger than MSE because MAE measures the absolute error rather than the squared error. Nevertheless, this value is still relatively small, indicating that the test error of the model is not large in most cases. MAE = 0.35: this value is larger than MSE because MAE measures the absolute error rather than the squared error. Nevertheless, this value is still relatively small, indicating that the test error of the model is not large in most cases.
[0160] In summary, the method for quantitatively identifying the color change of test strips based on improved MSRCR and random forest of the present invention is applicable to the quantitative study of the color change of test strips in the cable thermal stability test. The improved MSRCR algorithm can adaptively enhance the image and improve the recognition stability, while the random forest algorithm accurately tests the degree of color change of the test strip by integrating multiple decision trees, showing good performance. By introducing machine vision technology, the present invention realizes the automation of test strip color change recognition, significantly improves the efficiency of inspection work, and effectively solves the problems of low inspection efficiency, strong subjectivity and insufficient data utilization in traditional methods. The present invention not only provides a new and efficient method for quantitatively analyzing the color change of test strips in the cable thermal stability test, but also provides a useful reference for similar problems in other fields.
Claims
1. A quantitative identification method for test paper color change based on improved MSRCR and random forest, characterized in that: It includes the following steps: S1. The camera collects photos of the test paper in different states as target images. The photos of the test paper in different states include photos of the test paper with different light intensities, different discoloration states, and different discoloration areas, as well as photos of the same test paper taken at different positions and angles; S2. Each target image is enhanced and normalized by the improved multi-scale adaptive gain MSRCR. The specific steps are as follows: S2.
1. Enhance the edge information in the image through the SSR algorithm, remove the low-frequency illumination part in the original image, and leave the high-frequency component corresponding to the original image: As the incident light shines on the reflective object, it is reflected by the reflective object to form reflected light that enters the human eye. The final image is expressed as the following formula: (1) In formula (1), R(x, y) represents the reflective property of the object, that is, the intrinsic property of the image; L(x, y) represents the incident light image, the original image is S(x, y), the reflected image is R(x, y), and the brightness image is L(x, y), and formula (2) is obtained: (2) In formula (2), F(x, y) is the center surround function, which is convolved with S(x, y) and expressed as: (3) In formula (3), E is the Gaussian surround scale, is the scale, and its value must satisfy: (4) S2.2, based on the SSR algorithm, the Gaussian center surround function is added to develop a multi-scale MSR algorithm, the formula is: (5) In formula (5), K is the number of Gaussian center surround functions, Represents the scale parameter corresponding to the Gaussian center surround function, 0< <1, and + +……+ =1; S2.
3. Based on MSR, a color restoration factor C is added to adjust the defect of color distortion caused by contrast enhancement in local areas of the image. The formula is: (6) (7) (8) in, Indicates The color restoration factor of each channel is used to adjust the color ratio of the three channels in the RGB color channel; Indicates An image with multiple channels; f represents the mapping function of the color space; β is the gain constant; α is the controlled nonlinear intensity; S2.4, Adaptive color restoration factor to improve the MSRCR algorithm, assuming Indicates The image of channels, =R,G,B, design the following adaptive color restoration factor C, the steps are: Step S2.4.1: Use the Sobel operator to calculate the gradient of the image in the x and y directions and , calculate the gradient magnitude: (9) Step S2.4.2, define a function f(G(x, y)) related to the gradient amplitude G(x, y), and use this function to redefine the adaptive color restoration factor C, C is expressed as: (10) Wherein, the exp part of formula (10) is a Gaussian function, which is used to adjust the color restoration factor according to the gradient amplitude G(x,y); and are the mean and standard deviation of the gradient magnitude G(x,y) respectively; the second half of formula (10) is the contrast enhancement part, which is used to enhance the color restoration factor according to the local contrast, and L(x,y) is the local brightness component of the image; and are the mean and standard deviation of the local brightness L(x,y); Step S2.4.3, using the new adaptive color restoration factor C of formula (10) to replace formula (7) as the new color factor of MSRCR, calculating the Laplacian operator Laplacian of the image, obtaining the high-frequency component, enhancing the high-frequency component, and then adding the enhanced high-frequency component back to the original image, adjusting the sharpening intensity to increase the clarity of the image; S3, based on the Euclidean distance and Otsu dichotomy method, the test paper positioning and target area segmentation of the enhanced image are completed. The specific steps are as follows: Automatically locate the target test paper area in a large scene image: traverse the Euclidean distances of each pixel (R, G, B) in the image with the pure red point (255, 0, 0) and the pure yellow point (255, 255, 0). The Euclidean distance of the test paper pixel color is significantly different from the Euclidean distance of the surrounding environment pixels. Therefore, the Euclidean distances corresponding to these pixels are binary classified using the Otsu maximum inter-class variance method, and the smaller one is taken as the target pixel point; and the Otsu method is continued to be used for binary classification in the target pixel points to carefully separate the color-changing red area and the yellow area; then, continue to traverse the Euclidean coordinate distances between these pixels, and cluster the closest distances on the same test paper into a group; finally, a minimum inner orthogonal rectangle is used to realize the frame selection of the target area, and the framed area is the test paper positioning area; S4. Use the Euclidean distance in the RGB color space to analyze the degree to which the color of the pixels in the test paper positioning area gradually changes from yellow to red. The Euclidean distance in the RGB color space is defined as: (15) , are the RGB values of color A respectively; , are the RGB values of color B respectively; According to formula (15), the following five features are extracted for each target image: Feature 1: The average Euclidean distance Ly between the target image and pure yellow (255,255,0); Feature 2: The average Euclidean distance Lr of the target image from pure red (255, 0, 0); Feature 3: Parameter P=Ly / Lr; Feature 4: The closest Euclidean distance Lymin between the target image and pure yellow (255,255,0); Feature 5: The closest Euclidean distance Lrmin between the target image and pure red (255,0,0); S5. An experienced experimenter scores the color change degree of the test paper photo collected in step S1. If the experimenter believes that there is no color change, it is scored as 0 points. If the test paper changes color, it is scored on a scale of 1 to 10. The scoring is based on: the color change degree visually observed by the experimenter and the area ratio of the visually discolored area to the entire test paper. The higher the color change degree and the larger the area ratio of the discolored area to the entire test paper, the higher the score value; The scoring result target and all the features extracted in step S4 are made into a data set; S6. Random forest training and testing: The data set is imported into the random forest algorithm for training, and the trained random forest algorithm is used to perform color change test on the test paper photos to be identified.
2. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 1, characterized in that: In step S2.2, there are three Gaussian center surround functions, namely, high, medium, and low Gaussian center surround functions of different scales, K=3, and =1 / 3; =15; =80; =200.
3. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 1, characterized in that: In step S5, the color change degree of the test paper photo is scored by experienced experimenters, and finally the average of the scores of three experimenters is taken as the final score of the test paper photo.
4. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 1, characterized in that: The dataset is in csv format.
5. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 1, characterized in that: In step S1, the number of test paper photos in different states collected is 1000. The data set obtained in step S5 contains 1000 test paper photo data, of which 700 test paper photo data are used for random forest training and 300 test paper photo data are used for random forest testing.
6. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 1, characterized in that: In step S6, during random forest training, the random forest parameters are optimized as follows: S6.1, Weighted Random Forest: Give higher weights to important features, making them more likely to be selected during the tree construction process. When calculating the mean square error (MSE) of the child nodes, the contribution of each feature is weighted. The formula is: (16) in, It is a feature The weight of is the actual value, is the test value; S6.2, Weighted Information Splitting Criterion: When using the random forest algorithm in a regression task, if it is known that some features are highly correlated with the target parameter, while other features are not highly correlated, optimization is performed by the following steps: In weighted information gain, the contribution of each feature is multiplied by its weight; for each candidate split, the reduction in impurity before and after the split is calculated and multiplied by the weight of the feature, expressed as: (17) in, is the weight of feature A, Impurity(D) is the impurity of the dataset, is the data set of the child node v after splitting; Follow the above steps to optimize the regression forest parameters and import the dataset for training.
7. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 6, characterized in that: In the data set, the importance of the five features is ranked as follows: P ≥ Lrmin ≥ Ly ≥ Lymin ≥ Lr.
8. The method for quantitatively identifying color change of test paper based on improved MSRCR and random forest according to claim 7, characterized in that: In the data set, the importance of the five features is: P is 0.35, Lrmin is 0.25, Ly is 0.2, Lymin is 0.1, and Lr is 0.1.
Citation Information
Patent Citations
A remote sensing image retrieval method based on nonlinear dimension reduction and sparse representation
CN109815357A
Immunochromatography concentration detection method and system based on machine learning
CN112071423A