Underground vision enhancement method and device combined with target recognition

By combining the downhole visual enhancement method of target recognition, using defog operation and light transmittance matrix fusion technology, the problem of degradation of surveillance video image quality in the mine environment is solved, and image clarity and target recognition accuracy are improved.

CN120339114APending Publication Date: 2025-07-18CHINA COAL RES INST +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510258913.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The quality of surveillance videos in the mine environment will decline, resulting in the cover or distortion of key information, affecting security management.

Method used

Combined with the downhole visual enhancement method of target recognition, through defog operation, pretreatment, object detection and light transmittance matrix fusion, image clarity is improved and moisture dust and dirt influence is removed.

Benefits of technology

It improves the recognition accuracy of the target recognition model under complex operating conditions, ensures the reliability and clarity of the output images, and solves the problem of image blurring and occlusion in the mine environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339114A_ABST
    Figure CN120339114A_ABST
Patent Text Reader

Abstract

The invention provides an underground vision enhancement method and device combined with target recognition. The method comprises the following steps: carrying out defogging operation on an original first image in a video stream to obtain a defogged second image and a first light transmittance matrix; preprocessing the second image to obtain a preprocessed third image, and inputting the third image into a target detection model to output a target detection matrix; obtaining a fourth image according to the target detection matrix and the second image; according to the first light transmittance matrix and the second light transmittance matrix, pixel value fusion is carried out on a historical fifth image and a historical fourth image, an enhanced target image is obtained, and the second light transmittance matrix is the light transmittance matrix corresponding to the fifth image. Image defogging and target detection are combined, and the image is enhanced, so that the quality and definition of the image are improved; moreover, the defogged image is used as the recognition input of the target detection model, thereby improving the recognition accuracy of the target recognition model under a complex working condition, and guaranteeing the reliability of an output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of coal mine safety technology, and particularly to a downhole vision enhancement method and device combining target recognition. Background Art

[0002] In the safe production of mines, the monitoring system has become the core tool to ensure the smooth progress of production and the safety of personnel. The clarity of downhole monitoring videos directly affects the identification of potential safety hazards and the effectiveness of equipment monitoring. However, the complex environmental conditions inside the mine, such as coal dust, muddy water, fog, and direct strong light underground, often cause dirt to accumulate on the surface of the camera lens, resulting in a significant decrease in the quality of the monitoring video, frequent problems such as blurred, blocked, and reflective video images, and seriously hindering the normal operation of the monitoring system. In this case, key information such as the operating state of downhole equipment and personnel activities may be obscured or distorted, unable to provide effective decision-making support for production management. Therefore, how to improve the video clarity in the harsh mine environment and remove the influence of water vapor, dust, and dirt on the monitoring screen has become one of the key issues in mine safety management. Summary of the Invention

[0003] The purpose of this application is to solve at least one of the technical problems in the related art to some extent.

[0004] To this end, the first object of this application is to propose a downhole vision enhancement method combining target recognition, which can improve the clarity of the monitoring screen in the mine environment, and can remove the influence of water vapor, dust, and dirt, thereby improving the safety of underground operations.

[0005] The second object of this application is to propose a downhole vision enhancement device combining target recognition.

[0006] The third object of this application is to propose an electronic device.

[0007] The fourth object of this application is to propose a computer-readable storage medium.

[0008] The fifth object of this application is to propose a computer program product.

[0009] To achieve the above object, the first aspect embodiment of this application proposes a downhole vision enhancement method combining target recognition, including:

[0010] Performing defogging operation on the original first image in the video stream to obtain a defogged second image and a first transmittance matrix;

[0011] Preprocessing the second image to obtain a preprocessed third image, and inputting the third image into a target detection model to output a target detection matrix;

[0012] Based on the target detection matrix and the second image, a fourth image is obtained;

[0013] Based on the first light transmittance matrix and the second light transmittance matrix, pixel value fusion is performed on the historical fifth image and the fourth image to obtain an enhanced target image, where the second light transmittance matrix is the light transmittance matrix corresponding to the fifth image.

[0014] To achieve the above object, an embodiment of the second aspect of the present application proposes an underground vision enhancement device combined with target recognition, including:

[0015] A defogging module for performing defogging operations on the original first image in the video stream to obtain a defogged second image and a first light transmittance matrix;

[0016] A model detection module for preprocessing the second image to obtain a preprocessed third image, and inputting the third image into a target detection model to output a target detection matrix;

[0017] A first fusion model for obtaining a fourth image based on the target detection matrix and the second image;

[0018] A second fusion module for performing pixel value fusion on the historical fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix to obtain an enhanced target image, where the second light transmittance matrix is the light transmittance matrix corresponding to the fifth image.

[0019] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor; and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory so that the processor can execute the underground vision enhancement method combined with target recognition described in the first aspect embodiment above.

[0020] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, on which a computer program is stored, and the computer instructions are used to make the computer execute the underground vision enhancement method combined with target recognition described in the above embodiment of one aspect.

[0021] To achieve the above object, an embodiment of the fifth aspect of the present application proposes a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the underground vision enhancement method combined with target recognition described in the above embodiment of one aspect.

[0022] The downhole vision enhancement method and device combining target recognition provided by this application can combine image dehazing and target detection, and use the dehazed image as the recognition input of the target detection model, improving the recognition accuracy of the target recognition model under complex working conditions and ensuring the reliability of the output image. Moreover, subsequent image enhancement is performed after target recognition, and during the subsequent image enhancement process, the prior information contained in the light transmittance matrix is used to quickly determine the relative clarity of different positions of the image in combination with historical images, and the clearer image is used as the model output.

[0023] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0025] Figure 1 is a schematic flowchart of a downhole vision enhancement method combining target recognition provided by an embodiment of this application;

[0026] Figure 1A is a schematic diagram of the image after dehazing provided by an embodiment of this application;

[0027] Figure 2 is a schematic flowchart of a process for obtaining a target image provided by an embodiment of this application;

[0028] Figure 3 is a schematic flowchart of another process for obtaining a target image provided by an embodiment of this application;

[0029] Figure 4 is a schematic flowchart of another downhole vision enhancement method combining target recognition provided by an embodiment of this application;

[0030] Figure 4A is a schematic diagram of the processing result of the image provided by an embodiment of this application;

[0031] Figure 5 is a schematic flowchart of another downhole vision enhancement method combining target recognition provided by an embodiment of this application;

[0032] Figure 6 is a schematic structural diagram of a downhole vision enhancement device combining target recognition provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.

[0034] The downhole vision enhancement method and device combined with target recognition according to embodiments of the present application will be described below with reference to the accompanying drawings.

[0035] Figure 1 It is a schematic flowchart of a downhole vision enhancement method combined with target recognition provided in an embodiment of the present application. As Figure 1 shown, the downhole vision enhancement method combined with target recognition may include but is not limited to the following steps:

[0036] S101, perform a defogging operation on the original first image in the video stream to obtain a defogged second image and a first transmittance matrix.

[0037] In some embodiments, the video stream may be video data collected downhole. Optionally, cameras may be deployed downhole to monitor the downhole situation through the cameras.

[0038] In some embodiments, the downhole camera may upload the collected video stream to the ground server through the downhole ring network. The video stream is analyzed and recognized by the ground server.

[0039] In some embodiments, in the complex working conditions of the mine shaft, during the operation of equipment such as mining machines, water mist and the like will be generated. The generated water mist and the like will block the downhole scene, and the generated water mist and the like also have extremely high characteristics of dynamic changes in range, shape, and concentration. If the first image in the video stream is directly used for recognition, the recognition result will often be relatively poor. In order to improve the accuracy of the recognition result, the dark channel method may be used to perform a defogging operation on the original first image in the video, and a defogged second image is obtained to achieve image enhancement of the first image, so as to provide a clearer image for subsequent target monitoring.

[0040] In some embodiments, the dark channel method is based on a prior observation result, that is, the local minimum values of the R, G, and B channels in areas with relatively light fog should be close to 0. Let the original fog-free image be J(x) and the fog-containing image be I(x):

[0041] J dark (x) = min {R,G,B} (min y∈θ(x) (J(y))) (1)

[0042] Except for the sky area, in the remaining areas J dark(x) has an extremely low intensity, approaching 0. This is the dark channel prior method. Among them, θ(x) is a region containing the pixel point x, and the other pixel points in this region are the neighboring pixel points of the pixel point x. y ∈ θ(x) means that the pixel point x is any pixel point in this region.

[0043] Based on this prior formula, the foggy image affected by fog is obtained as shown in formula (2):

[0044] I(x) = J(x) * t(x) + A * (1 - t(x)) (2)

[0045] Among them, t(x) is the transmittance matrix, A represents the illumination intensity of the underground monitoring scene, and I(x) is the original first image affected by fog.

[0046] Optionally, the average value of the first n brightest pixels in the image is determined, where n is a hyperparameter set by humans. w is the proportion of removing illumination, usually set to 0.95 by humans.

[0047] Optionally, the calculation formula of the light transmittance matrix is as follows in formula (3):

[0048]

[0049] In some embodiments, based on the above formula (1), the de-fogged second image J(x) processed by the dark channel method can be obtained, and the conversion formula of the second image is as follows in formula (4):

[0050]

[0051] In the complex working conditions of the underground mine, different from the relatively uniform fog or haze in the atmosphere, the water mist generated by equipment such as mining machines will completely block the scene, and at the same time has extremely high characteristics of dynamic changes in range, shape, and concentration. Based on the dark channel method, the first image is de-fogged and enhanced to remove the scene blocked by water mist, etc., as Figure 1A shown.

[0052] In some embodiments, while performing the de-fogging operation on the first image, the light transmittance matrix of the first image can be determined based on the image information of the first image. Optionally, obtain the illumination intensity of the underground monitoring scene, and determine the set proportion of removing illumination. Further, according to the illumination intensity and the proportion of removing illumination, combined with the pixel point information in the first image, determine the first light transmittance matrix of the first image.

[0053] S102. Preprocess the second image to obtain the preprocessed third image, and input the third image into the target detection model to output the target detection matrix.

[0054] In some embodiments, after receiving a video stream including a second image, preprocessing may be performed on the second image in the video stream. That is to say, it is necessary to preprocess second images with different resolutions or formats.

[0055] In some embodiments, the preprocessing performed on the second image may include, but is not limited to: cropping, normalization operations. Optionally, after obtaining the second image, a cropping operation may be performed on the second image according to the recognition requirements to meet the processing requirements of the subsequent model. Further, downsampling and normalization processing may be performed on the cropped second image to ensure that the input of the model meets the expected standards.

[0056] In some embodiments, in order to further process the second based on the dark channel output result to enhance the defogging effect, the preprocessed third image may be input into a pre-trained object detection model, and the object detection model performs object detection on the third image to identify the objects included in the third image.

[0057] In some embodiments, the object detection model may identify common objects underground in a mine, such as: mining machines, coal mine baffles, conveyor belts, etc. The input of the object detection model is the image after defogging, which improves the inference efficiency and the anti-interference effect against fog.

[0058] In some embodiments, the object detection model may be a YOLO object detection model.

[0059] In some embodiments, the object detection model may output the object detection region of the detected object and the category of the object detection region. It can be understood that the object detection region is the pixel coordinate range of the detected object.

[0060] In some embodiments, the object detection model may output an object detection matrix. It can be understood that the object detection matrix may include the detected object detection region and the corresponding category.

[0061] In some embodiments, the detected object detection region and the corresponding category output by the object detection model may be used as a subsequent discrimination basis or as part of the final output information.

[0062] S103, obtain a fourth image according to the object detection matrix and the second image.

[0063] In some embodiments, the detection result of each pixel point in the second image in the object detection matrix may be determined, and the detection result may be marked at the pixel point to obtain a fourth image. Optionally, the detection result may include the object type corresponding to the pixel points within the object detection region.

[0064] In some embodiments, according to the coordinate information of each pixel point in the second image, further, based on the coordinate information of the pixel point, mapping is performed to a target detection matrix to determine the target detection area corresponding to the pixel point, and the target type marked by the target detection area is associated with the pixel point.

[0065] It can be understood that the target detection matrix includes multiple target detection areas and the target type corresponding to each target detection area. Further, the target type is mapped into the second image according to the coordinate information corresponding to the target detection area, and then a fourth image carrying the target detection result is obtained. That is to say, the target detection area and the corresponding target type can be marked in the fourth image, and the target detection area can be understood as the detection box output by the target detection model.

[0066] S104, according to the first transmittance matrix and the second transmittance matrix, perform pixel value fusion on the historical fifth image and the fourth image to obtain an enhanced target image.

[0067] In some embodiments, the fifth image is the enhanced target image output at the previous moment. It should be noted that the target image is the image to be finally output. It can be understood that the second transmittance matrix is the transmittance matrix corresponding to the fifth image.

[0068] In some embodiments, according to the first transmittance matrix and the second transmittance matrix, determine the target pixel value of each pixel point from the fifth image and the fourth image, and determine the enhanced target image according to the target pixel value of each pixel point.

[0069] In some embodiments, according to the first transmittance matrix, the second transmittance matrix, and a set transmittance threshold, generate a mask matrix. Further, according to the mask matrix, perform pixel value fusion on the fifth image and the fourth image to obtain an enhanced target image.

[0070] In the embodiments of the present application, by using the prior information contained in the transmittance matrix and combining historical images to quickly determine the relative clarity of different positions of the image, and taking the clearer image as the model output. It is simpler than other methods such as deep learning models, greatly increasing the inference efficiency. Moreover, in the above steps, joint dehazing and target detection can effectively solve the characteristics of "high concentration" and "high dynamics" of the fog in the mine, ensuring the image output quality under harsh working conditions.

[0071] In the embodiments of the present application, image dehazing and object detection can be combined, and the dehazed image can be used as the recognition input of the object detection model to improve the recognition accuracy of the object recognition model under complex working conditions and ensure the reliability of the output image. Moreover, subsequent image enhancement is performed after object recognition, and during the subsequent image enhancement process, the prior information contained in the light transmittance matrix is used to quickly discriminate the relative sharpness of different positions of the image in combination with historical images, and the clearer image is used as the model output.

[0072] Based on the above embodiments, the process of step S103 of performing fusion processing on the historical fifth image and fourth image according to the first light transmittance matrix and the second light transmittance matrix to obtain an enhanced target image can be explained:

[0073] Figure 2 It is a schematic flowchart of a process for obtaining a target image provided in the embodiments of the present application. As Figure 2 shown, the process for obtaining a target image may include but is not limited to the following steps:

[0074] S201, determine the target pixel value of each pixel point from the fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix.

[0075] S202, determine the enhanced target image according to the target pixel value of each pixel point.

[0076] In some embodiments, for any pixel point i, the first light transmittance and the second light transmittance corresponding to the pixel point i are respectively determined from the first light transmittance matrix and the second light transmittance matrix. Further, according to the first light transmittance and the second light transmittance corresponding to the pixel point i and the set light transmittance threshold, it is determined whether the pixel point i satisfies the pixel update condition.

[0077] In response to the pixel point i not satisfying the pixel update condition, determine the pixel value of the pixel point i in the fifth image as the pixel value of the pixel point i in the target image; in response to the pixel point i satisfying the pixel update condition, determine the pixel value of the pixel point i in the fourth image as the target pixel value of the pixel point i.

[0078] In some embodiments, it is determined whether the first light transmittance of the pixel point i is greater than the second light transmittance of the pixel point i, and it is determined whether the first light transmittance of the pixel point i is greater than the set light transmittance threshold.

[0079] In response to the first light transmittance of the pixel point i being greater than the second light transmittance of the pixel point i and the first light transmittance of the pixel point i being greater than the set light transmittance threshold, determine that the pixel point i satisfies the pixel update condition.

[0080] In response to the first light transmittance of pixel point i being less than or equal to the second light transmittance of pixel point i, and / or the first light transmittance of pixel point i being less than or equal to a set light transmittance threshold, it is determined that pixel point i does not meet the pixel update condition.

[0081] In the embodiments of the present application, by using the prior information contained in the light transmittance matrix and combining historical images, the relative clarity of different positions of the image is quickly discriminated, and the clearer image is used as the model output. It is simpler than other methods such as deep learning models, greatly increasing the inference efficiency. Moreover, in the above steps, joint dehazing and target detection can effectively solve the characteristics of "high concentration" and "high dynamics" of the fog in the mine, ensuring the image output quality under harsh working conditions.

[0082] Figure 3 It is a schematic flowchart of another process for obtaining a target image provided in the embodiments of the present application. As Figure 3 shown, the process for obtaining the target image may include but is not limited to the following steps:

[0083] S301, generate a mask matrix according to the first light transmittance matrix, the second light transmittance matrix, and a set light transmittance threshold.

[0084] S302, perform pixel value fusion on the fifth image and the fourth image according to the mask matrix to obtain an enhanced target image.

[0085] In some embodiments, for any pixel point i, the first light transmittance and the second light transmittance corresponding to pixel point i are respectively determined from the first light transmittance matrix and the second light transmittance matrix. And according to the first light transmittance and the second light transmittance corresponding to pixel point i, and a set light transmittance threshold, the value of pixel point i in the mask matrix is judged; further, according to the value of pixel point i in the mask matrix, a mask matrix is generated.

[0086] In some embodiments, according to the first light transmittance and the second light transmittance corresponding to pixel point i, and a set light transmittance threshold, it is judged whether pixel point i meets the pixel update condition. In response to pixel point i not meeting the pixel update condition, it is determined that the value of pixel point i is 0, which is used to indicate that the pixel value of pixel point i in the fifth image is the target pixel value of pixel point i; in response to pixel point i meeting the pixel update condition, it is determined that the value of pixel point i is 1, which is used to indicate that the pixel value of pixel point i in the fourth image is the target pixel value of pixel point i.

[0087] In some embodiments, it is determined whether the first light transmittance of pixel point i is greater than the second light transmittance of pixel point i, and it is determined whether the first light transmittance of pixel point i is greater than a set light transmittance threshold; in response to the first light transmittance of pixel point i being greater than the second light transmittance of pixel point i and the first light transmittance of pixel point i being greater than the set light transmittance threshold, it is determined that pixel point i meets the pixel update condition; in response to the first light transmittance of pixel point i being less than or equal to the second light transmittance of pixel point i, and / or the first light transmittance of pixel point i being less than or equal to the set light transmittance threshold, it is determined that pixel point i does not meet the pixel update condition.

[0088] In the embodiments of the present application, by using the prior information contained in the light transmittance matrix and combining historical images, the relative clarity of different positions of the image is quickly discriminated, and the clearer image is used as the model output. It is simpler than other methods such as deep learning models, greatly increasing the inference efficiency. Moreover, in the above steps, joint dehazing and target detection can effectively solve the characteristics of "high concentration" and "high dynamics" of the fog in the mine, ensuring the image output quality under harsh working conditions.

[0089] In the above embodiment, for each pixel point i, it is also necessary to determine whether pixel point i is within the target detection area detected by the target detection model.

[0090] In response to pixel point i being within the target detection area, the pixel value of pixel point i in the fourth image is determined as the target pixel value of pixel point i in the target image.

[0091] In response to pixel point i not being within the target detection interval, the pixel values of the fifth image and the fourth image are fused according to the first light transmittance matrix and the second light transmittance matrix of pixel point i to determine the target pixel value of pixel point i in the target image. That is to say, when pixel point i is not within the target detection interval, the target pixel value of pixel point i in the target image can be determined according to the embodiments shown in Figure 2 and Figure 3 to determine the target pixel value of pixel point i in the target image.

[0092] Figure 4 Another flow chart of the downhole vision enhancement method combined with target recognition provided in the embodiments of the present application. As shown in Figure 4 shown, the downhole vision enhancement method combined with target recognition may include but is not limited to the following steps:

[0093] S401, collect the original first image in the video stream.

[0094] S402, perform dehazing enhancement on the first image to obtain the dehazed second image.

[0095] S403, preprocess the second image to obtain the third image, where the preprocessing includes cropping, downsampling, and normalization processing.

[0096] S404. Input the third image into the target detection model to output a target detection matrix.

[0097] S405. Obtain a fourth image based on the target detection matrix and the second image.

[0098] S406. Determine whether the pixel meets the update condition.

[0099] In response to the pixel meeting the update condition, execute step S406.

[0100] In response to the pixel not meeting the update condition, execute step S407.

[0101] S407. Update the pixel value of the pixel point using the pixel value of the fourth image.

[0102] S408. Determine the historical image.

[0103] S409. Determine the historical pixel value in the historical image.

[0104] S410. Update the pixel value of the pixel point using the historical pixel value.

[0105] S411. Obtain the enhanced target image.

[0106] S413. Output the image to the front - end display.

[0107] S413. Use the output target image as the historical image for the next moment.

[0108] As Figure 4A shown is a schematic diagram of the result after processing the image by the downhole vision enhancement method combining target recognition provided in the embodiment of the present application. From top to bottom, they are the original first image, the defogged second image corresponding to the dark channel enhancement method, and the target image.

[0109] In the embodiment of the present application, image defogging and target detection can be combined, and the defogged image is used as the recognition input of the target detection model, improving the recognition accuracy of the target recognition model in complex working conditions and ensuring the reliability of the output image. Moreover, subsequent image enhancement is performed after target recognition, and during the subsequent image enhancement process, the prior information contained in the transmittance matrix is utilized to quickly determine the relative clarity of different positions of the image in combination with the historical image, and the clearer image is used as the model output.

[0110] Figure 5 A schematic flow diagram of a pixel update process provided in the embodiment of the present application. As Figure 5 shown, the pixel update process may include but is not limited to the following steps:

[0111] S501, Determine whether the pixel is a pixel within the target detection area.

[0112] In response to the pixel being a pixel within the non - target detection area, step S502 can be executed;

[0113] In response to the pixel being a pixel within the target detection area, step S505 can be executed.

[0114] S502, Determine whether the first light transmittance corresponding to the pixel is greater than the second light transmittance corresponding to the pixel.

[0115] In response to the first light transmittance corresponding to the pixel being greater than the second light transmittance corresponding to the pixel, step S503 can be executed;

[0116] In response to the first light transmittance corresponding to the pixel being less than or equal to the second light transmittance corresponding to the pixel, step S505 can be executed.

[0117] S503, Determine whether the first light transmittance corresponding to the pixel is greater than the set light transmittance threshold.

[0118] In response to the first light transmittance corresponding to the pixel being greater than the light transmittance threshold, step S504 can be executed;

[0119] In response to the first light transmittance corresponding to the pixel being less than or equal to the light transmittance threshold, step S505 can be executed.

[0120] S504, Update the pixel value of the pixel to the pixel value of the pixel in the fourth image.

[0121] S505, Continue to use the pixel value of the pixel in the fifth image for the pixel value.

[0122] Such as Figure 5As shown, when a certain area in the fourth image is completely obscured by temporary thick fog, the historical image (i.e., the fifth image) will be used as the image output for this area at the current moment. That is to say, the pixel values of the pixel points in the historical image will be used as the pixel values of the pixel points in this area at the current moment. In the embodiments of the present application, in order to ensure reliability when facing key targets and moving objects, the pixel values of the pixel points within the areas of the key targets and moving objects output by the target detection model in the fourth image can be forcibly updated to ensure the timeliness and reliability of the monitoring images. The second discrimination is to confirm that the updated image is clearer than the historical image; otherwise, the historical image will still be used for output. The third discrimination targets the characteristics of high dynamic range of fog and fluctuating light transmittance matrix. To ensure that in high-obscuration working conditions, even if the image at this moment is clearer than the historical image, it must meet the minimum quality requirements to be updated and overwrite the historical clear image. In the embodiments of the present application, based on the above steps, the final output can be a mask matrix with the same resolution as the original image. The value corresponding to the pixel point in the mask matrix is 1 or 0, where 1 indicates using the pixel value in the current fourth image as the final pixel value of the pixel point, and 0 indicates using the pixel value in the historical fifth image as the final pixel value of the pixel point.

[0123] In the embodiments of the present application, by using the prior information contained in the light transmittance matrix and combining with the historical image to quickly discriminate the relative clarity of different positions of the image, the clearer image is used as the model output. It is simpler than other methods such as deep learning models, greatly increasing the inference efficiency. Moreover, in the above steps, joint defogging and target detection can effectively solve the characteristics of "high concentration" and "high dynamic range" of the fog in the mine, ensuring the image output quality under harsh working conditions.

[0124] Figure 6 It is a schematic structural diagram of an underground vision enhancement device combining target recognition provided by the embodiments of the present application. As Figure 6 shown, the underground vision enhancement device 600 combining target recognition includes: a defogging module 601, a model detection module 602, a fusion processing module 603, and an output module 604.

[0125] The defogging module 601 is used to perform defogging operations on the original first image in the video stream to obtain a defogged fourth image and a first light transmittance matrix;

[0126] The model detection module 602 is used to preprocess the fourth image to obtain a preprocessed third image, and input the third image into the target detection model to output a target detection matrix;

[0127] The first fusion module 603 is used to obtain a fourth image according to the target detection matrix and the second image

[0128] The second fusion module 604 is configured to perform pixel value fusion on the historical fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix to obtain an enhanced target image, where the second light transmittance matrix is the light transmittance matrix corresponding to the fifth image.

[0129] In some embodiments, the second fusion module 604 is further configured to:

[0130] Determine the target pixel value of each pixel point from the pixel values of the fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix;

[0131] Determine the target image according to the target pixel value of each pixel point.

[0132] In some embodiments, the second fusion module 604 is further configured to:

[0133] For any pixel point i, determine the first light transmittance and the second light transmittance corresponding to the pixel point i from the first light transmittance matrix and the second light transmittance matrix respectively, where i is an integer greater than or equal to 1;

[0134] Judge whether the pixel point i satisfies the pixel update condition according to the first light transmittance and the second light transmittance corresponding to the pixel point i and a set light transmittance threshold;

[0135] In response to the pixel point i not satisfying the pixel update condition, determine the pixel value of the pixel point i in the fifth image as the target pixel value of the pixel point i;

[0136] In response to the pixel point i satisfying the pixel update condition, determine the pixel value of the pixel point i in the fourth image as the target pixel value of the pixel point i.

[0137] In some embodiments, the second fusion module 604 is further configured to:

[0138] Generate a mask matrix according to the first light transmittance matrix, the second light transmittance matrix, and a set light transmittance matrix;

[0139] Perform pixel value fusion on the fifth image and the fourth image according to the mask matrix to obtain an enhanced target image.

[0140] In some embodiments, the second fusion module 604 is further configured to:

[0141] For any pixel point i, determine the first light transmittance and the second light transmittance corresponding to the pixel point i from the first light transmittance matrix and the second light transmittance matrix respectively;

[0142] Based on the first light transmittance and the second light transmittance corresponding to the pixel point i, and the set light transmittance threshold, determine the value of the pixel point i in the mask matrix;

[0143] Generate the mask matrix according to the value of the pixel point i in the mask matrix.

[0144] In some embodiments, the second fusion module 604 is further configured to:

[0145] Based on the first light transmittance and the second light transmittance corresponding to the pixel point i, and the set light transmittance threshold, determine whether the pixel point i meets the pixel update condition;

[0146] In response to the pixel point i not meeting the pixel update condition, determine that the value of the pixel point i is 0, which is used to indicate that the pixel value of the pixel point i in the fifth image is the target pixel value of the pixel point i.

[0147] In response to the pixel point i meeting the pixel update condition, determine that the value of the pixel point i is 1, which is used to indicate that the pixel value of the pixel point i in the fourth image is the target pixel value of the pixel point i.

[0148] In some embodiments, the second fusion module 604 is further configured to:

[0149] Judge whether the first light transmittance is greater than the second light transmittance and whether the first light transmittance is greater than the set light transmittance threshold;

[0150] In response to the first light transmittance being greater than the second light transmittance and the first light transmittance being greater than the light transmittance threshold, determine that the pixel point i meets the pixel update condition;

[0151] In response to at least one of the conditions that the first light transmittance is less than or equal to the second light transmittance and the first light transmittance is less than or equal to the light transmittance threshold being satisfied, determine that the pixel point i does not meet the pixel update condition.

[0152] In some embodiments, the second fusion module 604 is further configured to:

[0153] Judge whether the pixel point i is within the target detection area detected by the target detection model;

[0154] In response to the pixel point i being within the target detection area, determine the pixel value of the pixel point i in the fourth image as the target pixel value of the pixel point i in the target image;

[0155] In response to the pixel point i not being within the target detection interval, the target pixel value of the pixel point i in the target image is determined by fusing the pixel values of the fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix of the pixel point i.

[0156] In some embodiments, the first fusion module 603 is further configured to:

[0157] Determine the recognition information of each pixel point in the second image in the target detection matrix, and mark the recognition information at the pixel point to obtain the fourth image.

[0158] In the embodiments of the present application, image dehazing and target detection can be combined, and the dehazed image is used as the recognition input of the target detection model to improve the recognition accuracy of the target recognition model in complex working conditions and ensure the reliability of the output image. Moreover, subsequent image enhancement is performed after target recognition, and during the subsequent image enhancement process, the prior information contained in the light transmittance matrix is used to quickly determine the relative clarity of different positions of the image in combination with historical images, and the clearer image is used as the model output.

[0159] It should be noted that the foregoing explanation of the embodiment of the downhole vision enhancement method combined with target recognition also applies to an embodiment of a downhole vision enhancement device combined with target recognition, and will not be repeated here.

[0160] To implement the above embodiments, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0161] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are used to implement the method provided in the foregoing embodiments when executed by a processor.

[0162] To implement the above embodiments, the present application also provides a computer program product including a computer program, and the computer program implements the method provided in the foregoing embodiments when executed by a processor.

[0163] The collection, storage, use, processing, transmission, provision, and application of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0164] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0165] This application is expected to provide an implementation for users to selectively block the use or access of personal information data. That is, this application is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0166] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0167] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0168] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of this application includes additional implementations, where the functions can be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0169] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0170] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0171] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above-described embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0172] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing module, can exist separately physically for each unit, or two or more units can be integrated in one module. The above-mentioned integrated module can be implemented in the form of hardware, or can be implemented in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0173] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An underground vision enhancement method combined with target recognition, characterized in that The method includes: Performing defogging operation on the original first image in the video stream to obtain a defogged second image and a first transmittance matrix; Preprocessing the second image to obtain a preprocessed third image, and inputting the third image into a target detection model to output a target detection matrix; Obtaining a fourth image according to the target detection matrix and the second image; Performing pixel value fusion on the historical fifth image and the fourth image according to the first transmittance matrix and a second transmittance matrix to obtain an enhanced target image, where the second transmittance matrix is the transmittance matrix corresponding to the fifth image.

2. The method according to claim 1, wherein The performing pixel value fusion on the historical fifth image and the fourth image according to the first transmittance matrix and the second transmittance matrix to obtain an enhanced target image includes: Determining the target pixel value of each pixel point from the pixel values of the fifth image and the fourth image according to the first transmittance matrix and the second transmittance matrix; Determining the target image according to the target pixel value of each pixel point.

3. The method according to claim 2, characterized in that, The determining the target pixel value of each pixel point from the pixel values of the fifth image and the fourth image according to the first transmittance matrix and the second transmittance matrix includes: For any pixel point i, determining the first transmittance and the second transmittance corresponding to the pixel point i from the first transmittance matrix and the second transmittance matrix respectively, where i is an integer greater than or equal to 1; Judging whether the pixel point i meets the pixel update condition according to the first transmittance and the second transmittance corresponding to the pixel point i and a set transmittance threshold; In response to the pixel point i not meeting the pixel update condition, determining the pixel value of the pixel point i in the fifth image as the target pixel value of the pixel point i; In response to the pixel point i meeting the pixel update condition, determining the pixel value of the pixel point i in the fourth image as the target pixel value of the pixel point i.

4. The method according to claim 1, wherein The performing pixel value fusion on the historical fifth image and the fourth image according to the first transmittance matrix and the second transmittance matrix to obtain an enhanced target image includes: Generating a mask matrix according to the first transmittance matrix, the second transmittance matrix, and a set transmittance matrix; Performing pixel value fusion on the fifth image and the fourth image according to the mask matrix to obtain an enhanced target image.

5. The method according to claim 4, wherein The generating a mask matrix according to the first transmittance matrix, the second transmittance matrix, and a set transmittance matrix includes: For any pixel point i, determining the first transmittance and the second transmittance corresponding to the pixel point i from the first transmittance matrix and the second transmittance matrix respectively; Judging the value of the pixel point i in the mask matrix according to the first transmittance and the second transmittance corresponding to the pixel point i and a set transmittance threshold; Generating the mask matrix according to the value of the pixel point i in the mask matrix.

6. The method according to claim 5, characterized in that, The judging the value of the pixel point i in the mask matrix according to the first transmittance and the second transmittance corresponding to the pixel point i and a set transmittance threshold includes: Based on the first light transmittance and the second light transmittance corresponding to the pixel point i, and the set light transmittance threshold, determine whether the pixel point i meets the pixel update condition; In response to the pixel point i not meeting the pixel update condition, determine that the value of the pixel point i is 0, which is used to indicate that the pixel value of the pixel point i in the fifth image is the target pixel value of the pixel point i. In response to the pixel point i meeting the pixel update condition, determine that the value of the pixel point i is 1, which is used to indicate that the pixel value of the pixel point i in the fourth image is the target pixel value of the pixel point i.

7. The method according to claim 3 or 6, characterized in that, The determining whether the pixel point i meets the pixel update condition according to the first light transmittance and the second light transmittance corresponding to the pixel point i, and the set light transmittance threshold includes: Determine whether the first light transmittance is greater than the second light transmittance, and determine whether the first light transmittance is greater than the set light transmittance threshold; In response to the first light transmittance being greater than the second light transmittance and the first light transmittance being greater than the light transmittance threshold, determine that the pixel point i meets the pixel update condition; In response to at least one of the conditions that the first light transmittance is less than or equal to the second light transmittance and the first light transmittance is less than or equal to the light transmittance threshold being satisfied, determine that the pixel point i does not meet the pixel update condition.

8. The method according to claim 3 or 5, characterized in that The method further includes: Determine whether the pixel point i is within the target detection area detected by the target detection model; In response to the pixel point i being within the target detection area, determine the pixel value of the pixel point i in the fourth image as the target pixel value of the pixel point i in the target image; In response to the pixel point i not being within the target detection interval, then perform pixel value fusion on the fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix of the pixel point i, and determine the target pixel value of the pixel point i in the target image.

9. The method according to any one of claims 1-6, characterized in that, The obtaining of the fourth image according to the target detection matrix and the second image includes: Determine the recognition information of each pixel point in the second image in the target detection matrix, and mark the recognition information at the pixel point to obtain the fourth image.

10. An underground vision enhancement device combined with target recognition, characterized in that, The method includes: A defogging module, which is used to perform defogging operations on the original first image in the video stream to obtain a defogged second image and a first light transmittance matrix; A model detection module, which is used to preprocess the second image to obtain a preprocessed third image, and input the third image into the target detection model to output a target detection matrix; A first fusion model, which is used to obtain a fourth image according to the target detection matrix and the second image; A second fusion module, which is used to perform pixel value fusion on the historical fifth image and the fourth image according to the first light transmittance matrix and the second light transmittance matrix to obtain an enhanced target image, where the second light transmittance matrix is the light transmittance matrix corresponding to the fifth image.

Citation Information

Patent Citations

  • Traffic monitoring video real-time defogging method based on moving targets

    CN103747213A

  • Operation video real-time defogging enhancement method based on atmospheric scattering model

    CN110910319A

  • Two-stage dense fog image defogging method based on reasoning

    CN115187474A

  • Label classification method and device for multimedia data, equipment and medium

    CN117540306A

  • Image defogging method and device and electronic equipment

    CN117635492A