An X-ray security inspection system and method integrating visible light images, depth images, and terahertz images
Through the multi-sensor fusion system, the problem of insufficient model generalization capability and image quality affected by focal length in terahertz security check is solved, and the target recognition and part analysis with high accuracy are achieved, which improves the security check effect.
Patent Information
- Application Number
- CN202010413169.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-05-15
AI Technical Summary
The existing terahertz security technology has problems such as insufficient generalization capabilities of the model, impact of the focal length, and inaccurate target position information, resulting in a decrease in recognition accuracy.
A multi-sensor fusion system is adopted, including synchronization, registration and fusion of visible light, depth and terahertz images. The visible light image is used to segment the human body Mask to filter the terahertz image background, and combined with depth images to perform focal length filtering and target part analysis to improve image quality and recognition accuracy.
It improves the target recognition accuracy of terahertz security, enhances the generalization ability of the model, provides specific location information of the target object in the human body, and reduces the false alarm rate.
Smart Images

Figure CN113673548B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security inspection technology, and more particularly to a system and method that integrates multiple sensor information, which can improve the accuracy of security inspection solutions and improve the display effect. Background Art
[0002] Terahertz imaging technology can be used to achieve non-cooperative, fast, and non-intrusive security inspection of passengers, and has been widely used in scenarios such as customs and civil aviation airports. However, this technology currently has the following defects:
[0003] 1. The current mainstream method for detecting and identifying suspicious objects in terahertz images is a deep learning model, but the effect of the deep learning model is highly dependent on the image data used for training. In addition, the models used on actual devices are all pre-trained. However, in reality, the usage scenarios of the devices are very complex and diverse. In particular, the background of the images used for training the model may be very different from the background of the images collected when the device is actually used. This may cause the recognition accuracy of the pre-trained suspicious object detection model to drop significantly when actually used, that is, the generalization ability of the model is not high.
[0004] 2. Terahertz cameras commonly used to obtain terahertz images have a certain focal length range. When a person is within the focal length range, the image quality of the obtained terahertz image is good, and the accuracy of target recognition based on this terahertz image is relatively high. However, when a person is outside the focal length range, the image quality of the obtained terahertz image will deteriorate, and the accuracy of target recognition based on this terahertz image may also drop. When using a terahertz camera for large-scale security inspection, there are usually people within and outside the imaging focal length range in the field of view of the terahertz camera. When directly performing target recognition based on such a large-scale terahertz image, due to the presence of people outside the focal length, the accuracy of target recognition may drop.
[0005] 3. Since terahertz images are grayscale images and lack details, it is difficult to accurately divide human body parts based on such images. Therefore, for the detected target, information related to which specific part of the human body the target is hidden in cannot be accurately provided.
[0006] All of the above problems pose challenges to the popularization of terahertz security inspection technology. And solutions that rely solely on terahertz images do not provide good effects. Therefore, the present disclosure adopts a method of fusing multiple different sensor information, and uses the complementary characteristics between different information to effectively solve the above defects of pure terahertz security inspection technology. Summary of the Invention
[0007] To address the above deficiencies of terahertz security inspection technology, the present disclosure provides a security inspection system and method that fuses information from multiple sensors. The system and method can utilize the complementary features between different types of information to improve the accuracy of target recognition.
[0008] In a first aspect, the present disclosure provides a security inspection system, comprising:
[0009] A plurality of sensor modules configured to collect visible light images, depth images, and terahertz images of a scene including one or more target objects;
[0010] An image synchronization module configured to synchronize the collected visible light images, depth images, and terahertz images;
[0011] An image registration module configured to register the synchronized visible light images, depth images, and terahertz images;
[0012] An image fusion module configured to fuse the registered visible light images, depth images, and / or terahertz images; and
[0013] An identification module configured to identify target items carried by the target objects based on the fused images.
[0014] In one embodiment, the plurality of sensor modules may be configured to collect the visible light images, the depth images, and the terahertz images at the same moment.
[0015] In one embodiment, the plurality of sensor modules may be configured to collect visible light video streams, depth video streams, and terahertz video streams at different frame rates. And, the synchronization module may be configured to: select an image from one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and for each of the other two video streams, determine an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the selected image.
[0016] In one embodiment, the plurality of sensor modules may be configured to collect visible light video streams, depth video streams, and terahertz video streams at different frame rates. And, the synchronization module may be configured to: select an image from one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and for each of the other two video streams: determine an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the selected image; obtain the first N images and the last N images centered on the determined image; and select the image with the highest similarity to the selected image from the first N images, the determined image, and the last N images.
[0017] In one embodiment, the image fusion module may be configured to: segment the registered visible light image to obtain a mask of the target object in the registered visible light image; and use the mask of the target object to filter the background in the registered terahertz image.
[0018] In one embodiment, the image fusion module may be configured to: extract a mask of the target object located outside the focal length range of the terahertz image from the registered depth image, use the mask of the target object to remove the corresponding target object in the registered visible light image, and segment the visible light image after removing the corresponding target object to obtain a mask of the remaining target objects in the visible light image; and use the mask of the remaining target objects to filter the registered terahertz image.
[0019] In one embodiment, the recognition module is further configured to identify the target item carried by the target object based on the registered terahertz image; and the image fusion module may further be configured to: perform part analysis on the target object in the registered visible light image; and map the terahertz image including the target item to the visible light image after the part analysis of the target object.
[0020] In a second aspect, the present disclosure provides a security inspection method, including:
[0021] Collect visible light images, depth images, and terahertz images of a scene including one or more target objects;
[0022] Synchronize the collected visible light images, depth images, and terahertz images;
[0023] Register the synchronized visible light images, depth images, and terahertz images;
[0024] Fuse the registered visible light images, depth images, and / or terahertz images; and
[0025] Identify the target item carried by the target object based on the fused image.
[0026] In one embodiment, the collecting step may include: collecting the visible light image, the depth image, and the terahertz image at the same moment.
[0027] In one embodiment, the step of acquisition may include: acquiring a visible light video stream, a depth video stream, and a terahertz video stream at different frame rates. And, the step of synchronization may include: selecting an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is acquired at the lowest frame rate; and for each of the other two video streams, determining an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image.
[0028] In one embodiment, the step of acquisition may include: acquiring a visible light video stream, a depth video stream, and a terahertz video stream at different frame rates. And, the step of synchronization may include: selecting an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is acquired at the lowest frame rate; and for each of the other two video streams: determining an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image; acquiring the first N images and the last N images centered on the determined image; and selecting an image with the highest similarity to the one image among the first N images, the determined image, and the last N images.
[0029] In one embodiment, the step of fusion may include: segmenting the registered visible light image to obtain a mask of the target object in the registered visible light image; and using the mask of the target object to filter the background in the registered terahertz image.
[0030] In one embodiment, the step of fusion may include: extracting a mask of the target object in the registered depth image that is outside the focal range of the terahertz image, using the mask of the target object to remove the corresponding target object in the registered visible light image, and segmenting the visible light image after removing the corresponding target object to obtain a mask of the remaining target objects in the visible light image; and using the mask of the remaining target objects to filter the registered terahertz image.
[0031] In one embodiment, the step of recognition may include: recognizing the target item carried by the target object based on the registered terahertz image. And, the step of fusion may include: performing part parsing on the target object in the registered visible light image; and mapping the terahertz image containing the target item to the visible light image that has undergone part parsing of the target object.
[0032] In a third aspect, the present disclosure provides a computer-readable storage medium storing instructions which, when executed by a processor of a security inspection system, cause the processor to execute the method according to the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A block diagram of a security inspection system according to an embodiment of the present disclosure is shown;
[0034] Figure 2 A flowchart of a security inspection method according to an embodiment of the present disclosure is shown;
[0035] Figure 3 An example layout diagram of a plurality of sensor modules is shown;
[0036] Figure 4 An example process flow of an image synchronization process according to an embodiment of the present disclosure is shown;
[0037] Figure 5 An example of an image fusion process according to an embodiment of the present disclosure is shown;
[0038] Figure 6 Another example of an image fusion process according to an embodiment of the present disclosure is shown; and
[0039] Figure 7 Yet another example of an image fusion process according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are only for illustrative purposes and are not intended to limit the present invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to those of ordinary skill in the art that the present invention does not have to be practiced with these specific details. In other instances, well-known circuits, materials, or methods have not been described in detail to avoid obscuring the present invention.
[0041] Throughout the specification, references to "an embodiment", "embodiments", "an example" or "examples" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present disclosure. Thus, the phrases "in an embodiment", "in embodiments", "an example" or "examples" that appear throughout the specification do not necessarily all refer to the same embodiment or example. Additionally, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Further, those of ordinary skill in the art will understand that the drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0042] Figure 1 FIG. 100 is a block diagram of a security inspection system according to an embodiment of the present disclosure. As Figure 1 shown, the security inspection system 100 may include a plurality of sensors 1011 - 101 n , an image synchronization module 102, an image registration module 103, an image fusion module 104, and an identification module 105. The plurality of sensors 1011 - 101 n may be configured to collect visible light, depth, and terahertz images or video streams of a scene including one or more target objects. Thus, in some embodiments, the plurality of sensors 1011 - 101 n may at least include a visible light camera, a depth camera, and a terahertz camera. Herein, the target object may refer to a human body, etc., but the present disclosure is not limited thereto.
[0043] The visible light, depth, and terahertz images or video streams collected by the plurality of sensors 1011 - 101 n are transmitted to the image synchronization module 102 for image synchronization. In some embodiments, if the plurality of sensors 1011 - 101 n can be controlled to collect images at the same moment, then the collected visible light, depth, and terahertz images or video streams are synchronized. In this case, the synchronization module 102 may not perform additional processing on the collected visible light, depth, and terahertz images or video streams, but directly provide them to the image registration module 103. However, in some embodiments, the frame rates of the plurality of sensors 1011 - 101 n may be different or the plurality of sensors 1011 - 101 nImaging acquisition is performed at the same time. In this case, the image synchronization module 103 needs to process the acquired visible light, depth, and terahertz images or video streams to make them synchronized or approximately synchronized. Specifically, the image synchronization module 103 needs to match the frames in the other video streams that are most consistent with the video with the lowest frame rate, that is, find the frames of multiple sensors 1011 - 101 n whose video frames are closest to the actual imaging time. This process is image synchronization. Only the image registration transformation after synchronization is meaningful.
[0044] In actual use, visible light and depth images are generally acquired synchronously by the corresponding cameras themselves, so these two videos generally do not require additional synchronization operations. When the visible light camera is controlled by one control unit (which can also be a processor) for acquisition, and the terahertz camera is controlled by another control unit (which can also be a processor) for acquisition, these two video streams are independent of each other, and the frame rate of the visible light camera is generally set higher than that of the terahertz camera. For this situation, it is necessary to synchronize the terahertz image and the visible light image.
[0045] In the present disclosure, for the situation where it is impossible to control multiple sensors 1011 - 101 n (such as visible light cameras, depth cameras, and terahertz cameras) to perform synchronous acquisition at the same frame rate, a method of performing synchronization in post-processing is proposed. In some embodiments, this post-processing can be executed in the image synchronization module 102. Taking the synchronization of two different types of images (such as terahertz images and visible light images) as an example, control units A and B of two cameras that respectively control the acquisition of these two images are set in the image synchronization module 102, and one of the control units A and B is set as the main control unit, and the other is set as the secondary control unit. The main control unit accesses the secondary control unit through a network or other means and obtains the video stream acquired on the secondary control unit. The main control unit records and maintains the system time of the main control unit when each frame of image is obtained from the secondary control unit (the time series is A1, A2,... A n , where n represents the number of the nth frame of image obtained), and binds this system time to the image one by one. At the same time, the main control unit also records and maintains the system time of the main control unit when each frame of image is obtained by the camera it controls (the time series is B1, B2,... B m , where m represents the number of the mth frame of image obtained), and binds this time to the acquired image one by one. Due to the transmission delay, the system time of the main control unit when each frame of image is obtained from the secondary control unit is actually later than the actual shooting time of the camera on the secondary control unit. This transmission delay can be represented as ΔT, and its value can be determined by experimental attempts. Therefore, A rj = Aj -ΔT can represent the time when the j-th picture is actually taken. For the convenience of description, assume that the camera frame rate controlled by the control unit A is relatively low. Then, for the j-th image, in the time series B1,..., B m Find an image captured at a time closest to its acquisition time A rj as the image synchronized with it. Alternatively, in some embodiments, for the j-th image, in the time series B1,..., B m Find an image captured at a time closest to its acquisition time A rj as the reference matching image. Taking the position of the reference matching image as the center point, match N frames before and after it respectively, and select the frame with the highest similarity to image A among these 2N + 1 frames rj as the optimal synchronization frame. Repeat this operation for each frame of image A rm Finally, the synchronization between the two video streams can be achieved.
[0046] The synchronized images can be provided to the image registration module 103. Since the multiple sensors 1011 - 101 n have different types, the significance of image registration is to make the pixel points at the same coordinates in the multiple different registered images point to the same imaging area of the target object (for example, a human body). There are many mature technologies for the registration of visible light and depth images, and many commercially available visible light cameras with depth information have already been pixel-level registered. Therefore, it will not be elaborated here. Although terahertz images are quite different from the general images we usually see, the terahertz camera is similar to the ordinary camera in the imaging optical path. Therefore, the registration of terahertz camera images and visible light images can also be carried out in the same way as that of ordinary cameras, such as the binocular calibration method, etc. Therefore, the technical details of the registration of terahertz images and visible light images will not be elaborated in this article.
[0047] After being registered by the image registration module 103, the visible light image, depth image, and terahertz image are pixel-aligned. The pixel-aligned visible light image, depth image, and terahertz image can be provided to the image fusion module 104 for image fusion. The first problem solved by image fusion is that there may be a large difference between the background of the terahertz image used in the training of the target detection model and the background of the terahertz image obtained during the actual use of the device, resulting in a reduction in the performance of the pre-trained target detection model. Using the image segmentation algorithm of deep learning, the human portrait Mask can be well segmented from the visible light, and the visible light image has been pixel-aligned with the terahertz image. Therefore, using the human body image Mask segmented from the visible light as a template to map the terahertz image can filter out the background area outside the human body in the terahertz image. By this means, it can be ensured that the background of the image used when pre-training the terahertz image target detection model is consistent with the background of the image detected during the actual use of the model, improving the scene adaptation ability of the target detection model. Since the depth image represents the distance from each point in space to the camera, the distance values of the human body area from the camera are relatively concentrated, and the human body area segmentation can also be achieved through the clustering algorithm. Therefore, replacing the above-mentioned visible light with the depth image after clustering can also achieve the background filtering of the terahertz image. However, the depth image is sensitive to the color of the human body's clothes. For example, the black clothes area may be regarded as the background due to its strong light absorption ability, resulting in a large missing area in the clustered human body area, making the mapped terahertz image also have obvious missing areas. Therefore, whether to use the depth image or the visible light image for the background filtering of the terahertz image can be determined according to the actual situation.
[0048] The second problem solved by image fusion is that in a terahertz image, some people may be within the imaging focal length range while others are outside it. Although the human portrait Mask segmented using visible light can filter the background of the terahertz image well, the images of the human body outside the imaging focal length will be very blurred, and target detection in this case will also increase the false alarm rate of the image. Since the depth image obtains the distance from the target to the camera, and the working distance of the depth camera can cover the imaging focal length range of the terahertz device, the distance information of the target in the depth image can be used to filter the human portraits outside the imaging focal length in the terahertz image. When the human body is far from the depth camera, the depth camera has a weak imaging ability for the edge parts of the human body. Compared with the human body Mask segmented from visible light, the edge parts will be missing in the human body Mask obtained by clustering and segmenting the data in the depth image. Therefore, if the human body Mask obtained directly from the depth image is used to post-process the Mask image segmented from visible light to filter the human body outside the focal length range, the edge parts of the Mask that need to be filtered will not be completely filtered, and finally the display effect of the terahertz image will be affected. When using a deep learning algorithm to segment human portraits from visible light, if the portrait part in the original image is severely missing, the remaining portrait part will be treated as the background. Using this feature, the present disclosure proposes an image fusion method, which can completely filter the human portraits outside the focal length range and completely retain the human portrait area within the focal length range. The image fusion method can be implemented via the image fusion module 104. The specific approach is to first use the depth image to segment and retain some of the human body Mask outside the focal length range; use the Mask outside the focal length range to remove the corresponding human portrait area in the visible light image; use the image segmentation algorithm of deep learning to segment the visible light image with some areas removed for human body Mask segmentation, so that the human body pixels connected to the removed areas will be treated as the background, and other unprocessed human portraits will be normally extracted as the human body Mask.
[0049] The third problem that image fusion needs to solve is that the position information of the alarm target in the terahertz image is not accurate enough. Since the terahertz image is a grayscale image with less detailed information, directly using the human body part parsing algorithm to parse the human body parts in the terahertz image has a poor effect. Because the registration of visible light and terahertz images has been achieved, the human body part parsing algorithm can be used to parse the human portrait in the visible light, and then map the pixel positions of the target items in the terahertz image to the visible light image, and finally obtain the alarm information about the specific parts of the human body where the target items are located. By determining the position information of the detected target items on the human body, the false alarm rate of target item detection can be further reduced, and more abundant security inspection information can also be provided.
[0050] In some embodiments, based on the images fused by the image fusion module 104, a target object (e.g., a human body) can be identified via an identification module (e.g., a neural network model) to determine the target item (e.g., a contraband) carried by the target object.
[0051] Figure 2 FIG. 200 shows a flowchart of a security inspection method 200 according to an embodiment of the present disclosure.
[0052] For the sake of convenience in explaining and describing the security inspection method 200, here, Rgb-A is used to represent a visible light camera or its image, Depth-B is used to represent a depth camera or its image, and Tera-C is used to represent a terahertz camera or its image. Regarding this security inspection method, Figure 3 FIG. 3 shows an example layout diagram of several sensors (e.g., cameras), but the applicability of the present invention is not limited to a specific layout scheme.
[0053] The method 200 includes, at step 201, respectively collecting a visible light image / video stream, a depth image / video stream, and a terahertz image / video stream through, for example, a visible light camera (e.g., an Rgb-A camera), a depth camera (e.g., a Depth-B camera), and a terahertz camera (e.g., a Tera-C camera).
[0054] Next, at step 202, the collected Rgb-A images, Depth-B images, and Tera-C images are synchronized. When the three cameras can be controlled by the same processor to collect images at the same time, the obtained image / video streams are directly synchronized, and this situation is not specifically discussed here. In some cases, the terahertz camera is independently controlled by a control unit to collect images, while the visible light and depth images are controlled by another control unit to collect images. This application mainly describes the video synchronization method in this case. Assume that the control unit for controlling Rgb-A and Depth-B is Processor-A, and the control unit for controlling Tera-C is Processor-B. Processor-A and Processor-B communicate with each other, and Processor-A serves as the main control unit, and Processor-B serves as the secondary control unit. The main control unit accesses the secondary control unit through a network or other means to obtain the terahertz images collected on the secondary control unit. Assume that every time the main control unit collects a visible light and a depth image, it obtains the system time (t11, t12,..., t1) of the current main control unit. j), where j represents the j-th acquired image. Since the visible light image and the depth image are hardware-synchronized, only one time series is required. The system time can be accurate to milliseconds. Similarly, the main control unit records the system time (t21, t22,..., t2) of the current main control unit each time Tera-C successfully acquires a terahertz image by accessing the secondary control unit. i ), where i represents the image sequence number. Assume that the terahertz camera has a lower frame rate than the visible light camera, and the acquisition control programs of the terahertz camera and the visible light camera images are independent of each other. Due to the transmission delay in the communication between the main control unit and the secondary control unit, the actual system time t2 i has a time delay of ΔT from the actual shooting time of the i-th picture. Since the transmission process and the shooting process are not completely stable and invariant, ΔT is an uncertain value, and its value can be determined by observing experiments, and this value does not need to be very accurate. The time obtained by subtracting ΔT from the recorded system time when acquiring terahertz can be approximately used as the shooting time of the terahertz image.
[0055] Maintain three queues on the main control unit, which are used to separately store the visible light, depth images, and the corresponding time series (t11, t12,..., t1 j ,...). For the i-th acquired terahertz image, first calculate the actual acquisition time of this terahertz image using t2 i - ΔT, and then find the time t1 m in the visible light time series that is closest to this actual acquisition time. The visible light / depth image obtained at this t1 m can be used as the image synchronized with the terahertz image. Alternatively, with t1 m as the center point, select N frames before and after it respectively, where N can be set as the ratio of the frame rate of the visible light image to the frame rate of the terahertz image. Cluster and binarize the selected 2N + 1 visible light / depth images, that is, set the pixels belonging to the human body part to 1 and the pixels in the area outside the human body to 0. Calculate the similarity between the current terahertz image and the 2N + 1 binarized visible light / depth images respectively. Some conventional methods for calculating the correlation between grayscale images can be used for the measurement of similarity, such as the mutual information method (MI) of images. Select a visible light / depth image with the highest mutual information with the current terahertz image as the synchronized image output. Clear the content before the optimal matching item in the above-mentioned three maintained queues, and then perform the same synchronization operation on the next acquired terahertz image t2 i+1 - ΔT. Figure 4 shows the flowchart of the video synchronization process.
[0056] Next, at step 203, the synchronized images are registered. In this article, it is assumed that the two cameras, Rgb-A and Depth-B, have already been image-registered by the hardware provider, and these two cameras are synchronized in hardware when collecting images. Therefore, only the image registration operation between Rgb-A and Tera-C needs to be performed. The image registration between Tera-C and Rgb-A can be completed according to the registration method of a conventional camera, such as the Surf registration method. The specific operation is that a person stands at M different distance positions within the focal range of Tera-C to collect still images, and ensure that Rgb-A captures the same images. Manually extract the human body in the terahertz image and the human body contour in the visible light image, and use the surf method to solve the registration transformation parameters between the two contour images.
[0057] Then, at step 204, the registered images are image-fused. Image fusion can solve three small problems. The first problem is to use the visible light image to segment the human body Mask to filter the background area in the terahertz image, so as to ensure that the background distribution of the training samples used to train the terahertz image target recognition model is consistent with the image background distribution detected by the target recognition model during actual use, thereby improving the environmental adaptability of the target recognition model. Figure 5 As an example of the image fusion process according to an embodiment of the present disclosure, where (a) is the registered terahertz image, (b) is the original visible light image after registration, (c) is the human body Mask obtained by segmenting the registered visible light image, and (d) is the image obtained by filtering the background of the terahertz image using the human body Mask in (c).
[0058] The second problem solved by image fusion is to completely filter the human body images outside the focal range in the terahertz image while retaining the integrity of the human body images within the focal range. To ensure the integrity of the segmented human body Mask, visible light image segmentation has an advantage over depth image clustering, but there is no depth information in the visible light image and distance filtering cannot be performed. To perform distance filtering, a depth image must be used, but using only the depth image cannot ensure the integrity of the segmented human body Mask. This article proposes to combine the visible light image and the depth image to achieve both ensuring the integrity of the segmented human body Mask and completely filtering the human body Mask outside the focal range. The specific approach is as follows: First, extract the human body Mask in the depth image outside the distance range, then use this Mask to remove the corresponding human body part in the corresponding visible light image, and finally use the visible light image segmentation algorithm to segment the human body image of the visible light image after removing the corresponding human body part. Figure 6Another example of the image fusion process according to an embodiment of the present disclosure is shown. Among them, (a) is the registered terahertz image, (b) is the corresponding human body Mask in the registered depth image, where the red color represents the human body pixel points outside the focal length range, (c) is the image obtained by removing the corresponding human body part in the visible light image using the human body Mask outside the focal length range in the depth image, and (d) is the result of performing human body image Mask segmentation on the visible light image after removing the corresponding human body part and mapping it to the terahertz image.
[0059] The third problem solved by image fusion is to provide alarm information about the hiding location of the target in the terahertz image. The specific approach is to perform target recognition and human body part analysis on the registered terahertz image and visible light image respectively, and then map the target position in the terahertz image to the image after human body analysis, so that the information about the hiding position of the target item in the terahertz image on the human body can be obtained. Figure 7 Another example of the image fusion process according to an embodiment of the present disclosure is shown. Among them, (a) is the registered terahertz image and the detection result of the target item, and (b) is the result of performing human body image analysis on the human body image in the registered visible light image and mapping the terahertz image containing the target item in (a) to the visible light image after human body image analysis.
[0060] Finally, method 200 may further include, at step 205, identifying the target item carried by the target object based on the fused image.
[0061] Beneficial effects
[0062] It can be seen from the specific embodiments that using the security inspection solution provided by the present disclosure has the following advantages: 1. Using this solution can keep the background of the terahertz image clean, so as to ensure that the distribution of the image background used in the terahertz image target detection model is consistent during the training and actual detection stages, and improve the generalization ability of the target detection model; 2. This solution can effectively alleviate the degradation of the image display quality caused by the human body being outside the imaging focal length, which helps to improve the accuracy of target item detection; 3. This solution can provide information related to the specific position of the target item on the human body.
[0063] The detailed description above has set forth numerous embodiments by way of illustration using schematic diagrams, flowcharts, and / or examples. In cases where such schematic diagrams, flowcharts, and / or examples contain one or more functions and / or operations, those skilled in the art will appreciate that each function and / or operation in such a schematic diagram, flowchart, or example can be implemented individually and / or jointly by a variety of structures, hardware, software, firmware, or substantially any combination thereof. In one embodiment, several portions of the subject matter of the embodiments of the present disclosure can be implemented by application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and the like. However, those skilled in the art will recognize that some aspects of the embodiments disclosed herein can be equivalently implemented, in whole or in part, in an integrated circuit, as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or substantially any combination of the above, and those skilled in the art will be capable of designing the circuitry and / or writing the software and / or firmware code based on the present disclosure. Additionally, those skilled in the art will recognize that the mechanisms of the subject matter of the present disclosure can be distributed as a variety of program products, and that the exemplary embodiments of the subject matter of the present disclosure apply regardless of the specific type of signal bearing medium actually used to carry out the distribution. Examples of signal bearing media include, but are not limited to: recordable media such as floppy disks, hard disk drives, compact discs (CDs), digital versatile discs (DVDs), digital magnetic tapes, computer memories, and the like; and transmission media such as digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).
[0064] While the present disclosure has been described with reference to several exemplary embodiments, it is to be understood that the terms used are illustrative and exemplary, rather than restrictive. Since the present disclosure can be embodied in many forms without departing from the spirit or essential characteristics thereof, it should be understood that the above embodiments are not limited to any of the foregoing details, but rather should be construed broadly within the spirit and scope defined by the appended claims, and accordingly, all variations and modifications that fall within the scope of the claims or their equivalents are intended to be covered by the appended claims.
Claims
1. An security inspection system that integrates visible light images, depth images, and terahertz images, comprising: Multiple sensor modules configured to collect visible light images, depth images, and terahertz images of a scene including one or more target objects; An image synchronization module configured to synchronize the collected visible light images, depth images, and terahertz images; An image registration module configured to register the synchronized visible light images, depth images, and terahertz images; An image fusion module configured to fuse the registered visible light images, depth images, and / or terahertz images; And An identification module configured to identify the target items carried by the target objects based on the fused images, Wherein, the image fusion module is configured to: Extract the Mask of the target objects in the registered depth image that are outside the focal range of the terahertz image, Use the Mask of the target objects to remove the corresponding target objects in the registered visible light image, and Segment the visible light image after removing the corresponding target objects to obtain the Mask of the remaining target objects in the visible light image; And Use the Mask of the remaining target objects to filter the registered terahertz image.
2. The security inspection system according to claim 1, wherein The multiple sensor modules are configured to: collect the visible light image, the depth image, and the terahertz image at the same moment.
3. The security inspection system according to claim 1, wherein, The multiple sensor modules are configured to: collect visible light video streams, depth video streams, and terahertz video streams at different frame rates, and Wherein, the synchronization module is configured to: Select an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and For each of the other two video streams, select an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image.
4. The security inspection system according to claim 1, wherein The multiple sensor modules are configured to: collect visible light video streams, depth video streams, and terahertz video streams at different frame rates, and Wherein, the synchronization module is configured to: Select an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and For each of the other two video streams: Determine an image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image; Obtain the first N images and the last N images centered on the determined image; and Select the image with the highest similarity to the one image among the first N images, the determined image, and the last N images.
5. The security inspection system according to any one of claims 1-4, wherein, The image fusion module is configured to: Segment the registered visible light image to obtain the Mask of the target objects in the registered visible light image; and Use the Mask of the target objects to filter the background in the registered terahertz image.
6. The security inspection system according to any one of claims 1-4, wherein, The identification module is further configured to: identify the target items carried by the target objects based on the registered terahertz image; and Wherein, the image fusion module is further configured to: Perform part analysis on the target object in the registered visible light image; and Map the terahertz image containing the target item to the visible light image after the part analysis of the target object.
7. A security inspection method that fuses visible light images, depth images, and terahertz images, including: Collect visible light images, depth images, and terahertz images of a scene containing one or more target objects; Synchronize the collected visible light images, depth images, and terahertz images; Register the synchronized visible light images, depth images, and terahertz images; Fuse the registered visible light images, depth images, and / or terahertz images; And Identify the target item carried by the target object based on the fused image, wherein the fusion step includes: Extract the mask of the target object in the registered depth image that is outside the focal length range of the terahertz image, Use the mask of the target object to remove the corresponding target object in the registered visible light image, and Segment the visible light image after removing the corresponding target object to obtain the mask of the remaining target objects in the visible light image; and Use the mask of the remaining target objects to filter the registered terahertz image.
8. The security inspection method according to claim 7, wherein, The collection step includes: Collect the visible light image, the depth image, and the terahertz image at the same time.
9. The security inspection method according to claim 7, wherein, The collection step includes: Collect visible light video streams, depth video streams, and terahertz video streams at different frame rates, and wherein the synchronization step includes: Select an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and For each of the other two video streams, select the image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image.
10. The security inspection method according to claim 7, wherein, The collection step includes: Collect visible light video streams, depth video streams, and terahertz video streams at different frame rates, and wherein the synchronization step includes: Select an image in one of the visible light video stream, the depth video stream, and the terahertz video stream that is collected at the lowest frame rate; and For each of the other two video streams: Determine the image in the corresponding video stream whose acquisition time is closest to the acquisition time of the one image; Obtain the first N images and the last N images centered on the determined image; and Select the image with the highest similarity to the one image among the first N images, the determined image, and the last N images.
11. The security inspection method according to any one of claims 7-10, wherein, The fusion step includes: Segment the registered visible light image to obtain the mask of the target object in the registered visible light image; and Use the mask of the target object to filter the background in the registered terahertz image.
12. The security inspection method according to any one of claims 7-10, wherein, The identification step includes: Identify the target item carried by the target object based on the registered terahertz image; and wherein the fusion step includes: Perform part analysis on the target object in the registered visible light image; and Map the terahertz image containing the target item to the visible light image after the part analysis of the target object has been performed.
13. A computer-readable storage medium storing instructions that, when executed by a processor of a security inspection system, cause the processor to perform the method according to any one of claims 7-12.
Citation Information
Patent Citations
Image processing apparatus and method thereof
CN109492714A