Liquid recognition method based on multimodal fusion images of security inspection X-ray machine
By acquiring and fusing specific areas of X-ray images and atomic number images, a multimodal fusion image that is not affected by the shape of the container is formed, which solves the problem of inaccurate liquid type recognition results and achieves higher recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202411723917.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the prior art, since liquids are usually contained in containers of different shapes, there is a large difference between high-energy and low-energy images and atomic number images, which affects the accuracy of liquid type recognition results.
By scanning the container containing the liquid to be tested, high-energy X-ray images, low-energy X-ray images and pseudo-color X-ray images are obtained, the center line and the mid-vertical line of the container are determined, and the multimodal fusion image is determined based on these areas. The high-energy X-ray image, low-energy X-ray image and effective atomic number image are fused to form a multimodal fusion image that is not affected by the shape of the container and is used for liquid type identification.
The accuracy of liquid type recognition is improved, the influence of container shape on recognition results is reduced, and recognition efficiency is improved.
Smart Images

Figure CN119741572B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security inspection technology, and in particular to a liquid identification method based on multimodal fusion images of a security inspection X-ray machine. Background Art
[0002] In the field of modern security testing, liquid detection is an important task, especially in public places such as airports and stations, where rapid and accurate identification of liquid hazardous materials is particularly important.
[0003] In the existing technology, high-energy images and low-energy images are generated by X-ray irradiation of the target object, and an atomic number image is generated from the high-energy and low-energy images. Then, a liquid target area mask is obtained through a segmentation algorithm, and the high- and low-energy images and atomic number images of the mask area are input into a pre-trained liquid model to achieve the purpose of liquid classification.
[0004] However, since liquids are usually stored in containers, and the shapes of different containers vary greatly, even if bottles of different shapes contain the same type of liquid, the high- and low-energy images and atomic number images obtained will be different. This difference will affect the results of liquid type identification, resulting in inaccurate liquid type identification results. Summary of the Invention
[0005] The present invention provides a liquid identification method based on multimodal fusion images of a security inspection X-ray machine, which is used to solve the defect of inaccurate liquid type identification results in the prior art and achieve the purpose of improving the accuracy of liquid type identification results.
[0006] The present invention provides a liquid identification method based on multimodal fusion images of a security inspection X-ray machine, comprising:
[0007] Scan the container containing the liquid to be tested to obtain high-energy X-ray images, low-energy X-ray images and pseudo-color X-ray images;
[0008] Determining a first-type cross region corresponding to the container in the pseudo-color X-ray image, where the first-type cross region includes a region where a center line of the container is located and a region where a perpendicular bisector of the center line is located;
[0009] determining an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image;
[0010] Based on the first type of cross region, determining a second type of cross region in the high-energy X-ray image, a third type of cross region in the low-energy X-ray image, and a fourth type of cross region in the effective atomic number image;
[0011] determining a multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region;
[0012] The liquid type of the liquid to be detected is identified based on the multimodal fusion image.
[0013] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, the step of determining a first-type cross region corresponding to the container in the pseudo-color X-ray image includes:
[0014] Inputting the pseudo-color X-ray image into an image segmentation model to obtain a mask region representing the container in the pseudo-color X-ray image output by the image segmentation model;
[0015] determining a centerline of the container in the mask area;
[0016] Extending the center line by a first preset number of pixels in a direction perpendicular to the center line to obtain an area where the center line is located;
[0017] Extending the perpendicular bisector by a second preset number of pixels in a direction perpendicular to the perpendicular bisector to obtain an area where the perpendicular bisector is located;
[0018] The first type of cross area is determined based on the area where the center line is located and the area where the perpendicular bisector is located.
[0019] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, the pseudo-color X-ray image is input into an image segmentation model to obtain a mask area representing the container in the pseudo-color X-ray image output by the image segmentation model, including:
[0020] Inputting the pseudo-color X-ray image into the backbone network of the image segmentation model to obtain at least two image features output by the backbone network;
[0021] Inputting each of the image features into the fusion network of the image segmentation model to obtain fusion features output by the fusion network;
[0022] Inputting the fused features into the target detection network and the mask extraction network of the image segmentation model respectively, to obtain the container features output by the target detection network and the mask features corresponding to the container output by the mask extraction network;
[0023] The mask region is obtained based on the container feature and the mask feature.
[0024] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, determining the center line of the container in the mask area includes:
[0025] performing a distance transform on the pixel value of each first pixel point in the mask region to obtain a transformed region, wherein the pixel value of each second pixel point in the transformed region is used to represent the minimum distance between the corresponding first pixel point in the mask region and a target boundary of the mask region, where the target boundary is the boundary closest to the corresponding pixel point;
[0026] The center line is determined based on a second pixel point corresponding to a maximum pixel value in the transformed area.
[0027] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, the multimodal fusion image is determined based on the second type of cross area, the third type of cross area, and the fourth type of cross area, including:
[0028] For each third pixel point in the second-type cross area, determining a normalized high and low energy coefficient based on a pixel value of the third pixel point and a pixel value of a fourth pixel point corresponding to the third pixel point in the third-type cross area;
[0029] Determining a pixel value of each pixel in the first channel based on the normalized high and low energy coefficients and the number of image data bits corresponding to the effective atomic number image;
[0030] determining a pixel value of a corresponding pixel point in a second channel based on a pixel value of each pixel point in the first channel and the number of bits of image data corresponding to the effective atomic number image;
[0031] The multimodal fusion image is determined based on the pixel value of each pixel in the first channel, the pixel value of the corresponding pixel in the second channel, and the pixel value of the corresponding pixel in the effective atomic number image.
[0032] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, determining the normalized high and low energy coefficients based on the pixel value of the third pixel point and the pixel value of the fourth pixel point corresponding to the third pixel point in the third type cross area includes:
[0033] Determine a first maximum value between a preset value and the pixel value of the third pixel point, wherein the preset value is used to represent the maximum value corresponding to normalization;
[0034] Determine a second maximum value between the preset value and the pixel value of the fourth pixel point;
[0035] The normalized high and low energy coefficients are determined based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high energy X-ray image, and the number of image data bits corresponding to the low energy X-ray image.
[0036] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, determining the normalized high- and low-energy coefficients based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high-energy X-ray image, and the number of image data bits corresponding to the low-energy X-ray image includes:
[0037] The normalized high and low energy coefficients are determined based on the following formula (1):
[0038] (1)
[0039] in, represents the normalized high and low energy coefficients, A represents the maximum pixel value of the pixel point under the image data bit number corresponding to the high energy X-ray image, B represents the first maximum value, C represents the maximum pixel value of the pixel point under the image data bit number corresponding to the low energy X-ray image, and D represents the second maximum value.
[0040] According to the present invention, a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine is provided, wherein the liquid type of the liquid to be detected is identified based on the multimodal fusion image, including:
[0041] Inputting the multimodal fusion image into a feature extraction module of an image classification model, dividing the multimodal fusion image into blocks by the feature extraction module, encoding each obtained image block, and performing feature extraction on each encoded image block to obtain each multimodal feature;
[0042] Each of the multimodal features is input into a classification module of the image classification model to obtain an identification result of the liquid type output by the classification module.
[0043] According to the present invention, a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine is provided, wherein a container containing a liquid to be detected is scanned to obtain a pseudo-color X-ray image, comprising:
[0044] Scan the container containing the liquid to be tested to obtain an original X-ray image;
[0045] The original X-ray image is colored using an image coloring algorithm to obtain the pseudo-color X-ray image.
[0046] According to a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine provided by the present invention, the multimodal fusion image is determined based on the second type of cross area, the third type of cross area, and the fourth type of cross area, including:
[0047] Rotating the second-type cross region, the third-type cross region, and the fourth-type cross region respectively to obtain a rotated second-type cross region, a rotated third-type cross region, and a rotated fourth-type cross region;
[0048] scaling the rotated second-type cross region, the rotated third-type cross region, and the rotated fourth-type cross region based on preset sizes to obtain scaled second-type cross region, scaled third-type cross region, and scaled fourth-type cross region;
[0049] The scaled second-type cross region, the scaled third-type cross region, and the scaled fourth-type cross region are fused to obtain the multimodal fusion image.
[0050] The present invention also provides a liquid identification device based on multimodal fusion images of a security inspection X-ray machine, comprising:
[0051] A scanning module is used to scan the container containing the liquid to be tested to obtain a high-energy X-ray image, a low-energy X-ray image and a pseudo-color X-ray image;
[0052] a determination module, configured to determine a first-type cross region corresponding to the container in the pseudo-color X-ray image, wherein the first-type cross region includes a region where a center line of the container is located and a region where a perpendicular bisector of the center line is located;
[0053] The determination module is further configured to determine an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image;
[0054] The determining module is further configured to determine, based on the first type of cross region, a second type of cross region in the high-energy X-ray image, a third type of cross region in the low-energy X-ray image, and a fourth type of cross region in the effective atomic number image;
[0055] The determining module is further configured to determine a multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region;
[0056] An identification module is used to identify the liquid type of the liquid to be detected based on the multimodal fusion image.
[0057] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements any of the above-described methods for liquid identification based on multimodal fusion images of a security inspection X-ray machine.
[0058] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine as described in any of the above is implemented.
[0059] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for liquid identification based on multimodal fusion images of a security inspection X-ray machine.
[0060] An embodiment of the present invention provides a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine. The method scans a container containing a liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image. A first-type cross area corresponding to the container in the pseudo-color X-ray image is determined, where the first-type cross area includes an area where the center line of the container and an area where the perpendicular bisector of the center line are located. An effective atomic number image is determined based on the high-energy X-ray image and the low-energy X-ray image. Based on the first-type cross area, a second-type cross area in the high-energy X-ray image, a third-type cross area in the low-energy X-ray image, and a fourth-type cross area in the effective atomic number image are determined. After a multimodal fusion image is determined based on the second-type cross area, the third-type cross area, and the fourth-type cross area, the liquid type of the liquid to be detected is identified based on the multimodal fusion image. Since the multimodal fusion image is obtained by fusing the cross-like regions extracted from the high-energy X-ray image, the low-energy X-ray image and the effective atomic number image, the multimodal fusion image is independent of the shape of the container. Therefore, when the liquid type of the liquid to be detected is identified by the multimodal fusion image, the identification result will no longer be affected by the shape of the container, thereby improving the accuracy of liquid type identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0062] Figure 1 A schematic flow chart of a liquid identification method based on multimodal fusion images of a security inspection X-ray machine provided in an embodiment of the present invention.
[0063] Figure 2 A schematic diagram of a high-energy X-ray image provided by an embodiment of the present invention.
[0064] Figure 3 A schematic diagram of a low-energy X-ray image provided by an embodiment of the present invention.
[0065] Figure 4 A schematic diagram of a pseudo-color X-ray image provided by an embodiment of the present invention.
[0066] Figure 5 A schematic diagram of the centerline of a container provided in an embodiment of the present invention.
[0067] Figure 6 A schematic diagram of a first type of cross region corresponding to a container provided in an embodiment of the present invention.
[0068] Figure 7 A schematic diagram of an effective atomic number image provided by an embodiment of the present invention.
[0069] Figure 8 A schematic diagram of a first type of cross region provided in an embodiment of the present invention.
[0070] Figure 9 This is a schematic diagram of the second type of cross area, the third type of cross area, and the fourth type of cross area determined based on the first type of cross area.
[0071] Figure 10 Schematic diagram of a scaled second-type cross region, a scaled third-type cross region, and a scaled fourth-type cross region provided by an embodiment of the present invention.
[0072] Figure 11 A schematic diagram of mask region segmentation provided by an embodiment of the present invention.
[0073] Figure 12 A schematic diagram of mask area distance transformation provided by an embodiment of the present invention.
[0074] Figure 13 This is a flowchart of a liquid identification method based on multimodal fusion images of a security inspection X-ray machine provided by an embodiment of the present invention.
[0075] Figure 14 A schematic diagram of the structure of a liquid identification device based on multimodal fusion images of a security inspection X-ray machine provided in an embodiment of the present invention.
[0076] Figure 15 A schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0077] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0078] In the prior art, a liquid recognition model is pre-trained, and after collecting high-energy and low-energy images of the container containing the liquid to be detected, an atomic number map is generated through the high-energy and low-energy images, and a liquid target area mask in the high-energy, low-energy and atomic number images is obtained through a segmentation algorithm. The high-energy, low-energy and atomic number images of the liquid target area mask are input into the pre-trained liquid recognition model to obtain the liquid type of the liquid to be detected output by the liquid recognition model. However, in actual applications, the shapes of containers are diverse, and the collected high-energy and low-energy images, as well as the atomic number maps obtained based on the high-energy and low-energy images, are all related to the shape of the container. Even if the container contains the same type of liquid, the atomic number maps obtained will be different when the container shape is different. Therefore, the pre-trained liquid recognition model cannot be applied to containers of all shapes, and the liquid type determined in the above manner is not accurate enough.
[0079] In response to the above problems, an embodiment of the present invention provides a liquid identification method based on the multimodal fusion image of a security inspection X-ray machine. Considering that no matter what the shape of the container is, it is usually cylindrical, and the high-energy image, low-energy image and effective atomic number image obtained for the cylinder all have the situation where the features of the area where the center line of the container and the area where the perpendicular bisector of the center line are located gradually change, therefore, no matter what the shape of the container is, the area where the center line and the area where the perpendicular bisector of the center line are located can be used to replace the overall liquid area. In this way, it is equivalent to normalizing containers of different shapes. When liquid type identification is performed based on the area where the center line and the area where the perpendicular bisector of the center line are located, the identification result will no longer be affected by the shape of the bottle, thereby improving the accuracy of liquid type identification.
[0080] The following combination Figures 1 to 13 The liquid identification method based on the multimodal fusion image of the security inspection X-ray machine provided in an embodiment of the present invention is described. The execution subject of this method can be an electronic device such as a security inspection machine, a computer or a server, or a specially designed intelligent device, or a liquid identification device based on the multimodal fusion image of the security inspection X-ray machine set in the electronic device or intelligent device. The liquid identification device based on the multimodal fusion image of the security inspection X-ray machine can be implemented by software, hardware or a combination of the two. This method can be applied to various scenarios where security inspections are required, such as bus stations, railway stations, subway stations or high-speed railway stations, to distinguish whether the liquid to be detected is a safe liquid such as water and beverages, or a flammable liquid such as alcohol and gasoline, or a chemical liquid such as sulfuric acid.
[0081] Figure 1 A flow chart of a liquid identification method based on multimodal fusion images of a security inspection X-ray machine provided in an embodiment of the present invention is shown as follows: Figure 1 As shown, the method includes:
[0082] Step 101: Scan a container containing a liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image.
[0083] In this step, Figure 2 A schematic diagram of a high-energy X-ray image provided by an embodiment of the present invention, Figure 3 A schematic diagram of a low-energy X-ray image provided by an embodiment of the present invention, Figure 4 A schematic diagram of a pseudo-color X-ray image provided by an embodiment of the present invention, such as Figure 2-4 As shown, when the container containing the liquid to be inspected passes through the security X-ray machine, the security X-ray machine uses high energy to scan the container and obtains the following Figure 2 The high-energy X-ray image Hi shown in the figure is obtained after the security inspection X-ray machine scans the container with low energy. Figure 3 The low-energy X-ray image Lo shown in FIG. 1 can be colored by an artificial coloring algorithm to obtain the following: Figure 4 The pseudo-color X-ray image RGB is shown.
[0084] Combine Figure 2-Figure 4 It can be seen that the color of the image data near the center axis of the container is darker, and the farther the image data deviates from the center axis, the lighter the color. The darker the image color, the fewer X-rays pass through, and the lighter the image color, the more X-rays pass through.
[0085] Step 102: Determine a first-type cross region corresponding to the container in the pseudo-color X-ray image. The first-type cross region includes a region where the center line of the container is located and a region where the perpendicular bisector of the center line is located.
[0086] In this step, Figure 5 A schematic diagram of the center line of a container provided in an embodiment of the present invention, such as Figure 5 As shown in FIG, by performing image recognition on the pseudo-color X-ray image, the liquid container mask area (mask) can be identified, and based on the position of each pixel point in the liquid container mask area (mask), the center line of the container can be determined, as shown in FIG. Figure 5 The center line of the container can also be understood as the central axis of the container.
[0087] Figure 6 A schematic diagram of the first type of cross region corresponding to the container provided in the embodiment of the present invention is shown as follows: Figure 6As shown, the area where the centerline is located can be obtained by expanding the area by a preset number of pixels in the left-right direction of the container, using the centerline as a reference. Furthermore, the perpendicular bisector of the centerline can be determined and, using the perpendicular bisector as a reference, expanding the area by a preset number of pixels in the vertical direction of the container to obtain the area where the perpendicular bisector is located. This preset number can be set based on actual conditions or experience, for example, based on the size of the container; the larger the container, the larger the preset number.
[0088] Step 103: Determine an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image.
[0089] In this step, due to the different atomic numbers of different elements, the attenuation characteristics of X-rays are also different. Elements with higher atomic numbers have a stronger attenuation effect on X-rays, while elements with lower atomic numbers have a weaker attenuation effect on X-rays. Therefore, the high-energy X-ray image and low-energy X-ray image obtained by X-rays passing through a container containing the liquid to be tested can be used to obtain X-ray attenuation information of different energies. Specifically, Figure 7 A schematic diagram of an effective atomic number image provided by an embodiment of the present invention, such as Figure 7 As shown, the relationship between the known attenuation coefficient and atomic number, as well as the ratio of the attenuation coefficients under high-energy X-ray images and low-energy X-ray images, can be used to determine the effective atomic number of each material area. The effective atomic number of each material area can be presented in the form of an image to obtain an effective atomic number image.
[0090] Step 104: Based on the first type of cross region, determine the second type of cross region in the high energy X-ray image, the third type of cross region in the low energy X-ray image, and the fourth type of cross region in the effective atomic number image.
[0091] In this step, Figure 8 A schematic diagram of the first type of cross region provided in an embodiment of the present invention, Figure 9 This is a schematic diagram of the second type of cross area, the third type of cross area, and the fourth type of cross area determined based on the first type of cross area, as shown in FIG. Figure 8-Figure 9 As shown, Figure 8 The outline of the first type of cross region in the pseudo-color X-ray image RGB is used as a template to extract the second type of cross region at the same position in the high-energy X-ray image Hi, the third type of cross region at the same position in the low-energy X-ray image Lo, and the fourth type of cross region at the same position in the effective atomic number image. Figure 9 From left to right, the second type of cross area, the third type of cross area, and the fourth type of cross area are shown in sequence.
[0092] Step 105: Determine a multimodal fusion image based on the second type of cross region, the third type of cross region, and the fourth type of cross region.
[0093] In this step, after obtaining the image blocks corresponding to the second, third, and fourth cross regions, the image data from these three image blocks are combined to produce a three-channel multimodal fusion image. It should be understood that this multimodal fusion image contains key information from the high-energy X-ray image, the low-energy X-ray image, and the effective atomic number image that can be used to determine the type of liquid being tested. The multimodal fusion image is equivalent to unifying or normalizing the image features corresponding to containers of different shapes.
[0094] For example, based on the above embodiments, due to the various placement positions of the containers, the shapes or angles of the second, third, and fourth cross regions used to characterize the container features are not uniform. To prevent errors in the liquid type recognition results due to angle differences or position differences, in this embodiment, the second, third, and fourth cross regions need to be corrected. When determining a multimodal fusion image based on the second, third, and fourth cross regions, the second, third, and fourth cross regions can be rotated to obtain rotated second, third, and fourth cross regions, respectively. The rotated second, third, and fourth cross regions are scaled based on a preset size to obtain scaled second, third, and fourth cross regions, respectively. The scaled second, third, and fourth cross regions are then fused to obtain a multimodal fusion image.
[0095] Specifically, Figure 10 Schematic diagram of the scaled second type cross region, the scaled third type cross region, and the scaled fourth type cross region provided in an embodiment of the present invention, as shown in FIG. Figure 10As shown, the second type of cross area, the third type of cross area and the fourth type of cross area are rotated respectively to ensure that the orientations of the rotated second type of cross area, the rotated third type of cross area and the rotated fourth type of cross area are consistent. The rotation process can align the features in each cross area with the features used in the training of the image classification model. Furthermore, the rotated second type of cross area, the rotated third type of cross area and the rotated fourth type of cross area are scaled based on the preset size to uniformly scale the different cross-like areas to the same size. For example, after the above-mentioned rotation and scaling process, the scaled second type of cross area, the scaled third type of cross area and the scaled fourth type of cross area of the same width and direction will be obtained, wherein the standard width can be, for example, 224×224 pixels. The scaled second type of cross area, the scaled third type of cross area and the scaled fourth type of cross area are then fused to obtain a three-channel multimodal fusion image. Among them, Figure 10 From left to right, the second type of cross area after scaling, the third type of cross area after scaling, and the fourth type of cross area after scaling are shown in sequence.
[0096] In this embodiment, by performing correction processing such as rotation and scaling on the second-type cross area, the third-type cross area, and the fourth-type cross area, the orientation and size of different cross-type areas can be unified. When performing image fusion, only the features of the pixel points at the same position need to be fused, and there will be no pixel confusion, thereby improving the accuracy of the multimodal fusion image.
[0097] Step 106: Identify the liquid type of the liquid to be detected based on the multimodal fusion image.
[0098] In this step, the multimodal fusion image is input into a pre-trained liquid classification model to obtain the liquid type recognition result output by the liquid classification model. The liquid classification model is obtained by training the initial liquid classification model based on the sample multimodal fusion image and liquid type labels.
[0099] An embodiment of the present invention provides a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine. The method scans a container containing a liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image. A first-type cross area corresponding to the container in the pseudo-color X-ray image is determined, where the first-type cross area includes an area where the center line of the container and an area where the perpendicular bisector of the center line are located. An effective atomic number image is determined based on the high-energy X-ray image and the low-energy X-ray image. Based on the first-type cross area, a second-type cross area in the high-energy X-ray image, a third-type cross area in the low-energy X-ray image, and a fourth-type cross area in the effective atomic number image are determined. After a multimodal fusion image is determined based on the second-type cross area, the third-type cross area, and the fourth-type cross area, the liquid type of the liquid to be detected is identified based on the multimodal fusion image. Since the multimodal fusion image is obtained by fusing the cross-like regions extracted from the high-energy X-ray image, the low-energy X-ray image and the effective atomic number image, the multimodal fusion image is independent of the shape of the container. Therefore, when the liquid type of the liquid to be detected is identified by the multimodal fusion image, the identification result will no longer be affected by the shape of the container, thereby improving the accuracy of liquid type identification.
[0100] Exemplarily, based on the above embodiment, when determining the first type of cross area corresponding to the container in the pseudo-color X-ray image, the pseudo-color X-ray image can be input into the image segmentation model to obtain a mask area representing the container in the pseudo-color X-ray image output by the image segmentation model, and the center line of the container in the mask area is determined, and the center line is extended by a first preset number of pixels in the vertical direction of the center line to obtain the area where the center line is located, and then the perpendicular bisector is extended by a second preset number of pixels in the vertical direction of the perpendicular bisector to obtain the area where the perpendicular bisector is located, and then the first type of cross area is determined based on the area where the center line is located and the area where the perpendicular bisector is located.
[0101] Specifically, the image segmentation model is a neural network model or any other model capable of image segmentation. After the pseudo-color X-ray image is input into the image segmentation model, the image segmentation model can generate a mask region of the liquid container in the pseudo-color X-ray image, that is, Figure 5 The mask region representing the container is shown in the green polygonal area.
[0102] In one possible implementation, when a pseudo-color X-ray image is input into an image segmentation model to obtain a mask region representing a container in the pseudo-color X-ray image output by the image segmentation model, the pseudo-color X-ray image can be input into a backbone network of the image segmentation model to obtain at least two image features output by the backbone network, and each image feature is input into a fusion network of the image segmentation model to obtain a fusion feature output by the fusion network, and the fusion feature is respectively input into a target detection network and a mask extraction network of the image segmentation model to obtain container features output by the target detection network and mask features corresponding to the container output by the mask extraction network, and then a mask region is obtained based on the container features and the mask features.
[0103] Specifically, Figure 11 A schematic diagram of mask region segmentation provided by an embodiment of the present invention is shown in FIG. Figure 11 As shown, when performing image segmentation on pseudo-color X-ray images, this embodiment uses an instance segmentation algorithm that combines a segmentation head and a detection head. After the pseudo-color X-ray image is input into the image segmentation model, feature extraction is performed through the backbone network to obtain at least two image features. The at least two image features are then input into the fusion network of the image segmentation model. Feature pyramid networks (FPN) are used to fuse feature maps of different scales to obtain fused features. The fused features are input into the target detection network and mask extraction network of the image segmentation model. While the target detection network outputs the detection box corresponding to the container, the mask extraction network can also output the mask features corresponding to the container. When extracting the mask region of the container, the container features corresponding to the container's detection box and the mask features corresponding to the container are merged, and then passed through a three-layer dynamic convolutional neural network to finally generate a segmented image and obtain the mask region of the container.
[0104] In this embodiment, the backbone network can extract multi-scale image features, which can improve the accuracy of subsequent liquid recognition based on these image features. Furthermore, by fusing multi-scale image features, a more comprehensive fused feature can be obtained. Based on the fused feature, container features and mask features can be extracted in parallel, improving the efficiency of determining the mask region.
[0105] After determining the mask area representing the container in the pseudo-color X-ray image, the skeleton of the container can be extracted through the image thinning algorithm to form a Figure 5 The centerline of the container in the mask area shown by the medium blue line.
[0106] Furthermore, if Figure 6As shown, the generated center line is extended by a first preset number of pixels on both sides in a direction perpendicular to the center line to form a center line strip mask, that is, the strip area where the center line is located. The first preset number can be set based on experience or the size of the container, for example, it can be set to 32. At the same time, a perpendicular median strip mask can also be constructed on the perpendicular median of the center line of the container. For example, the perpendicular median of the center line is extended by a second preset number of pixels on both sides in a direction perpendicular to the perpendicular median to form the strip area where the perpendicular median is located. The value of the second preset number can be the same as or different from the value of the first preset number. For example, it can also be set to 32. It should be understood that when the perpendicular median of the center line is extended, there is a gradient change in the liquid characteristics in the area of the perpendicular median of the container, which can characterize the characteristics of the liquid in the entire container.
[0107] By combining the long strip area where the center line is located and the long strip area where the perpendicular bisector is located, the first type of cross area in the pseudo-color X-ray image can be obtained.
[0108] In this embodiment, after the center line of the container in the mask area is determined, the center line is extended by a first preset number of pixels in the vertical direction of the center line, and the perpendicular bisector of the center is extended by a second preset number of pixels in the vertical direction of the perpendicular bisector to obtain a first type of cross area. The features in the first type of cross area can characterize the overall characteristics of the liquid and are not related to the shape of the container. Therefore, the liquid type of the liquid can be determined based on the features in the first type of cross area. This not only eliminates the interference of the container shape, but also reduces the amount of calculated data and improves the efficiency of liquid recognition.
[0109] Exemplarily, based on the above embodiments, when determining the center line of the container in the mask area, it is necessary to perform a distance transformation on the pixel values of each first pixel point in the mask area to obtain a transformation area. The pixel values of each second pixel point in the transformation area are used to represent the minimum distance between the corresponding first pixel point in the mask area and the target boundary of the mask area. The target boundary is the boundary closest to the corresponding pixel point, and the center line is determined based on the second pixel point corresponding to the maximum pixel value in the transformation area.
[0110] Specifically, the mask area representing the container is determined from the pseudo-color X-ray image, usually a binary image. Then, an image thinning algorithm is used to extract the center line or central axis of the mask area, which is also called a skeleton extraction algorithm. The skeleton removes some points in the original container image while still maintaining the original structural information of the container. Figure 12 A schematic diagram of mask area distance transformation provided by an embodiment of the present invention is shown in FIG. Figure 12As shown, a distance transform is first performed on the pixel values of each first pixel in the masked area. In the resulting transformed area image, the pixel value of each second pixel represents the minimum distance between the corresponding first pixel in the original masked area and the target boundary of the masked area. The target boundary is the boundary closest to the corresponding pixel, which can also be understood as the background. This distance can be Euclidean distance, Manhattan distance, or checkerboard distance, etc.
[0111] After obtaining the transformed region, based on the pixel values of each second pixel in the transformed region, the second pixel farthest from the target boundary (i.e., the second pixel corresponding to the maximum pixel value in the transformed region) is retained. The skeleton formed by combining these second pixel points is the centerline of the container. Alternatively, the second pixel farthest and second-to-farthest from the target boundary (i.e., the second pixel corresponding to the maximum and second-to-largest pixel values in the transformed region) can be retained and combined to obtain the centerline of the container.
[0112] In this embodiment, by performing a distance transformation on the pixel values of each first pixel point in the mask area, and determining the center line of the container based on the second pixel point corresponding to the maximum pixel value in the obtained transformed area, the center line can be determined quickly and accurately based on the above-mentioned skeleton extraction algorithm.
[0113] For example, based on the above embodiments, when determining the multimodal fusion image based on the second type of cross region, the third type of cross region, and the fourth type of cross region, the following method can be used:
[0114] For each third pixel point in the second type of cross area, the normalized high and low energy coefficients are determined based on the pixel value of the third pixel point and the pixel value of the fourth pixel point corresponding to the third pixel point in the third type of cross area, the pixel value of each pixel point in the first channel is determined based on the normalized high and low energy coefficients and the number of image data bits corresponding to the effective atomic number image, and after the pixel value of the corresponding pixel point in the second channel is determined based on the pixel value of each pixel point in the first channel and the number of image data bits corresponding to the effective atomic number image, a multimodal fusion image is determined based on the pixel value of each pixel point in the first channel, the pixel value of the corresponding pixel point in the second channel, and the pixel value of the corresponding pixel point in the effective atomic number image.
[0115] Specifically, since the number of image data bits of the collected high-energy X-ray images and low-energy X-ray images is different from that of the commonly used image channels, for example, the image data bits of high-energy X-ray images and low-energy X-ray images are both 16 bits, and the pixel value range of each pixel is 0~65535, while the image data of commonly used image channels (such as RGB images and effective atomic number images) is 8 bits, and the pixel value range of the pixel is 0~255. Therefore, normalization processing is required when performing multimodal data fusion on the second-type cross area, the third-type cross area, and the fourth-type cross area.
[0116] For each third pixel in the second cross region of the high-energy X-ray image, a normalized high- and low-energy coefficient can be determined based on the pixel value of the third pixel and the pixel value of a fourth pixel corresponding to the third pixel in the third cross region of the low-energy X-ray image. In one possible implementation, when determining the normalized high- and low-energy coefficients, a first maximum value between a preset value and the pixel value of the third pixel can be determined, where the preset value is used to represent the maximum value corresponding to normalization, and a second maximum value between the preset value and the pixel value of the fourth pixel can be determined. The normalized high- and low-energy coefficients are determined based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high-energy X-ray image, and the number of image data bits corresponding to the low-energy X-ray image.
[0117] Among them, the preset value can be 1, for example. Based on the number of image data bits corresponding to the high-energy X-ray image, the maximum value of the pixel value range of the pixel point in the high-energy X-ray image can be determined. Based on the number of image data bits corresponding to the low-energy X-ray image, the maximum value of the pixel value range of the pixel point in the low-energy X-ray image can be determined. Therefore, based on the first maximum value, the second maximum value, the maximum value of the pixel value range of the pixel point in the high-energy X-ray image, and the maximum value of the pixel value range of the pixel point in the low-energy X-ray image, the normalized high and low energy coefficients are determined.
[0118] Specifically, the normalized high and low energy coefficients can be determined based on the following formula (1):
[0119] (1)
[0120] in, represents the normalized high and low energy coefficients, A represents the maximum pixel value of the pixel under the image data bit number corresponding to the high energy X-ray image, B represents the first maximum value, C represents the maximum pixel value of the pixel under the image data bit number corresponding to the low energy X-ray image, and D represents the second maximum value.
[0121] Taking the preset value as 1 and the image data bits of both high-energy X-ray images and low-energy X-ray images as an example, the above-mentioned normalized high and low energy coefficients can be obtained by the following formula (2):
[0122] (2)
[0123] in, Represents the pixel value of the third pixel (i, j), Indicates the pixel value of the fourth pixel (i, j).
[0124] In this embodiment, the normalized high and low energy coefficients are determined by the pixel value of the third pixel point and the pixel value of the fourth pixel point corresponding to the third pixel point in the third type of cross area, so that the number of image data bits of the second type of cross area and the third type of cross area is normalized to the same number of image data bits as the commonly used image channel, eliminating the differences between different images, facilitating the subsequent fusion of multimodal data, and improving the accuracy of data fusion.
[0125] After determining the normalized high and low energy coefficients, the pixel value of each pixel in the first channel is determined by the normalized high and low energy coefficients and the number of image data bits corresponding to the effective atomic number image. Taking the image data bit number corresponding to the effective atomic number image as 8 bits and the maximum pixel value of the corresponding pixel as 255 as an example, the pixel value of the pixel point (i, j) in the first channel can be determined by the following formula (3): :
[0126] (3)
[0127] The pixel value of the pixel point (i, j) in the second channel is determined by the following formula (4): :
[0128] (4)
[0129] in, The value is 10. The value is 1.1.
[0130] In addition, the pixel value of the third channel of the multimodal fusion image is the pixel value of the corresponding pixel point in the effective atomic number image.
[0131] In this embodiment, by fusing the second type of cross area, the third type of cross area and the fourth type of cross area, determining the multimodal fusion image through the pixel values of different channels, and integrating the pixel values of multiple channels, more comprehensive image information can be provided, which helps to improve the accuracy of liquid type recognition.
[0132] Exemplarily, based on the above embodiments, when the liquid type of the liquid to be detected is identified based on the multimodal fusion image, the multimodal fusion image can be input into the feature extraction module of the image classification model, the multimodal fusion image can be divided into blocks by the feature extraction module, the obtained image blocks are encoded, and the features of each encoded image block are extracted to obtain each multimodal feature, and each multimodal feature is input into the classification module of the image classification model to obtain the recognition result of the liquid type output by the classification module.
[0133] Specifically, the image classification model uses the Vision Transformer (ViT) model. The entire Vision Transformer can be divided into two parts: a feature extraction module and a classification module. In the feature extraction module, ViT is used for feature extraction. Its corresponding regions in multimodal images are block segmentation plus position embedding and a Transformer Encoder. Block segmentation plus position embedding primarily partitions the input multimodal fusion image, dividing the image into blocks of a certain size. This is typically achieved by introducing a convolutional layer using a sliding window, such as a 16×16 convolution. The divided image blocks are then combined into an image sequence. Each image block is also encoded with a position and a class token. The class token is associated with the annotated image category during training. After obtaining the image sequence, it is passed to the Transformer Encoder for feature extraction, resulting in multimodal features. This is the unique multi-head self-attention structure of the Transformer, which uses the self-attention mechanism to focus on the importance of each image block.
[0134] In the classification module, the multimodal features extracted by the Transformer Encoder are input into a Multilayer Perceptron (MLP) classifier for classification, resulting in liquid type recognition results. These liquid types include safe liquids such as water and beverages, flammable and explosive liquids such as alcohol and gasoline, and hazardous chemicals such as sulfuric acid.
[0135] In this embodiment, the multimodal fusion image is segmented and encoded through the feature extraction module of the image classification model, which can capture the key information in the multimodal fusion image in more detail. Moreover, the local features in the multimodal fusion image can be identified through block processing, and the expression ability of the features can be enhanced through encoding, thereby improving the accuracy of the classification results when classifying liquid types through various multimodal features.
[0136] Illustratively, based on the above embodiments, when scanning a container containing the liquid to be detected to obtain a pseudo-color X-ray image, the container containing the liquid to be detected can be scanned to obtain an original X-ray image, and the original X-ray image can be colored using an image coloring algorithm to obtain a pseudo-color X-ray image.
[0137] Specifically, when an X-ray scanner is used to scan a container containing a liquid to be inspected, the resulting raw X-ray image is typically displayed in grayscale, where different grayscale values represent different material densities. When colorizing the raw X-ray image using an image colorization algorithm, different areas within the image are assigned colors based on the grayscale information within the image. This process enhances the visual quality of the image, making it easier to distinguish between different materials and structures within the pseudo-color X-ray image.
[0138] In this embodiment, the original X-ray image is colored by an image coloring algorithm, and the resulting pseudo-color X-ray image retains the density information of the original X-ray image, and the color change makes the image more intuitive, thereby improving the accuracy of the detected first-type cross area.
[0139] Figure 13 The flow chart of the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine provided in the embodiment of the present invention is as follows: Figure 13 As shown, after obtaining the high-energy X-ray image Hi, the low-energy X-ray image Lo, and the pseudo-color X-ray image RGB of the container containing the liquid to be detected, the pseudo-color X-ray image RGB is segmented to obtain the mask area to which the liquid container belongs in the pseudo-color X-ray image RGB. The center line of the container is extracted from the mask area, and then the skeleton extraction algorithm is used to determine the first type of cross area in the pseudo-color X-ray image RGB.
[0140] In addition, the effective atomic number image Zeff can be determined based on the high-energy X-ray image Hi and the low-energy X-ray image Lo, and based on the first-type cross area, the second-type cross area in the high-energy X-ray image Hi, the third-type cross area in the low-energy X-ray image Lo, and the fourth-type cross area in the effective atomic number image Zeff can be determined.
[0141] Multi-channel data fusion is performed on the second, third and fourth cross regions to obtain a multimodal fusion image, which is then input into the image classification model to obtain the liquid classification result.
[0142] An embodiment of the present invention provides a liquid identification method based on a multimodal fusion image of a security inspection X-ray machine. The method scans a container containing a liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image. A first-type cross area corresponding to the container in the pseudo-color X-ray image is determined, where the first-type cross area includes an area where the center line of the container and an area where the perpendicular bisector of the center line are located. An effective atomic number image is determined based on the high-energy X-ray image and the low-energy X-ray image. Based on the first-type cross area, a second-type cross area in the high-energy X-ray image, a third-type cross area in the low-energy X-ray image, and a fourth-type cross area in the effective atomic number image are determined. After a multimodal fusion image is determined based on the second-type cross area, the third-type cross area, and the fourth-type cross area, the liquid type of the liquid to be detected is identified based on the multimodal fusion image. Since the multimodal fusion image is obtained by fusing the cross-like regions extracted from the high-energy X-ray image, the low-energy X-ray image and the effective atomic number image, the multimodal fusion image is independent of the shape of the container. Therefore, when the liquid type of the liquid to be detected is identified by the multimodal fusion image, the identification result will no longer be affected by the shape of the container, thereby improving the accuracy of liquid type identification.
[0143] The following describes the liquid identification device based on the multimodal fusion image of the security inspection X-ray machine provided by the present invention. The liquid identification device based on the multimodal fusion image of the security inspection X-ray machine described below and the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine described above can be referenced to each other.
[0144] Figure 14 The structural diagram of the liquid identification device based on the multimodal fusion image of the security inspection X-ray machine provided in the embodiment of the present invention is shown in FIG. Figure 14 As shown, the liquid identification device 1400 based on the multimodal fusion image of the security inspection X-ray machine includes:
[0145] Scanning module 11, used to scan the container containing the liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image and a pseudo-color X-ray image;
[0146] a determination module 12, configured to determine a first type of cross region corresponding to the container in the pseudo-color X-ray image, wherein the first type of cross region includes a region where the center line of the container is located and a region where the perpendicular bisector of the center line is located;
[0147] The determining module 12 is further configured to determine an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image;
[0148] The determining module 12 is further configured to determine, based on the first type of cross region, a second type of cross region in the high-energy X-ray image, a third type of cross region in the low-energy X-ray image, and a fourth type of cross region in the effective atomic number image;
[0149] The determining module 12 is further configured to determine a multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region;
[0150] The identification module 13 is configured to identify the liquid type of the liquid to be detected based on the multimodal fusion image.
[0151] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0152] Inputting the pseudo-color X-ray image into an image segmentation model to obtain a mask region representing the container in the pseudo-color X-ray image output by the image segmentation model;
[0153] determining a centerline of the container in the mask area;
[0154] Extending the center line by a first preset number of pixels in a direction perpendicular to the center line to obtain an area where the center line is located;
[0155] Extending the perpendicular bisector by a second preset number of pixels in a direction perpendicular to the perpendicular bisector to obtain an area where the perpendicular bisector is located;
[0156] The first type of cross area is determined based on the area where the center line is located and the area where the perpendicular bisector is located.
[0157] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0158] Inputting the pseudo-color X-ray image into the backbone network of the image segmentation model to obtain at least two image features output by the backbone network;
[0159] Inputting each of the image features into the fusion network of the image segmentation model to obtain fusion features output by the fusion network;
[0160] Inputting the fused features into the target detection network and the mask extraction network of the image segmentation model respectively, to obtain the container features output by the target detection network and the mask features corresponding to the container output by the mask extraction network;
[0161] The mask region is obtained based on the container feature and the mask feature.
[0162] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0163] performing a distance transform on the pixel value of each first pixel point in the mask region to obtain a transformed region, wherein the pixel value of each second pixel point in the transformed region is used to represent the minimum distance between the corresponding first pixel point in the mask region and a target boundary of the mask region, where the target boundary is the boundary closest to the corresponding pixel point;
[0164] The center line is determined based on a second pixel point corresponding to a maximum pixel value in the transformed area.
[0165] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0166] For each third pixel point in the second-type cross area, determining a normalized high and low energy coefficient based on a pixel value of the third pixel point and a pixel value of a fourth pixel point corresponding to the third pixel point in the third-type cross area;
[0167] Determining a pixel value of each pixel in the first channel based on the normalized high and low energy coefficients and the number of image data bits corresponding to the effective atomic number image;
[0168] determining a pixel value of a corresponding pixel point in a second channel based on a pixel value of each pixel point in the first channel and the number of bits of image data corresponding to the effective atomic number image;
[0169] The multimodal fusion image is determined based on the pixel value of each pixel in the first channel, the pixel value of the corresponding pixel in the second channel, and the pixel value of the corresponding pixel in the effective atomic number image.
[0170] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0171] Determine a first maximum value between a preset value and the pixel value of the third pixel point, wherein the preset value is used to represent the maximum value corresponding to normalization;
[0172] Determine a second maximum value between the preset value and the pixel value of the fourth pixel point;
[0173] The normalized high and low energy coefficients are determined based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high energy X-ray image, and the number of image data bits corresponding to the low energy X-ray image.
[0174] In an exemplary embodiment, the determining module 12 is specifically configured to:
[0175] The normalized high and low energy coefficients are determined based on the following formula (1):
[0176] (1)
[0177] in, represents the normalized high and low energy coefficients, A represents the maximum pixel value of the pixel point under the image data bit number corresponding to the high energy X-ray image, B represents the first maximum value, C represents the maximum pixel value of the pixel point under the image data bit number corresponding to the low energy X-ray image, and D represents the second maximum value.
[0178] In an exemplary embodiment, the identification module 13 is specifically configured to:
[0179] Inputting the multimodal fusion image into a feature extraction module of an image classification model, dividing the multimodal fusion image into blocks by the feature extraction module, encoding each obtained image block, and performing feature extraction on each encoded image block to obtain each multimodal feature;
[0180] Each of the multimodal features is input into a classification module of the image classification model to obtain an identification result of the liquid type output by the classification module.
[0181] In an exemplary embodiment, the scanning module 11 is specifically configured to:
[0182] Scan the container containing the liquid to be tested to obtain an original X-ray image;
[0183] The original X-ray image is colored using an image coloring algorithm to obtain the pseudo-color X-ray image.
[0184] In an exemplary embodiment, the determination module 12 is specifically configured to:
[0185] Rotating the second-type cross region, the third-type cross region, and the fourth-type cross region respectively to obtain a rotated second-type cross region, a rotated third-type cross region, and a rotated fourth-type cross region;
[0186] scaling the rotated second-type cross region, the rotated third-type cross region, and the rotated fourth-type cross region based on preset sizes to obtain scaled second-type cross region, scaled third-type cross region, and scaled fourth-type cross region;
[0187] The scaled second-type cross region, the scaled third-type cross region, and the scaled fourth-type cross region are fused to obtain the multimodal fusion image.
[0188] The device of this embodiment can be used to execute the method of any embodiment in the embodiment of the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine. Its specific implementation process and technical effects are similar to those in the embodiment of the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine. For details, please refer to the detailed introduction in the embodiment of the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine, which will not be repeated here.
[0189] Figure 15 A schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 15 As shown, the electronic device may include: a processor (processor) 1510 , a communication interface (Communications Interface) 1520 , a memory (memory) 1530 and a communication bus 1540 , wherein the processor 1510 , the communication interface 1520 , and the memory 1530 communicate with each other via the communication bus 1540 . The processor 1510 can call the logic instructions in the memory 1530 to execute a liquid identification method based on the multimodal fusion image of the security inspection X-ray machine, the method including: scanning a container containing the liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image; determining a first-type cross area corresponding to the container in the pseudo-color X-ray image, the first-type cross area including the area where the center line of the container is located and the area where the perpendicular bisector of the center line is located; determining an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image; determining a second-type cross area in the high-energy X-ray image, a third-type cross area in the low-energy X-ray image, and a fourth-type cross area in the effective atomic number image based on the first-type cross area; determining a multimodal fusion image based on the second-type cross area, the third-type cross area, and the fourth-type cross area; and identifying the liquid type of the liquid to be detected based on the multimodal fusion image.
[0190] Furthermore, the logic instructions in the aforementioned memory 1530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0191] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine provided by the above-mentioned methods, the method including: scanning a container containing a liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image; determining a first type of cross area corresponding to the container in the pseudo-color X-ray image, the first type of cross area including an area where the centerline of the container is located and an area where the perpendicular bisector of the centerline is located; determining an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image; based on the first type of cross area, determining a second type of cross area in the high-energy X-ray image, a third type of cross area in the low-energy X-ray image, and a fourth type of cross area in the effective atomic number image; determining a multimodal fusion image based on the second type of cross area, the third type of cross area, and the fourth type of cross area; and identifying the liquid type of the liquid to be detected based on the multimodal fusion image.
[0192] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the liquid identification method based on the multimodal fusion image of the security inspection X-ray machine provided by the above-mentioned methods, the method comprising: scanning a container containing the liquid to be detected to obtain a high-energy X-ray image, a low-energy X-ray image, and a pseudo-color X-ray image; determining a first type of cross area corresponding to the container in the pseudo-color X-ray image, the first type of cross area including an area where the center line of the container is located and an area where the perpendicular bisector of the center line is located; determining an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image; determining a second type of cross area in the high-energy X-ray image, a third type of cross area in the low-energy X-ray image, and a fourth type of cross area in the effective atomic number image based on the first type of cross area; determining a multimodal fusion image based on the second type of cross area, the third type of cross area, and the fourth type of cross area; and identifying the liquid type of the liquid to be detected based on the multimodal fusion image.
[0193] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0194] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A liquid identification method based on multimodal fusion images of security inspection X-ray machines, characterized in that: include: Scan the container containing the liquid to be tested to obtain high-energy X-ray images, low-energy X-ray images and pseudo-color X-ray images; Determining a first-type cross region corresponding to the container in the pseudo-color X-ray image, where the first-type cross region includes a region where a center line of the container is located and a region where a perpendicular bisector of the center line is located; determining an effective atomic number image based on the high-energy X-ray image and the low-energy X-ray image; Based on the first type of cross region, determining a second type of cross region in the high-energy X-ray image, a third type of cross region in the low-energy X-ray image, and a fourth type of cross region in the effective atomic number image; determining a multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region; Identifying the liquid type of the liquid to be detected based on the multimodal fusion image; The determining of the multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region includes: For each third pixel point in the second-type cross area, determining a normalized high and low energy coefficient based on a pixel value of the third pixel point and a pixel value of a fourth pixel point corresponding to the third pixel point in the third-type cross area; Determining a pixel value of each pixel in the first channel based on the normalized high and low energy coefficients and the number of image data bits corresponding to the effective atomic number image; determining a pixel value of a corresponding pixel point in a second channel based on a pixel value of each pixel point in the first channel and the number of bits of image data corresponding to the effective atomic number image; determining the multimodal fusion image based on a pixel value of each pixel in the first channel, a pixel value of a corresponding pixel in the second channel, and a pixel value of a corresponding pixel in the effective atomic number image; The determining of the normalized high and low energy coefficients based on the pixel value of the third pixel and the pixel value of a fourth pixel corresponding to the third pixel in the third type cross area includes: Determine a first maximum value between a preset value and the pixel value of the third pixel point, wherein the preset value is used to represent the maximum value corresponding to normalization; Determine a second maximum value between the preset value and the pixel value of the fourth pixel point; The normalized high and low energy coefficients are determined based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high energy X-ray image, and the number of image data bits corresponding to the low energy X-ray image.
2. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 1 is characterized in that: Determining a first type of cross region corresponding to the container in the pseudo-color X-ray image includes: Inputting the pseudo-color X-ray image into an image segmentation model to obtain a mask region representing the container in the pseudo-color X-ray image output by the image segmentation model; determining a centerline of the container in the mask area; Extending the center line by a first preset number of pixels in a direction perpendicular to the center line to obtain an area where the center line is located; Extending the perpendicular bisector by a second preset number of pixels in a direction perpendicular to the perpendicular bisector to obtain an area where the perpendicular bisector is located; The first type of cross area is determined based on the area where the center line is located and the area where the perpendicular bisector is located.
3. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 2 is characterized in that: Inputting the pseudo-color X-ray image into an image segmentation model to obtain a mask region representing the container in the pseudo-color X-ray image output by the image segmentation model includes: Inputting the pseudo-color X-ray image into the backbone network of the image segmentation model to obtain at least two image features output by the backbone network; Inputting each of the image features into the fusion network of the image segmentation model to obtain fusion features output by the fusion network; Inputting the fused features into the target detection network and the mask extraction network of the image segmentation model respectively, to obtain the container features output by the target detection network and the mask features corresponding to the container output by the mask extraction network; The mask region is obtained based on the container feature and the mask feature.
4. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 2 is characterized in that: Determining the center line of the container in the mask area includes: performing a distance transform on the pixel value of each first pixel point in the mask region to obtain a transformed region, wherein the pixel value of each second pixel point in the transformed region is used to represent the minimum distance between the corresponding first pixel point in the mask region and a target boundary of the mask region, where the target boundary is the boundary closest to the corresponding pixel point; The center line is determined based on a second pixel point corresponding to a maximum pixel value in the transformed area.
5. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 1 is characterized in that: The determining of the normalized high- and low-energy coefficients based on the first maximum value, the second maximum value, the number of image data bits corresponding to the high-energy X-ray image, and the number of image data bits corresponding to the low-energy X-ray image includes: The normalized high and low energy coefficients are determined based on the following formula (1): Among them, R(i, j) represents the normalized high and low energy coefficients, A represents the maximum pixel value of the pixel point under the number of image data bits corresponding to the high-energy X-ray image, B represents the first maximum value, C represents the maximum pixel value of the pixel point under the number of image data bits corresponding to the low-energy X-ray image, and D represents the second maximum value.
6. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 1 is characterized in that: The identifying the liquid type of the liquid to be detected based on the multimodal fusion image includes: Inputting the multimodal fusion image into a feature extraction module of an image classification model, dividing the multimodal fusion image into blocks by the feature extraction module, encoding each obtained image block, and performing feature extraction on each encoded image block to obtain each multimodal feature; Each of the multimodal features is input into a classification module of the image classification model to obtain an identification result of the liquid type output by the classification module.
7. The liquid identification method based on multimodal fusion images of security inspection X-ray machine according to claim 1 is characterized in that: Scan the container containing the liquid to be tested to obtain a pseudo-color X-ray image, including: Scan the container containing the liquid to be tested to obtain an original X-ray image; The original X-ray image is colored using an image coloring algorithm to obtain the pseudo-color X-ray image.
8. The liquid identification method based on multimodal fusion images of a security inspection X-ray machine according to any one of claims 1 to 7, characterized in that: The determining of the multimodal fusion image based on the second-type cross region, the third-type cross region, and the fourth-type cross region includes: Rotating the second-type cross region, the third-type cross region, and the fourth-type cross region respectively to obtain a rotated second-type cross region, a rotated third-type cross region, and a rotated fourth-type cross region; scaling the rotated second-type cross region, the rotated third-type cross region, and the rotated fourth-type cross region based on preset sizes to obtain scaled second-type cross region, scaled third-type cross region, and scaled fourth-type cross region; The scaled second-type cross region, the scaled third-type cross region, and the scaled fourth-type cross region are fused to obtain the multimodal fusion image.
Citation Information
Patent Citations
Article category identification method, device and equipment based on X-ray security inspection equipment
CN115081469A
Image fusion method and device, server and storage medium
CN115601276A