CT Object Recognition Method and Apparatus for Security Inspection
By generating 2D dimension reduction views from 3D CT data and expanding recognition results to 3D, the method addresses limitations in recognizing complex objects, achieving improved accuracy and real-time performance in security inspections.
Patent Information
- Application Number
- JP2024504228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-27
- Filing Date
- 2022-07-08
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2042-07-08
Smart Images

Figure 0007675277000001 
Figure 0007675277000002 
Figure 0007675277000003
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of security inspection CT, and in particular to a method and apparatus for recognizing objects in security inspection CT. [Background technology]
[0002] Currently, in the field of security inspection, CT devices are often used to recognize objects such as prohibited goods. When using a security inspection CT device to recognize objects, the prior art mainly uses CT reconstruction technology to obtain a three-dimensional tomographic image containing material attribute information, divides the three-dimensional image into a number of suspect objects, and performs statistics and classification of the material attributes of the suspect objects.
[0003] However, while the above-mentioned conventional techniques have good performance in recognizing prohibited items that have strong discrimination in terms of material attributes, such as explosives and poisons, they have obvious limitations in recognizing objects that have strong three-dimensional shape features and relatively complex material compositions and physical attributes.
[0004] In order to overcome such limitations, Patent Document 1 proposes a CT detection method and device that recognizes three-dimensional tomographic images and two-dimensional images of an object, respectively, and obtains recognition results for explosives using the former and recognition results for other prohibited items using the latter.
[0005] Patent document 1: 109975335A Summary of the Invention
[0006] As described above, Patent Document 1 attempts to improve the above limitations, but the inventors of the present invention have found through research that Patent Document 1 still has the following technical problems.
[0007] (1) If the object is projected only along the direction perpendicular to the moving direction (Z direction) of the object during the detection process, the projection area is too small in some object orientations, and the representation of the shape information is incomplete, making it impossible to accurately recognize the object. In addition, if the object is projected in this way, it may be occluded by other objects, in which case the shape information of the object will also be lost, making it impossible to accurately recognize the object.
[0008] (2) In this method, two-dimensional projection images obtained from three-dimensional tomographic images may be used to recognize other prohibited items except explosives, but the recognition operation is limited to a two-dimensional plane, and recognition results cannot be obtained in three-dimensional space. Since the amount of information in two-dimensional images is significantly lower than that in three-dimensional data, the advantages of security inspection CT devices cannot be fully utilized unless two-dimensional recognition results can be effectively integrated. [Means for solving the problem]
[0009] The present application provides a CT object recognition method and device for security inspection, which can improve the recognition effect for three-dimensional objects.
[0010] A first aspect of the present application provides a security inspection CT object recognition method, performing dimensionality reduction on the three-dimensional CT data to generate a plurality of two-dimensional reduced-dimensional views; performing object recognition on a plurality of two-dimensional views, including a plurality of two-dimensional reduced dimensionality views, to obtain a set of two-dimensional semantic descriptions of the object; and performing dimensional expansion on the set of two-dimensional semantic descriptions to obtain a three-dimensional recognition result of the object.
[0011] In the above-mentioned security inspection CT object recognition method, performing dimensional expansion on the two-dimensional semantic description set and obtaining a three-dimensional recognition result of the object includes mapping the two-dimensional semantic description set into a three-dimensional space by back projection to obtain a three-dimensional probability map, and performing feature extraction on the three-dimensional probability map to obtain a three-dimensional recognition result of the object.
[0012] In the above-mentioned security inspection CT object recognition method, mapping the two-dimensional semantic description set into three-dimensional space by back projection and obtaining a three-dimensional probability map includes mapping from the two-dimensional semantic description set into three-dimensional space by voxel-driven or pixel-driven, obtaining a semantic feature matrix, and compressing the semantic feature matrix into a three-dimensional probability map.
[0013] In the above-mentioned security inspection CT object recognition method, voxel driving includes matching each voxel in the 3D CT data to a pixel in each 2D view, querying and accumulating the 2D semantic description information corresponding to the pixel, and generating a semantic feature matrix; pixel driving includes that each pixel in the 2D view corresponds to a straight line in the 3D CT data, traversing each pixel in each 2D view or each pixel of the region of interest given by the 2D semantic description set, and propagating the 2D semantic description information corresponding to the pixel along the straight line into the three-dimensional space to generate a semantic feature matrix.
[0014] In the above-mentioned security inspection CT object recognition method, in the voxel driving or pixel driving, the correspondence between the voxel and the pixel is obtained by a mapping function or a lookup table.
[0015] In the above-mentioned security inspection CT object recognition method, performing feature extraction on the three-dimensional probability map and obtaining a three-dimensional recognition result of the object includes adopting at least one or a combination of an image processing method, a classical machine learning method, and a deep learning method to perform feature extraction on the three-dimensional probability map, thereby obtaining a three-dimensional image semantic description set as the three-dimensional recognition result.
[0016] In the above-mentioned CT object recognition method for security inspection, performing feature extraction on the 3D probability map and obtaining a 3D recognition result for the object includes: binarizing the 3D probability map to obtain a 3D binary image; performing connected region analysis on the 3D binary image to obtain connected regions; and generating a 3D image semantic description set for the connected regions.
[0017] In the above security inspection CT object recognition method, the connected region analysis includes performing connected component labeling on the 3D binary image, and performing a mask operation on each label region to obtain a connected region.
[0018] In the above-mentioned CT object recognition method for security inspection, generating a 3D image semantic description set for the connected region includes extracting all probability values in the connected region, performing principal component analysis to obtain an analysis set, and taking the analysis set as the object effective voxel region to perform statistics on the 3D image semantic description set.
[0019] In the above-mentioned security inspection CT object recognition method, the three-dimensional image semantic description set includes category information and / or confidence level for each unit of one or more of voxels, three-dimensional regions of interest, and three-dimensional CT images, or the three-dimensional image semantic description set includes at least one of category information, object position information, and confidence level for each unit of three-dimensional regions of interest and / or three-dimensional CT images.
[0020] In the above security inspection CT object recognition method, the position information includes a three-dimensional bounding box.
[0021] In the above-mentioned security inspection CT object recognition method, the two-dimensional semantic description set includes category information and / or confidence level for each unit of a pixel, a region of interest, and / or a two-dimensional image, or the two-dimensional semantic description set includes at least one of category information, confidence level, and object location information for each unit of a region of interest and / or a two-dimensional image.
[0022] In the above-mentioned security inspection CT object recognition method, performing object recognition for each of the multiple two-dimensional views includes employing at least one or a combination of an image processing method for two-dimensional images, a classical machine learning method, and a deep learning method to perform object recognition.
[0023] In the above-mentioned security inspection CT object recognition method, performing dimensionality reduction on the 3D CT data to generate multiple 2D dimensionality reduced views includes setting multiple directions for the 3D CT data and performing projection or rendering according to the multiple directions.
[0024] In the above security inspection CT object recognition method, the multiple directions are arbitrary directions and are not limited to directions perpendicular to the traveling direction of the object in the detection process.
[0025] In the above security inspection CT object recognition method, the plurality of two-dimensional views further includes a two-dimensional DR image, and the two-dimensional DR image is obtained by a DR imaging device.
[0026] In the above security inspection CT object recognition method, the three-dimensional recognition result is projected onto a two-dimensional DR image, and further output as the recognition result of the two-dimensional DR image.
[0027] A second aspect of the present application provides a security inspection CT object recognition device, which includes a dimensionality reduction module that performs dimensionality reduction on 3D CT data to generate a plurality of 2D dimensionality reduced views, a 2D recognition module that performs object recognition on a plurality of 2D views including the plurality of 2D dimensionality reduced views to obtain a set of 2D semantic descriptions of the object, and a dimensionality expansion module that performs dimensionality expansion on the set of 2D semantic descriptions to obtain a 3D recognition result of the object.
[0028] A third aspect of the present application provides a computer-readable storage medium having stored thereon a program for causing a computer to perform dimensionality reduction on three-dimensional CT data to generate a plurality of two-dimensional reduced views; perform object recognition on a plurality of two-dimensional views including the plurality of two-dimensional reduced views to obtain a set of two-dimensional semantic descriptions of the object; and perform dimensionality expansion on the set of two-dimensional semantic descriptions to obtain a three-dimensional recognition result of the object.
[0029] As described above, in the present application, a plurality of two-dimensional dimension-reduced views are generated by performing dimensionality reduction from three-dimensional CT data, and a plurality of two-dimensional views including the plurality of two-dimensional dimension-reduced views are used to perform object recognition, to obtain a two-dimensional semantic description set, and then dimensional expansion is performed on the two-dimensional semantic description set to obtain a three-dimensional recognition result, that is, the three-dimensional recognition is first performed by reducing the dimension from three dimensions to two dimensions, and then the three-dimensional result is generated by performing dimensional expansion, so that the object having relatively complex material composition and physical attributes and shape characteristics can be effectively recognized through the recognition based on two dimensions, and the two-dimensional recognition result can be effectively integrated to provide a three-dimensional recognition result with rich information. Therefore, the recognition effect for the object can be improved, and the real-time requirement of security inspection can be met. [Brief description of the drawings]
[0030] [Figure 1] FIG. 1 is a flowchart showing a security inspection CT object recognition method according to the first embodiment. [Diagram 2] FIG. 2 is a flowchart showing a specific example of the dimensionality reduction process. [Diagram 3] FIG. 3 is a flowchart showing a specific example of the dimension expansion process. [Figure 4] FIG. 4 is a flow chart showing a specific example of three-dimensional feature extraction. [Diagram 5] FIG. 5 is a flowchart showing a security inspection CT object recognition method according to the second embodiment. [Figure 6]FIG. 6 is a schematic diagram showing an example of a security inspection CT object recognition device according to the third embodiment. [Figure 7] FIG. 7 is a schematic diagram showing another example of the security inspection CT object recognition device according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] Hereinafter, exemplary embodiments or examples of the present invention will be described in detail with reference to the accompanying drawings. Although the accompanying drawings show exemplary embodiments of the present invention, it should be understood that the present invention can be realized in various forms and should not be limited by the embodiments or examples described herein. On the contrary, these embodiments or examples are provided for a clear understanding of the present invention.
[0032] The terms "first", "second", and the like in the specification and claims of this application are not intended to describe a particular order or sequence, but are intended to distinguish between similar objects. As will be understood, the data used in this manner may be interchanged as appropriate, such that the embodiments or examples of this application described herein may be performed in an order other than that shown or described. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion, e.g., a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units explicitly recited, and may include other steps or units not explicitly shown. Identical or similar reference numerals indicate components having the same or similar functions.
[0033] First Embodiment As a first embodiment of the present application, a security inspection CT object recognition method is provided. Fig. 1 is a flowchart showing the security inspection CT object recognition method according to the first embodiment. The security inspection CT object recognition method is applied to a security inspection CT system, and may be executed, for example, in a security inspection CT device or a server connected to the security inspection CT device.
[0034] As shown in Fig. 1, in step S10, a dimensionality reduction process is performed, that is, dimensionality reduction is performed on the 3D CT data to generate multiple 2D reduced-dimensional views.
[0035] Specifically, as shown in FIG. 2, step S10 may include steps S11 and S12.
[0036] In step S11, a plurality of directions are set for the three-dimensional CT data. Here, the plurality of directions is not limited to a specific direction such as a direction perpendicular to the traveling direction of the object in the detection process, but may be any direction.
[0037] In addition, simultaneously with setting the multiple directions, or before or after setting the multiple directions, certain pre-processing operations of the 3D volume data, such as filtering of invalid voxels and pre-calculation of geometric parameters required for projection or rendering, can be selectively performed, thereby improving the speed of subsequent processing.
[0038] At S12, multiple projections or renderings are performed according to multiple directions to obtain multiple 2D reduced dimension views.
[0039] As an example, a ray or light beam can be projected based on a CT image slice sequence, and one ray or light beam can be fired from each pixel of the image along a specific direction, and the ray or light beam can traverse the entire image sequence. In the process, the image sequence can be sampled to obtain attribute or color information. At the same time, by the time the ray or light beam traverses the entire image sequence, attribute or color values can be accumulated based on a certain model, and the finally obtained attribute or color value can be a two-dimensional view after dimensionality reduction.
[0040] In the present application, by obtaining a two-dimensional dimensionality reduced view by projecting according to an arbitrary direction, it is possible to avoid performing dimensionality reduction only along a specific direction, for example, to avoid performing dimensionality reduction only along a direction perpendicular to the moving direction of the object in the detection process, thereby solving the problems that exist when performing dimensionality reduction only along a specific direction, that is, (1) in some object placement postures, the area of the object after dimensionality reduction is too small, and the representation of shape information is incomplete, making it impossible to accurately recognize the object, and (2) the object is occluded by other objects, resulting in the loss of shape information of the object and the inability to accurately recognize the target object.
[0041] In step S20, a 2D recognition process is performed, that is, object recognition is performed on the multiple 2D views to obtain a set of 2D semantic descriptions of the object, where the multiple 2D views include the multiple 2D reduced dimension views obtained in step S10 described above.
[0042] Specifically, as an object recognition method for a two-dimensional view, at least one of image processing methods used for two-dimensional images, classic machine learning methods, and deep learning methods, or a combination thereof, can be adopted.
[0043] For example, a two-dimensional view is input into a neural network model and a set of two-dimensional semantic descriptions is obtained as output.
[0044] Specifically, a target detection neural network based on deep learning can be used to detect the two-dimensional position of an object. The convolutional neural network used in the target detection task is a typical structure of deep learning in computer vision tasks, and such a convolutional neural network has characteristics such as local connection, weight sharing, and spatial resampling. Due to these characteristics, the convolutional neural network has a certain degree of translational scaling invariance. Here, the two-dimensional semantic description set includes category information and / or confidence for one or more of a pixel, a region of interest, and a two-dimensional image, or the two-dimensional semantic description set includes at least one of category information, confidence, and object position information for a region of interest and / or a two-dimensional image. Here, the category information indicates a category to which the object belongs, for example, a gun, a knife, etc. The position information may include a center coordinate, a bounding box, etc. The confidence represents the magnitude of the possibility that the object exists, and may be a normalized scalar or vector.
[0045] In other words, the two-dimensional semantic description set includes at least one of the following information: category information, confidence level, etc. that a pixel belongs to an object; category information, position information, confidence level, etc. of an object included in a region of interest; category information, position information, confidence level, etc. of an object included in a two-dimensional image. Here, the at least one information may be information included in one set, or may be information included in different sets.
[0046] The description set of two-dimensional semantic information may also include category information, confidence level, location information, as well as other semantic information such as object posture, number of objects, etc.
[0047] The object recognition method for the two-dimensional view of the present application is not particularly limited as long as it can obtain the above-mentioned two-dimensional semantic description set based on the two-dimensional view.
[0048] As described above, in this application, the 2D recognition result of the object is represented in the form of a 2D semantic description set, and this 2D semantic description set is input to step S30 as an input and dimensionally expanded to 3D, so that the 2D recognition result can be integrated into 3D. In addition, the form of the 2D semantic description set is flexible, so that the contained information can be enriched.
[0049] In step S30, a dimension expansion process is performed, that is, dimension expansion is performed on the two-dimensional semantic description set to obtain a three-dimensional recognition result of the object.
[0050] Specifically, as shown in FIG. 3, step S30 may include steps S31 and S32.
[0051] In step S31, the two-dimensional semantic description set is mapped into a three-dimensional space by backprojection to obtain a three-dimensional probability map. Backprojection can be considered as the inverse process of projection.
[0052] Optionally, the backprojection process may be realized in a manner such as voxel-driven or pixel-driven. In particular, the semantic feature matrix may be obtained by voxel-driven or pixel-driven, and the semantic feature matrix may be compressed into a three-dimensional probability map.
[0053] Here, voxel driving includes matching each voxel in the 3D CT data to a pixel in each 2D view, querying and accumulating the 2D semantic description information corresponding to the pixel, and generating a semantic feature matrix.
[0054] The voxel to pixel correspondence can be constructed as a mapping function or look-up table to improve computation speed.
[0055] As described above, according to the voxel drive, each voxel in the 3D CT data is traversed, its semantic feature matrix is obtained sequentially, and finally the semantic feature matrix is compressed to obtain a 3D probability map.
[0056] Voxel driving allows parallel calculations to be performed for each voxel, which increases the calculation speed and improves the real-time nature of security inspections.
[0057] Pixel driving involves traversing each pixel in each two-dimensional view or each pixel in a region of interest along a straight line, propagating the two-dimensional semantic description information corresponding to the pixel along the straight line to a three-dimensional space, and generating a semantic feature matrix, where the region of interest is given by the set of two-dimensional semantic descriptions, and the correspondence between voxels and pixels may be obtained by a mapping function or a look-up table.
[0058] As described above, for each pixel in multiple 2D views, its semantic feature matrix is obtained sequentially according to pixel driving, and finally the semantic feature matrix is compressed to obtain a 3D probability map.
[0059] Pixel driving also enables parallel calculations to be performed for each pixel, improving the calculation speed and contributing to improving the real-time nature of security inspections.
[0060] As can be seen from the above description of voxel driving and pixel driving, the semantic feature matrix is generated from the two-dimensional semantic description information based on its spatial correspondence, and is a matrix obtained by digitizing and aggregating based on the two-dimensional semantic description information. For example, for category information in the two-dimensional semantic description set, a semantic feature matrix can be obtained for each category of each object. For example, it can be assumed that if the object belongs to the category, the corresponding value in the semantic feature matrix is 1, and if the object does not belong to the category, the corresponding value in the semantic feature matrix is 0. In addition, semantic feature matrices can also be obtained for other semantic information in the two-dimensional semantic description set by a similar method.
[0061] Typical methods for compressing a semantic feature matrix include weighted average, principal component analysis, etc. In this case, the input is a semantic feature matrix and the output is a probability map.
[0062] As an example, assuming there are two two-dimensional views, there are two two-dimensional semantic description sets, among which the semantic information in a pixel (or region of interest or two-dimensional image) is expressed as a numerical value of 1 or 0, and can be mapped to a three-dimensional space using a back projection method to generate a corresponding semantic feature matrix in the three-dimensional space, the values in the matrix are vectors consisting of 0 or 1, and a weighted average method can be used to calculate the probability map value of a voxel in the corresponding three-dimensional space, for example, the semantic feature matrix value of a voxel is v=[0, 1], and if the weights are the same, the probability map value of the voxel is 0.5. When compressing the semantic feature matrix corresponding to all object categories, the dimension of the output probability map value is determined by the quantity of object categories. The method of obtaining the probability map value described here is an example, and the probability map value may be obtained in other ways. For example, the weights may be different, and the probability map value may be obtained by weighting the semantic feature matrix value with different weights.
[0063] As another example, one or more vectors in the three-dimensional semantic feature matrix may be used as input variables in a principal component analysis, and such input variables may be subjected to principal component analysis to obtain output variables as principal components, which may then be normalized to obtain probability map values corresponding to the voxels.
[0064] The above-mentioned calculation method not only ensures the real-time calculation, but also effectively integrates the two-dimensional recognition results, improving the final recognition effect.
[0065] In step S32, feature extraction is performed on the 3D probability map to obtain a 3D recognition result of the object.
[0066] Specifically, feature extraction is performed on the 3D probability map by employing at least one or a combination of image processing methods, classical machine learning methods, and deep learning methods, and a set of 3D image semantic descriptions is obtained as the 3D recognition result.
[0067] As an example, a 3D probability map is input to a deep learning model as input, and 3D recognition results such as confidence and 3D bounding box are obtained as output. The deep learning model used here can adopt techniques such as classification neural networks or target detection networks with a small number of layers. By adopting such techniques, the amount of information contained in the original 3D CT data is effectively simplified and abstracted after being processed by the above steps, which is closer to the final goal of contraband recognition, and a 3D semantic description set can be quickly and accurately extracted by applying a simple feature extraction method.
[0068] Here, the 3D image semantic description set includes category information and / or confidence level for one or more of a voxel, a 3D region of interest, and a 3D CT image, or the 3D image semantic description set includes at least one of category information, object position information, and confidence level for a 3D region of interest and / or a 3D CT image. The object position information in the 3D CT image may include a 3D bounding box.
[0069] In other words, the 3D image semantic description set includes at least one of the following information: category information, confidence level, etc. that a certain voxel belongs to an object; category information, position information, confidence level, etc. of an object included in a certain 3D region of interest (VOI); category information, position information, confidence level, etc. of an object included in a certain 3D CT image. Here, the at least one information may be information included in one set, or may be information included in different sets.
[0070] Since the three-dimensional image semantic description set is generated from a three-dimensional probability map generated based on the two-dimensional semantic description set, the types of semantic information contained in the three-dimensional image semantic description set and the types of semantic information contained in the two-dimensional semantic description set are consistent or mutually convertible.
[0071] In this application, by using the above-mentioned dimension expansion process, dimension expansion is performed on the two-dimensional semantic description set to obtain the three-dimensional recognition result of the object, thereby solving the problem that the amount of information is significantly reduced when performing two-dimensional recognition only by dimension reduction, and while adopting two-dimensional recognition, the loss of information amount is reduced, and it is possible to achieve both real-timeness and accuracy of security inspection.
[0072] As another example of step S32, for example, an image processing method may be adopted, and step S32 may include steps S321-S323 as shown in FIG.
[0073] In step S321, the three-dimensional probability map is binarized to obtain a three-dimensional binary image.
[0074] In step S322, connected regions are obtained by performing connected region analysis on the 3D binary image.
[0075] As an example, connected regions can be obtained by labeling a three-dimensional binary image using connected component labeling and performing a mask operation for each labeled region.
[0076] In step S323, a set of 3D image semantic descriptions is generated for the connected regions.
[0077] In this case, the 3D image semantic description set may include a 3D bounding box, which can provide the spatial boundary of an object in the 3D image and more intuitively indicate the object's position, extent, orientation, shape, etc., and is advantageous for the accuracy with which security inspectors can determine whether an object is dangerous.
[0078] For example, all probability values in the connected region can be extracted, and a principal component analysis can be performed to obtain an analysis set, and the analysis set can be taken as the object effective voxel region. A set of 3D image semantic descriptions can be statistically calculated for the effective voxel region, which can further improve the accuracy of 3D recognition.
[0079] In the first embodiment, a plurality of two-dimensional dimension-reduced views are generated by performing dimensionality reduction from three-dimensional CT data, a plurality of two-dimensional views including the plurality of two-dimensional dimension-reduced views are used to perform object recognition, a two-dimensional semantic description set is obtained, and a three-dimensional recognition result is obtained by performing dimensional expansion on the two-dimensional semantic description set, that is, after performing dimensionality reduction from three dimensions to two dimensions for recognition, a three-dimensional result is further expanded, so that the object having relatively complex material composition and physical attributes and shape characteristics can be effectively recognized by the two-dimensional recognition, and the two-dimensional recognition result can be effectively integrated to provide a three-dimensional recognition result with rich information, thereby improving the recognition effect of the object and meeting the real-time requirement of security inspection.
[0080] <Second embodiment> As a second embodiment of the present application, another security inspection CT object recognition method is provided. Fig. 5 is a flowchart showing the security inspection CT object recognition method according to the first embodiment.
[0081] The difference between the second embodiment and the first embodiment is that in the second embodiment, object recognition is performed using not only a two-dimensional reduced image generated by reducing the dimensions of three-dimensional CT data, but also two-dimensional DR data.
[0082] Specifically, in step S20, the multiple two-dimensional views further include a two-dimensional DR image, and object recognition is also performed on the two-dimensional DR image to obtain a two-dimensional semantic description set of the object. Here, the two-dimensional DR image is obtained by a DR imaging device arranged separately from the security inspection CT equipment. The two-dimensional DR image is an image of the same security inspection object as the three-dimensional CT data.
[0083] In the second embodiment, as shown in Fig. 5, before step S20, step S40 may be further included. In step S40, a two-dimensional DR image is obtained from a DR imaging device and is one of a plurality of two-dimensional views. This step S40 may be performed in parallel with step S10.
[0084] In this case, in step S30, not only is dimensional expansion performed on the 2D semantic description set of the 2D dimension-reduced image, but also dimensional expansion is performed on the 2D semantic description set of the 2D DR image, thereby obtaining a 3D recognition result.
[0085] A two-dimensional DR image is a two-dimensional image whose principles and properties are different from those of a two-dimensional reduced image generated by reducing the dimensions of three-dimensional CT data. By using such two-dimensional DR images for object recognition, the amount of information used for recognition can be increased, thereby improving the accuracy of recognition.
[0086] In the second embodiment, as shown in Fig. 5, after step S30, step S50 may be optionally included. In step S50, the 3D recognition result generated in step S30 is projected onto a 2D DR image, and is further output as a recognition result of the 2D DR image.
[0087] Due to working habits and needs, some security inspectors would like to confirm the recognition results using 2D DR images. However, if the recognition results of 2D DR images are directly used, when the object in the DR image is seriously occluded or has a special placement posture, the object information will be incomplete, which will affect the recognition accuracy. In contrast, 3D recognition results can be more accurate and reliable by effectively integrating the semantic information of several 2D views. Therefore, by projecting the 3D recognition results onto 2D DR images and outputting them as recognition results, the work needs of security inspectors to confirm the recognition results using 2D DR images can be met and the accuracy of the recognition results can be improved.
[0088] The results of steps S30 and S50 may be output simultaneously.
[0089] In this case, by comparing and verifying the three-dimensional recognition results with the recognition results in the two-dimensional DR image, it is advantageous for security inspectors to more accurately determine whether or not the object is a dangerous object.
[0090] <Third embodiment> As a third embodiment of the present application, a security inspection CT object recognition device is provided. Fig. 6 is a schematic diagram showing the security inspection CT object recognition device according to the first embodiment.
[0091] As shown in FIG. 6, the security inspection CT object recognition device 100 according to this embodiment includes a dimension reduction module 10, a two-dimensional recognition module 20, and a dimension expansion module 30.
[0092] The dimension reduction module 10 reduces the dimension of the three-dimensional CT data to generate a plurality of two-dimensional reduced views, i.e., can execute the process of step S10 in the first and second embodiments.
[0093] The 2D recognition module 20 performs object recognition on a plurality of 2D views to obtain a set of 2D semantic descriptions of the object, where the plurality of 2D views includes the plurality of 2D dimension-reduced views, i.e., can execute the process of step S20 in the first and second embodiments.
[0094] The dimension expansion module 30 performs dimension expansion on the two-dimensional semantic description set to obtain a three-dimensional recognition result of the object, that is, the process of step S30 in the first and second embodiments can be executed.
[0095] The specific processes of the dimension reduction module 10, the two-dimensional recognition module 20, and the dimension expansion module 30 can be referred to in the first and second embodiments described above, and therefore will not be repeated here.
[0096] 7, the security inspection CT object recognition device 100 may further include a DR image acquisition module 40 that acquires a two-dimensional DR image from a DR imaging device and converts it into one of a plurality of two-dimensional views. That is, the DR image acquisition module 40 can execute the process of step S40 in the second embodiment.
[0097] The security inspection CT object recognition device 100 may further include a DR output module 50 that projects the three-dimensional recognition result generated by the dimension extension module 30 onto a two-dimensional DR image and outputs the two-dimensional DR image as a recognition result. That is, the DR output module 50 can execute the process of step S50 in the second embodiment.
[0098] In the present application, the security inspection CT object recognition device 100 may be realized in hardware, or in software modules executed on one or more processors, or in a combination thereof.
[0099] For example, the security inspection CT object recognition device 100 may be realized in the form of a combination of software and hardware by any appropriate electronic device, such as a desktop computer equipped with a processor, a tablet computer, a smartphone, a server, etc. For example, the security inspection CT object recognition device 100 may be a control computer of a security inspection CT system, or a server connected to a security inspection CT scanning device in the security inspection CT system.
[0100] Furthermore, the security inspection CT object recognition device 100 may be realized in the form of a software module in any appropriate electronic device, such as a desktop computer, a tablet computer, a smartphone, a server, etc. For example, it may be a software module installed in a control computer of a security inspection CT system, or a software module installed in a server connected to a security inspection CT scanning device in the security inspection CT system.
[0101] The processor of the security inspection CT object recognition device 100 can execute the security inspection CT object recognition method described below.
[0102] The security inspection CT object recognition device 100 may further include a memory (not shown) and a communication module (not shown).
[0103] The memory of the security inspection CT object recognition device 100 may store steps for executing a security inspection CT object recognition method described later, and data for performing security inspection CT object recognition, etc. The memory may be, for example, a ROM (Read Only Memory image), a RAM (Random Access Memory), etc. The memory has a storage space for program codes for executing any steps in the security inspection CT object recognition method. When these program codes are read and executed by a processor, the security inspection CT object recognition method is executed. These program codes may be read from one or more computer program products or written into one or more computer program products. These computer program products include program code carriers, such as hard disks, compact discs (CDs), memory cards, and flexible disks. Such computer program products are usually portable or fixed storage units. The program codes for executing any steps in the above method may be downloaded via a network. The program codes may be, for example, compressed in an appropriate format.
[0104] The communication module in the security inspection CT object recognition device 100 can support establishing a direct (e.g., wired) communication channel or a wireless communication channel between the security inspection CT object recognition device 100 and an external electronic device, and performing communication through the established communication channel. For example, the communication module receives three-dimensional CT data, etc. from a CT scanner device via a network.
[0105] In addition, the security inspection CT object recognition device 100 may further include an output unit such as a display, a microphone, and a speaker to output the object recognition result.
[0106] The above-described security inspection CT object recognition device 100 can provide the same effects as those of the first and second embodiments.
[0107] Although the embodiments and specific examples of the present invention have been described above with reference to the drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and all such modifications and variations are included within the scope defined by the claims. [Explanation of symbols]
[0108] 100 Security inspection CT object recognition device 10 Dimensionality Reduction Module 20 2D Recognition Module 30 Dimensional Extension Module 40 DR Image Acquisition Module 50 DR Output Module
Claims
1. A security inspection CT object recognition method, comprising: performing dimensionality reduction on 3D CT data to generate a plurality of 2D reduced-dimensional views; performing object recognition on a plurality of two-dimensional views including the plurality of two-dimensional reduced dimensionality views to obtain a set of two-dimensional semantic descriptions of the object; performing a voxel-driven or pixel-driven mapping from the set of two-dimensional semantic descriptions to a three-dimensional space to obtain a semantic feature matrix; compressing the semantic feature matrix into a three-dimensional probability map; and performing feature extraction on the three-dimensional probability map to obtain a three-dimensional recognition result of the object. A security inspection CT object recognition method comprising the steps of:
2. The voxel drive is Corresponding each voxel in the 3D CT data to a pixel in each of the 2D views, and querying and accumulating 2D semantic description information corresponding to the pixel to generate the semantic feature matrix; The pixel drive is each pixel in the two-dimensional views corresponds to a straight line in the three-dimensional CT data, traversing each pixel in each two-dimensional view or each pixel in a region of interest given by the set of two-dimensional semantic descriptions, and propagating two-dimensional semantic description information corresponding to the pixel along the straight line into a three-dimensional space to generate the semantic feature matrix.
2. The security inspection CT object recognition method according to claim 1 .
3. In the voxel driving or pixel driving, a correspondence relationship between the voxels and the pixels is obtained by a mapping function or a lookup table.
3. The security inspection CT object recognition method according to claim 2.
4. performing feature extraction on the 3D probability map to obtain a 3D recognition result of the object; and performing feature extraction on the 3D probability map by employing an image processing method or a machine learning method, or a combination thereof, to obtain a set of 3D image semantic descriptions, which is the 3D recognition result.
2. The security inspection CT object recognition method according to claim 1 .
5. performing feature extraction on the 3D probability map to obtain a 3D recognition result of the object; binarizing the three-dimensional probability map to obtain a three-dimensional binary image; performing a connected region analysis on the three-dimensional binary image to obtain connected regions; generating a set of 3D image semantic descriptions for the connected regions.
5. The security inspection CT object recognition method according to claim 4.
6. The connected region analysis includes: performing connected component labeling on the three-dimensional binary image, and performing a masking operation on each labeled region to obtain the connected region; 6. The security inspection CT object recognition method according to claim 5.
7. Generating a set of 3D image semantic descriptions for the connected regions includes: extracting all probability values in the connected regions, performing a principal component analysis to obtain an analysis set, and using the analysis set to generate statistics on the 3D image semantic description.
6. The security inspection CT object recognition method according to claim 5.
8. The set of 3D image semantic descriptions includes category information and / or confidence level for one or more of a voxel, a 3D region of interest, and a 3D CT image; Alternatively, the 3D image semantic description set includes at least one of category information, object position information, and confidence level for each 3D region of interest and / or each 3D CT image.
5. The security inspection CT object recognition method according to claim 4.
9. The position information includes a three-dimensional bounding box.
9. The security inspection CT object recognition method according to claim 8.
10. The set of two-dimensional semantic descriptions includes category information and / or confidence for one or more of a pixel, a region of interest, and a two-dimensional image; Alternatively, the two-dimensional semantic description set includes at least one of category information, confidence level, and object position information for each region of interest and / or each two-dimensional image.
2. The security inspection CT object recognition method according to claim 1 .
11. Performing object recognition on each of the plurality of two-dimensional views includes: This includes employing image processing methods or machine learning methods for two-dimensional images, or a combination of both, to perform object recognition.
2. The security inspection CT object recognition method according to claim 1 .
12. Performing dimensionality reduction on the three-dimensional CT data to generate a plurality of two-dimensional reduced-dimensional views, Setting a plurality of directions for the three-dimensional CT data; and projecting or rendering according to said plurality of directions.
2. The security inspection CT object recognition method according to claim 1 .
13. The multiple directions may be any directions and are not limited to directions perpendicular to the direction of travel of the object during the detection process. The security inspection CT object recognition method according to claim 12 .
14. the plurality of two-dimensional views further comprises a two-dimensional DR image; The two-dimensional DR image is obtained by a DR imaging device. The security inspection CT object recognition method according to any one of claims 1 to 13.
15. The three-dimensional recognition result is projected onto the two-dimensional DR image, and is output as a recognition result of the two-dimensional DR image. The security inspection CT object recognition method according to claim 14 .
16. A security inspection CT object recognition device, a dimensionality reduction module that performs dimensionality reduction on the 3D CT data to generate a plurality of 2D reduced-dimensional views; a 2D recognition module that performs object recognition on a plurality of 2D views, including the plurality of 2D reduced dimensionality views, to obtain a set of 2D semantic descriptions of the object; a dimensional expansion module for performing a voxel-driven or pixel-driven mapping from the two-dimensional semantic description set to a three-dimensional space, obtaining a semantic feature matrix, compressing the semantic feature matrix into a three-dimensional probability map, performing feature extraction on the three-dimensional probability map, and obtaining a three-dimensional recognition result of the object. A security inspection CT object recognition device characterized by the above.
17. On the computer, performing dimensionality reduction on the 3D CT data to generate a plurality of 2D reduced dimensionality views; performing object recognition on a plurality of two-dimensional views including the plurality of two-dimensional reduced dimensionality views to obtain a set of two-dimensional semantic descriptions of the object; performing a voxel-driven or pixel-driven mapping from a set of two-dimensional semantic descriptions to a three-dimensional space, obtaining a semantic feature matrix, compressing the semantic feature matrix into a three-dimensional probability map, performing feature extraction on the three-dimensional probability map, and obtaining a three-dimensional recognition result of the object. A computer-readable storage medium.
Citation Information
Patent Citations
Composite object partitioning method and system {COMPOUNDOBJECTSEPARATION}
JP2014508954A
Image display method
JP2015227872A