Three-view x-ray security image recognition method, device, equipment and medium

CN122454347BActive Publication Date: 2026-09-08HUNAN SUKE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610918649.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-08
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

[0003]本申请的目的在于提供一种三视角X光安检图像识别方法、装置、设备及介质,以解决现有技术中单、双视角成像维度不足易产生识别歧义,三视角识别缺乏几何约束注意力机制,跨视角特征交互与融合效果不佳、安检识别精度低的技术问题

Benefits of technology

[0008] Beneficial effects: The three-view X-ray security inspection image recognition method, device, equipment and medium of this application solve the technical problems of insufficient single and dual-view imaging dimensions which easily lead to recognition ambiguity, lack of geometric constraint attention mechanism in three-view recognition, poor cross-view feature interaction and fusion effect and low security inspection recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454347B_ABST
    Figure CN122454347B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of security image recognition, and discloses a three-view X-ray security image recognition method, device, equipment and medium, which comprises the following steps: constructing a security geometric position code based on a security target geometric vector to optimize original image features and obtain security geometric reinforcement features; generating security supplementary features for representing cross-view attention of original view geometric constraints based on a security cross-view mask; fusing the security geometric reinforcement features and all the security supplementary features to obtain cross-view interaction features; generating three-view features suitable for geometric constraint cross-view attention; inputting the three-view features into a detection head to perform recognition processing and outputting a recognition result of the three-view X-ray security image. The application solves the technical problems of the prior art, such as insufficient single-view and double-view imaging dimensions, easy generation of recognition ambiguity, lack of geometric constraint attention mechanism in three-view recognition, poor cross-view feature interaction and fusion effect, and low security recognition precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security inspection image recognition technology, specifically a three-view X-ray security inspection image recognition method, device, equipment, and medium. Background Technology

[0002] X-ray security inspection equipment is a core device for screening prohibited items at airports, train stations, customs, logistics facilities, and large-scale events. Traditional single- and dual-view X-ray security inspections have limited imaging dimensions, and feature confusion easily occurs when inspected items overlap, are made of mixed materials, or have similar shapes, leading to significant identification ambiguity and a high risk of missed detections and misjudgments. Three-view X-ray security inspection equipment can acquire projection images from multiple directions, compensating for the shortcomings of single- and dual-view imaging. However, existing recognition schemes do not effectively utilize the geometric correlation of multi-view projections, making it difficult to accurately establish effective feature interactions between different perspectives. They also cannot rely on geometric constraints to filter out interference from irrelevant areas, resulting in poor effectiveness of multi-view feature fusion and ultimately reducing the accuracy of security inspection and recognition in complex item scenarios. Summary of the Invention

[0003] The purpose of this application is to provide a three-view X-ray security inspection image recognition method, device, equipment and medium to solve the technical problems of insufficient single and dual-view imaging dimensions which easily lead to recognition ambiguity, lack of geometric constraint attention mechanism in three-view recognition, poor cross-view feature interaction and fusion effect and low security inspection recognition accuracy.

[0004] To achieve the above objectives, this application provides a three-view X-ray security inspection image recognition method, which predefines any one of the three views corresponding to the three-view X-ray security inspection image as the original view, and the remaining two views as non-original views; including: Extract the original image features from the original viewpoint; construct the security inspection geometric position code based on the security inspection target geometric vector; optimize the original image features based on the security inspection geometric position code to obtain the security inspection geometric enhancement features from the original viewpoint; Based on the security check cross-view distance between the original view and any non-original view, a security check cross-view mask is constructed; based on the security check cross-view mask, supplementary security check features are generated to characterize the cross-view attention constrained by physical geometric relationships under the original view. For the original perspective, the security inspection geometric enhancement features and all security inspection supplementary features are integrated to obtain the cross-perspective interaction features of the original perspective; the cross-perspective interaction features of the three perspectives are spliced ​​together to generate three-perspective features that adapt to the physical geometric relationship constraints of cross-perspective attention. The three-view features are input into a preset detection head for recognition processing, and the recognition results of the three-view X-ray security inspection image are output.

[0005] To achieve the above objectives, this application also provides a three-view X-ray security inspection image recognition device, which applies the three-view X-ray security inspection image recognition method described above, including: The security inspection geometric enhancement feature module is configured to: extract the original image features from the original viewpoint; construct the security inspection geometric position code based on the security inspection target geometric vector; and optimize the original image features based on the security inspection geometric position code to obtain the security inspection geometric enhancement features from the original viewpoint. The security check supplementary feature module is configured to: construct a security check cross-view mask based on the security check cross-view distance between the original view and any non-original view; and generate security check supplementary features based on the security check cross-view mask to characterize the cross-view attention constrained by physical geometric relationships under the original view. The three-view feature module is configured to: for the original viewpoint, fuse the security inspection geometric enhancement features and all security inspection supplementary features to obtain the cross-view interaction features of the original viewpoint; and splice the cross-view interaction features of the three viewpoints to generate three-view features that adapt to the cross-view attention constrained by physical geometric relationships. The three-view recognition module is configured to input three-view features into a preset detection head for recognition processing and output the recognition results of the three-view X-ray security inspection image.

[0006] To achieve the above objectives, this application also provides a three-view X-ray security inspection image recognition device, including at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which calls the program instructions to execute the three-view X-ray security inspection image recognition method as described above.

[0007] To achieve the above objectives, this application also provides a three-view X-ray security inspection image recognition medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the three-view X-ray security inspection image recognition method as described above.

[0008] Beneficial effects: The three-view X-ray security inspection image recognition method, device, equipment and medium of this application solve the technical problems of insufficient single and dual-view imaging dimensions which easily lead to recognition ambiguity, lack of geometric constraint attention mechanism in three-view recognition, poor cross-view feature interaction and fusion effect and low security inspection recognition accuracy. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart of the three-view X-ray security inspection image recognition method provided in this embodiment; Figure 2 This is a structural block diagram of the three-view X-ray security inspection image recognition device provided in this embodiment.

[0011] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0012] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0013] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0014] Three-view X-ray security inspection equipment can simultaneously acquire projected images of the inspected object from three independent acquisition directions: front, side, and top. The three sets of images naturally possess geometric complementarity in spatial dimensions, which theoretically can effectively alleviate the information loss defects of single / dual-view imaging. However, the current mainstream recognition and fusion methods for three-view security inspection images are mostly direct feature stitching, fixed-weight weighted averaging, or post-hoc voting of recognition results. These shallow fusion strategies do not explore the inherent spatial geometric correspondence constraints between different views under the X-ray projection imaging mechanism, resulting in low utilization of effective correlation information between views.

[0015] Existing dual-view attention interaction schemes for security inspection images are only compatible with dual-optical-path imaging architectures and cannot adapt to the unique three-dimensional spatial geometric complementarity characteristics of three-view devices, making it difficult to fully explore the spatial correlation between the three sets of projected images. While the general Transformer global attention mechanism can achieve global feature interaction, it suffers from huge computational overhead and is prone to establishing incorrect feature matching in irrelevant areas without spatial projection correlation, making it unsuitable for the real-time hard requirements of recognition and inference speed in security inspection scenarios.

[0016] Based on the above-mentioned current state of technology, this embodiment discloses a three-view X-ray security inspection image recognition method, device, equipment and medium, aiming to solve the technical problems of insufficient single and dual-view imaging dimensions which easily lead to recognition ambiguity, lack of geometric constraint attention mechanism in three-view recognition, poor cross-view feature interaction and fusion effect, and low security inspection recognition accuracy.

[0017] The three-view X-ray security inspection image recognition method, apparatus, equipment and medium of this embodiment will now be described.

[0018] Reference Figure 1 , Figure 1 This is a flowchart of the three-view X-ray security inspection image recognition method provided in this embodiment.

[0019] For ease of explanation, this embodiment predefines any one of the three perspectives corresponding to the three-view X-ray security inspection image as the original perspective, and the remaining two perspectives as non-original perspectives. It should be noted that this is only a logical definition for clearly distinguishing each imaging perspective. In actual engineering deployment, it is not necessary to classify the perspectives as original and non-original. It is only necessary to assign independent labels to the three groups of perspective images to complete the distinction.

[0020] In this specific application, the three-view X-ray security inspection image recognition method is applicable to three-view X-ray security inspection equipment. The three views include view one, view two, and view three. For ease of description, the three views are respectively denoted as: in: This indicates perspective one; This represents a second perspective; This indicates a third perspective; This represents any one of the three perspectives. In a simple example, if... If the original perspective is used, then the remaining two perspectives are... and All of these are non-original perspectives.

[0021] A three-view X-ray security inspection machine images the same package being inspected, obtaining images from viewpoint one, viewpoint two, and viewpoint three. The three input images are denoted as follows: in: Represents an X-ray image from a specific viewpoint; This represents an X-ray image from a second viewpoint. This represents a three-view X-ray image. For uniform representation, the input image from any viewpoint is denoted as... ,Right now Indicates perspective The corresponding X-ray image, and This is the original input image. It should be noted that in practical applications, the X-ray image can be a grayscale image or a pseudo-color X-ray image. When the image is a pseudo-color image, its channel information typically includes color-coded information related to material properties.

[0022] Before proceeding with the subsequent steps, in practical applications, the three-view images need to be preprocessed. The preprocessing of the three-view images used in this embodiment will now be described.

[0023] The three images are subjected to size normalization, grayscale or pseudo-color normalization, noise suppression, and intensity correction, respectively. The preprocessing process can be represented as follows: in: Indicates perspective The image after preprocessing; Indicates perspective The original input image; This represents an image preprocessing function. In a specific application, the image preprocessing function... This includes the following operations: in: This indicates an image resizing operation; This represents the image pixel normalization operation. After preprocessing, the three images have a uniform input size and numerical range, which facilitates subsequent feature extraction and cross-view fusion.

[0024] Firstly, such as Figure 1 As shown, this embodiment discloses a three-view X-ray security inspection image recognition method, including: To ensure that features from different perspectives are in the same semantic space, this embodiment obtains basic features from three perspectives, namely the original image features, through a shared feature extraction network.

[0025] S10: Extract the original image features from the original viewpoint; construct the security inspection geometric position code based on the security inspection target geometric vector; optimize the original image features based on the security inspection geometric position code to obtain the security inspection geometric enhancement features from the original viewpoint.

[0026] In the specific application of this embodiment, extracting the original image features corresponding to the original viewpoint involves inputting the three preprocessed images into a shared feature extraction network to obtain the corresponding image features. This process can be represented as follows: in: Indicates perspective Basic image features; Indicates perspective The image after preprocessing; This represents a feature extraction network. In a specific application, the same feature extraction network is shared by three perspectives. This feature extraction network Convolutional neural networks, Transformer networks, or a hybrid structure of convolutional neural networks and Transformers can be used. Specifically, the first... If the layer feature map is used for subsequent cross-view interaction, it can be denoted as: in: Indicates perspective In the Feature map of the layer; Indicates the first The height of the layer feature map; Indicates the first Width of the layer feature map; Indicates the number of feature channels; Indicates the size is The real number tensor space. It should be noted that, for the sake of simplicity, this embodiment will use... Abbreviated as .

[0027] The following describes the security inspection geometric position coding and its application in this embodiment, namely, constructing security inspection geometric position coding based on the geometric vector of the security inspection target, and optimizing the original image features based on the security inspection geometric position coding to obtain the security inspection geometric enhancement features of the original viewpoint.

[0028] Specifically, the security inspection geometric position encoding is determined based on a preset geometric encoding function and combined with the security inspection target geometric vector; wherein, for any feature point in the original view, the security inspection target geometric vector includes normalized horizontal pixel coordinates, normalized vertical pixel coordinates, the view identification encoding of the original view, the three-dimensional ray direction corresponding to the feature point, the parameters of the ray entering the inspection volume, and the parameters of the ray leaving the inspection volume.

[0029] This embodiment is based on security inspection geometric position coding, which introduces the physical imaging geometric relationship of the three-view X-ray security inspection machine into the network, so that the network can understand the correspondence of ray projection between different viewpoints.

[0030] In the specific application of this embodiment, the geometric calibration parameters for each viewpoint are obtained. Specifically, the three-view security inspection machine is geometrically calibrated to obtain the X-ray source position, detector plane position, and detector direction parameters for each viewpoint. For any viewpoint perspective The corresponding geometric calibration parameters are expressed as follows: in: Indicates perspective The set of geometric calibration parameters; Indicates perspective The position of the X-ray source in the three-dimensional coordinate system of the equipment; Indicates perspective The detector plane reference point; Indicates perspective The transverse unit direction vector of the detector plane; Indicates perspective The longitudinal unit direction vector of the detector plane; Indicates perspective The horizontal pixel physical spacing; Indicates perspective The vertical pixel physical spacing; Indicates perspective The original image height; Indicates perspective The original image width.

[0031] The normal vector of the detector plane is defined as: in: Indicates perspective The detector plane normal vector; × represents the vector cross product operation.

[0032] The effective inspection volume of the security inspection machine is recorded as: in: Indicates the three-dimensional spatial area where the inspected package may exist; and Indicates the volume of the inspection. The boundary in the direction; and Indicates the volume of the inspection. The boundary in the direction; and Indicates the volume of the inspection. The boundary in a direction.

[0033] The feature map locations are mapped to the pixel coordinates of the original image. In this embodiment, the downsampling stride of the current feature map used for cross-view interaction relative to the original image is labeled as follows: The downsampling step size Indicates the first The downsampling factor of the layer feature map relative to the original image.

[0034] For perspective Any feature point in the feature map Its row and column coordinates in the feature map are: in: This represents a feature point in the feature map; Representing feature points Row coordinates in the feature map; Representing feature points The column coordinates in the feature map. At this point, the feature points... The corresponding pixel coordinates in the original image are: in: Indicates perspective Middle feature points The corresponding horizontal pixel coordinates of the original image; Indicates perspective Middle feature points The corresponding vertical pixel coordinates of the original image; This indicates the location of the center of the feature unit.

[0035] A 3D ray is obtained by backprojecting from pixel coordinates. Specifically, it is obtained by backprojecting from the pixel coordinates of the original image. Calculate the three-dimensional position of the original image pixels on the detector plane: in: Indicates perspective Middle feature points The corresponding three-dimensional points of the detector; Indicates perspective The detector plane reference point; This indicates the horizontal pixel coordinates of the detector reference point in the image; This indicates the vertical pixel coordinates of the detector reference point in the image; other symbols have the same meaning as described above. (From the X-ray source) and detector points The direction of the unit ray is obtained: in: Indicates perspective Middle feature points The corresponding unit ray direction; This represents the L2 norm of a vector. Therefore, it is clear that feature points... The corresponding three-dimensional ray is: in: Represents a three-dimensional point on a ray; Represents ray parameters, This indicates that the ray extends from the ray source towards the detector. Calculate feature points. Corresponding 3D X-ray and effective inspection volume The intersection of these parameters yields the parameters for the entry and exit of the X-rays into the inspection volume, including the parameters for the X-rays entering the inspection volume. Parameters of the volume of X-rays being examined. Therefore, the X-ray segment located within the effective inspection volume is: in: Indicates perspective Middle feature points Within the effective radiation range of the inspection volume .

[0036] In this embodiment, the above data processing is a data processing method for generating security check geometric location codes.

[0037] Regarding the geometric vector of the security check target, in this embodiment, for the viewpoint... Feature points in Construct the geometric vector of the security inspection target. The mathematical expression of this geometric vector is: in: Indicates perspective Middle feature points The geometric vector of the security inspection target; Represents the normalized horizontal pixel coordinates; Represents the normalized vertical pixel coordinates; Indicates perspective Viewpoint identifier encoding; This indicates the direction of the three-dimensional ray corresponding to the feature point; The parameter indicating the amount of radiation entering the inspection volume; This indicates the parameter representing the ray's departure from the inspection volume. In a specific example, the normalized lateral pixel coordinates of this embodiment... and normalized vertical pixel coordinates It is obtained based on the following two formulas: in: Indicates perspective The original image width; Indicates perspective The original image height.

[0038] For security check geometric location encoding, the encoding is determined based on a preset geometric encoding function and the geometric vector of the security target. In this embodiment, the geometric vector of the security target is input into the geometric encoding function to obtain the security check geometric location code: in: Indicates perspective Middle feature points Security check geometric location coding; This represents the preset geometric coding function, which can be a multilayer perceptron, a sine / cosine coding function, or a combination of both.

[0039] To optimize the original image features based on security check geometric position coding and obtain security check geometric enhancement features from the original viewpoint, in this embodiment, the security check geometric position coding is added to the corresponding original image features: in: Indicates perspective Middle feature points The original image features, i.e., the original image features; This represents the features after adding security check geometric location coding, i.e., security check geometric enhancement features.

[0040] S20: Construct a security check cross-view distance based on the original view and any non-original view; based on the security check cross-view mask, generate supplementary security check features to characterize cross-view attention constrained by physical geometric relationships under the original view.

[0041] Regarding the security check cross-view mask in this embodiment.

[0042] Specifically, the security check cross-view mask is determined based on the security check cross-view distance and a preset geometric tolerance scale on the feature map. For the original view, the calculation of the security check cross-view distance includes: obtaining the feature map center coordinates of any feature point in any non-original view; determining each three-dimensional sampling point in the original view, and obtaining the projection coordinates of each three-dimensional sampling point after projecting it onto the feature map of the corresponding feature point in the corresponding non-original view; and calculating each geometric distance based on the feature map center coordinates and each projection coordinate, and finding the minimum value to obtain the security check cross-view distance.

[0043] For ease of explanation, in this embodiment, it is referred to as and To represent different perspectives, specifically, for any two different perspectives and and ,in: For example: if the perspective As the original perspective, the non-original perspective is the perspective. ; if the original perspective for Then it is not the original perspective. Corresponding to or .

[0044] Regarding the determination of each 3D sampling point in the original viewpoint, the viewpoint Feature points Uniform sampling on its effective ray segment Three-dimensional points: in: Indicates perspective Middle feature points The Three-dimensional sampling points; Indicates perspective Location of the X-ray source; Indicates perspective Middle feature points The parameters of the X-ray entering the inspection volume; Indicates perspective Middle feature points The parameters of the ray leaving the inspected volume; Indicates perspective Middle feature points The direction of the unit ray; Indicates the number of sampling points; Indicates the sampling point number, and 3D sampling points Projected to viewpoint From the image plane, we obtain the two-dimensional projected coordinates: in: Represents the 3D point-to-viewpoint relationship. The projection function of the image plane; Representing a three-dimensional point Projected to viewpoint The horizontal pixel coordinates after; Representing a three-dimensional point Projected to viewpoint The vertical pixel coordinates are then shown. Corresponding to the feature map scale, the projected coordinates are: in: This represents the projection of 3D sampling points onto the viewpoint. The Projected coordinates of the layer feature map; Indicates the first The downsampling step size of the layer feature map.

[0045] Regarding obtaining the feature map center coordinates of any feature point in any non-original viewpoint, the viewpoint... Middle feature points The feature map coordinates are: in: Indicates perspective Middle feature points The coordinates of the center of the feature map; Representing feature points The row coordinates; Representing feature points The column coordinates.

[0046] At this point, based on the center coordinates of the feature map and each projected coordinate, the geometric distances are calculated and the minimum value is obtained to obtain the cross-view distance of the security check, i.e., the feature point is calculated. To view Middle feature points The minimum distance between the projected sampling point sets: in: Indicates perspective Middle feature points Perspective Middle feature points The distance between security checkpoints across different viewing angles; This indicates that the minimum value is taken among all sampled points; This represents the Euclidean distance.

[0047] At this point, the security check cross-view mask is determined based on the security check cross-view distance and the preset geometric tolerance scale on the feature map; that is, the security check cross-view mask is generated according to the geometric distance. in: Indicates perspective Middle feature points To view Middle feature points The value of the security check cross-view mask; This indicates the cross-view distance between the two security checkpoints; Indicates the preset perspective The The geometric tolerance scale on the layer feature map. Therefore, this embodiment is based on cross-view masking for security checks. The following constraint is imposed: if two feature points are geometrically closer, then Smaller A value close to 0 indicates weaker inhibition of attention; if two feature points are geometrically far apart, then... The larger the absolute value of the negative value, the stronger the inhibition of attention. In applications targeting three-view X-ray security inspection images, i.e., for three-view systems, geometric masks in six directions need to be constructed: Each of them Both are used to constrain the viewpoint To view Cross-perspective attention computation.

[0048] Specifically, the security check supplementary features are generated based on the security check cross-view mask and combined with attention interaction.

[0049] In the specific application of this embodiment, a three-view cross-view attention interaction is performed; specifically, in obtaining geometrically enhanced features... Cross-view mask for security checks Subsequently, cross-perspective attention interaction is performed on the three-view features, still combining the aforementioned different perspectives. and Explanation: For perspective To view For cross-perspective attention, the query matrix, key matrix, and value matrix are first computed: in: Indicates perspective The query matrix; Indicates perspective The key matrix; Indicates perspective The value matrix; Indicates perspective The geometrically enhanced features, i.e., the viewpoint The geometric enhancement features of security checks; Indicates perspective The geometrically enhanced features, i.e., the viewpoint The geometric enhancement features of security checks; This indicates a query for a linear mapping matrix; Represents a linear mapping matrix of the keys; The value is represented by a linear mapping matrix.

[0050] Thus, the supplementary security check features are obtained based on cross-view attention calculation, and the corresponding mathematical expression is: in: Indicates perspective From a perspective Obtained supplementary security inspection features from multiple perspectives; This represents the normalized exponential function; Key matrix transpose; Indicates the channel dimension of the query vector and key vector; This represents the scale normalization factor; Indicates perspective To view The geometric mask matrix, i.e., the security inspection cross-view mask.

[0051] S30: For the original viewpoint, integrate the security inspection geometric enhancement features and all security inspection supplementary features to obtain the cross-view interaction features of the original viewpoint; stitch together the cross-view interaction features of the three viewpoints to generate three-view features that adapt to the physical geometric relationship constraints of cross-view attention.

[0052] In the application of three-view X-ray security imaging, for each viewpoint, information from the other two views is simultaneously received. For example, viewpoint one... The cross-perspective interaction features are: in: This refers to the features of viewpoint 1 after cross-viewpoint interaction, i.e., cross-viewpoint interaction features. This represents the geometrically enhanced features of viewpoint one, i.e., the security inspection geometrically enhanced features; This indicates the supplementary security check features obtained from perspective 2 by perspective 1. This refers to the supplementary security check features obtained from viewpoint 3 from viewpoint 1. All supplementary security check features include... and ; This represents the fusion weight of information from perspective two to perspective one; This represents the fusion weight of information from viewpoint three with that from viewpoint one. Similarly, we can obtain the fusion weight of information from viewpoint two. and perspective three Post-interaction features: in: This indicates a second perspective. Features resulting from cross-perspective interaction, i.e., perspective two Cross-perspective interactive features; Perspective 3 Features resulting from cross-perspective interaction, i.e., perspective three Cross-perspective interactive features; This represents the fusion weight of information from viewpoint one to viewpoint two; This represents the fusion weight of information from perspective three to information from perspective two; This represents the fusion weight of information from viewpoint one to viewpoint three; This represents the fusion weight of the information from viewpoint two to viewpoint three. It should be noted that this embodiment... , , , , and These are learnable parameters generated by a gating network.

[0053] Thus, this embodiment provides the data foundation for a three-view X-ray security inspection image recognition technology solution based on geometrically constrained cross-view attention.

[0054] S40: Input the three-view features into the preset detection head to perform recognition processing, and output the recognition result of the three-view X-ray security inspection image.

[0055] In the specific application of this embodiment, the detection head settings are determined based on the task type, and the identification results may include, but are not limited to, hazardous material categories, identification confidence levels, target detection boxes, and target segmentation masks.

[0056] Unlike existing technologies, the three-view X-ray security inspection image recognition method in this embodiment targets X-ray images, which have severe color stacking. From the perspective of imaging principles, it is basically impossible to use conventional feature point recognition and matching methods to obtain depth information. Therefore, this embodiment enhances the feature association and expression between different perspectives of contraband based on the three-view approach, thereby constructing a more discriminative three-dimensional structural representation of contraband without relying on depth maps. Furthermore, this embodiment only relies on the X-ray images from each perspective and calculates the association attention through prior knowledge of the internal structure.

[0057] In the practical application of this embodiment, it was found that the identification of hazardous materials in X-ray images depends not only on shape and structure, but also on material density gradients, edge pseudocolor, grain texture, and subtle high-frequency variations. Traditional convolutional neural networks tend to weaken these high-frequency material textures during downsampling, resulting in insufficient ability to distinguish between material categories such as liquids, powders, plastics, and rubber. Therefore, this embodiment optimizes the design for high-frequency textures in three-view X-ray security inspection image recognition.

[0058] Specifically, the cross-view interaction features are transformed from the spatial domain to the frequency domain to obtain the frequency domain cross-view enhanced interaction features; the cross-view enhanced interaction features of the three views are stitched together to generate a three-view fusion feature with frequency domain enhancement and physical geometric relationship constraint on cross-view attention; the three-view fusion feature is input into the detection head to perform recognition processing, and the recognition result of the frequency domain enhanced three-view X-ray security inspection image is output.

[0059] It should be noted that, compared with the existing technology which relies more on the context and attention mechanism of the model itself, this embodiment focuses on designing a multi-view position encoding method for real physical spatial position relationships. This allows the model to more accurately and comprehensively consider the characteristics of a certain area under different views, and performs fusion enhancement in the frequency domain, making it more sensitive to material edges, density changes and subtle textures.

[0060] In the specific application of this embodiment, the cross-viewpoint enhanced interaction feature in the frequency domain is represented as: Combined with the aforementioned cross-perspective interaction features In practice, there are technical problems such as high computational load and failure to highlight important features. To address this, this embodiment introduces a lightweight fusion module.

[0061] Specifically, lightweight fusion is performed before the three-view feature input to the detection head is used for recognition processing and before the three-view fused feature input to the detection head is used for recognition processing.

[0062] Enhanced Interaction Features from a Cross-Perspective in this Embodiment This describes the lightweight fusion implemented in this embodiment.

[0063] The frequency domain enhancement cross-view interaction features from the three perspectives are then stitched together: in: This represents the features after stitching together three perspectives; This indicates splicing along the channel dimension; Perspective 1 Frequency domain enhancement of cross-view interaction features; Perspective 2 Frequency domain enhancement of cross-view interaction features; Perspective 3 Frequency domain enhancement of cross-view interaction features. In response to the need to reduce computational cost and highlight important features, a lightweight fusion module is introduced: in: This represents the three-view features after lightweight fusion; This indicates a lightweight fusion module. In one specific implementation, Includes channel-space attention and depthwise separable convolution: Among them: CSA ( ) represents the channel spatial attention module; DSConv This represents a depthwise separable convolution module. Thus, this embodiment achieves three-view information fusion with relatively low computational cost.

[0064] The frequency domain cross-viewpoint enhanced interactive features of this embodiment will now be described in detail.

[0065] Specifically, the step of performing spatial-frequency domain transformation on the cross-view interaction features to obtain frequency-domain enhanced cross-view interaction features includes: processing the cross-view interaction features using a preset frequency-domain enhancement weight map and preset frequency-domain enhancement intensity coefficients, combined with Fourier transform, to obtain frequency-domain enhanced spatial features; and optimizing the cross-view interaction features based on the frequency-domain enhanced spatial features to obtain frequency-domain enhanced cross-view interaction features.

[0066] In the specific application of this embodiment, in response to the identification of dangerous goods in X-ray security inspection images, it not only depends on shape information, but also on the actual situation of material edges, density changes and fine textures. This embodiment performs frequency domain enhancement on the features after cross-view interaction.

[0067] For any perspective Features after cross-perspective interaction Perform frequency domain texture enhancement.

[0068] First, perform a two-dimensional Fourier transform: in: Indicates perspective Frequency domain characteristics; Indicates Fast Fourier Transform; Indicates perspective Cross-perspective interactive features.

[0069] Enhancement of frequency domain features: in: This represents the enhanced frequency domain characteristics; This represents element-wise multiplication; This represents the preset frequency domain enhancement weighting map; This represents the preset frequency domain enhancement intensity coefficient.

[0070] In this embodiment, the frequency domain enhancement weight map It can be generated by a learnable function: in: This represents the preset frequency domain weighting generation function; The amplitude spectrum represents the frequency domain characteristics.

[0071] Then, the inverse Fourier transform is used to return to the spatial domain: in: This represents the spatial characteristics after frequency domain enhancement; This represents the inverse fast Fourier transform.

[0072] To avoid excessive noise amplification, this embodiment uses the residual form: in: The residual enhancement weights can be fixed values ​​or learnable parameters. To enhance the cross-view interaction features obtained in the frequency domain.

[0073] Now, we combine cross-perspective enhanced interaction features corresponding The recognition results of the three-view X-ray security inspection image output by performing recognition processing on the input detection head in this embodiment are illustrated with an example.

[0074] Three-view fusion features Input the detector head, output the hazardous materials identification result: in: This indicates the final security check and identification result; Indicates the detection output head; This indicates the fusion feature of three perspectives.

[0075] Depending on the task type, the recognition results It can include: in: Indicates the category of dangerous goods; Indicates the confidence level of identification; Represents the target detection bounding box; This represents the target segmentation mask.

[0076] When the task is a categorized task It may include only category and confidence; when the task is a detection task, It can include categories, confidence levels, and bounding boxes; when the task is a segmentation task... It can further include a segmentation mask.

[0077] The model training and loss function design of this embodiment will now be explained.

[0078] During the training phase, supervised learning is used to optimize the model. The total loss function is expressed as: in: This represents the total loss of the model; Indicates loss of the main task; This indicates cross-perspective consistency loss; Indicates frequency domain regularization loss; The weight representing the cross-perspective consistency loss; The weights represent the frequency domain regularization loss.

[0079] The primary task loss is determined based on the specific task. For example, for classification tasks, cross-entropy loss can be used: in: Indicates the number of categories; Indicates a category index; Indicates the true label in the category The value of ; This indicates that the model predicts a category. The probability of.

[0080] Cross-view consistency loss is used to constrain predictions of the same object from different viewpoints to remain consistent, and can be expressed as: in: Indicates perspective The prediction results or intermediate features are represented; Indicates perspective The prediction results or intermediate features are represented; This represents the squared Euclidean distance between two viewpoint predictions or features.

[0081] Frequency domain regularization loss is used to prevent the frequency domain enhancement module from over-amplifying noise, and can be expressed as: in: Indicates perspective Frequency domain enhancement weighting map; This represents the square norm of the frequency domain enhancement weights.

[0082] By jointly optimizing the above loss functions, the model can simultaneously possess the ability to identify hazardous materials, cross-perspective consistency, and stable frequency domain enhancement capabilities.

[0083] In summary, the three-view X-ray security inspection image recognition method of this embodiment includes: first, acquiring three-view X-ray images and extracting basic features of viewpoint 1, viewpoint 2, and viewpoint 3 through a shared feature extraction network to obtain three-view features; then, generating security inspection geometric position codes based on the X-ray source, detector plane, projection angle, and imaging geometric parameters of the security inspection machine, and constructing a security inspection cross-view mask; next, using a three-view cross-view attention module to achieve fully connected interaction of the three-view features under the constraints of the geometric mask; further, adaptively enhancing the high-frequency texture of the material through a frequency domain texture enhancement module; finally, using a channel spatial attention module and a depthwise separable convolution module for lightweight fusion, outputting a three-view fused feature map for detection, classification, or segmentation to complete the three-view X-ray security inspection image recognition. Thus, the three-view X-ray security inspection image recognition method of this embodiment provides a three-view X-ray security inspection image recognition technology based on geometrically constrained cross-view attention and frequency domain enhancement, to solve the problems of isolated viewpoint information, occlusion ambiguity, insufficient high-frequency material texture recognition, large global attention computation, and poor cross-device scene generalization ability in existing technologies.

[0084] Compared with ordinary three-view image stitching or voting methods, this embodiment utilizes ray projection geometry to establish cross-view position coding and geometric masks, achieving feature interaction under physical consistency constraints.

[0085] Compared with the dual-view attention method, this embodiment constructs a three-view fully connected cross-view attention for view 1, view 2 and view 3, which can make fuller use of the naturally complementary information of the three views.

[0086] Compared to ordinary global Transformer attention, this embodiment limits the attention search space through geometric masks, reducing invalid computations and erroneous associations.

[0087] Compared with the single-view frequency domain enhancement method, this embodiment embeds frequency domain high-frequency enhancement into a three-view geometric fusion framework, which can simultaneously utilize the complementary structure of multiple views and the texture features of X-ray materials.

[0088] Reference Figure 2 , Figure 2 The diagram shows the structure of the three-view X-ray security inspection image recognition device provided in this embodiment; in the diagram: 10, security inspection geometric enhancement feature module; 20, security inspection supplementary feature module; 30, three-view feature module; 40, three-view recognition module.

[0089] Secondly, such as Figure 2 As shown, this embodiment also discloses a three-view X-ray security inspection image recognition device, which applies the three-view X-ray security inspection image recognition method described above, including: The security inspection geometric enhancement feature module 10 is configured to: extract the original image features from the original viewpoint; construct the security inspection geometric position code based on the security inspection target geometric vector; optimize the original image features based on the security inspection geometric position code to obtain the security inspection geometric enhancement features from the original viewpoint. The security check supplementary feature module 20 is configured to: construct a security check cross-view mask based on the security check cross-view distance between the original view and any non-original view; and generate security check supplementary features based on the security check cross-view mask to characterize the cross-view attention constrained by physical geometric relationships under the original view. The three-view feature module 30 is configured to: for the original viewpoint, fuse the security inspection geometric enhancement features and all security inspection supplementary features to obtain the cross-view interaction features of the original viewpoint; and splice the cross-view interaction features of the three viewpoints to generate three-view features that adapt to the physical geometric relationship constraints of cross-view attention. The three-view recognition module 40 is configured to input three-view features into a preset detection head for recognition processing and output the recognition result of the three-view X-ray security inspection image.

[0090] Thirdly, this embodiment also discloses a three-view X-ray security inspection image recognition device, including at least one processor, at least one memory and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which calls the program instructions to execute the three-view X-ray security inspection image recognition method as described above.

[0091] Fourthly, this embodiment also discloses a three-view X-ray security inspection image recognition medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the three-view X-ray security inspection image recognition method as described above.

[0092] It should be noted that the three-view X-ray security inspection image recognition device, equipment, and medium of this embodiment correspond to the aforementioned three-view X-ray security inspection image recognition method. Therefore, any content not specifically described in the three-view X-ray security inspection image recognition device, equipment, and medium of this embodiment, including but not limited to functional definitions, working principles, and technical effects, can be referred to the description in the aforementioned three-view X-ray security inspection image recognition method, and will not be repeated here.

[0093] In summary, the three-view X-ray security inspection image recognition method, apparatus, equipment, and medium of this embodiment have at least the following technical effects: By utilizing the known geometric relationships of ray projection from a three-view security inspection machine, the feature correspondence between different views is explicitly constrained, thereby reducing invalid associations and erroneous matches.

[0094] By using three-view fully connected cross-view attention, full interaction is achieved between view 1, view 2 and view 3, mitigating recognition ambiguity caused by occlusion and overlap.

[0095] Frequency domain texture enhancement strengthens the high-frequency material texture, density gradient, and false-color edges in X-ray images, improving the ability to identify fine-grained materials such as liquids, powders, and plastics.

[0096] By using geometric masks, CBAM, and depthwise separable convolutions, computational load can be controlled and real-time performance improved.

[0097] Improve the model's generalization ability across different devices, scenarios, and placement postures by using cross-view consistency loss.

[0098] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.

[0099] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A three-view X-ray security inspection image recognition method, wherein any one of the three views corresponding to the three-view X-ray security inspection image is predefined as the original view, and the remaining two views are non-original views; characterized in that, include: Extract original image features from the original viewpoint; Based on the geometric vector of the security inspection target, a security inspection geometric position code is constructed. Based on the security inspection geometric position code, the original image features are optimized to obtain the security inspection geometric enhancement features from the original viewpoint. The security inspection geometric position encoding is determined based on a preset geometric encoding function and combined with the security inspection target geometric vector. For any feature point in the original view, the security inspection target geometric vector includes normalized horizontal pixel coordinates, normalized vertical pixel coordinates, the view identification encoding of the original view, the three-dimensional ray direction corresponding to the feature point, the parameters of the ray entering the inspection volume, and the parameters of the ray leaving the inspection volume. Based on the security check cross-view distance between the original view and any non-original view, a security check cross-view mask is constructed; based on the security check cross-view mask, supplementary security features are generated to characterize the cross-view attention constrained by physical geometric relationships under the original view; the security check cross-view mask is determined based on the security check cross-view distance and the geometric tolerance scale on the preset feature map; in: Indicates perspective Middle feature points To view Middle feature points The value of the security check cross-view mask; This indicates the cross-view distance between the two security checkpoints; Indicates the preset perspective The The geometric tolerance scale on the layer feature map; where, for the original viewpoint, the calculation of the security check cross-viewpoint distance includes: obtaining the feature map center coordinates of any feature point in any non-original viewpoint; determining each three-dimensional sampling point in the original viewpoint, and obtaining the projection coordinates of each three-dimensional sampling point after being projected onto the feature map of the corresponding feature point in the corresponding non-original viewpoint; based on the feature map center coordinates and each projection coordinate, calculating each geometric distance and finding the minimum value to obtain the security check cross-viewpoint distance; For the original perspective, the security inspection geometric enhancement features and all security inspection supplementary features are integrated to obtain the cross-perspective interaction features of the original perspective; the cross-perspective interaction features of the three perspectives are spliced ​​together to generate three-perspective features that adapt to the physical geometric relationship constraints of cross-perspective attention. The three-view features are input into a preset detection head for recognition processing, and the recognition results of the three-view X-ray security inspection image are output.

2. The three-view X-ray security inspection image recognition method according to claim 1, characterized in that, The security check supplementary features are generated based on security check cross-view masks and combined with attention interaction.

3. The three-view X-ray security inspection image recognition method according to claim 1, characterized in that, The cross-view interaction features are transformed from spatial to frequency domain to obtain the frequency domain cross-view enhanced interaction features. By splicing the cross-view enhanced interaction features of the three perspectives, a three-view fusion feature with frequency domain enhancement and physical geometric relationship constraint on cross-view attention is generated; The three-view fusion features are input into the detection head for recognition processing, and the recognition result of the frequency domain enhanced three-view X-ray security inspection image is output.

4. The three-view X-ray security inspection image recognition method according to claim 3, characterized in that, The step of performing spatial-frequency domain transformation on the cross-view interaction features to obtain frequency-domain enhanced cross-view interaction features includes: processing the cross-view interaction features using a preset frequency-domain enhancement weight map and preset frequency-domain enhancement intensity coefficients, combined with Fourier transform, to obtain frequency-domain enhanced spatial features; and optimizing the cross-view interaction features based on the frequency-domain enhanced spatial features to obtain frequency-domain enhanced cross-view interaction features.

5. The three-view X-ray security inspection image recognition method according to claim 3, characterized in that, Lightweight fusion is performed before the three-view feature input to the detection head for recognition processing and before the three-view fused feature input to the detection head for recognition processing.

6. A three-view X-ray security inspection image recognition device, employing the three-view X-ray security inspection image recognition method as described in any one of claims 1 to 5, characterized in that, include: The security inspection geometric enhancement feature module is configured to extract original image features from the original viewpoint. Based on the geometric vector of the security inspection target, a security inspection geometric position code is constructed. Based on the security inspection geometric position code, the original image features are optimized to obtain the security inspection geometric enhancement features from the original viewpoint. The security check supplementary feature module is configured to: construct a security check cross-view mask based on the security check cross-view distance between the original view and any non-original view; Based on the security inspection cross-view mask, supplementary security inspection features are generated to characterize cross-view attention constrained by physical geometric relationships under the original view. three The viewpoint feature module is configured to: for the original viewpoint, fuse the security inspection geometric enhancement features and all security inspection supplementary features to obtain the cross-viewpoint interaction features of the original viewpoint; By splicing together the cross-view interaction features of the three perspectives, a three-view feature adapted to cross-view attention constrained by physical geometric relationships is generated. The three-view recognition module is configured to input three-view features into a preset detection head for recognition processing and output the recognition results of the three-view X-ray security inspection image.

7. A three-view X-ray security inspection image recognition device, characterized in that, Includes at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the three-view X-ray security inspection image recognition method according to any one of claims 1 to 5.

8. A three-view X-ray security inspection image recognition medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the three-view X-ray security inspection image recognition method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Sparse view three-dimensional reconstruction method and system based on dynamic occlusion perception and multi-prior fusion

    CN120411412A

  • Characteristic fusion-based contraband detection method and system for dual-view X-ray image

    CN121170714A