Multi-view target detection and model training method, device, equipment, storage medium and program product
By introducing position coding features related to imaging angles into X-ray images, the problem of image feature mismatch in multi-view target detection is solved, and the detection precision and recognition accuracy are improved.
Patent Information
- Application Number
- CN202510905618.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-02
AI Technical Summary
X-ray images from different perspectives are difficult to match, which leads to feature mismatch and low recognition accuracy in multi-view target detection.
Position coding features related to imaging angles are introduced to compensate for the impact of imaging angle changes on image features and thus achieve image feature fusion.
The detection accuracy of multi-view target detection is improved, the feature mismatch problem of images with different viewpoints is solved, and the accuracy of target recognition is enhanced.
Smart Images

Figure CN120411485B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of X-ray imaging, and specifically to a multi-view target detection and model training method, device, equipment, storage medium and program product thereof. Background Art
[0002] X-ray imaging is a technology that uses X-rays to penetrate an object and generate images of its internal structures based on the differences in absorption of the rays by different tissues or materials. During imaging, X-rays are emitted by an emitter, penetrate the object, and reach a detector. The detector records the intensity of the transmitted rays, converts them into digital signals, and generates an image.
[0003] Given its non-invasive and penetrating nature, X-ray imaging is widely used in radiography, industrial inspection, and safety checks. Given the unique characteristics of X-ray imaging, multi-angle imaging can be performed using X-ray imaging equipment to obtain three-dimensional data of the object being measured. However, due to the principles of X-ray imaging, matching X-ray images from different perspectives is difficult. Therefore, effectively utilizing X-ray images from different perspectives is a pressing technical challenge for those skilled in the art. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a multi-perspective target detection and its model training method, device, equipment, storage medium and program product, which compensates for the influence of different imaging angles on multi-perspective target recognition by introducing position coding features related to imaging angles, thereby enabling target recognition based on multi-perspective fusion features.
[0005] In a first aspect, the present application provides a multi-view target detection method, which includes: determining multiple X-ray images of the same object to be detected acquired by an imaging device at different imaging viewpoints. For a target image in the multiple X-ray images, feature extraction is performed on the target image to determine the image features of the target image. Based on the imaging geometric parameters corresponding to the target image and the position information of the image features of the target image, the position coding features of the target image are determined to determine the position coding features of each X-ray image, wherein the imaging geometric parameters of the X-ray image include the imaging angle and the detection distance. For a first image and a second image in the multiple X-ray images, the perspective fusion features of at least one first image and the second image are determined based on the imaging geometric parameters, image features and position coding features of the first image and the imaging geometric parameters, image features and position coding features of the second image. The target detection result is determined by a target recognition model based on the perspective fusion features.
[0006] In a second aspect, the present application provides a multi-view target detection device. The multi-view target detection device includes an image acquisition module, a feature extraction module, a position encoding module, a feature fusion module, and a detection execution module. The image acquisition module is configured to determine multiple X-ray images of the same object to be detected, acquired by an imaging device at different imaging perspectives. The feature extraction module is configured to perform feature extraction on a target image in the multiple X-ray images to determine image features of the target image. The position encoding module is configured to determine position encoding features of the target image based on imaging geometric parameters corresponding to the target image and position information of the image features of the target image, thereby determining position encoding features of each X-ray image. The imaging geometric parameters of the X-ray image include imaging angle and detection distance. The feature fusion module is configured to determine, for a first image and a second image in the multiple X-ray images, a view fusion feature of at least one of the first and second images based on the imaging geometric parameters, image features, and position encoding features of the first image and the imaging geometric parameters, image features, and position encoding features of the second image. The detection execution module is configured to determine a target detection result using a target recognition model based on the view fusion feature.
[0007] In a third aspect, the present application provides a model training method based on multi-view target detection. The model training method includes: obtaining at least one set of training data, wherein the training data includes multiple X-ray images of the same object to be detected acquired by an imaging device at different imaging perspectives, and target identification labels in the multiple X-ray images. For target images in the multiple X-ray images in the training data, feature extraction is performed on the target images to determine image features of the target images. Position coding features of the target images are determined based on imaging geometric parameters corresponding to the target images and position information of the image features of the target images, thereby determining position coding features of each X-ray image in the training data. For a first image and a second image in the multiple X-ray images, perspective fusion features of at least one of the first image and the second image are determined based on the imaging geometric parameters, image features, and position coding features of the first image, and the imaging geometric parameters, image features, and position coding features of the second image. Target detection results are determined using a target recognition model based on the perspective fusion features, wherein the imaging geometric parameters of the X-ray images include imaging angle and detection distance. The target recognition model is trained based on the target identification labels and the target detection results.
[0008] In a fourth aspect, the present application provides a model training device based on multi-view target detection. The model training device includes: a data acquisition module, a target detection module, and a parameter iteration module. The data acquisition module is configured to acquire at least one set of training data, wherein the training data includes multiple X-ray images of the same object to be detected, acquired by an imaging device at different imaging perspectives, and target identification labels in the multiple X-ray images. The target detection module is configured to perform feature extraction on target images in the multiple X-ray images in the training data, determine image features of the target images, and determine position coding features of the target images based on imaging geometric parameters corresponding to the target images and position information of the image features of the target images, thereby determining position coding features of each X-ray image in the training data. For a first image and a second image in the multiple X-ray images, a perspective fusion feature is determined for at least one of the first image and the second image based on the imaging geometric parameters, image features, and position coding features of the first image, and the imaging geometric parameters, image features, and position coding features of the second image. A target detection result is determined using a target recognition model based on the perspective fusion feature, wherein the imaging geometric parameters of the X-ray images include imaging angle and detection distance. The parameter iteration module is configured to train the target recognition model based on the target identification labels and the target detection results.
[0009] In a fifth aspect, the present application provides an electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multi-view object detection method described in the first aspect or the model training method based on multi-view object detection described in the third aspect.
[0010] In a sixth aspect, the present application provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a processor to implement the multi-view target detection method described in the first aspect or the model training method based on multi-view target detection described in the third aspect when executed.
[0011] In the seventh aspect, the present application provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer implements the multi-view target detection method described in the first aspect or the model training method based on multi-view target detection described in the third aspect.
[0012] The present application provides a multi-view target detection and its model training method, apparatus, equipment, storage medium and program product. These methods address the technical problem that X-ray images at different viewpoints are difficult to match during target recognition based on multi-view X-ray imaging. By introducing position coding features related to the imaging angle into the image features of the X-ray image, the position information contained in the feature image of the X-ray image is corrected based on the imaging angle, thereby eliminating the influence of the viewpoint change on the X-ray image. Furthermore, the image features of the X-ray images at different viewpoints can be fused based on the aforementioned position coding features, so that target recognition and classification can be performed based on the fused features, thereby improving the detection accuracy of the multi-view target detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 This is a schematic diagram of a multi-view imaging process provided by some embodiments of the present application.
[0015] Figure 2 This is a system module diagram of a multi-view target detection device provided in some embodiments of the present application.
[0016] Figure 3 This is an exemplary flowchart of the multi-view target detection method provided in some embodiments of the present application.
[0017] Figure 4 This is an exemplary flowchart of a method for determining position coding features provided in some embodiments of the present application.
[0018] Figure 5 This is an exemplary flowchart of a method for determining first-person perspective fusion features provided in some embodiments of the present application.
[0019] Figure 6 This is a system module diagram of a model training device provided in some embodiments of the present application.
[0020] Figure 7 This is a data diagram of the target recognition model training process provided by some embodiments of the present application.
[0021] Figure 8 This is an exemplary flowchart of the target recognition model training method provided in some embodiments of the present application.
[0022] Among them, 100, multi-view target detection system; 110, imaging device; 120, processing device; 130, imaging space; 200, multi-view target detection device; 210, image acquisition module; 220, feature extraction module; 230, position encoding module; 240, feature fusion module; 250, detection execution module; 260, model training module; 600, model training device; 610, data acquisition module; 620, target detection module; 630, parameter iteration module; 710, target recognition model to be trained; 720, training data; 721, multiple X-ray images; 722, target recognition label; 731, target detection result. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] Application Overview:
[0025] Based on the aforementioned background technology, in practical applications, X-ray imaging is often used for target detection. For example, in the field of security inspection, X-rays can be used to check for prohibited items in luggage. Another example is that in the medical field, X-rays can be used to detect nodules and tumors inside the breast.
[0026] Taking into account the principles of X-ray imaging, imaging equipment can perform multi-view target detection on the object to be detected to simulate human stereoscopic vision. In this case, the imaging equipment often includes multiple imaging devices, each of which consists of a set of emitters and detectors positioned opposite each other. In multi-view target detection, the detection system often forms an open space with at least one opening. The imaging devices in the imaging equipment are often arranged perpendicular to the opening (possibly at a certain angle), encompassing the open space within the imaging range.
[0027] For example, in security inspection equipment, the open space is often configured as a box with openings on opposite sides. A conveyor is located between the two openings to transport the object to be inspected (also known as the imaging object). Multi-view security inspection equipment includes at least two sets of imaging devices, which can be located at different locations in the open space, such as the top and side, to obtain X-ray images of the object to be inspected from different perspectives.
[0028] To further illustrate this situation, this application also provides an application scenario diagram of a multi-view target detection system ( Figure 1 ).
[0029] like Figure 1 As shown, a multi-view object detection system 100 (security inspection system) may include two imaging devices 110 (referred to as a first imaging device 111 and a second imaging device 112 ), a processing device 120 , and an imaging space 130 . During object detection, an object to be detected may be placed within the imaging space 130 (e.g., transported via a conveyor belt). The imaging device 110 may collect imaging data of the object within the imaging space 130 based on the principles of X-ray imaging, and then generate an X-ray image through the processing device 120 .
[0030] by Figure 1 For example, the X-rays released by the aforementioned first imaging device 111 can be distributed in the vertical direction in the imaging space 130 to obtain a top-illuminated image of the object to be measured. The X-rays released by the second imaging device 112 can be distributed in the horizontal direction in the imaging space 130 to obtain a side-illuminated image of the object to be measured, thereby observing the object to be measured from multiple angles, simulating the stereoscopic vision of humans when observing objects, and thus improving recognition accuracy. Therefore, the aforementioned multi-view imaging can provide X-ray images of the imaging object at different perspectives. In the field of security inspection, multi-view imaging can obtain more morphological and structural information of the object to be measured. In particular, experienced security inspectors rely on the information complementarity between multi-angle images to judge the true shape and structure of the object, thereby effectively identifying complexly stacked and severely obscured contraband.
[0031] In conventional scenarios, the imaging postures of each group of imaging devices in a multi-view target detection system are fixed, and the mapping of the X-ray images they collect to the imaging space is also fixed. Multi-view image fusion can be achieved by calibrating this fixed mapping relationship.
[0032] However, in the actual multi-view target detection process, the imaging angles of imaging devices with different perspectives often change. Existing detection models are mostly designed for single-view images and lack mechanisms for multi-view information alignment and complementary modeling. In the case of multi-view input, common practices include simply splicing, averaging, or feeding two-view images into parallel networks before fusing the outputs. These methods are unable to properly handle issues such as spatial offset, occlusion differences, and texture distortion caused by changes in imaging angles between different perspectives, which can easily lead to feature mismatches or distraction, further causing false detections and missed detections.
[0033] In addition, methods such as those based on the twin Faster R-CNN structure and introducing a bidirectional attention module mainly perform matching at the feature level in actual processing logic. Although this type of method has alleviated some feature alignment problems to a certain extent, it still does not explicitly model the geometric relationship between the two perspectives in space, resulting in limited fusion effects. Especially in X-ray images, due to the penetrating nature of imaging, it is very common for objects of different categories to be highly superimposed or partially occluded. If we only rely on image-level semantic features and ignore their consistency in space, it is difficult to effectively identify contraband with complex structures or blurred boundaries. Therefore, there is an urgent need to design a lightweight and efficient feature fusion mechanism that can perceive spatial structural constraints to truly achieve deep collaborative understanding of dual-perspective information.
[0034] Furthermore, given that the emitter is a point light source, the detection distance between the radiation emitted by the emitter and the detector varies with the angle of the radiation. Due to the mechanism of X-ray imaging, changes in detection distance will lead to changes in the X-ray image, which can further lead to inconsistencies in image features.
[0035] In response to the technical problem that X-ray images at different perspectives are difficult to match during target recognition based on multi-perspective X-ray imaging, the present application provides a multi-perspective target detection and its model training method, device, equipment, storage medium and program product. By introducing position coding features related to the imaging angle into the image features of the X-ray image, the position information contained in the feature image of the X-ray image is corrected based on the imaging angle, thereby eliminating the influence of the imaging perspective change on the X-ray image, and then the image features of the X-ray images at different perspectives can be fused based on the aforementioned position coding features, so that the target recognition and classification can be based on the fused features, thereby improving the detection accuracy of the multi-perspective target detection method.
[0036] The target detection method provided in the present application is not only applicable to imaging devices with multiple imaging devices, but also X-ray images of the same object to be detected acquired at different time sequences or with a single light source and multiple viewing angles can be regarded as images at different viewing angles and the method provided in the present application can be similarly applied.
[0037] The foregoing Figure 1 The processing device 120 shown in the figure can serve as an execution device for the multi-view target detection method provided in this application. After acquiring multiple X-ray images, it performs image feature fusion and target recognition based on the multi-view target detection method provided in this application. Specifically, the processing device 120 (or electronic device) can include, at the implementation level, at least one processor and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multi-view target detection method provided in this application.
[0038] The processor can process data and / or information obtained from other devices or system components. The processor can execute program instructions based on these data, information and / or processing results to perform one or more functions described in this application. In some embodiments, the processor may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-core processing device). By way of example only, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic circuit (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), a microprocessor, etc., or any combination thereof.
[0039] The memory can be used to store data and / or instructions. The memory may include one or more storage components, each of which may be an independent device or part of another device. In some embodiments, the memory may include random access memory (RAM), read-only memory (ROM), mass storage, removable memory, volatile read-write memory, etc. or any combination thereof. Exemplarily, the mass storage may include a magnetic disk, an optical disk, a solid-state disk, etc. In some embodiments, the memory may be implemented on a cloud platform. Data refers to a digital representation of information and may include various types, such as binary data, text data, image data, video data, etc. Instructions refer to programs that can control a device or component to perform a specific function.
[0040] In some embodiments, the aforementioned memory can be implemented through a computer-readable storage medium. By storing computer instructions in the computer-readable storage medium, when the computer instructions are called by the processor, the processor can implement the multi-view target detection method provided in this application when executing the computer instructions.
[0041] It should be noted that the processing device 120 can be connected to the imaging device 110 via wired or wireless communication to execute the methods provided herein based on the corresponding data. Furthermore, some of the methods herein may involve model training, which can be performed by the processing device 120 or another computing device. The trained model is stored in the memory of the processing device 120 to facilitate execution of the multi-view object detection methods provided herein.
[0042] In some embodiments, the multi-perspective target detection method provided in the present application can be characterized as a multi-perspective target detection device, which can be written into a storage medium / called by a processor as a program product / code constructed as a multi-functional module to implement the multi-perspective target detection method provided in the present application.
[0043] To further illustrate the aforementioned multi-view target detection device, the present application also provides a system module diagram of a multi-view target detection device ( Figure 2 ).
[0044] like Figure 2 As shown, the multi-view target detection apparatus 200 may include an image acquisition module 210 , a feature extraction module 220 , a position encoding module 230 , a feature fusion module 240 and a detection execution module 250 .
[0045] The image acquisition module 210 can determine multiple X-ray images of the same object to be tested acquired by the imaging device at different imaging angles. For more information on acquiring X-ray images, please refer to step S310 and its related description, which will not be repeated here.
[0046] The feature extraction module 220 can be used to extract features of a target image from a plurality of X-ray images and determine image features of the target image. For more information about image feature extraction, please refer to step S320 and its related description, which will not be repeated here.
[0047] Position encoding module 230 can be configured to determine position encoding features of the target image based on imaging geometric parameters corresponding to the target image and position information of image features of the target image, thereby determining position encoding features of each X-ray image. The imaging geometric parameters of the X-ray image include imaging angle and detection distance. For more information on position encoding, see step S330 and its related description, and are not further elaborated here.
[0048] Feature fusion module 240 can be configured to determine, for a first image and a second image in the plurality of X-ray images, a perspective fusion feature of at least one of the first image and the second image based on the imaging geometry parameters, image features, and position coding features of the first image, and the imaging geometry parameters, image features, and position coding features of the second image. For more information on feature fusion, see step S340 and its related description, and are not further elaborated here.
[0049] The detection execution module 250 can be used to determine the target detection result through the target recognition model based on the view fusion feature. For more information about target detection, please refer to step S350 and its related description, which will not be repeated here.
[0050] Furthermore, considering that the aforementioned object detection methods (particularly their feature-based object detection steps / modules) often rely on machine learning models for implementation, the aforementioned multi-view object detection apparatus 200 may, in special cases, further include a model training module 260 to provide a suitable machine learning model. The model training module 260 may be configured to iterate the parameters of the various machine learning models or related parameters involved in the aforementioned process based on training samples and sample labels to provide suitable model parameters / algorithm parameters.
[0051] It should be noted that, considering that model training requires high computing power, the aforementioned model training module 260 and the aforementioned other modules can often be integrated into different electronic devices (such as directly as a separate model training device, see Figure 6 In some special cases (such as model fine-tuning), the aforementioned model training module 260 and the aforementioned other modules can be integrated into one electronic device.
[0052] It should be noted that the module names involved in the embodiments of the present application can be defined as other names, as long as the functions of each module can be achieved, and there is no specific restriction on the names of the modules. In addition, it should be understood that the aforementioned device 200 is embodied in the form of a functional module. The term "module" here can refer to an application-specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a merged logic circuit and / or other suitable components that support the described functions. At the practical application level, the aforementioned module can serve as the execution subject of the corresponding step, and the subject omitted in the relevant description of the subsequent steps may be the aforementioned device / module.
[0053] The following combination Figures 3 to 5 The multi-view target detection method provided in this application is described in detail.
[0054] Exemplary multi-view object detection method:
[0055] Figure 3 This is an exemplary flow chart of a multi-view target detection method shown in some embodiments of the present application. Figure 3 The illustrated process may be executed by the aforementioned processing device 120 .
[0056] like Figure 3 As shown, P300 may include the following steps:
[0057] S310: Determine multiple X-ray images of the same object to be tested acquired by an imaging device at different imaging viewing angles.
[0058] S320 , for a target image among the multiple X-ray images, perform feature extraction on the target image to determine image features of the target image.
[0059] S330 , determining a position coding feature of the target image based on imaging geometric parameters corresponding to the target image and position information of image features of the target image, so as to determine position coding features of each X-ray image.
[0060] S340. For a first image and a second image among a plurality of X-ray images, determine a perspective fusion feature of at least one of the first image and the second image based on the imaging geometric parameters, image features and position coding features of the first image and the imaging geometric parameters, image features and position coding features of the second image.
[0061] S350: Determine the target detection result through the target recognition model based on the view fusion feature.
[0062] In the aforementioned S310, the object to be tested may refer to an object undergoing X-ray scanning, or may also be referred to as the object to be tested. Different imaging perspectives may refer to the imaging perspectives (e.g., top-down perspective, side-view perspective, etc.) of different imaging devices (i.e., a combination of emitters and detectors) within an imaging device. Correspondingly, multiple X-ray images may refer to images captured by each imaging device after the object to be tested enters the X-ray imaging device. These images are generally X-ray images captured by different imaging devices to capture the internal structure of the object to be tested from multiple perspectives.
[0063] It should be noted that if a penetrating image imaging technology similar to the aforementioned X-ray image imaging principle appears in subsequent technologies, the technical solution provided in this application can also be adaptively applied.
[0064] In some embodiments, S310 can be implemented by data transmission between imaging devices with different viewing angles within an imaging device and a processing device. In practical applications, this process is often real-time. That is, the X-ray images captured by each imaging device can be synchronously transmitted to the processing device, allowing the processing device to execute S310 and subsequent steps in real time to determine the individual items contained within the object under test at that moment.
[0065] Furthermore, considering that the models involved in P300 often also require the execution of the aforementioned P300 process during training, the aforementioned P300 can, in some cases, be performed by the electronic device used for model training. Given the computational requirements of model training, the electronic device used for model training is often different from the electronic device that executes P300 under normal operating conditions. Generally, after the electronic device used for model training completes model training, the trained model is burned into the electronic device that executes P300 under normal operating conditions, enabling it to execute the aforementioned P300.
[0066] In the aforementioned process, imaging geometric parameters may refer to the relevant imaging parameters (generally referred to as imaging extrinsics) of each imaging device in the imaging device when acquiring X-ray images. In actual X-ray imaging, considering that changes in imaging angles can significantly affect the aforementioned X-ray images, the imaging geometric parameters may generally include at least the imaging angles of each imaging device when acquiring X-ray images at different viewing angles.
[0067] Furthermore, considering that the emitter of the imaging device is a point light source, the detection distance between the X-rays emitted by the emitter and the detector will vary with the angle of the rays. Therefore, the aforementioned imaging geometric parameters may further include a detection distance reflecting the distance between the emitter and the detector at a corresponding imaging angle. The detection distance is generally associated with the aforementioned imaging angle.
[0068] In the aforementioned S320, the target image may refer to a single X-ray image among the multiple X-ray images. That is, the aforementioned S320 and S330 may reflect a general processing method for any X-ray image among the multiple X-ray images. In practical applications, there is generally no additional restriction on the selection of the target image.
[0069] Image features can refer to the results of feature extraction from the target image. For example, image features can be represented as feature maps, feature vectors, semantic vectors, and other forms. Specifically, this can be performed using a feature extraction algorithm based on image features. Given that the target image is similar to images captured by traditional cameras at the data level, existing feature extraction methods can be used for feature extraction.
[0070] Overall, image feature extraction can generally be divided into two categories:
[0071] ① Feature extraction based on local regions of interest: This is when extracting features, rather than processing the entire image, local features are extracted from the local region of interest (ROI) to determine the image features corresponding to the region of interest. When performing this type of image feature extraction, it is often necessary to first perform target segmentation (recognition only, not classification) to delineate the region of interest.
[0072] ② Global feature extraction: This involves performing feature extraction processing (such as convolution) on the entire image to determine the image features for the entire image. Considering the subsequent need to configure position encoding features for image features, the image features obtained from global feature extraction are generally presented as feature maps.
[0073] In the aforementioned S330, the position coding feature may be a compensation of the aforementioned image feature based on the imaging geometric parameters to eliminate the impact of changes in the imaging geometric parameters on the image feature. The position coding feature may be characterized as a recalculation of the position parameters of the image feature based on the imaging geometric parameters (primarily the imaging angle).
[0074] For the image features constructed based on the aforementioned local region of interest, the position coding features constructed in S330 can be used to modify the positional relationship between the image features and the X-ray image. The region of interest is typically represented as a framed area in the X-ray image, which can be characterized by one or more coordinates and a region size. The aforementioned position coding features can modify the coordinates and size used to identify the region of interest.
[0075] For the image features determined by the aforementioned global feature extraction, the position coding features constructed in S330 can also be used to recalibrate the coordinates of each pixel in the image. That is, the coordinate data of each pixel point (or other basic units, such as pixel blocks, voxels, etc.) in the image features can be corrected based on the imaging angle. The specific calculation process can be found in Figure 4 and its related descriptions.
[0076] During the coordinate reconstruction process, the position-encoding feature can account for the effects of imaging angle (and detection distance) on image features through extrinsic correction factors. These correction factors are typically presented as correction values for each coordinate / dimension. These extrinsic correction factors can be represented as output values based on imaging angle and detection distance. Their calculation typically relies on a trained model (such as a multilayer perceptron) to reflect the effects of different imaging angles on image features.
[0077] Therefore, based on the aforementioned steps, each X-ray image can be processed with reference to the aforementioned target image processing method to determine the image features and position coding features of each X-ray image.
[0078] In the aforementioned S340, the first image and the second image may refer to two different images that need to be fused and matched among the multiple X-ray images. Similar to the aforementioned target image, the selection of the first image and the second image is generally not limited. In practical applications, considering the principle of X-ray imaging, the multiple X-ray images of the same object to be tested acquired by an imaging device often appear as X-ray images of the same object to be tested under two viewing angles (for example, a top view angle and a side view angle), that is, the aforementioned first image and second image can be directly determined. In the case of acquiring more than two X-ray images, when determining the view angle fusion feature, any two X-ray images can be selected to construct all combinations that can be used for image fusion. In addition, an X-ray image based on a reference view angle (such as the top view angle / side view angle in each view angle) can also be combined with other X-ray images to obtain the fusion feature when the other X-ray images are fused into the reference view angle X-ray image.
[0079] The perspective fusion feature of the first image and the second image can be understood as the image feature that fuses the information of the first image and the second image. Based on common feature fusion methods, the aforementioned perspective fusion features can theoretically include three categories:
[0080] ① The fusion feature that takes the image features of the first image as the main body and loads the characteristic information of the image features of the second image can also be recorded as the perspective fusion feature of the first image based on the second image (hereinafter referred to as the first perspective fusion feature).
[0081] ② The fusion feature that takes the image features of the second image as the main body and loads the characteristic information of the image features of the first image can also be recorded as the perspective fusion feature of the second image based on the first image (hereinafter referred to as the second perspective fusion feature).
[0082] ③ For the fusion features that are indistinguishable from the first image and the second image, they can generally be processed through synchronous feature extraction (such as treating different images as different channels of the same image and then performing convolution processing), splicing and other operations to generate an overall fusion feature.
[0083] When actually executing the aforementioned S340, the matching relationship between the first image and the second image may be determined first, and then the image features may be fused. For example, taking the image features of the aforementioned local region of interest as an example, the features of the two images may be matched based on the aforementioned position coding features and the semantic information of the features themselves, and the two image features corresponding to the matching results may be fused.
[0084] In some embodiments, for the aforementioned global convolution results, when performing feature fusion, a query-key matching method (i.e., a query-key matching algorithm) can be used to model the corresponding regions across viewpoints. Specifically, attention can be used to perform fault-tolerant modeling of spatial structural differences. This eliminates the need for manual labeling of the same object in different images for the corresponding training data, reducing training difficulty and cost, and improving redundancy. For more information on query-key matching, see Figure 5 and its related descriptions.
[0085] It should be noted that the processing logic of the aforementioned global convolution result in the aforementioned S340 is matching and fusion, which only needs to meet this requirement. Query-Key matching is only a preferred embodiment, and other algorithms that can achieve this requirement can also execute the aforementioned S340.
[0086] In the aforementioned S350, the target recognition model is generally a machine learning model capable of performing target detection based on image features. It can generally identify the types of various objects from the original image (such as the aforementioned X-ray image) using input feature data (such as the aforementioned view fusion features) (generally presented as multiple object identification boxes in the image and the classification probabilities of the local images within those boxes, with the classification with the highest probability generally being considered the type corresponding to the local image). Accordingly, the target detection result is generally determined by inputting the aforementioned determined view fusion features into the target recognition model. The target recognition model generally includes a classifier (such as an SVM), which can be used to map low-dimensional, nonlinearly separable data into a high-dimensional space using a kernel function (such as an RBF or polynomial kernel) to determine the probabilities of the corresponding objects belonging to different types.
[0087] Based on the aforementioned image fusion, multiple target frames containing the identified objects and classification information of the objects in the target frames can be output in the aforementioned S350. The specific calculation process can be referred to the existing technology, and this application will not be repeated here.
[0088] In some embodiments, taking into account the various types of feature fusion, the aforementioned S350 can be executed based on the actual situation of feature fusion. For example, if the aforementioned feature fusion outputs only one set of features, the aforementioned S350 can determine a target recognition result based on a single feature. For example, if the aforementioned feature fusion outputs multiple sets of features, the aforementioned S350 can determine multiple sets of target recognition results based on the multiple sets of features, and then perform comprehensive consideration. Among them, the specific algorithm of comprehensive consideration can be adjusted based on actual needs. For example, in the identification of prohibited items in security inspections, for prohibited items, if they are detected in a set of features, it is determined that prohibited items exist.
[0089] Taking the simultaneous determination of the first-perspective fusion feature and the second-perspective fusion feature as an example, when executing the aforementioned S350, the first-perspective fusion feature can be input into the target recognition model to determine the first detection result; the second-perspective fusion feature can be input into the target recognition model to determine the second detection result; and the target detection result can be determined based on the first detection result and the second detection result.
[0090] Furthermore, considering that there may be more than two imaging angles, if feature fusion is performed for any combination in S340, similar operations may be performed on all fusion results. If multiple fusion features based on a certain perspective are determined by fusion of other perspectives, similar operations may also be used during target recognition. When features from other perspectives are migrated to the same perspective, the fusion of target detection results may include the fusion of target recognition frames.
[0091] Therefore, in order to address the technical problem that X-ray images at different perspectives are difficult to match during target recognition in multi-perspective X-ray imaging, the aforementioned multi-perspective target detection method P300 can introduce position coding features related to the imaging angle into the image features of the X-ray image, so that the position information contained in the feature image of the X-ray image is corrected based on the imaging angle, thereby eliminating the influence of the imaging angle change on the fusion of X-ray images of different perspectives, and further enabling the image features of the X-ray images of different perspectives to be fused based on the aforementioned position coding features, so that target recognition and classification can be performed based on the fused features, thereby improving the detection accuracy of the multi-perspective target detection method.
[0092] Exemplary Position Encoding Feature Determination Method:
[0093] In practical applications, X-ray imaging generally displays objects based on penetration intensity, essentially creating a grayscale image. However, in practical applications (especially security inspections), images are often colored based on penetration intensity to facilitate inspection by personnel. Accordingly, the aforementioned X-ray images are configured as multi-channel image data based on penetration intensity.
[0094] When performing global convolution feature extraction, the convolution kernel can have a certain cross-channel processing capability (i.e., multi-channel convolution kernel), extracting high-dimensional information from the X-ray image to obtain image features. In other words, image features can be configured as the global convolution result of the X-ray image, which can specifically include multiple channel features of the same size (actually image data).
[0095] As mentioned above, the spatial distribution of pixels within the global convolution can be adjusted by position coding features. To further illustrate this process, the present application also provides an exemplary flow chart of a method for determining position coding features ( Figure 4 ).
[0096] like Figure 4 As shown, process P400 may include the following steps:
[0097] S410 , normalizing the position information of the image features of the target image to determine normalized coordinate features.
[0098] S420: Determine position coding features based on normalized coordinate features and imaging geometric parameters.
[0099] For the convenience of explanation, for the X-ray image I, the solution after feature extraction can be recorded as image feature F The feature extraction process can be understood as The image size of the X-ray image I can be recorded as H0×W0, which is generally presented as a grayscale image or a multi-channel image (such as a traditional RGB image based on penetration intensity). The image features can satisfy .Right now F The image size is H×W and the number of channels is C.
[0100] It should be noted that the aforementioned P400 processes image data simultaneously across all channels, effectively a mapping relationship based on pixel position. Therefore, the position-encoding feature can also be represented as multidimensional data of size H×W to represent the results of each pixel mapping. In actual processing, the position-encoding feature is generally presented as two-dimensional pixels (i.e., two channels). If the X-ray image contains spatial information (such as an integrated depth detector to identify the depth of the imaging object surface), the position-encoding feature can be represented as three-dimensional data (the processing process is similar to that of the present application).
[0101] In the aforementioned S410, the normalized coordinate feature may also refer to the result of converting the pixel positions in each channel of the image feature to the range [0,1]. Normalized coordinates can avoid the influence of too large a range on coordinate processing. Specifically, the normalization process can be performed based on the size of the aforementioned image feature, that is, for the image feature F The coordinates of the point (i, j) in the image after the processing in S410 can be (i / H, j / W), which can be represented as P ij , which can also be recorded as the basic position encoding vector.
[0102] In S420, considering that the effect of imaging angle on the image is often nonlinear, the process of determining the position encoding features based on the normalized coordinate features and imaging geometric parameters can be performed using a spatial mapping model. The spatial mapping model is a machine learning model with learnable parameters. Its parameters can be determined by collaborative training with the object recognition model to reflect the nonlinear effect of imaging angle on the image.
[0103] Considering that the imaging geometric parameters of the imaging device are mainly characterized by the imaging angle and the detection distance, the imaging angle and the detection distance can be used as an input of the spatial mapping model, thereby generating a position coding feature in combination with the normalized coordinate features determined above.
[0104] In actual calculations, the normalized coordinate features can be used as input to the spatial mapping model, allowing the spatial mapping model to output correction values for each position to form the position coding feature. Alternatively, the normalized coordinate features can be omitted from the spatial mapping model, allowing the spatial mapping model to directly output the overall correction value for each position based on the imaging angle and detection distance, and then the position coding feature is calculated using the normalized coordinate features.
[0105] In the case where the normalized coordinate features are not input into the spatial mapping model, the result output by the spatial mapping model based on the imaging angle and the detection distance can be recorded as the external parameter correction factor, recorded as Q. In some embodiments, when the imaging angle and the detection distance are input into the spatial mapping model, the sine value and the cosine value can be used for normalization representation. The above processing process can be recorded as , among which, MLP geo θ can represent the spatial mapping model, θ can represent the imaging angle, and d can represent the detection distance. In practical applications, the spatial mapping model can be implemented using a small multi-layer perceptron. For example, a two-layer FFN can be used to implement the aforementioned spatial mapping model. Furthermore, if θ is a three-dimensional angle, then in the aforementioned process, θ can be resolved into two two-dimensional values in space and then input into the spatial mapping model.
[0106] Specifically, the aforementioned imaging angle can be described as needed. For example, the aforementioned imaging angle can be described based on a baseline imaging angle of the imaging device. The baseline imaging angle of the imaging device can reflect the angular deviation between the current X-ray direction and the direction of the ray when the emitter and detector in the imaging device are perpendicular. For another example, the aforementioned imaging angle can be directly described as the current angle value between the emitter and the detector.
[0107] Therefore, when determining the position coding feature, the normalized coordinate feature can be superimposed with the external parameter correction factor, so the position coding feature E here is ij =P ij + Q ij , where Q ij The value of the extrinsic parameter correction factor Q at pixel (i, j) can be obtained.
[0108] In addition, in order to limit the influence of the model, the aforementioned superposition process can be a weighted process of the external parameter correction factor, that is, the position coding feature E ij =P ij +αQ ij , where α is a scaling parameter that reflects the effect of the extrinsic correction factor on the encoding feature. For example, the scaling parameter can be set to 0.1.
[0109] If the normalized coordinate features are not input into the spatial mapping model, they can be processed directly through the spatial mapping model, that is, .
[0110] It should be noted that the form of the final position coding feature can be set according to actual needs. ij, which can be represented as a 2-dimensional vector to represent the corrected coordinates of the corresponding pixel point; it can also be represented as a C-dimensional vector, using a single non-coordinate value to represent the location or offset of the pixel points of each channel; it can also be represented as an n×C multi-dimensional vector, representing the corrected position of the pixels of each channel in a multi-dimensional manner (e.g., 2D). In addition, if the aforementioned X-ray image is a 3D image, the aforementioned position encoding feature can also be represented as a 3D vector.
[0111] Furthermore, the value of the detection distance itself varies greatly. In order to narrow the effect on the result, the detection distance can be normalized during the aforementioned spatial correction. Therefore, the aforementioned S420 can further include the following steps:
[0112] S421 . Based on the maximum detection distances of the plurality of X-ray images, normalize the detection distances corresponding to the target image to determine a normalized distance of the target image.
[0113] S422: Determine an extrinsic parameter correction factor based on the normalized distance of the target image and the imaging geometric parameters.
[0114] S423. Superimpose the normalized coordinate features and the external parameter correction factors to determine the position coding features.
[0115] In the aforementioned S421 and S422, the detection distance can refer to the vertical distance between the detector and the emitter, or it can refer to the distance from the emitter to the sensor after the angle changes. Considering that the imaging angle has been introduced into the model, the processing of the detection distance can also be adjusted according to the situation.
[0116] The normalization in the aforementioned S421 is similar to the above normalization, that is, mapping the detection distance to [0,1], which can be specifically represented as Based on the superposition processing of S423, the processing of S422 can be characterized as follows: .
[0117] In the aforementioned S430 , the normalized coordinate features of the aforementioned image features and the external parameter correction factors may be superimposed to determine the position coding features.
[0118] In addition, the aforementioned spatial mapping model can also process each position during processing, thereby directly outputting the E of each position. ij .at this time .
[0119] Therefore, based on the above steps, the influence of the imaging angle on the X-ray image can be transferred to the image features for subsequent feature fusion.
[0120] Furthermore, in subsequent processing, the image features and the position coding features can be fused into one feature, or processed as two separate features. Taking the fusion of two features as an example, . Based on the above description, Can be added as an additional channel After the features of That is, the composite feature of the target image can be determined based on the image feature and the position coding feature of the target image, wherein the image feature and the position coding feature are configured as channel data of different channels of the composite feature.
[0121] Example feature fusion process:
[0122] Based on the aforementioned global convolution results, query-key matching can be used for feature matching. Query-key matching is the core process in the attention mechanism, used to calculate the similarity or correlation between the query and a set of keys, thereby determining which values should be focused on.
[0123] In query-key matching, the query represents the user input (such as a search keyword), the key represents the attribute to be matched (such as a document feature), and the value represents the actual content (such as a search result). The matching goal is to assign weights to the query through similarity calculations, and then aggregate the weighted values to generate the final output. This process is widely used in natural language processing (such as the Transformer model), recommendation systems, and search engines.
[0124] Query-Key matching involves three core matrices, all of which are generated by linear transformation of the input embedding vector:
[0125] The query matrix (Q) represents the set of query vectors, and its dimensions are usually n × d , where n is the number of queries and d is the feature dimension.
[0126] The key matrix (K) represents a set of key vectors, typically of dimension m × d, where m is the number of keys. Keys represent attributes to be matched (such as document features) and are used to compare similarity with the query.
[0127] The Value matrix (V) represents a vector set of values, typically of dimension m × d. Value is the actual content (such as text or data), which is weighted and output based on the query-key similarity weight.
[0128] These three matrices are obtained by linearly transforming the input embedding vector X through the weight matrix:
[0129] Q=X⋅W Q , K=X⋅W K , V=X⋅W V , where W Q , W K , W V It is a learnable weight matrix with the same size, ensuring that the Q, K, and V dimensions are compatible.
[0130] After determining the above three matrices based on the input content, the query-key similarity can be determined, which is generally performed through the dot product, that is, S=Q·K T , and then normalized and passed through the SoftMax function to output its corresponding weight matrix. That is, the weight matrix A = SoftMax(S / d), where d is the dimension of the key, to prevent the dot product value from being too large and causing the gradient to disappear (called scaled dot product attention).
[0131] Subsequently, the normalized weight matrix A can be multiplied by the Value matrix (V) to obtain the final output, that is, Output = A·V.
[0132] The present application uses the aforementioned relationship to process the matching of image features of two X-ray images. Specifically, for the first image and the second image, when the image features of the second image are transferred to the first image, the image features of the first image (and position coding features, such as ) as the query, and the image features of the second image (and position encoding features, such as ) as the key and value, thereby determining, through query-key matching, the image features of the second image that can be fused with the features of the first image when the image features of the first image are "queried" (i.e., the cross-view fusion features from the second image). In other words, a query key matching algorithm can be used to determine the cross-view fusion features from the second image. The first-view fusion features are then determined based on the image features and position coding features of the first image and the cross-view fusion features from the second image. The query sequence of the query key matching algorithm is determined based on the image features of the first image, and the key sequence and value sequence of the query key matching algorithm are determined based on the image features of the first image.
[0133] In some embodiments, considering that image features are often represented as multidimensional pixel matrices, and the aforementioned queries, keys, and values are often represented as multiple feature points, filtering and / or conversion can be performed during configuration. For example, multiple feature points can be selected from the image features. In another example, the multidimensional pixel matrix can be converted into a sequence based on the aforementioned position coding features (e.g., the C-dimensional position coding features).
[0134] To further illustrate the query-key matching process of the first image and the second image, the present application also provides an exemplary flow chart of a method for determining a first perspective fusion feature ( Figure 5 ). As mentioned above, the first perspective fusion feature is the perspective fusion feature of the first image based on the second image, and the second perspective fusion feature can be processed based on the same logical correspondence.
[0135] like Figure 5 As shown, process P500 may include the following steps:
[0136] S510 : Determine a first query sequence, a second key sequence, and a second value sequence based on the query weight matrix, the key weight matrix, and the value weight matrix of the query key matching algorithm, as well as the query sequence, the key sequence, and the value sequence.
[0137] S520. Determine the attention encoding value based on the first query sequence and the second key sequence.
[0138] S530. Determine a cross-view attention weight from the first image to the second image based on the attention encoding value.
[0139] S540: Determine a cross-view fusion feature of the second image based on the cross-view attention weight and the second value sequence.
[0140] S550: Determine a first-view fusion feature based on the image feature and position coding feature of the first image and the cross-view fusion feature from the second image.
[0141] In the aforementioned S510, the weight matrix, key weight matrix and value weight matrix are important parameters that need to be trained in Query-Key matching (i.e. the aforementioned three matrices). In this application, their identifiers are consistent with the aforementioned ones, also W Q , W K , W V .
[0142] In practical applications, the aforementioned weight matrix, key weight matrix, and value weight matrix are often operators with pre-configured parameters. The parameters of these three matrices can be determined through collaborative training with the target recognition model. For details, see the matrix parameter determination algorithm in related art and will not be elaborated here.
[0143] Taking into account that the embedding vector X of the aforementioned Query-Key matching process is often presented as a d-dimensional vector, the corresponding embedding vectors can also be configured based on the aforementioned first image and second image before performing the aforementioned P500. Taking composite features as an example, for the composite features of the first image and the composite features of the second image that have been determined, the composite features of the first image can be converted into a query sequence, and the composite features of the second image can be converted into a key sequence and a value sequence. The specific conversion process is often to take one or more pixels as elements of a vector to achieve the conversion from matrix to vector, which will not be elaborated here. In addition, considering actual needs, it is also possible not to construct an embedding vector based on composite features, but only to construct an embedding vector based on image features, and introduce position coding features in the subsequent processing process. This solution can also make the finally determined cross-view fusion features involve the aforementioned position coding features.
[0144] The first query sequence is configured as the dot product of the image features of the first image, the position encoding features and the query weight matrix. That is, the first query sequence Q1= ⋅W Q Similarly, in the aforementioned S530, the second key sequence is configured as the dot product of the image feature of the second image, the position encoding feature and the key weight matrix. That is, the second key sequence K2= ⋅W K .
[0145] In the aforementioned S520, the attention encoding value can represent the similarity between the image features of the first image and the image features of the second image. Based on the aforementioned Query-Key matching, the attention encoding value includes the dot product similarity between the first query sequence and the second key sequence. That is, the attention encoding value = , where d2 is the dimension of the key vector (i.e., key sequence). Similarly, in S530, the attention encoding value can be processed by SoftMax to determine the cross-view attention weight from the first image to the second image. .
[0146] In the aforementioned S540, the cross-view fusion features from the second image .
[0147] Thus, the aforementioned S510 to S530 can determine the cross-view fusion features from the second image determined based on the query of the first image. In the aforementioned S540, the image features of the first image can be fused with the aforementioned cross-view fusion features to determine the first-view fusion features for object recognition.
[0148] The fusion of the two features mentioned above can be achieved by direct superposition or other superposition methods. For example, feature fusion can be achieved by residual connection and Layer Norm. Therefore, the first-person perspective fusion feature can be characterized as follows:
[0149] .
[0150] It should be noted that the composite feature is used in the above process , image features can also be used To achieve the superposition of image content. The following is an example of the second perspective fusion feature:
[0151] Similar to the above, the following parameters can be determined in sequence when determining the second perspective fusion feature:
[0152] Second query sequence Q2= ⋅W Q 、First key sequence K1= ⋅W K Cross-view attention weights from the second image to the first image Then the cross-view fusion features from the first image .
[0153] Therefore, the second perspective fusion feature can be characterized as follows: .
[0154] In some embodiments, in order to further characterize the relative spatial relationship between the first image and the second image, when executing the aforementioned S550, a relative position relationship (relative position encoding) can be further introduced to simultaneously construct a cross-view attention weight based on the attention encoding value and the relative position encoding.
[0155] Therefore, the aforementioned S530 may include the following sub-steps:
[0156] S531. For a first feature point in the first image and a second feature point in the second image, determine a relative position coding value based on the position coding feature value of the first feature point and the position coding feature value of the second feature point to determine the relative position coding of the first image and the second image.
[0157] S532: Input the attention encoding value and the relative position encoding into the activation function to determine the cross-view attention weight from the first image to the second image.
[0158] Similar to the aforementioned position encoding process, S531 can also be implemented using a machine learning model (herein referred to as a relative space model) to account for the nonlinear effects of imaging geometric parameters on the imaging process. Specifically, in S531, the position encoding feature values of the first feature point and the second feature point can be input into the relative space model to output a relative position code.
[0159] Among them, the relative space model is similar to the aforementioned spatial mapping model. It is also a machine learning model with learnable parameters, and its parameters can be determined by collaborative training with the target recognition model.
[0160] When executing the aforementioned S531, similar to the aforementioned first image and second image, the aforementioned first feature point and second feature point can also be represented as a feature point group arbitrarily selected from the two feature point sequences after the image features are converted into feature point sequences. Among them, the first feature point of the first image can be recorded as , and its corresponding position encoding value can be Similarly, the second feature point of the second image The positional encoding value is .
[0161] Then the relative position code value in the aforementioned S531 is Wherein, θ1 is the imaging angle of the imaging device corresponding to the first image, and θ2 is the imaging angle of the imaging device corresponding to the second image.
[0162] In some embodiments, when the aforementioned data is input into the relative spatial model, it can be calculated by difference. That is, when executing the aforementioned S531, the position difference between the first feature point and the second feature point can be determined based on the position encoding feature value of the first feature point and the position encoding feature value of the second feature point. Then, the difference in imaging geometric parameters between the first image and the second image can be determined. Finally, the relative position encoding value is determined based on the position difference and the imaging angle difference.
[0163] Specifically, .in, Indicates the X coordinate difference between the first feature point and the second feature point in the position encoding feature value, is the Y coordinate difference, is the imaging angle difference between the first image and the second image, is the difference between the normalized detection distances of the first image and the second image. It should be noted that all or part of the above parameters can also be directly input in the form of original values instead of differences.
[0164] Specifically, the specific representation of the aforementioned imaging angle difference can be determined in accordance with the description method of the actual imaging process to reflect the imaging angles or differences of each imaging device for the current imaging space. For example, the imaging angle difference can be directly represented as the angle between the X-ray of the imaging device corresponding to the first image and the X-ray of the imaging device corresponding to the second image in the imaging space (which can be represented by a two-dimensional value or by directly calculating the one-dimensional angle between the two in space (i.e., the angle between the two X-rays in the plane where the two X-rays are located). Figure 1In this example, the angle between the two viewing angles is 90°. For another example, during the imaging process, the imaging devices may rotate synchronously to align the object. To simplify the computational complexity in this case, the difference in imaging angles can be directly represented as the imaging angle of one of the imaging devices.
[0165] In the aforementioned S552, based on the relative position encoding value determined above, the weight matrix determination process can be adjusted as a parameter. Thus, the cross-view attention weight from the first image to the second image is It can be calculated as follows:
[0166] The corresponding cross-view attention weights from the second image to the first image It should be noted that in this step, d1 and d2 are the number of feature points of the image features of the first image and the second image, respectively, rather than the aforementioned detection distance.
[0167] Based on the above steps, image feature matching and fusion can be achieved based on query-key matching. In particular, the above position encoding and imaging angle can be introduced during image matching and fusion to reflect the impact of imaging posture on image matching.
[0168] Furthermore, in order to enhance the matching capability of the aforementioned matching process, the loss function can be constructed based on the aforementioned relative position encoding, introducing sparsity constraints, encouraging the concentration of attention distribution, and improving the reliability of matching. Figure 8 and its related descriptions.
[0169] Therefore, based on the aforementioned feature fusion process, the perspective fusion feature determined in this application can be determined by the aforementioned position coding feature, so that the influence of different perspectives on the imaging image can be taken into account during feature fusion, so that both the similarity of the features themselves and their spatial correlation can be considered during feature fusion, thereby improving the representation ability of the fused perspective fusion feature for the image semantics and spatial relationships, and thus enabling the target recognition process based on perspective fusion features to make full use of the semantic information and spatial relationships of different perspectives, thereby improving the accuracy of the target detection results.
[0170] Example model training process:
[0171] Taking into account that the aforementioned target detection process involves multiple machine learning models, the present application also provides a model training method and device based on multi-view target detection to train the various machine learning models and related trainable parameters involved in the aforementioned process.
[0172] Figure 6 This is a system module diagram of a model training device provided in some embodiments of the present application.
[0173] like Figure 6 As shown, the model training device 600 may include a data acquisition module 610, a target detection module 620 and a parameter iteration module 630.
[0174] The data acquisition module 610 can be used to acquire at least one set of training data, wherein the training data includes multiple X-ray images of the same object to be detected acquired by the imaging device at different imaging angles and target identification labels in the multiple X-ray images, and imaging geometric parameters include imaging angle and detection distance. For more information about training data, please refer to Figure 7 The related descriptions are not repeated here.
[0175] The target detection module 620 can be used to extract features of target images from multiple X-ray images in the training data, determine the image features of the target images, and determine the position coding features of the target images based on the imaging geometric parameters corresponding to the target images and the position information of the image features of the target images, so as to determine the position coding features of each X-ray image in the training data. For the first image and the second image in the multiple X-ray images, the perspective fusion features of at least one first image and the second image are determined based on the imaging geometric parameters, image features and position coding features of the first image and the imaging geometric parameters, image features and position coding features of the second image. The target detection result is determined by the target recognition model based on the perspective fusion features. For the target detection process, please refer to the aforementioned P300 and its related descriptions, which will not be repeated here.
[0176] The parameter iteration module 630 can be used to train the target recognition model based on the target recognition label and the target detection result. For more information about parameter iteration, please refer to Figure 7 The relevant contents will not be elaborated here.
[0177] To further illustrate the above process, this application also provides a data diagram of the target recognition model training process ( Figure 7 ).
[0178] like Figure 7 As shown, the model training method provided in this application can train the target recognition model to be trained 710 to obtain a target recognition model with configured parameters. The trained target recognition model can execute the aforementioned P300 to perform target recognition based on the perspective fusion feature and determine the target detection result.
[0179] like Figure 7As shown, the aforementioned training process can be performed based on multiple sets of training data 720. Among them, for a set of training data 720, it may include multiple X-ray images 721 of the same object to be measured acquired by the imaging device at different imaging perspectives and target identification labels 722 in the multiple X-ray images. Among them, the multiple X-ray images 721 can refer to the aforementioned description and will not be repeated here. The target identification label 722 can be an annotation of the recognition result in the X-ray image (such as manual annotation). Among them, considering the association between the multiple X-ray images 721, the aforementioned target identification label 722 can achieve unified annotation of the same object in multiple X-ray images 721 through some features (such as ID) when forming the target identification frame and marking the type.
[0180] The training process based on the aforementioned multiple sets of training data 720 can be characterized as an inference process (i.e., the aforementioned P300) and a parameter iteration process. Specifically, the aforementioned multi-view object detection method can be first executed based on the multiple X-ray images 721 in the training data 720 to determine the target detection results 731 for the target recognition model 710 to be trained. Iterations are then performed based on the differences between the target detection results 731 and the target recognition labels 722. The process of determining the target detection results 731 is similar to the aforementioned P300 and its related description and will not be further elaborated here.
[0181] When iterating the target recognition model 710 to be trained based on the difference between the target detection result 731 and the target recognition label 722, it is generally performed through a corresponding loss function. During the training process of the target recognition model, the loss function is generally configured as a classification loss and a regression loss.
[0182] Classification loss is used to measure the difference between the target category predicted by the model and the actual category. The core is to evaluate the matching degree of probability distribution. Cross Entropy Loss is generally used to calculate the difference between the predicted probability and the actual label. In this application, it can be expressed as .
[0183] Regression loss is used to optimize the target position (such as bounding box coordinates, size, angle), the core of which is to reduce the geometric deviation between the predicted value and the true value. GIoU Loss can be used for calculation. For details, please refer to the common training process of the target recognition model, which will not be described here. In this application, it can be recorded as .
[0184] In some embodiments, given that the aforementioned P300 process may involve multiple unlabeled machine learning models or trainable parameters (such as the aforementioned spatial mapping model, relative spatial model, and various parameter matrices of the matching algorithm), the training process can also be collaboratively trained based on the training process of the aforementioned object recognition model. That is, the object recognition model and its related models can be iterated simultaneously based on the aforementioned loss function. The specific parameter iteration principles can be found in related art and are not detailed here.
[0185] In some embodiments, considering the representational meaning of the aforementioned relative space model, the training of the relative space model can be strengthened in the form of sparsity constraints during the actual training process. In view of this situation, the present application also provides an exemplary flow chart of the target recognition model training method ( Figure 8 ).
[0186] like Figure 8 As shown, process P800 may include the following steps:
[0187] S810: Determine the classification loss and regression loss of the target recognition model based on the target recognition label and the target detection result.
[0188] S820: Determine a set of relative position encoding values of multiple X-ray images in the training data when determining a view fusion feature.
[0189] S830. Determine a discrete attention loss based on a set of relative position encoding values.
[0190] S840: Determine the overall loss of the target recognition model based on the classification loss, regression loss, and attention discrete loss weights.
[0191] S850: Iterate the parameters of the target recognition model and the parameters of the related models of the target recognition model based on the overall loss.
[0192] In the aforementioned S810 , the classification loss and regression loss can be referred to the aforementioned description and will not be elaborated here.
[0193] In the aforementioned S820, the relative position encoding value set is the set of relative position encoding values during the training process. That is, the relative position encoding value set may refer to the relative position encoding value set output by the relative spatial model, or may refer to the set of values output by the aforementioned relative position encoding φ for any feature point group during the training process.
[0194] In the aforementioned S830, considering that the relative position encoding has no practical meaning, the attention discrete loss can be calculated based on the output value of the aforementioned relative position encoding. The attention discrete loss is the average value of the modulus length of each relative position encoding value in the relative position encoding value set. That is, the attention discrete loss .
[0195] Based on the three losses determined above, the final loss can be determined in the above S640 based on the three losses, which is generally presented in the form of weighted summation, that is, the final loss Among them, the weight coefficient , , It is used to balance the losses of different parts. Preferably, they can be 1, 2, and 0.1 respectively.
[0196] Therefore, based on the above loss calculation process, by introducing the discrete attention loss, a sparse constraint is introduced for the relative position encoding, thereby encouraging the concentration of attention distribution and improving the reliability of matching.
[0197] All of the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, and will not be described in detail here.
[0198] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0199] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0200] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0201] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0202] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0203] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program check codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0204] It should be noted that, in the description of this application, the terms "first," "second," "third," etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0205] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A multi-view target detection method, characterized in that: The multi-view target detection method comprises: Determining multiple X-ray images of the same object to be tested acquired by an imaging device at different imaging viewing angles, wherein the imaging geometric parameters of the X-ray images include an imaging angle and a detection distance; For a target image among the multiple X-ray images, performing feature extraction on the target image to determine image features of the target image; Normalizing the position information of the image features of the target image to determine normalized coordinate features; Normalizing the detection distance corresponding to the target image based on the maximum detection distance of the multiple X-ray images to determine the normalized distance of the target image; Determining an extrinsic parameter correction factor based on the normalized distance and imaging geometric parameters of the target image; Superimposing the normalized coordinate feature and the extrinsic parameter correction factor to determine a position coding feature; For a first image and a second image among the plurality of X-ray images, determining a perspective fusion feature of at least one of the first image and the second image based on imaging geometric parameters, image features, and position coding features of the first image and imaging geometric parameters, image features, and position coding features of the second image; The target detection result is determined by a target recognition model based on the perspective fusion feature.
2. The multi-view target detection method according to claim 1, wherein: The perspective fusion feature of the first image and the second image includes a first perspective fusion feature of the first image based on the second image, and determining the first perspective fusion feature from at least one perspective fusion feature of the first image and the second image determined based on imaging geometric parameters, image features, and position coding features of the first image and imaging geometric parameters, image features, and position coding features of the second image includes: Determining cross-view fusion features from the second image using a query key matching algorithm, wherein a query sequence of the query key matching algorithm is determined based on image features of the first image, and a key sequence and a value sequence of the query key matching algorithm are determined based on image features of the second image; The first view fusion feature is determined based on the image feature and position encoding feature of the first image and the cross-view fusion feature from the second image.
3. The multi-view target detection method according to claim 2, characterized in that: Before determining the cross-view fusion features from the second image using the query key matching algorithm, the multi-view object detection method further includes: determining a composite feature of the first image based on image features of the first image and the position coding feature, and determining a composite feature of the second image based on image features of the second image and the position coding feature, wherein the image features and the position coding feature are configured as channel data of different channels of the composite feature; The composite feature of the first image is converted into the query sequence, and the composite feature of the second image is converted into the key sequence and the value sequence.
4. The multi-view target detection method according to claim 2, wherein: The determining of the cross-view fusion features from the second image by using a query key matching algorithm includes: Determine a first query sequence, a second key sequence, and a second value sequence based on a query weight matrix, a key weight matrix, and a value weight matrix of the query key matching algorithm, the query sequence, the key sequence, and the value sequence; Determining an attention encoding value of the cross-view fusion feature of the second image based on the first query sequence and the second key sequence; determining a cross-view attention weight from the first image to the second image based on the attention encoding value; A cross-view fusion feature of the second image is determined based on the cross-view attention weight and the second value sequence.
5. The multi-view target detection method according to claim 4, characterized in that: The determining a cross-view attention weight from the first image to the second image based on the attention encoding value includes: For a first feature point in the first image and a second feature point in the second image, determining a relative position coding value based on a position coding feature value of the first feature point and a position coding feature value of the second feature point, so as to determine a relative position coding of the first image and the second image; The attention encoding value and the relative position encoding are input into an activation function to determine a cross-view attention weight from the first image to the second image.
6. The multi-view target detection method according to claim 5, characterized in that: The determining of the relative position coding value based on the position coding feature value of the first feature point and the position coding feature value of the second feature point includes: determining a position difference between the first feature point and the second feature point based on the position coding feature value of the first feature point and the position coding feature value of the second feature point; determining a difference in imaging geometric parameters between the first image and the second image; The relative position encoding value is determined based on the position difference and the imaging geometric parameter difference.
7. The multi-view target detection method according to claim 2, wherein: The perspective fusion feature of the first image and the second image further includes a second perspective fusion feature of the second image based on the first image, and determining the target detection result through the target recognition model based on the perspective fusion feature includes: Inputting the first perspective fusion feature into the target recognition model to determine a first detection result; Inputting the second perspective fusion feature into the target recognition model to determine a second detection result; The target detection result is determined based on the first detection result and the second detection result.
8. A multi-view target detection device, characterized in that: The multi-view target detection device comprises: An image acquisition module, configured to determine multiple X-ray images of the same object to be tested acquired by an imaging device at different imaging viewing angles, wherein the imaging geometric parameters of the X-ray images include imaging angles and detection distances; a feature extraction module, configured to extract features of a target image from the plurality of X-ray images and determine image features of the target image; a position encoding module, configured to normalize the position information of the image features of the target image to determine a normalized coordinate feature; normalize the detection distance corresponding to the target image based on the maximum detection distance of the plurality of X-ray images to determine a normalized distance of the target image; determine an extrinsic parameter correction factor based on the normalized distance of the target image and imaging geometric parameters; and superimpose the normalized coordinate feature and the extrinsic parameter correction factor to determine a position encoding feature; a feature fusion module, configured to determine, for a first image and a second image among the plurality of X-ray images, a viewing angle fusion feature of at least one of the first image and the second image based on imaging geometric parameters, image features, and position coding features of the first image and imaging geometric parameters, image features, and position coding features of the second image; A detection execution module is used to determine a target detection result through a target recognition model based on the perspective fusion feature.
9. A model training method based on multi-view target detection, characterized in that: The model training method includes: Acquire at least one set of training data, wherein the training data includes a plurality of X-ray images of the same object to be tested acquired by an imaging device at different imaging perspectives and target identification labels in the plurality of X-ray images; For a target image among the multiple X-ray images in the training data, feature extraction is performed on the target image to determine the image features of the target image, and the position information of the image features of the target image is normalized to determine the normalized coordinate features; based on the maximum detection distance of the multiple X-ray images, the detection distance corresponding to the target image is normalized to determine the normalized distance of the target image; an extrinsic parameter correction factor is determined based on the normalized distance and imaging geometric parameters of the target image; the normalized coordinate features and the extrinsic parameter correction factor are superimposed to determine the position coding features to determine the position coding features of each X-ray image in the training data; for a first image and a second image among the multiple X-ray images, a perspective fusion feature of at least one of the first image and the second image is determined based on the imaging geometric parameters, image features and position coding features of the first image and the imaging geometric parameters, image features and position coding features of the second image; a target detection result is determined through a target recognition model based on the perspective fusion feature, wherein the imaging geometric parameters of the X-ray image include the imaging angle and the detection distance; The target recognition model is trained based on the target recognition label and the target detection result.
10. The model training method according to claim 9, characterized in that: The training of the target recognition model based on the target recognition label and the target detection result includes: Determining the classification loss and regression loss of the target recognition model based on the target recognition label and the target detection result; Determining a set of relative position coding values of a plurality of X-ray images in the training data when determining the view fusion feature, wherein the relative position coding values in the set of relative position coding values are determined based on a difference in imaging geometric parameters between two X-ray images in the plurality of X-ray images and a position difference in position coding feature values of a combination of feature points in the two X-ray images; Determining a discrete attention loss based on the relative position encoding value set, wherein the discrete attention loss is an average of the modulus lengths of the relative position encoding values in the relative position encoding value set; Determine the overall loss of the target recognition model based on the classification loss, the regression loss, and the attention discrete loss weighted; Parameters of the target recognition model and parameters of related models of the target recognition model are iterated based on the overall loss.
11. A model training device based on multi-view target detection, characterized in that: The model training device comprises: a data acquisition module, configured to acquire at least one set of training data, wherein the training data includes a plurality of X-ray images acquired by an imaging device at different imaging perspectives and target identification labels in the plurality of X-ray images; a target detection module configured to extract features of target images from a plurality of X-ray images in the training data, determine image features of the target images, and normalize position information of the image features of the target images to determine normalized coordinate features; normalize the detection distance corresponding to the target images based on the maximum detection distance of the plurality of X-ray images to determine the normalized distance of the target images; determine an extrinsic parameter correction factor based on the normalized distance and imaging geometric parameters of the target images; superimpose the normalized coordinate features and the extrinsic parameter correction factor to determine position coding features to determine position coding features of each X-ray image in the training data; determine, for a first image and a second image in the plurality of X-ray images, a perspective fusion feature of at least one of the first image and the second image based on the imaging geometric parameters, image features, and position coding features of the first image and the imaging geometric parameters, image features, and position coding features of the second image; and determine a target detection result using a target recognition model based on the perspective fusion feature, wherein the imaging geometric parameters of the X-ray images include an imaging angle and a detection distance; A parameter iteration module is used to train the target recognition model based on the target recognition label and the target detection result.
12. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-view target detection method described in any one of claims 1 to 7 or the model training method based on multi-view target detection described in any one of claims 9 and 10.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable the processor to implement the multi-view target detection method according to any one of claims 1 to 7 or the model training method based on multi-view target detection according to any one of claims 9 and 10 when executed.
14. A computer program product, characterized in that The computer program product includes computer program code, and when the computer program code is run on a computer, the computer implements the multi-view target detection method according to any one of claims 1 to 7 or the model training method based on multi-view target detection according to any one of claims 9 and 10.
Citation Information
Patent Citations
Binocular distance measurement method and system based on depth stereo matching algorithm
CN117635730A
Wide-angle video monitoring system, method and equipment based on multiple cameras and medium
CN120238632A