Target detection network training and detection method, device, terminal and storage medium

By building an initial detection network and iterative training, the hazardous product detection problem of X-ray security machines in different ray sources is solved, effective detection of scanned images of different ray sources is achieved, and the scope of application of the target detection network is expanded.

CN114462487BActive Publication Date: 2025-07-29ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111624203.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-07-29
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

The existing target detection network is difficult to be suitable for the detection of hazardous goods by X-ray security machines of different ray sources, resulting in poor detection results.

Method used

By building an initial detection network, using the source domain and the target domain image set for training, building a loss function, and iteratively training the target detection network, including the initial feature extraction module, the image feature domain classification network and the target feature domain classification network, to adapt to the scanned images of different ray sources.

Benefits of technology

The target detection network is implemented to detect dangerous goods on scanned images of different ray sources, expand the scope of application, and save time and labor costs for real information labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462487B_ABST
    Figure CN114462487B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for training and detecting a target detection network, a terminal, and a storage medium. The method for training the target detection network includes: obtaining a source domain image set and a target domain image set; respectively detecting a first scanned image and a second scanned image through an initial detection network to obtain target prediction information of the first scanned image, a first prediction domain classification label, a second prediction domain classification label, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image; constructing a loss function based on the prediction information and the corresponding labeled ground truth information of the same first scanned image and the prediction information and the corresponding labeled ground truth information of the same second scanned image; and iteratively training the initial detection network based on the loss function to obtain a target detection network. The present application can improve the accuracy of the target detection network by training the target detection network with the source domain image set and the target domain image set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a method, device, terminal and storage medium for training and detecting a target detection network. Background Art

[0002] X-ray security inspection machines are widely used in important places that require security inspections, such as railways, airports, docks, government agencies, logistics, etc. With the development of computer technology and intelligent image processing technology, using target detection algorithms to assist security inspectors in identifying dangerous goods can effectively improve the security inspection efficiency and reduce missed reports and false reports caused by the visual fatigue of security inspectors. According to different scenario requirements, the ray sources used by X-ray security inspection machines are not exactly the same, and the colors of the objects imaged by different ray sources will also vary. For existing detection technologies, it is difficult to effectively adapt the pictures imaged by different ray sources to the same algorithm, which requires the target detection algorithm to be applicable to the dangerous goods detection tasks of X-ray security inspection machines with different ray sources. Summary of the Invention

[0003] The main technical problem to be solved by the present application is to provide a method, device, terminal and storage medium for training and detecting a target detection network, so as to solve the problem that it is difficult for the target detection network in the prior art to detect dangerous goods in scanned images corresponding to different ray sources.

[0004] To solve the above technical problem, the first technical solution adopted by the present application is: to provide a method for training a target detection network, the method for training a target detection network includes: obtaining a source domain image set and a target domain image set, the source domain image set includes multiple first scanned images labeled with target true information, a first true domain classification label and a second true domain classification label; the target domain image set includes multiple second scanned images labeled with the first true domain classification label and the second true domain classification label and not labeled with target information; the first scanned image and the second scanned image are respectively scanned based on different ray sources; detecting the first scanned image and the second scanned image respectively through a constructed initial detection network to obtain target prediction information, a first prediction domain classification label, a second prediction domain classification label of the first scanned image, and a first prediction domain classification label and a second prediction domain classification label of the second scanned image; constructing a loss function based on the target prediction information of the same first scanned image and the labeled target true information, the first prediction domain classification label of the same first scanned image and the labeled first true domain classification label and the second prediction domain classification label and the labeled second true domain classification label, and the first prediction domain classification label of the same second scanned image and the labeled first true domain classification label and the second prediction domain classification label and the labeled second true domain classification label; and iteratively training the initial detection network based on the loss function to obtain a target detection network.

[0005] Among them, the steps of constructing the initial detection network include: constructing the initial detection network based on the initial object detection network, the initial image feature domain classification network, and the initial object feature domain classification network; among them, the initial object detection network includes an initial feature extraction module, an initial feature aggregation module, and an initial object prediction module connected in sequence; the initial feature extraction module is connected to the initial image feature domain classification network, and the initial feature aggregation module is connected to the initial object feature domain classification network.

[0006] Among them, iteratively training the initial detection network based on the loss function to obtain the object detection network includes: correcting the weights in the initial object detection network, the initial image feature domain classification network, and the initial object feature domain classification network in the initial detection network based on the loss function to obtain the object detection network, the image feature domain classification network, and the object feature domain classification network; removing the image feature domain classification network and the object feature domain classification network, and retaining the object detection network.

[0007] Among them, the first scanned image and the second scanned image are scanned images corresponding to X-rays from different ray sources; the object prediction information includes the object prediction position and the object prediction category; before obtaining the object prediction information, the first prediction domain classification label, the second prediction domain classification label of the first scanned image, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image by detecting the first scanned image and the second scanned image respectively through the constructed initial detection network, it also includes: adjusting the sizes of the first scanned image and the second scanned image to a first preset size; and performing normalization processing on the first scanned image and the second scanned image to obtain the corresponding preprocessed scanned images.

[0008] Among them, a first gradient reversal layer is provided between the initial feature extraction module and the initial image feature domain classification network; and / or a second gradient reversal layer is provided between the initial feature aggregation module and the initial object feature domain classification network.

[0009] Among them, obtaining the object prediction information, the first prediction domain classification label, the second prediction domain classification label of the first scanned image, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image by detecting the first scanned image and the second scanned image respectively through the constructed initial detection network includes: extracting features from the preprocessed scanned image through the initial feature extraction module to obtain the corresponding image features; performing domain adaptation on the image features through the initial image feature domain classification network to obtain the first prediction domain classification label of the preprocessed scanned image.

[0010] Among them, the initial feature extraction module includes multiple sequentially connected feature extraction layers, and the last feature extraction layer is connected to the initial image feature domain classification network; the initial feature extraction module extracts features from the preprocessed scanned image to obtain corresponding image features, including: the last feature extraction layer extracts features from the feature map output by the previous feature extraction layer to obtain image features.

[0011] Among them, the initial image feature domain classification network includes a first image feature extraction unit, a second image feature extraction unit, a pooling layer, and a fully connected layer that are sequentially connected; among them, the first image feature extraction unit includes a cascaded first convolutional layer and a first activation function layer, and the second image feature extraction unit is multiple and sequentially connected, and the second image feature extraction unit includes a cascaded second convolutional layer and a second activation function layer; the initial image feature domain classification network performs domain adaptation on the image features to obtain the first predicted domain classification label of the preprocessed scanned image, including: after the first convolutional layer extracts features from the image features, a first feature map is obtained; the first activation function layer performs non-linear activation processing on the first feature map to obtain a second feature map; the feature map corresponding to the second convolutional layer in the last second image feature extraction unit is fused with the feature maps corresponding to the second convolutional layer in at least one previous second image feature extraction unit to obtain a third feature map; the second activation function layer in the last second image feature extraction unit performs non-linear activation processing on the third feature map to obtain a fourth feature map; the pooling layer adjusts the size of the fourth feature map; the fully connected layer determines the first predicted domain classification label of the first scanned image / second scanned image based on the fourth feature map with adjusted size.

[0012] Among them, the second convolutional layer includes at least two convolutional kernels, and the second convolutional layer is a depthwise separable convolutional layer.

[0013] Among them, the initial image feature domain classification network performs domain adaptation on the image features to obtain the first predicted domain classification label of the preprocessed scanned image, and further includes: adjusting the size of the first feature map to a second preset size.

[0014] Among them, the pooling layer includes a global average pooling layer and a global max pooling layer; adjusting the size of the fourth feature map through the pooling layer includes: the global average pooling layer adjusts the size of the fourth feature map to obtain a fifth feature map; the global max pooling layer adjusts the size of the fourth feature map to obtain a sixth feature map; the fully connected layer determines the first predicted domain classification label of the first scanned image / second scanned image based on the fourth feature map with adjusted size, including: the fully connected layer performs feature fusion on the fifth feature map and the sixth feature map to determine the first predicted domain classification label of the first scanned image / second scanned image.

[0015] Among them, the initial detection network is constructed to detect the first scanned image and the second scanned image respectively, and the target prediction information of the first scanned image, the classification labels of the first prediction domain, the classification labels of the second prediction domain, and the classification labels of the first prediction domain and the second prediction domain of the second scanned image are obtained. It also includes: the initial feature aggregation module aggregates the image features to obtain target features; the initial target feature domain classification network performs domain adaptation on the target features to obtain the classification labels of the second prediction domain of the preprocessed scanned image.

[0016] Among them, the initial feature aggregation module aggregates the image features to obtain target features, including: the initial feature aggregation module adjusts the size of the image features to obtain multiple feature maps with different sizes; the multiple feature maps are fused to obtain the target features corresponding to the first scanned image / second scanned image.

[0017] Among them, the initial target feature domain classification network includes cascaded first target feature extraction unit, second target feature extraction unit and third target feature extraction unit; the first target feature extraction unit includes cascaded first feature extraction layer and first activation layer, the second target feature extraction unit includes cascaded second feature extraction layer and second activation layer, and the third feature extraction unit includes cascaded third feature extraction layer and output layer; the initial target feature domain classification network performs domain adaptation on the target features to obtain the classification labels of the second prediction domain of the preprocessed scanned image, including: the first feature extraction layer in the first target feature extraction unit extracts features from the target features corresponding to the first scanned image / second scanned image to obtain the first target feature map; the first activation layer performs non-linear activation on the first target feature map to obtain the second target feature map; the target feature map corresponding to the second feature extraction layer in the last second target feature extraction unit is fused with the target feature maps corresponding to the second feature extraction layers in at least one previous second target feature extraction unit to obtain the third target feature map; the second activation layer in the last second target feature extraction unit performs non-linear activation processing on the third target feature map to obtain the fourth target feature map; the third feature extraction layer in the third feature extraction unit extracts features from the fourth target feature map to obtain the fifth target feature map; the output layer determines the classification labels of the second prediction domain of the first scanned image / second scanned image based on the fifth target feature map.

[0018] Among them, the second feature extraction layer includes at least two convolutional kernels, and the second feature extraction layer is a depthwise separable convolutional layer.

[0019] Among them, the initial detection network is constructed to detect the first scanned image and the second scanned image respectively, and the target prediction information, the first prediction domain classification label, the second prediction domain classification label of the first scanned image, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image are obtained. It also includes: the initial target prediction module performs target detection on the target features corresponding to the first scanned image / second scanned image, and obtains the target prediction information containing the target in the preprocessed scanned image.

[0020] Among them, based on the target prediction information of the same first scanned image and the labeled target ground truth information, the first prediction domain classification label of the same first scanned image and the labeled first ground truth domain classification label and the second prediction domain classification label and the labeled second ground truth domain classification label, and the first prediction domain classification label of the same second scanned image and the labeled first ground truth domain classification label and the second prediction domain classification label and the labeled second ground truth domain classification label, a loss function is constructed, including: the loss function is:

[0021]

[0022] In the formula: L det is the target information loss function of the target detection network; is the domain classification label loss function of the image feature domain classification network; is the domain classification label loss function of the target feature domain classification network; λ is the weighting coefficient.

[0023] Among them, based on the target prediction information of the same first scanned image and the labeled target ground truth information, the first prediction domain classification label of the same first scanned image and the labeled first ground truth domain classification label and the second prediction domain classification label and the labeled second ground truth domain classification label, and the first prediction domain classification label of the same second scanned image and the labeled first ground truth domain classification label and the second prediction domain classification label and the labeled second ground truth domain classification label, a loss function is constructed, including: the domain classification label loss function of the image feature domain classification network is:

[0024]

[0025] In the formula: is the domain classification label loss function of the image feature domain classification network; M is the total number of the first scanned image and the second scanned image; n + 1 is the total number of the first prediction domain classification labels; Imc is the sign function, when the first ground truth domain classification label of the mth scanned image is equal to c, the value of this function is 1, otherwise the value is 0, is the prediction probability that the mth scanned image belongs to the category c, where the first prediction domain classification label indicates whether the scanned image belongs to the source domain image set or the target domain image set.

[0026] Among them, a loss function is constructed based on the target prediction information and the labeled target ground truth information of the same first scan image, the first predicted domain classification label and the labeled first ground truth domain classification label of the same first scan image, the second predicted domain classification label and the labeled second ground truth domain classification label, and the first predicted domain classification label and the labeled second ground truth domain classification label of the same second scan image, including: The domain classification label loss function of the target feature domain classification network is:

[0027]

[0028] In the formula, is the domain classification label loss function of the target feature domain classification network; represents the probability that the result output by the target feature domain classification network of the m-th scan image is predicted as the second predicted domain class label at the coordinate.

[0029] Among them, the first scan image and the second scan image are respectively detected by the constructed initial detection network to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scan image, and the first predicted domain classification label and the second predicted domain classification label of the second scan image. Before that, it also includes: pre-training the initial target detection network in the initial detection network.

[0030] Among them, pre-training the initial target detection network in the initial detection network includes: detecting the first scan image through the initial target detection network to obtain the target prediction information of the first scan image; constructing an initial loss function through the target prediction information and the target ground truth information corresponding to the same target of the first scan image; pre-training the initial target detection network based on the initial loss function.

[0031] To solve the above technical problems, the second technical solution adopted by this application is: to provide an object detection method, the object detection method includes: obtaining an image to be detected; the image to be detected includes at least one target object; preprocessing the image to be detected to obtain a preprocessed image; performing object detection on the preprocessed image through an object detection network to obtain object information of the target object; among them, the object detection network is obtained through the above object detection network training method.

[0032] Among them, the object information includes the target position and the target category; preprocessing the image to be detected to obtain a preprocessed image includes: adjusting the size of the image to be detected to a preset size; and performing normalization processing on the image to be detected to obtain the corresponding preprocessed image.

[0033] To solve the above technical problems, the third technical solution adopted by this application is: to provide a training device, which includes: a sample acquisition module, configured to acquire a source domain image set and a target domain image set. The source domain image set includes multiple first scanned images labeled with target real information, a first real domain classification label, and a second real domain classification label; the target domain image set includes multiple second scanned images labeled with the first real domain classification label and the second real domain classification label and without labeled target information. The first scanned images and the second scanned images are scanned respectively based on different ray sources; a prediction module, configured to detect the first scanned images and the second scanned images respectively through a constructed initial detection network, to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned images, and the first predicted domain classification label and the second predicted domain classification label of the second scanned images; a function construction module, configured to construct a loss function based on the target prediction information of the same first scanned image and the labeled target real information, the first predicted domain classification label of the same first scanned image and the labeled first real domain classification label, the second predicted domain classification label and the labeled second real domain classification label, and the first predicted domain classification label of the same second scanned image and the labeled first real domain classification label, and the second predicted domain classification label and the labeled second real domain classification label; a processing module, configured to iteratively train the initial detection network based on the loss function to obtain a target detection network.

[0034] To solve the above technical problems, the fourth technical solution adopted by this application is: to provide a target detection device, which includes: an image acquisition module, configured to acquire an image to be detected; the image to be detected includes at least one target object; a preprocessing module, configured to preprocess the image to be detected to obtain a preprocessed image; a detection module, configured to perform target detection on the preprocessed image through a target detection network to obtain target object information of the target object; wherein, the target detection network is obtained through the above target detection network training method.

[0035] To solve the above technical problems, the fifth technical solution adopted by this application is: to provide a terminal, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. The processor is configured to execute program data to implement the steps in the above image quality assessment method.

[0036] To solve the above technical problems, the sixth technical solution adopted by this application is: to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the above image quality assessment method.

[0037] The beneficial effects of the present application are as follows: Different from the prior art, a method, device, terminal, and storage medium for training and detecting a target detection network are provided. The method for training the target detection network includes: obtaining a source domain image set and a target domain image set. The source domain image set includes multiple first scanned images labeled with target true information, as well as a first true domain classification label and a second true domain classification label. The target domain image set includes multiple second scanned images labeled with the first true domain classification label and the second true domain classification label but without target information. The first scanned images and the second scanned images are scanned based on different radiation sources respectively. Detecting the first scanned images and the second scanned images respectively through the constructed initial detection network to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned images, and the first predicted domain classification label and the second predicted domain classification label of the second scanned images. Constructing a loss function based on the target prediction information and the labeled target true information of the same first scanned image, the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same first scanned image, and the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same second scanned image. Iteratively training the initial detection network based on the loss function to obtain the target detection network. The present application predicts the scanned images included in the source domain image set and the target domain image set through the initial detection network, and trains the initial detection network based on the first predicted domain classification label and the second predicted domain classification label corresponding to the first scanned images and the second scanned images respectively and the first true domain classification label and the second true domain classification label corresponding to the first scanned images and the second scanned images respectively, saving the time cost and labor cost of labeling the true information of the second scanned images. The obtained target detection network can detect dangerous goods in scanned images collected by different radiation sources, making the applicable range of the target detection network wider. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a flowchart of the method for training the target detection network provided by the present application;

[0040] Figure 2 It is a flowchart of a specific embodiment of the method for training the target detection network provided by the present application;

[0041] Figure 3Yes Figure 2 Schematic flow diagram of a specific embodiment of step S203 in the provided object detection network training method;

[0042] Figure 4 Block diagram of a specific embodiment of the object detection network training method provided by the present application;

[0043] Figure 5 Block diagram of a specific embodiment of the initial object detection network provided by the present application;

[0044] Figure 6 Block diagram of a specific embodiment of the initial image feature domain classification network provided by the present application;

[0045] Figure 7 Block diagram of a specific embodiment of the initial object feature domain classification network provided by the present application;

[0046] Figure 8 Block diagram of a specific embodiment of the second convolutional layer or the second feature extraction layer provided by the present application;

[0047] Figure 9 Schematic flow diagram of the object detection method provided by the present application;

[0048] Figure 10 Schematic flow diagram of a specific embodiment of the object detection method provided by the present application;

[0049] Figure 11 Schematic block diagram of the object detection network training device provided by the present application;

[0050] Figure 12 Schematic block diagram of the object detection device provided by the present application;

[0051] Figure 13 Schematic block diagram of an implementation manner of the terminal provided by the present application;

[0052] Figure 14 Schematic block diagram of an implementation manner of the computer-readable storage medium provided by the present application. Detailed implementation manners

[0053] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0054] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0055] In this text, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in this text, the character " / " generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "multiple" in this text means two or more than two.

[0056] To enable those skilled in the art to better understand the technical solution of this application, the following further describes in detail a target detection network training method and a target detection method provided by this application in conjunction with the accompanying drawings and specific implementation manners.

[0057] Please refer to Figure 1 , Figure 1 is a schematic flowchart of the target detection network training method provided by this application. In this embodiment, a target detection network training method is provided, and this target detection network training method includes the following steps.

[0058] S11: Obtain a source domain image set and a target domain image set.

[0059] Specifically, the source domain image set includes multiple first scanned images labeled with target true information, a first true domain classification label, and a second true domain classification label; the target domain image set includes multiple second scanned images labeled with the first true domain classification label and the second true domain classification label and without labeled target information. The first scanned image and the second scanned image are scanned images of X-rays corresponding to different ray sources. In this embodiment, based on different ray sources, it means that the voltage, current, or energy spectrum is different, and the penetration power of the X-rays is different, resulting in different scanned images. The target true information includes the target true position and the target true category.

[0060] S12: Detect the first scanned image and the second scanned image respectively through the constructed initial detection network to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned image, and the first predicted domain classification label and the second predicted domain classification label of the second scanned image.

[0061] Specifically, an initial detection network is constructed based on an initial target detection network, an initial image feature domain classification network, and an initial target feature domain classification network; among them, the initial target detection network includes an initial feature extraction module, an initial feature aggregation module, and an initial target prediction module connected in sequence; the initial feature extraction module is connected to the initial image feature domain classification network, and the initial feature aggregation module is connected to the initial target feature domain classification network. Among them, a first gradient reversal layer is provided between the initial feature extraction module and the initial image feature domain classification network; and / or a second gradient reversal layer is provided between the initial feature aggregation module and the initial target feature domain classification network.

[0062] In one embodiment, the sizes of the first scanned image and the second scanned image are adjusted to a first preset size; and the first scanned image and the second scanned image are normalized to obtain corresponding preprocessed scanned images.

[0063] In one embodiment, the initial feature extraction module extracts features from the preprocessed scanned images to obtain corresponding image features; the initial image feature domain classification network performs domain adaptation on the image features to obtain the first predicted domain classification label of the preprocessed scanned images.

[0064] In one embodiment, the initial feature extraction module includes a plurality of sequentially connected feature extraction layers, and the last feature extraction layer is connected to the initial image feature domain classification network; the last feature extraction layer extracts features from the feature map output by the previous feature extraction layer to obtain image features.

[0065] In one embodiment, the initial image feature domain classification network includes a first image feature extraction unit, a second image feature extraction unit, a pooling layer, and a fully connected layer that are sequentially connected; wherein, the first image feature extraction unit includes a cascaded first convolutional layer and a first activation function layer, the second image feature extraction units are multiple and sequentially connected, and the second image feature extraction unit includes a cascaded second convolutional layer and a second activation function layer; after the first convolutional layer extracts features from the image features, a first feature map is obtained; after the first activation function layer performs non-linear activation processing on the first feature map, a second feature map is obtained; the feature map corresponding to the second convolutional layer in the last second image feature extraction unit is fused with the feature maps corresponding to the second convolutional layers in at least one previous second image feature extraction unit to obtain a third feature map; the second activation function layer in the last second image feature extraction unit performs non-linear activation processing on the third feature map to obtain a fourth feature map; the pooling layer adjusts the size of the fourth feature map; the fully connected layer determines the first predicted domain classification label of the first scanned image / second scanned image based on the fourth feature map with the adjusted size. Wherein, the second convolutional layer includes at least two convolutional kernels, and the second convolutional layer is a depthwise separable convolutional layer. Preferably, the first feature map is adjusted to a second preset size.

[0066] In one embodiment, the pooling layer includes a global average pooling layer and a global max pooling layer. The global average pooling layer adjusts the size of the fourth feature map to obtain a fifth feature map; the global max pooling layer adjusts the size of the fourth feature map to obtain a sixth feature map; the fully connected layer performs feature fusion on the fifth feature map and the sixth feature map to determine the first predicted domain classification label of the first scanned image / second scanned image.

[0067] In one embodiment, the initial feature aggregation module aggregates the image features to obtain target features; the initial target feature domain classification network performs domain adaptation on the target features to obtain the second predicted domain classification label of the preprocessed scanned image. Specifically, the initial feature aggregation module adjusts the size of the image features to obtain multiple feature maps with different sizes; the multiple feature maps are fused to obtain the target features corresponding to the first scanned image / second scanned image.

[0068] In one embodiment, the initial target feature domain classification network includes a cascaded first target feature extraction unit, a second target feature extraction unit, and a third target feature extraction unit; the first target feature extraction unit includes a cascaded first feature extraction layer and a first activation layer, the second target feature extraction unit includes a cascaded second feature extraction layer and a second activation layer, and the third feature extraction unit includes a cascaded third feature extraction layer and an output layer; the first feature extraction layer in the first target feature extraction unit extracts features from the target features corresponding to the first scanned image / second scanned image to obtain a first target feature map; the first activation layer non-linearly activates the first target feature map to obtain a second target feature map; the target feature map corresponding to the second feature extraction layer in the last second target feature extraction unit is fused with the target feature maps corresponding to the second feature extraction layers in at least one previous second target feature extraction unit to obtain a third target feature map; the second activation layer in the last second target feature extraction unit non-linearly activates the third target feature map to obtain a fourth target feature map; the third feature extraction layer in the third feature extraction unit extracts features from the fourth target feature map to obtain a fifth target feature map; the output layer determines the second predicted domain classification label of the first scanned image / second scanned image based on the fifth target feature map. Among them, the second feature extraction layer includes at least two convolutional kernels, and the second feature extraction layer is a depthwise separable convolutional layer.

[0069] In one embodiment, the initial target prediction module performs target detection on the target features corresponding to the first scanned image / second scanned image to obtain target prediction information including the target in the preprocessed scanned image.

[0070] S13: Construct a loss function based on the target prediction information and the annotated target ground truth information of the same first scanned image, the first predicted domain classification label and the annotated first true domain classification label and the second predicted domain classification label and the annotated second true domain classification label of the same first scanned image, and the first predicted domain classification label and the annotated first true domain classification label and the second predicted domain classification label and the annotated second true domain classification label of the same second scanned image.

[0071] Specifically, the loss function is:

[0072]

[0073] where: L det is the target information loss function of the target detection network; is the domain classification label loss function of the image feature domain classification network; is the domain classification label loss function of the target feature domain classification network; λ is the weighting coefficient.

[0074] In one embodiment, the domain classification label loss function of the image feature domain classification network is:

[0075]

[0076] where: is the domain classification label loss function of the image feature domain classification network; M is the total number of the first scanned image and the second scanned image; n + 1 is the total number of the first predicted domain classification labels; I mc is the sign function, which takes the value of 1 when the first true domain classification label of the m-th scanned image is equal to c, otherwise it takes the value of 0, is the predicted probability that the m-th scanned image belongs to the category c, where the first predicted domain classification label indicates whether the scanned image belongs to the source domain image set or the target domain image set.

[0077] In one embodiment, the domain classification label loss function of the target feature domain classification network is:

[0078]

[0079] where, is the domain classification label loss function of the target feature domain classification network; represents the probability that the result output by the target feature domain classification network of the m-th scanned image is predicted as the second predicted domain class label at the coordinate.

[0080] S14: Iteratively train the initial detection network based on the loss function to obtain the target detection network.

[0081] Specifically, correct the weights in the initial target detection network, the initial image feature domain classification network, and the initial target feature domain classification network in the initial detection network based on the loss function to obtain the target detection network, the image feature domain classification network, and the target feature domain classification network; remove the image feature domain classification network and the target feature domain classification network, and retain the target detection network.

[0082] In one embodiment, before step S12, the initial target detection network in the initial detection network is pre-trained. Specifically, the first scan image is detected by the initial target detection network to obtain the item prediction information of the first scan image; the initial target detection network is pre-trained by the loss value between the item prediction information of the first scan image and the target ground truth information.

[0083] The target detection network training method provided in this embodiment includes: the source domain image set includes multiple first scan images labeled with target ground truth information, a first true domain classification label, and a second true domain classification label; the target domain image set includes multiple second scan images labeled with a first true domain classification label and a second true domain classification label and without labeled target information; the first scan image and the second scan image are scanned based on different radiation sources respectively; the first scan image and the second scan image are detected respectively by the constructed initial detection network to obtain the target prediction information, a first predicted domain classification label, a second predicted domain classification label of the first scan image, and the first predicted domain classification label and the second predicted domain classification label of the second scan image; a loss function is constructed based on the target prediction information and the labeled target ground truth information of the same first scan image, the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same first scan image, and the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same second scan image; the initial detection network is iteratively trained based on the loss function to obtain the target detection network. In this application, the initial detection network is used to predict the scan images included in the source domain image set and the target domain image set, and the initial detection network is trained based on the first predicted domain classification label and the second predicted domain classification label corresponding to the first scan image and the second scan image respectively and the first true domain classification label and the second true domain classification label corresponding to the first scan image and the second scan image respectively, saving the time cost and labor cost of labeling the ground truth information of the second scan image. The obtained target detection network can detect dangerous goods for scan images collected by different radiation sources, making the applicable range of the target detection network wider.

[0084] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a specific embodiment of the target detection network training method provided by this application. In this embodiment, a target detection network training method is provided, and the target detection network training method includes the following steps.

[0085] S201: Obtain a source domain image set and a target domain image set.

[0086] Specifically, the source domain image set includes multiple first scanned images labeled with the target's true information, the first true domain classification label, and the second true domain classification label; the target domain image set includes multiple second scanned images labeled with the first true domain classification label and the second true domain classification label but without the target's true information. The targets in the first scanned images and the second scanned images are the same and are collected based on different radiation sources. Among them, the first scanned images included in the source domain image set are all images that have been labeled with the target's true information. Among them, the target's true information includes the target's true position and the target's true category. Among them, the first scanned images and the second scanned images are scanned images collected by an X-ray security inspection machine, and the radiation sources of the first scanned images and the second scanned images are different. Among them, the radiation sources for scanning the first scanned images and the second scanned images can be X-rays with different voltages, currents, or energy spectra. The penetration powers of the X-rays are different, and the obtained scanned images are different..

[0087] In a specific embodiment, the same target or different targets are scanned by different radiation sources to obtain at least two different scanned images. One of the scanned images is labeled, and the labeled scanned image is used as the first scanned image. The true information of all the targets included in the first scanned image, as well as the first true domain classification label and the second true domain classification label of the first scanned image, are labeled. The scanned images obtained by other radiation sources are used as the second scanned images, and the true information of the targets included in the second scanned images is not labeled. The first true domain classification label and the second true domain classification label of the second scanned images are labeled.

[0088] The first scanned image labeled with the target's true information, the first true domain classification label, and the second true domain classification label is the source data; the second scanned image labeled with the first true domain classification label and the second true domain classification label is the target data. In this embodiment, the dangerous goods included in the first scanned images and the second scanned images are used as the targets. For example, when the first scanned images and the second scanned images contain a knife, the true category of the target is labeled as a cutter.

[0089] In an alternative embodiment, the first true domain classification label corresponding to the source data is a one-hot matrix with a dimension of (n + 1, 1); where n is the number of different radiation sources. Specifically, the first true domain classification label corresponding to the source data is [1, 0, 0, …, 0], and the first true domain classification label corresponding to the target data collected by the first other radiation source is [0, 1, 0, …, 0]. And so on, the first true domain classification label corresponding to the target data collected by the nth other radiation source is [0, 0, 0, …, 1].[[]END]]

[0090] In an optional embodiment, the second true domain classification label corresponding to the target data is a matrix with a dimension of (n + 1, 256); where n is the number of different ray sources. Specifically, in the classification label of the second true domain classification label corresponding to the source data, the first dimension is all 1, and the other dimensions are all 0. In the second true domain classification label corresponding to the target data, the second true domain classification label corresponding to the target data collected by the first other ray source has all 1s in the second dimension and all 0s in the other dimensions. And so on, the second true domain classification label corresponding to the target data collected by the nth other ray source has all 1s in the nth dimension and all 0s in the other dimensions.

[0091] S202: Preprocess the first scanned image and the second scanned image.

[0092] Specifically, adjust the sizes of the first scanned image and the second scanned image to a first preset size. In a specific embodiment, the first preset size is 512 * 512 * 3. In other optional embodiments, the first preset size can also be set according to the actual situation. In a specific embodiment, interpolate the first scanned image and the second scanned image into images with a size of 512 * 512 * 3 in terms of length and width.

[0093] In an embodiment, according to the size relationship between the first scanned image and the first scanned image after size adjustment, correct the true position of the target included in the first scanned image. That is, use the coordinate position of the target in the first scanned image after size adjustment as the true position of the target.

[0094] In an embodiment, perform normalization processing on the first scanned image after size adjustment and the second scanned image after size adjustment to obtain the preprocessed scanned image corresponding to the first scanned image and the preprocessed scanned image corresponding to the second scanned image.

[0095] S203: Train the initial target detection network to obtain a pre-trained target detection network.

[0096] Specifically, in order to improve the accuracy and precision of the initial target detection network in detecting the target, and to minimize the workload of training the initial detection network by means of domain adaptation technology in the subsequent steps, and to accelerate the training speed, the initial target detection network can be pre-trained in advance to preliminarily correct the weights of each module in the initial target detection network.

[0097] In one embodiment, an initial target detection network is trained using a source domain image set. The initial target detection network has a three-layer structure, including an initial feature extraction module, an initial feature aggregation module, and an initial target prediction module connected in sequence. In this embodiment, the initial target detection network can directly adopt the YOLO-v4 (You Only Look Once-v4) network structure without modifying the YOLO-v4 network structure. In other alternative embodiments, any single-stage target detection network structure can also be used, as long as the backbone network part and the output part of the target detection network are known.

[0098] Please refer to Figure 3 , Figure 3 is Figure 2 a schematic flowchart of a specific embodiment of step S203 in the target detection network training method provided.

[0099] Among them, the specific steps of training the initial target detection network to obtain a pre-trained target detection network are as follows.

[0100] S2031: Detect the first scanned image using the initial target detection network to obtain target prediction information of the first scanned image.

[0101] Specifically, perform target prediction on multiple first scanned images in the source domain image set using the initial target detection network respectively to obtain target prediction information of all targets included in the first scanned image. Among them, the target prediction information includes target prediction positions and target prediction categories.

[0102] S2032: Construct an initial loss function using the target prediction information corresponding to the same target in the first scanned image and the target ground truth information.

[0103] Specifically, construct an initial loss function based on the target ground truth position and the target prediction position, and the target ground truth category and the target prediction category corresponding to the same target in the first scanned image. In this embodiment, in order to accelerate the training speed of the initial target detection network, the optimizer used can be SGD (Stochastic Gradient Descent).

[0104] S2033: Train the initial target detection network based on the initial loss function to obtain a pre-trained target detection network.

[0105] Specifically, pre-train the initial target detection network based on the loss value between the target ground truth position and the target prediction position and the loss value between the target ground truth category and the target prediction category in the first scanned image to obtain a pre-trained target detection network, so as to improve the detection accuracy of the pre-trained target detection network. Among them, the learning rate of the initial target detection network is set to 0.0001.

[0106] In a specific embodiment, the prediction result of the initial target detection network is backpropagated, and the weights of the initial target detection network are corrected according to the loss value fed back by the initial loss function. In an alternative embodiment, the parameters of the initial target detection network can also be corrected to implement the training of the initial target detection network. In this embodiment, the parameters in the initial feature extraction module, the initial feature aggregation module, and the initial target prediction module are corrected to obtain the pre-trained feature extraction module, the pre-trained feature aggregation module, and the pre-trained target prediction module.

[0107] The first scanned image containing the target is input into the initial target detection network, and the initial target detection network predicts the target position and target category. When the weighted sum of the loss value between the target true position and the target predicted position and the loss value between the target true category and the target predicted category is less than a preset threshold, which can be set by oneself, such as 1%, 5%, etc., the training of the initial target detection network is stopped and the pre-trained target detection network is obtained.

[0108] In another alternative embodiment, after training the initial target detection network a preset number of times with the first scanned images included in the source domain image set, the pre-trained target detection network is obtained. In this embodiment, the preset number of times is 100 times. It can also be set according to the actual situation by oneself.

[0109] S204: Construct an initial detection network based on the pre-trained target detection network, the initial image feature domain classification network, and the initial target feature domain classification network.

[0110] In this embodiment, domain adaptation technology is added to the pre-trained target detection network, and the pre-trained target detection network can be trained by using the second scanned images without labeled target true information in the target domain image set.

[0111] Specifically, the structure of the pre-trained target detection network obtained after training is the same as that of the initial target detection network, and only the parameters of each module included in the initial target detection network are corrected. Therefore, the pre-trained target detection network includes a pre-trained feature extraction module, a pre-trained feature aggregation module, and a pre-trained target prediction module connected in sequence.

[0112] Please refer to Figure 4 and Figure 5 , Figure 4 which is the structural block diagram of a specific embodiment of the target detection network training method provided by this application; Figure 5 which is the structural block diagram of a specific embodiment of the initial target detection network provided by this application.

[0113] In this embodiment, the pre-trained feature extraction module is a backbone network, which can be a CSP-DarkNet structure and is mainly used for extracting image features. For example, the pre-trained feature extraction module is a CSPDarknet53 structure, as Figure 4 . The pre-trained feature aggregation module includes an SPP (Spatial Pyramid Pooling) network structure and a PANet structure, which are mainly used for fusing the extracted features to obtain image features. The SPP network structure stacks feature maps after using max pooling with different scales. In this embodiment, there are 3 max pooling layers. Among them, the SPP network structure is connected to the CSP-DarkNet structure, and the PANet structure is respectively connected to the SPP network structure and the CSP-DarkNet structure. The pre-trained target prediction module can be a YOLO Head, and the YOLO Head is connected to the PANet structure, and there are three YOLO Heads, which are mainly used for target prediction based on the aggregated image features. Among them, the three YOLO Heads are three branches output by the FPN (Feature Pyramid Networks), as Figure 5 .

[0114] Since the overall brightness and / or style of the scanned images corresponding to different ray sources are different, it is necessary to perform domain adaptation on the global features of the scanned images. Therefore, the initial image feature domain classification network is connected to the pre-trained feature extraction module, so that the global features of the scanned images corresponding to different ray sources can achieve distribution alignment. In a specific embodiment, the backbone network part of the target detection network is determined, and the initial image feature domain classification network is connected to the last layer of the backbone network. In this embodiment, a first gradient reversal layer is provided between the pre-trained feature extraction module and the initial image feature domain classification network. Specifically, a first gradient reversal layer is provided between the last "Resblock_body" of the CSP-DarkNet structure and the initial image feature domain classification network.

[0115] Since the color and / or size of the same target in the scanned images corresponding to different radiation sources are different, it is necessary to perform domain adaptation on the target features of the scanned images. Therefore, the pre-trained feature aggregation module is connected to the initial target feature domain classification network, so that the target features of the scanned images corresponding to different radiation sources can achieve distribution alignment. In a specific embodiment, the output part of the target detection network is determined, and the initial target feature domain classification network is connected to the first two layers of the output layer. Specifically, the initial target feature domain classification network is connected to the "Conv" before YOLO Head in the YOLO-v4 network structure. In this embodiment, a second gradient reversal layer is provided between the pre-trained feature aggregation module and the initial target feature domain classification network. Specifically, a second gradient reversal layer is provided between the "Conv" before YOLO Head in the YOLO-v4 network structure and the initial target feature domain classification network.

[0116] Among them, both the first gradient reversal layer and the second gradient reversal layer are used to confuse the source domain image set and the target domain image set, so as to reduce the difference between the first scanned image and the second scanned image, make the first scanned image and the second scanned image have almost the same marginal distribution, and then improve the classification accuracy and precision of the initial image feature domain classification network and the initial target feature domain classification network.

[0117] In this embodiment, the initial detection network is trained by the source domain image set and the target domain image set to correct the parameters in the pre-trained target detection network, the initial image feature domain classification network and the initial target feature domain classification network, so that the trained target detection network can detect dangerous goods in the scanned images corresponding to different radiation sources, and then improve the applicability of the target detection network.

[0118] In a specific embodiment, the initial image feature domain classification network can be an image global feature domain classifier. The initial target feature domain classification network can be an image local feature domain classifier.

[0119] S205: Perform target prediction on the first scanned image and the second scanned image through the pre-trained target detection network.

[0120] Specifically, the multiple first scanned images and multiple second scanned images whose sizes are adjusted in step S202 are input into the pre-trained target detection network, and the pre-trained target detection network performs target detection on the multiple first scanned images and multiple second scanned images respectively to obtain the target prediction information in each first scanned image. Among them, the target detection information includes the target prediction position and the target prediction category.

[0121] The pre-trained feature extraction module extracts features from the pre-processed scanned images to obtain corresponding image features. The pre-trained feature extraction module includes multiple sequentially connected feature extraction layers. The last feature extraction layer is connected to the initial image feature domain classification network. The last feature extraction layer extracts features from the feature map output by the previous feature extraction layer to obtain image features. In this embodiment, there are five feature extraction layers, and the last feature extraction layer is connected to the initial image feature domain classification network. Among them, the sizes of the convolutional kernels of the five feature extraction layers are different. Specifically, the output features of the last "Resblock_body" in the CSP-DarkNet structure in the YOLO-v4 network structure are input into the initial image feature domain classification network. Since the last "Resblock_body" is the last layer of the backbone network of YOLO-v4, the features it extracts contain relatively rich information. Therefore, the output features of this layer are adapted to the image feature domain.

[0122] In a specific embodiment, the pre-trained feature extraction module extracts features from the pre-processed first scanned image and the pre-processed second scanned image obtained in step S202 respectively, to obtain the image features corresponding to the first scanned image and the image features corresponding to the second scanned image, so that the initial image feature domain classification network predicts the first predicted domain classification label of the first scanned image and the first predicted domain classification label of the second scanned image according to the image features output by the pre-trained feature extraction module.

[0123] The pre-trained feature aggregation module aggregates the image features to obtain target features. Among them, the feature aggregation module aggregates the image features to obtain target features; among them, the sizes of multiple target features are different. In a specific embodiment, the image features output by the last "Resblock_body" in the CSP-DarkNet structure are used as the input of the SPP network structure. The SPP network structure performs maximum pooling of 5*5, 9*9, and 13*13 on the image features to increase the receptive field of the object detection network. The three pooled features corresponding to the image features are fused to obtain target features. In another specific embodiment, the PANet structure performs upsampling or downsampling processing on the image features output by at least one "Resblock_body" before the last "Resblock_body" in the CSP-DarkNet structure and the fused image features output by the SPP network structure respectively to obtain corresponding target features. In a specific embodiment, the SPP network structure and the PANet structure in the pre-trained feature aggregation module are connected to the initial target feature domain classification network, so that the initial target feature domain classification network predicts the second predicted domain classification label of the first scanned image and the second predicted domain classification label of the second scanned image according to the target features output by the SPP network structure and the PANet structure respectively, asFigure 5 。

[0124] In a specific embodiment, the pre-trained feature aggregation module performs feature extraction and pooling processing on the image features corresponding to the first scanned image and the image features corresponding to the second scanned image respectively, to obtain the target features corresponding to the first scanned image and the target features corresponding to the second scanned image.

[0125] The pre-trained target prediction module performs target detection on the target features to obtain target prediction information including the target in the pre-processed scanned image. Specifically, the YOLO Head in the YOLO-v4 network structure predicts the target prediction information in the scanned image based on the target features corresponding to the scanned image output by the pre-trained feature aggregation module.

[0126] In a specific embodiment, the YOLO Head predicts the target prediction information including the target in the first scanned image according to the target features corresponding to the first scanned image. Specifically, the target prediction position and the target prediction category of the target are predicted.

[0127] S206: The initial image feature domain classification network performs domain adaptation on the image features to obtain the first predicted domain classification label of the pre-processed scanned image.

[0128] Please refer to Figure 6 , Figure 6 which is the structural block diagram of a specific embodiment of the initial image feature domain classification network provided by this application.

[0129] Specifically, the initial image feature domain classification network includes a first image feature extraction unit, a second image feature extraction unit, a pooling layer, and a fully connected layer that are connected in sequence. Among them, there is one first image feature extraction unit, and there are three second image feature extraction units, and the three second image feature extraction units are connected in sequence, and the first second image feature extraction unit is connected to the first image feature extraction unit.

[0130] In this embodiment, the first image feature extraction unit includes a cascaded first convolutional layer and a first activation function layer, and the second image feature extraction unit includes a cascaded second convolutional layer and a second activation function layer. That is to say, there are multiple second convolutional layers and second activation function layers, and the second convolutional layer and the second activation function layer are spaced apart from each other.

[0131] In a specific embodiment, the first convolutional layer in the first image feature extraction unit extracts image features corresponding to the first scanned image / second scanned image to obtain a corresponding first feature map. In an alternative embodiment, in order to reduce the workload and reduce the differences between different image features, the size of the first feature map is adjusted to a second preset size. Specifically, the size of the first feature map is reduced to (16, 16, 1024), and they are concatenated by channels and input together into the first activation function layer in the first image feature extraction unit. The first activation function layer performs a non-linear activation process on the first feature map to obtain a second feature map. The second feature map is input into the second convolutional layer of the first second image feature extraction unit. The second convolutional layer extracts features from the second feature map to obtain a corresponding feature map, and the second activation function layer performs an activation process on the feature map corresponding to the second convolutional layer to obtain a feature map corresponding to the second activation function layer; the second second image feature extraction unit sequentially performs feature extraction and activation processes on the feature map output by the first second image feature extraction unit to obtain a feature map corresponding to the second second image feature extraction unit. The second convolutional layer in the third second image feature extraction unit extracts features from the feature map corresponding to the second second image feature extraction unit to obtain a corresponding feature map. In order to fuse the feature data in the shallow feature map and avoid the degradation of the deep network, the feature map output by the second convolutional layer in the first second image feature extraction unit is fused with the feature map output by the second convolutional layer in the third second image feature extraction unit to obtain a third feature map; the second activation function layer in the third second image feature extraction unit performs a non-linear activation process on the third feature map to obtain a fourth feature map.

[0132] The size of the obtained fourth feature map is adjusted by a pooling layer. In this embodiment, the pooling layer includes a global average pooling layer and a global maximum pooling layer. Specifically, to reduce the computational amount of the fully connected layer, the size of the fourth feature map is adjusted by the global average pooling layer to obtain a fifth feature map. In order to enrich the feature information, the size of the fourth feature map is adjusted by the global maximum pooling layer to obtain a sixth feature map, which is cascaded with the fifth feature map.

[0133] The pooled feature maps are connected by a fully connected layer, and the first prediction domain classification label of the scanned image is determined based on the connected feature maps. In a specific embodiment, the fifth feature map and the sixth feature map are fused by the fully connected layer, and then the first prediction domain classification label of the scanned image is predicted.

[0134] In a specific embodiment, the initial image feature domain classification network obtains the first predicted domain classification label corresponding to the first scanned image / second scanned image based on the image features corresponding to the first scanned image / second scanned image output by the pre-trained feature extraction module. Specifically, the output dimension of the initial image feature domain classification network is (n + 1, 1).

[0135] S207: The initial target feature domain classification network performs domain adaptation on the target features to obtain the second predicted domain classification label of the preprocessed scanned image.

[0136] Please refer to Figure 7 , Figure 7 which is the structural block diagram of a specific embodiment of the initial target feature domain classification network provided by the present application.

[0137] Specifically, the initial target feature domain classification network includes a cascaded first target feature extraction unit, a second target feature extraction unit, and a third target feature extraction unit. Among them, the first target feature extraction unit and the third target feature extraction unit are each one. The second target feature extraction unit is three, and the three second target feature extraction units are connected in sequence. The first second target feature extraction unit is connected to the first target feature extraction unit, and the third second target feature extraction unit is connected to the third target feature extraction unit.

[0138] In this embodiment, the first target feature extraction unit includes a cascaded first feature extraction layer and a first activation layer. The second target feature extraction unit includes a cascaded second feature extraction layer and a second activation layer. The third feature extraction unit includes a cascaded third feature extraction layer and an output layer. That is to say, the second feature extraction layer and the second activation layer are each multiple, and the third feature extraction layer and the third activation layer are arranged at intervals of each other.

[0139] In a specific embodiment, the first feature extraction layer in the first target feature extraction unit extracts the target features corresponding to the first scan image / second scan image to obtain a first target feature map, and the first activation layer performs non-linear activation on the first target feature map to obtain a second target feature map. The second target feature map is input into the first second target feature extraction unit. The second feature extraction layer in the second target feature extraction unit extracts the features of the second feature map to obtain a corresponding target feature map, and the second activation layer performs non-linear activation on the target feature map extracted by the previous second feature extraction layer to obtain the target feature map corresponding to the second activation layer. The second second target feature extraction unit and the third feature extraction unit sequentially perform target feature extraction and non-linear activation processing on the target feature map output by the previous feature extraction unit. Among them, in order to fuse the feature data in the shallow target feature map and avoid the degradation of the deep network, the target feature map output by the second feature extraction layer in the first second target feature extraction unit is fused with the target feature map output by the second feature extraction layer in the third second target feature extraction unit to obtain a third target feature map; the second activation layer in the third second target feature extraction unit performs non-linear activation on the third target feature map to obtain a fourth target feature map. The fourth target feature map is input into the third target feature extraction unit, and the third feature extraction layer extracts the target features of the fourth target feature map to obtain a fifth target feature map; the output layer determines and outputs the second prediction domain classification label corresponding to the scan image based on the fifth target feature map.

[0140] In a specific embodiment, the initial target feature domain classification network obtains the second prediction domain classification label corresponding to the first scan image / second scan image based on the target features corresponding to the first scan image / second scan image output by the pre-trained feature aggregation module. Specifically, the output dimension of the initial target feature domain classification network is (n + 1, 256).

[0141] In an alternative embodiment, in order to improve the initial image feature domain classification network and the initial target feature domain classification network, and enable the initial image feature domain classification network and the initial target feature domain classification network to have a larger receptive field and stronger information expression ability, a multi-scale feature fusion mechanism is added to the second convolutional layer and the second feature extraction layer. Specifically, the second convolutional layer and the second feature extraction layer each include at least two different convolutional kernels. The feature maps output by the previous level are respectively extracted by different convolutional kernels and output, and then the sub-feature maps output by each convolutional kernel are fused to obtain the feature map / target feature map corresponding to the second convolutional layer / second target feature map.

[0142] Please refer to Figure 8 , Figure 8 which is the structural block diagram of a specific embodiment of the second convolutional layer or the second feature extraction layer provided in this application.

[0143] In this embodiment, the second convolutional kernel and the second feature extraction layer each include three different convolutional kernels with sizes of 3*3, 5*5, and 7*7 respectively, and the number of channels of all three convolutional kernels is 128.

[0144] In another alternative embodiment, in order to shorten the running time of the initial detection network, the second convolutional layer and / or the second feature extraction layer use depthwise separable convolutional layers.

[0145] S208: Construct a loss function based on the target prediction information and the labeled target ground truth information of the same first scanned image, the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same first scanned image, and the first predicted domain classification label and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label of the same second scanned image.

[0146] Specifically, construct a loss function based on the target prediction information and the target ground truth information of the same first scanned image, the first predicted domain classification label of the first scanned image and the first true domain classification label labeled for the same first scanned image, the second predicted domain classification label of the first scanned image and the second true domain classification label labeled for the same first scanned image, the first predicted domain classification label of the second scanned image and the first true domain classification label labeled for the same second scanned image, and the second predicted domain classification label of the second scanned image and the second true domain classification label labeled for the same second scanned image.

[0147] In a specific embodiment, the loss function is as follows:

[0148]

[0149] In the formula: L det is the target information loss function of the target detection network; is the domain classification label loss function of the image feature domain classification network; is the domain classification label loss function of the target feature domain classification network; λ is the weighting coefficient.

[0150] In this embodiment, in order to make the training of the initial detection network more stable, λ is set to 0.8.

[0151] Construct the target information loss function L of the pre-trained target detection network based on the target prediction information corresponding to the detected target in the first scanned image obtained from the pre-trained target detection network and the target ground truth information corresponding to the target labeled in the corresponding first scanned image det , and the target information loss function of the pre-trained target detection network includes a target position loss function and a target category loss function.

[0152] Based on the first predicted domain classification label of the first scanned image predicted by the initial image feature domain classification network and the first true domain classification label annotated for the same first scanned image, and the first predicted domain classification label of the second scanned image and the first true domain classification label annotated for the same second scanned image, construct the domain classification label loss function of the initial image feature domain classification network

[0153] In a specific embodiment, the domain classification label loss function of the initial image feature domain classification network is as follows:

[0154]

[0155] In the formula: is the domain classification label loss function of the initial image feature domain classification network; M is the total number of the first scanned image and the second scanned image; n + 1 is the total number of the first predicted domain classification labels; I mc is the sign function. When the first true domain classification label / second true domain classification label of the m-th scanned image is equal to c, the value of this function is 1, otherwise the value is 0. is the predicted probability that the m-th scanned image belongs to the category c.

[0156] Based on the second predicted domain classification label of the first scanned image predicted by the initial target feature domain classification network and the second true domain classification label annotated in the same first scanned image, and the second predicted domain classification label of the second scanned image and the second true domain classification label annotated for the same second scanned image, construct the domain classification label loss function of the initial target feature domain classification network

[0157] In a specific embodiment, the domain classification label loss function of the target feature domain classification network is as follows:

[0158]

[0159] In the formula, is the domain classification label loss function of the initial target feature domain classification network; represents the probability that the result output by the target feature domain classification network of the m-th scanned image is predicted as the second predicted domain class label at the coordinate.

[0160] S209: Train the initial detection network based on the loss function to obtain the detection network.

[0161] Specifically, correct the weights in the pre-trained target detection network, the initial image feature domain classification network, and the initial target feature domain classification network in the initial detection network based on the loss function to obtain the detection network. Among them, the detection network includes a target detection network, an image feature domain classification network, and a target feature domain classification network.

[0162] In a specific embodiment, based on the error value between the target prediction information and the target true information of the same first scanned image, the error value between the first predicted domain classification label of the first scanned image and the first true domain classification label annotated for the same first scanned image, the error value between the second predicted domain classification label of the first scanned image and the second true domain classification label annotated for the same first scanned image, the error value between the first predicted domain classification label of the second scanned image and the first true domain classification label annotated for the same second scanned image, and the error value between the second predicted domain classification label of the second scanned image and the second true domain classification label annotated for the same second scanned image, the initial detection network is trained to obtain a detection network.

[0163] In a specific embodiment, the prediction result of the initial detection network is backpropagated, and the weights of the initial detection network are corrected according to the loss value fed back by the loss function. In an alternative embodiment, the parameters of the initial detection network can also be corrected to achieve the training of the initial detection network.

[0164] In a specific embodiment, when updating the parameters of each module in the initial detection network, after each round of training is completed, first update the gradient generated by the pre-trained object detection network, and then sequentially update the gradients generated by the initial image feature domain classification network and the initial target feature domain classification network. At this time, the parameters of the pre-trained object detection network, the initial image feature domain classification network, and the initial target feature domain classification network in this round are updated. After that, continuously repeat the iterative training until the network converges.

[0165] Input the first scanned image and the second scanned image containing the target into the initial detection network, and the initial detection network predicts the target position, target category, the first predicted domain classification label corresponding to the scanned image, and the second predicted domain classification label. When the weighted sum of the error value between the target prediction information and the target true information of the same first scanned image, the error value between the first predicted domain classification label of the first scanned image and the first true domain classification label annotated for the same first scanned image, the error value between the second predicted domain classification label of the first scanned image and the second true domain classification label annotated for the same first scanned image, the error value between the first predicted domain classification label of the second scanned image and the first true domain classification label annotated for the same second scanned image, and the error value between the second predicted domain classification label of the second scanned image and the second true domain classification label annotated for the same second scanned image is less than a preset threshold, and the preset threshold can be set by oneself, such as 1%, 5%, etc., then stop the training of the initial detection network and obtain a detection network.

[0166] In another alternative embodiment, after the initial detection network is trained a preset number of times with the first scanned image included in the source domain image set and the second scanned image included in the target domain image set, the detection network can be obtained. In this embodiment, to make the training process more stable and further improve the training speed, the preset number of times is 10 times. It can also be set according to the actual situation.

[0167] S210: Remove the image feature domain classification network and the target feature domain classification network, and retain the target detection network.

[0168] Specifically, in order to enable the global features and local features of the scanned images corresponding to different radiation sources to achieve aligned distribution and make the training process more stable, the image feature domain classification network and the target feature domain classification network are used to assist the target detection network in training, so that the target detection network can detect the targets in the scanned images corresponding to different radiation sources. In this embodiment, the image feature domain classification network and the target feature domain classification network are used to assist the target detection network in training, taking into account the domain matching of the image global features and the image local features.

[0169] After the detection network is stable, separate the image feature domain classification network and the target feature domain classification network in the detection network from the target detection network, and only retain the target detection network.

[0170] In this embodiment, the trained target detection network has the same network structure as the YOLO-v4 network. Without modifying the known YOLO-v4 network structure, by using the image feature domain classification network and the target feature domain classification network to assist the target detection network in training, the detection accuracy of the target detection network for dangerous goods in the scanned images under different radiation source security inspection machines can be improved.

[0171] The target detection network training method provided in this embodiment includes: obtaining a source domain image set and a target domain image set, where the source domain image set includes multiple first scanned images labeled with target true information, as well as a first true domain classification label and a second true domain classification label; the target domain image set includes multiple second scanned images labeled with a first true domain classification label and a second true domain classification label and without target information; the first scanned images and the second scanned images are scanned based on different radiation sources respectively; detecting the first scanned images and the second scanned images respectively through the constructed initial detection network to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned images, and the first predicted domain classification label and the second predicted domain classification label of the second scanned images; constructing a loss function based on the target prediction information of the same first scanned image and the labeled target true information, the first predicted domain classification label of the same first scanned image and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label, and the first predicted domain classification label of the same second scanned image and the labeled first true domain classification label and the second predicted domain classification label and the labeled second true domain classification label; iteratively training the initial detection network based on the loss function to obtain a target detection network. In this application, the initial detection network is used to predict the scanned images included in the source domain image set and the target domain image set, and the initial detection network is trained based on the first predicted domain classification label and the second predicted domain classification label corresponding to the first scanned image and the second scanned image respectively and the first true domain classification label and the second true domain classification label corresponding to the first scanned image and the second scanned image respectively, saving the time cost and labor cost of labeling the true information of the second scanned image. The obtained target detection network can detect dangerous goods in scanned images collected by different radiation sources, making the application range of the target detection network wider.

[0172] Please refer to Figure 9 , Figure 9 which is a schematic flowchart of the target detection method provided in this application. In this embodiment, a target detection method is provided, and the target detection method includes the following steps.

[0173] S31: Obtain an image to be detected.

[0174] Specifically, obtain the image to be detected through an X-ray security inspection machine. Among them, the radiation source in the X-ray security inspection machine can be an X-ray. The image to be detected is a scanned image including at least one target object. In other alternative embodiments, the scanned images corresponding to different radiation sources can also be obtained through other means, and the obtained scanned images are used as the images to be detected. Different radiation sources refer to different voltages, currents, or different energy spectra. The penetration power of X-rays is different, and the obtained scanned images are different.

[0175] S32: Preprocess the obtained image to be detected to obtain a preprocessed image.

[0176] Specifically, adjust the size of the image to be detected to a preset size; and perform normalization processing on the image to be detected to obtain a corresponding preprocessed image.

[0177] In one embodiment, adjust the size of the image to be detected to a preset size. Specifically, the preset size is 512*512*3. In other alternative embodiments, the preset size can also be set according to actual situations. In a specific embodiment, interpolate the image to be detected to a size of 512*512*3 in terms of length and width.

[0178] Perform normalization processing on the image to be detected after size adjustment to obtain a preprocessed image corresponding to the image to be detected. Specifically, convert the matrix of the image to be detected after size adjustment to between 0 and 1, thereby achieving normalization processing.

[0179] S33: Perform object detection on the preprocessed image through an object detection network to obtain object information.

[0180] Specifically, extract features from the preprocessed image through the feature extraction module in the object detection network obtained in the above embodiment to obtain a feature map; perform feature fusion on the feature map output by the feature extraction module through the feature fusion module to obtain a target feature map; perform object detection based on the target feature map corresponding to the feature fusion module through the target prediction module to detect the object information of the image to be detected. Among them, the object information includes the object position and the object category.

[0181] Please refer to Figure 10 , Figure 10 which is a schematic flowchart of a specific embodiment of the object detection method provided by this application.

[0182] In this embodiment, the object detection network detects the object included in the preprocessed image. In response to detecting that the object category of the object is a preset category, output the object whose object category is the same as the preset category to prompt that the object belongs to the item of the preset category. Specifically, the preset categories are knives, flammable substances, explosive substances, liquids, etc., and the items corresponding to the preset categories are dangerous items. In a specific embodiment, the image to be detected is an X-ray image, and the object detection network is a trained YOLO-v4 network structure.

[0183] The object detection method provided in this embodiment obtains a to-be-detected image, preprocesses the obtained to-be-detected image to obtain a preprocessed image, and performs object detection on the preprocessed image through an object detection network to obtain object information. In this embodiment, the object detection network is used to perform object detection on the scanned images corresponding to different radiation sources, so as to obtain the target position and target category of the object in the scanned image, and improve the accuracy of detecting dangerous goods in the scanned image.

[0184] Please refer to Figure 11 , Figure 11 which is a schematic block diagram of the object detection network training device provided in this application. In this embodiment, a training device 40 is provided, which includes a sample acquisition module 41, a prediction module 42, a function construction module 43, and a processing module 44.

[0185] The sample acquisition module 41 is used to acquire a source domain image set and a target domain image set. The source domain image set includes multiple first scanned images labeled with target true information, a first true domain classification label, and a second true domain classification label; the target domain image set includes multiple second scanned images labeled with a first true domain classification label and a second true domain classification label and without labeled object information; the first scanned image and the second scanned image are scanned based on different radiation sources respectively.

[0186] The prediction module 42 is used to perform detection on the first scanned image and the second scanned image respectively through the constructed initial detection network, and obtain the target prediction information, the first prediction domain classification label, the second prediction domain classification label of the first scanned image, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image.

[0187] The function construction module 43 is used to construct a loss function based on the target prediction information and the labeled target true information of the same first scanned image, the first prediction domain classification label and the labeled first true domain classification label and the second prediction domain classification label and the labeled second true domain classification label of the same first scanned image, and the first prediction domain classification label and the labeled first true domain classification label and the second prediction domain classification label and the labeled second true domain classification label of the same second scanned image.

[0188] The processing module 44 is used to perform iterative training on the initial detection network based on the loss function to obtain the object detection network.

[0189] In this embodiment, the initial detection network predicts the scanned images included in the source domain image set and the target domain image set, and trains the initial detection network based on the first predicted domain classification label and the second predicted domain classification label corresponding to the first scanned image and the second scanned image respectively, and the first true domain classification label and the second true domain classification label corresponding to the first scanned image and the second scanned image respectively, saving the time cost and labor cost of annotating the true information of the second scanned image. The obtained target detection network can detect dangerous goods in the scanned images collected by different ray sources, making the applicable range of the target detection network wider.

[0190] Please refer to Figure 12 , Figure 12 which is a schematic block diagram of the target detection device provided by this application. In this embodiment, a target detection device 50 is provided, which includes an image acquisition module 51, a preprocessing module 52, and a detection module 53.

[0191] The image acquisition module 51 is used to acquire the image to be detected; the image to be detected includes at least one target object. Among them, the image to be detected is a scanned image corresponding to X-rays of different ray sources; different ray sources refer to different voltages, currents, or different energy spectra, different penetrations of X-rays, and different obtained scanned images.

[0192] The preprocessing module 52 is used to preprocess the image to be detected to obtain a preprocessed image.

[0193] The detection module 53 is used to perform target detection on the preprocessed image through the target detection network to obtain the target object information of the target object; among them, the target detection network is obtained by the target detection network training method in the above embodiment.

[0194] The target detection device provided in this embodiment includes: an image acquisition module, which is used to acquire the image to be detected; the image to be detected includes at least one target object; a preprocessing module, which is used to preprocess the image to be detected to obtain a preprocessed image; a detection module, which is used to perform target detection on the preprocessed image through the target detection network to obtain the target object information of the target object. In this embodiment, the target detection network performs target detection on the scanned images corresponding to different ray sources, and then obtains the target position and target category of the target object in the scanned image, improving the accuracy of detecting dangerous goods in the scanned image.

[0195] Please refer to Figure 13 , Figure 13It is a schematic block diagram of an embodiment of the terminal provided by the present application. The terminal 70 in this embodiment includes: a processor 71, a memory 72, and a computer program stored in the memory 72 and executable on the processor 71. When the computer program is executed by the processor 71, it implements the target detection network training method and the target detection method in the above embodiments. To avoid repetition, details are not described herein one by one.

[0196] Please refer to Figure 14 , Figure 14 It is a schematic block diagram of an embodiment of the computer-readable storage medium provided by the present application. In an embodiment of the present application, a computer-readable storage medium 90 is further provided. The computer-readable storage medium 90 stores a computer program 901. The computer program 901 includes program instructions. When the processor executes the program instructions, it implements the target detection network training method and the target detection method provided by the embodiment of the present application.

[0197] Among them, the computer-readable storage medium 90 may be an internal storage unit of the computer device in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium 90 may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0198] The above are only embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for training an object detection network, characterized in that, The method for training the object detection network includes: Obtain a source domain image set and a target domain image set. The source domain image set includes multiple first scanned images labeled with object true information, as well as a first true domain classification label and a second true domain classification label. The target domain image set includes multiple second scanned images labeled with a first true domain classification label and a second true domain classification label and without labeled object information. The first scanned images and the second scanned images are scanned based on different ray sources respectively; Detect the first scanned images and the second scanned images respectively through the constructed initial detection network to obtain the object prediction information, first prediction domain classification label, second prediction domain classification label of the first scanned images, and the first prediction domain classification label and second prediction domain classification label of the second scanned images; The steps for constructing the initial detection network include: Construct the initial detection network based on an initial object detection network, an initial image feature domain classification network, and an initial object feature domain classification network. Among them, the initial object detection network includes an initial feature extraction module, an initial feature aggregation module, and an initial object prediction module connected in sequence. The initial feature extraction module is connected to the initial image feature domain classification network, and the initial feature aggregation module is connected to the initial object feature domain classification network; Construct a loss function based on the object prediction information of the same first scanned image and the labeled object true information, the first prediction domain classification label of the same first scanned image and the labeled first true domain classification label, the second prediction domain classification label and the labeled second true domain classification label, and the first prediction domain classification label of the same second scanned image and the labeled first true domain classification label, and the second prediction domain classification label and the labeled second true domain classification label; Iteratively train the initial detection network based on the loss function to obtain an object detection network.

2. The method for training the object detection network according to claim 1, wherein The iteratively training the initial detection network based on the loss function to obtain an object detection network includes: Correct the weights in the initial object detection network, the initial image feature domain classification network, and the initial object feature domain classification network in the initial detection network based on the loss function to obtain an object detection network, an image feature domain classification network, and an object feature domain classification network; Remove the image feature domain classification network and the object feature domain classification network, and retain the object detection network.

3. The method for training an object detection network according to claim 1, wherein The first scanned images and the second scanned images are scanned images corresponding to X-rays of different ray sources. The object prediction information includes an object prediction position and an object prediction category; Before the step of detecting the first scanned images and the second scanned images respectively through the constructed initial detection network to obtain the object prediction information, first prediction domain classification label, second prediction domain classification label of the first scanned images, and the first prediction domain classification label and second prediction domain classification label of the second scanned images, it further includes: Adjust the sizes of the first scanned image and the second scanned image to a first preset size; And perform normalization processing on the first scanned image and the second scanned image to obtain corresponding preprocessed scanned images.

4. The method for training an object detection network according to claim 1, wherein A first gradient reversal layer is provided between the initial feature extraction module and the initial image feature domain classification network; and / or a second gradient reversal layer is provided between the initial feature aggregation module and the initial target feature domain classification network.

5. The method for training a target detection network according to claim 3, wherein Detecting the first scanned image and the second scanned image respectively by the constructed initial detection network to obtain target prediction information of the first scanned image, a first prediction domain classification label, a second prediction domain classification label, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image, includes: Performing feature extraction on the preprocessed scanned image by the initial feature extraction module to obtain corresponding image features; Performing domain adaptation on the image features by the initial image feature domain classification network to obtain the first prediction domain classification label of the preprocessed scanned image.

6. The method for training an object detection network according to claim 5, wherein The initial feature extraction module includes a plurality of sequentially connected feature extraction layers, and the last feature extraction layer is connected to the initial image feature domain classification network; The initial feature extraction module performing feature extraction on the preprocessed scanned image to obtain corresponding image features, includes: The last feature extraction layer performs feature extraction on the feature map output by the previous feature extraction layer to obtain the image features.

7. The method for training an object detection network according to claim 5, wherein, The initial image feature domain classification network includes a first image feature extraction unit, a second image feature extraction unit, a pooling layer, and a fully connected layer that are sequentially connected; wherein, the first image feature extraction unit includes a cascaded first convolutional layer and a first activation function layer, the second image feature extraction units are multiple and sequentially connected, and the second image feature extraction unit includes a cascaded second convolutional layer and a second activation function layer; The initial image feature domain classification network performing domain adaptation on the image features to obtain the first prediction domain classification label of the preprocessed scanned image, includes: After performing feature extraction on the image features by the first convolutional layer, a first feature map is obtained; Performing non-linear activation processing on the first feature map by the first activation function layer to obtain a second feature map; Fusing the feature map corresponding to the second convolutional layer in the last second image feature extraction unit with the feature maps corresponding to the second convolutional layer in at least one previous second image feature extraction unit to obtain a third feature map; Performing non-linear activation processing on the third feature map by the second activation function layer in the last second image feature extraction unit to obtain a fourth feature map; Adjusting the size of the fourth feature map by the pooling layer; The fully connected layer determines the first prediction domain classification label of the first scanned image / the second scanned image based on the fourth feature map after size adjustment.

8. The method for training an object detection network according to claim 7, wherein, The second convolutional layer includes at least two convolutional kernels, and the second convolutional layer is a depthwise separable convolutional layer.

9. The method for training an object detection network according to claim 7, wherein the initial image feature domain classification network performs domain adaptation on the image features to obtain the first predicted domain classification label of the preprocessed scanned image, and further includes: Adjusting the size of the first feature map to a second preset size.

10. The method for training an object detection network according to claim 7, wherein, The pooling layer includes a global average pooling layer and a global max pooling layer; Adjusting the size of the fourth feature map through the pooling layer includes: The global average pooling layer adjusts the size of the fourth feature map to obtain a fifth feature map; The global max pooling layer adjusts the size of the fourth feature map to obtain a sixth feature map; The fully connected layer determines the first predicted domain classification label of the first scanned image / the second scanned image based on the fourth feature map with adjusted size, including: The fully connected layer performs feature fusion on the fifth feature map and the sixth feature map to determine the first predicted domain classification label of the first scanned image / the second scanned image.

11. The method for training an object detection network according to claim 5, wherein detecting the first scanned image and the second scanned image respectively through the constructed initial detection network to obtain the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned image, and the first predicted domain classification label and the second predicted domain classification label of the second scanned image, and further includes: The initial feature aggregation module performs feature aggregation on the image features to obtain target features; The initial target feature domain classification network performs domain adaptation on the target features to obtain the second predicted domain classification label of the preprocessed scanned image.

12. The method for training an object detection network according to claim 11, wherein the initial feature aggregation module performs feature aggregation on the image features to obtain target features, including: Adjusting the size of the image features through the initial feature aggregation module to obtain a plurality of feature maps with different sizes; Performing feature fusion on the plurality of feature maps to obtain the target features corresponding to the first scanned image / the second scanned image.

13. The method for training an object detection network according to claim 11, wherein The initial target feature domain classification network includes a cascaded first target feature extraction unit, a second target feature extraction unit, and a third target feature extraction unit; the first target feature extraction unit includes a cascaded first feature extraction layer and a first activation layer, the second target feature extraction unit includes a cascaded second feature extraction layer and a second activation layer, and the third target feature extraction unit includes a cascaded third feature extraction layer and an output layer; The initial target feature domain classification network performs domain adaptation on the target features to obtain the second predicted domain classification label of the preprocessed scanned image, including: Performing feature extraction on the target features corresponding to the first scanned image / the second scanned image through the first feature extraction layer in the first target feature extraction unit to obtain a first target feature map; Non-linearly activate the first target feature map through the first activation layer to obtain a second target feature map; Fuse the target feature map corresponding to the second feature extraction layer in the last second target feature extraction unit with the target feature maps corresponding to the second feature extraction layers in at least one previous second target feature extraction unit to obtain a third target feature map; Non-linearly activate the third target feature map through the second activation layer in the last second target feature extraction unit to obtain a fourth target feature map; Extract features from the fourth target feature map through the third feature extraction layer in the third feature extraction unit to obtain a fifth target feature map; The output layer determines the second prediction domain classification label of the first scan image / the second scan image based on the fifth target feature map.

14. The method for training an object detection network according to claim 13, wherein The second feature extraction layer includes at least two convolutional kernels, and the second feature extraction layer is a depthwise separable convolutional layer.

15. The method for training a target detection network according to claim 11, wherein The step of respectively detecting the first scan image and the second scan image through the constructed initial detection network to obtain the target prediction information, the first prediction domain classification label, and the second prediction domain classification label of the first scan image and the second scan image further includes: The initial target prediction module performs target detection on the target features corresponding to the first scan image / the second scan image to obtain the target prediction information of the preprocessed scan image containing the target.

16. The method for training a target detection network according to claim 1, wherein Constructing a loss function based on the target prediction information of the same first scan image and the annotated target ground truth information, the first prediction domain classification label of the same first scan image and the annotated first ground truth domain classification label and the second prediction domain classification label and the annotated second ground truth domain classification label, and the first prediction domain classification label of the same second scan image and the annotated first ground truth domain classification label and the second prediction domain classification label and the annotated second ground truth domain classification label includes: The loss function is: Where: L det is the target information loss function of the target detection network; is the domain classification label loss function of the image feature domain classification network; is the domain classification label loss function of the target feature domain classification network; λ is the weighting coefficient.

17. The method for training a target detection network according to claim 16, wherein Constructing a loss function based on the target prediction information of the same first scan image and the annotated target ground truth information, the first prediction domain classification label of the same first scan image and the annotated first ground truth domain classification label and the second prediction domain classification label and the annotated second ground truth domain classification label, and the first prediction domain classification label of the same second scan image and the annotated first ground truth domain classification label and the second prediction domain classification label and the annotated second ground truth domain classification label includes: The domain classification label loss function of the image feature domain classification network is: In the formula: is the domain classification label loss function of the image feature domain classification network; M is the total number of the first scan image and the second scan image; n + 1 is the total number of the first predicted domain classification labels; I mc is the sign function, when the first true domain classification label of the m-th scan image is equal to c, the value of this function is 1, otherwise the value is 0, is the predicted probability that the m-th scan image belongs to the category c, where the first predicted domain classification label indicates whether the scan image belongs to the source domain image set or the target domain image set.

18. The method for training a target detection network according to claim 16, wherein Construct a loss function based on the target prediction information and the labeled target ground truth information of the same first scanned image, the first predicted domain classification label and the labeled first ground truth domain classification label, and the second predicted domain classification label and the labeled second ground truth domain classification label of the same first scanned image, and the first predicted domain classification label and the labeled first ground truth domain classification label, and the second predicted domain classification label and the labeled second ground truth domain classification label of the same second scanned image, including: The domain classification label loss function of the target feature domain classification network is: In the formula, is the domain classification label loss function of the target feature domain classification network; represents the probability that the result output by the target feature domain classification network of the m-th scanned image is predicted as the second predicted domain class label at the coordinate.

19. The method for training a target detection network according to claim 1, wherein: Before obtaining the target prediction information, the first predicted domain classification label, the second predicted domain classification label of the first scanned image, and the first predicted domain classification label and the second predicted domain classification label of the second scanned image by respectively detecting the first scanned image and the second scanned image through the constructed initial detection network, further includes: Pre-train the initial target detection network in the initial detection network.

20. The method for training a target detection network according to claim 19, wherein: The pre-training of the initial target detection network in the initial detection network includes: Detect the first scanned image through the initial target detection network to obtain the target prediction information of the first scanned image; Construct an initial loss function through the target prediction information and the target ground truth information corresponding to the same target of the first scanned image; Pre-train the initial target detection network based on the initial loss function.

21. A target detection method, characterized in that, The target detection method includes: Obtain an image to be detected; the image to be detected includes at least one target object; Preprocess the image to be detected to obtain a preprocessed image; Perform target detection on the preprocessed image through a target detection network to obtain target object information of the target object; wherein, the target detection network is obtained by the method for training a target detection network according to any one of claims 1-20 above.

22. The object detection method according to claim 21, wherein The image to be detected is a scanned image corresponding to X-rays of different radiation sources; the target object information includes a target position and a target category; The preprocessing of the image to be detected to obtain a preprocessed image includes: Adjust the size of the image to be detected to a preset size; And perform normalization processing on the image to be detected to obtain the corresponding preprocessed image.

23. A training device, characterized in that, The training device includes: A sample acquisition module, configured to acquire a source domain image set and a target domain image set, the source domain image set includes multiple first scanned images labeled with target ground truth information, a first ground truth domain classification label, and a second ground truth domain classification label; the target domain image set includes multiple second scanned images labeled with a first ground truth domain classification label and a second ground truth domain classification label and without labeled target information; the first scanned image and the second scanned image are respectively scanned based on different radiation sources; A prediction module, configured to detect the first scanned image and the second scanned image respectively through the constructed initial detection network, so as to obtain the target prediction information of the first scanned image, the first prediction domain classification label, the second prediction domain classification label, and the first prediction domain classification label and the second prediction domain classification label of the second scanned image; The prediction module is further configured to construct the initial detection network based on the initial target detection network, the initial image feature domain classification network, and the initial target feature domain classification network; wherein, the initial target detection network includes an initial feature extraction module, an initial feature aggregation module, and an initial target prediction module connected in sequence; the initial feature extraction module is connected to the initial image feature domain classification network, and the initial feature aggregation module is connected to the initial target feature domain classification network; A function construction module, configured to construct a loss function based on the target prediction information of the same first scanned image and the labeled target ground truth information, the first prediction domain classification label of the same first scanned image and the labeled first ground truth domain classification label, the second prediction domain classification label and the labeled second ground truth domain classification label, and the first prediction domain classification label of the same second scanned image and the labeled first ground truth domain classification label, and the second prediction domain classification label and the labeled second ground truth domain classification label; A processing module, configured to perform iterative training on the initial detection network based on the loss function to obtain a target detection network.

24. A target detection device, characterized in that, The target detection device includes: An image acquisition module, configured to acquire an image to be detected; the image to be detected includes at least one target object; A preprocessing module, configured to preprocess the image to be detected to obtain a preprocessed image; A detection module, configured to perform target detection on the preprocessed image through a target detection network to obtain target object information of the target object; wherein, the target detection network is obtained by the target detection network training method according to any one of claims 1-20 above.

25. A terminal, characterized in that, The terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is configured to execute program data to implement the steps in the target detection network training method according to any one of claims 1 to 20; or implement the steps in the target detection method according to claim 21 or 22.

26. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps in the target detection network training method according to any one of claims 1 to 20; or implements the steps in the target detection method according to claim 21 or 22.

Citation Information

Patent Citations

  • Target detection network construction method and device and target detection method

    CN111274981A

  • Domain adaptive target detection method and system considering category semantic matching

    CN113807420A

  • Object detection device, object detection system and object detection method

    JP2022083513A