Method and device for cross-domain target detection and related product
By using the cross-domain object detection model for feature extraction and domain classification in cross-domain object detection, the performance degradation caused by domain differences in cross-domain detection is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411803632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-27
AI Technical Summary
Cross-domain object detection faces the problem of degradation in model performance due to differences in data distribution and feature representation between data sets or fields.
A cross-domain object detection method is adopted to obtain training sample images and their domain labels, and use the feature extraction network, image-level domain adaptive network, target object detection network and instance-level domain adaptive network in the cross-domain object detection model to perform feature extraction, domain classification and loss calculation, and finally update the model parameters to improve detection accuracy.
By eliminating the domain differences between the source domain and the target domain, the accuracy of cross-domain object detection is improved, and the generalization and transfer learning ability of the model are enhanced.
Smart Images

Figure CN120047657A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for cross-domain object detection. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section.
[0003] Cross-domain object detection refers to the task of performing object detection or image classification between different data sets or domains. Due to differences in data distribution, feature representation, etc. between different data sets or domains, cross-domain detection faces great challenges. For example, in the object detection task of autonomous driving, due to changes in weather conditions (such as foggy days), conventional object detection models may experience a significant drop in performance due to inconsistent feature distributions between training and test data. This is mainly because the model usually assumes that the data has the same feature distribution during training, but in actual applications, this assumption does not always hold. Therefore, a method is needed to reduce the domain difference between the source domain (such as sunny days) and the target domain (such as foggy days). Summary of the Invention
[0004] An object of the present disclosure is to provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for cross-domain object detection, which can reduce the inter-domain difference between the source domain and the target domain, thereby improving the accuracy of cross-domain object detection.
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or will be partially learned through the practice of the present disclosure.
[0006] Embodiments of the present disclosure provide a method for cross-domain object detection, including: obtaining a training sample image and a domain label of the training sample image, where the domain label of the training sample image is a source domain or a target domain; performing feature extraction processing on the training sample image through a feature extraction network in a cross-domain object detection model to obtain an image-level feature of the training sample image; wherein, the cross-domain object detection model further includes an image-level domain adaptation network, an object detection network, and an instance-level domain adaptation network; performing image-level domain classification on the image-level feature through the image-level domain adaptation network to determine an image-based domain classification result of the training sample image, and determining a first loss based on the domain label and the image-based domain classification result; performing object feature extraction on the image-level feature through the object detection network to obtain an instance-level feature of the training sample image; performing feature extraction processing on the instance-level feature through the instance-level domain adaptation network to obtain an object-based domain classification result of the training sample image, and determining a second loss based on the domain label and the object-based domain classification result; calculating a consistency loss between the image-based domain classification result and the object-based domain classification result to determine a third loss; updating parameters of the cross-domain object detection model based on the first loss, the second loss, and the third loss, so as to perform cross-domain object detection according to the cross-domain object detection model with updated parameters.
[0007] In some embodiments, the image-level domain adaptation network includes a multi-scale information aggregation sub-network, a first adversarial gradient reversal layer, and an image-level domain classifier; wherein, performing image-level domain classification on the image-level feature through the image-level domain adaptation network to determine an image-based domain classification result of the training sample image includes: performing multi-scale information extraction on the image-level feature through the multi-scale information aggregation sub-network to obtain multi-scale information aggregation features; performing feature extraction processing on the multi-scale information aggregation features through the first adversarial gradient reversal layer to obtain multi-scale information aggregation depth features; performing image-level domain classification on the image-level feature through the image-level domain classifier to determine an image-based domain classification result of the training sample image.
[0008] In some embodiments, when updating parameters of the cross-domain object detection model, the reverse propagation directions of parameters before and after the first adversarial gradient reversal layer are opposite.
[0009] In some embodiments, the first adversarial gradient reversal layer includes a reversal degree parameter, which is used to weight the propagation function gradient in the first adversarial gradient reversal layer; the method further includes: when the first loss is greater than a preset threshold, the reversal degree parameter is equal to a preset constant; when the first loss is less than or equal to the preset threshold, the reversal degree parameter is inversely proportional to the first loss.
[0010] In some embodiments, the multi-scale information aggregation sub-network includes multiple convolutional layers with different dilation rates; wherein, through the multi-scale information aggregation sub-network, multi-scale information extraction is performed on the image-level features to obtain multi-scale information aggregation features, including: respectively performing feature extraction on the image-level features through the multiple convolutional layers with different dilation rates to obtain feature maps with multiple different receptive fields; performing information aggregation processing on the feature maps with multiple different receptive fields to obtain the multi-scale information aggregation features.
[0011] In some embodiments, the instance-level domain adaptation network includes a second adversarial gradient reversal layer and an instance-level domain classifier, wherein, through the instance-level domain adaptation network, feature extraction processing is performed on the instance-level features to obtain an object-based domain classification result of the training sample image, and a second loss is determined based on the domain label and the object-based domain classification result, including: performing feature extraction processing on the instance-level features through the second adversarial gradient reversal layer to obtain instance-level depth features; wherein, when updating the parameters of the cross-domain target object detection model, the backpropagation directions of the parameters before and after the second adversarial gradient reversal layer are opposite; performing classification processing on the instance-level depth features through the instance-level domain classifier to obtain the object-based domain classification result.
[0012] In some embodiments, the target object detection network includes a region generation sub-network and a region of interest pooling sub-network; wherein, through the target object detection network, object feature extraction is performed on the image-level features to obtain instance-level features of the training sample image, including: performing region generation processing on the image-level features through the region generation sub-network to obtain multiple region bounding boxes; wherein, the region bounding boxes are used to frame the position of the target object; performing pooling processing on the multiple region bounding boxes through the region of interest pooling sub-network to obtain the instance-level features.
[0013] In some embodiments, if there are object bounding boxes and object labels in each object bounding box in the training sample image; wherein, the region generation sub-network performs region generation processing on the image-level feature to obtain a plurality of region bounding boxes, including: the region generation sub-network performs region generation on the image-level feature to obtain the plurality of region bounding boxes and the object existence probability in each region bounding box; the method further includes: inputting the instance-level feature into the class classifier and the bounding box regressor of the target object detection network to obtain the bounding box regression value of each region bounding box and the target recognition result in each region bounding box, wherein the bounding box regression value is used to measure the error between the region bounding box and the object bounding box; determining a fourth loss according to the object labels in each object bounding box and the object existence probability in each region bounding box; determining a fifth loss according to the object labels in each object bounding box and the target object recognition result in each region bounding box; determining a sixth loss according to the object bounding box, each region bounding box and the bounding box regression value; wherein, based on the first loss, the second loss and the third loss, parameter updating is performed on the cross-domain target object detection model, including: based on the first loss, the second loss, the third loss, the fourth loss, the fifth loss and the sixth loss, parameter updating is performed on the cross-domain target object detection model.
[0014] In some embodiments, the method further includes: obtaining an image to be predicted; performing feature extraction processing on the image to be predicted through a feature extraction network in a cross-domain target object detection model to obtain the image-level feature of the image to be predicted; performing region feature extraction on the image-level feature of the image to be predicted through the target object detection network to obtain the instance-level feature of the image to be predicted and a plurality of region bounding boxes; inputting the instance-level feature of the image to be predicted into the class classifier and the bounding box regressor of the target object detection network to obtain the bounding box regression value of each region bounding box in the image to be predicted and the target recognition result in each region bounding box.
[0015] An embodiment of the present disclosure provides an apparatus for cross-domain target detection, including: a sample acquisition module, a feature extraction module, a first loss determination module, an instance-level feature acquisition module, a second loss determination module, a third loss determination module, and a parameter update module.
[0016] Among them, the sample acquisition module is used to acquire a training sample image and the domain label of the training sample image, where the domain label of the training sample image is the source domain or the target domain; the feature extraction module can be used to perform feature extraction processing on the training sample image through the feature extraction network in the cross-domain object detection model to obtain the image-level feature of the training sample image; among them, the cross-domain object detection model further includes an image-level domain adaptation network, an object detection network, and an instance-level domain adaptation network; the first loss determination module can be used to perform image-level domain classification on the image-level feature through the image-level domain adaptation network to determine the image-based domain classification result of the training sample image, and determine the first loss based on the domain label and the image-based domain classification result; the instance-level feature acquisition module can be used to perform object feature extraction on the image-level feature through the object detection network to obtain the instance-level feature of the training sample image; the second loss determination module can be used to perform feature extraction processing on the instance-level feature through the instance-level domain adaptation network to obtain the object-based domain classification result of the training sample image, and determine the second loss based on the domain label and the object-based domain classification result; the third loss determination module can be used to calculate the consistency loss between the image-based domain classification result and the object-based domain classification result to determine the third loss; the parameter update module can be used to update the parameters of the cross-domain object detection model based on the first loss, the second loss, and the third loss, so as to perform cross-domain object detection according to the cross-domain object detection model after parameter update.
[0017] An embodiment of the present disclosure provides an electronic device, which includes: a memory and a processor; the memory is used to store computer program instructions; the processor calls the computer program instructions stored in the memory to implement the method for cross-domain object detection described in any one of the above.
[0018] An embodiment of the present disclosure provides a computer-readable storage medium, on which computer program instructions are stored, implementing the method for cross-domain object detection described in any one of the above.
[0019] An embodiment of the present disclosure provides a computer program product or a computer program, which includes computer program instructions, and the computer program instructions are stored in a computer-readable storage medium. Reading the computer program instructions from the computer-readable storage medium, the processor executes the computer program instructions to implement the above method for cross-domain object detection.
[0020] The method, apparatus, electronic device, computer-readable storage medium, and computer program product for cross-domain object detection provided by the embodiments of the present disclosure can eliminate the domain difference between the source domain and the target domain through the first loss between the domain label and the image-based domain classification result, the second loss between the domain label and the object-based domain classification result, and the third loss (consistency loss) between the image-based domain classification result and the object-based domain classification result, thereby improving the object detection accuracy of the object detection network.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 A schematic diagram of a scenario to which the method for cross-domain object detection or the apparatus for cross-domain object detection according to the embodiments of the present disclosure can be applied is shown.
[0024] Figure 2 It is a flowchart of a method for cross-domain object detection shown according to an exemplary embodiment.
[0025] Figure 3 It is a schematic diagram of an object detection model shown according to an exemplary embodiment.
[0026] Figure 4 It is a flowchart of an image-level domain classification method shown according to an exemplary embodiment.
[0027] Figure 5 The image-level domain adaptive network based on multi-scale adversarial mining shown can be the image-level domain adaptive network in the present application.
[0028] Figure 6 It is a schematic diagram of an image-level domain adaptive network shown according to an exemplary embodiment.
[0029] Figure 7 It is a flowchart of a multi-scale information extraction method shown according to an exemplary embodiment.
[0030] Figure 8 It is a flowchart of a gradient reversal method of an adversarial gradient reversal layer shown according to an exemplary embodiment.
[0031] Figure 9 It is a schematic diagram of a coefficient relationship shown according to an exemplary embodiment.
[0032] Figure 10 It is a flowchart of a method for obtaining an object-based domain classification result shown according to an exemplary embodiment.
[0033] Figure 11 It is a flowchart of a method for extracting instance-level features shown according to an exemplary embodiment.
[0034] Figure 12 It is a flowchart of a method for cross-domain object detection shown according to an exemplary embodiment.
[0035] Figure 13 It is a flowchart of a cross-domain object detection method shown according to an exemplary embodiment.
[0036] Figure 14 It is a block diagram of a device for cross-domain object detection shown according to an exemplary embodiment.
[0037] Figure 15 It shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repeated description will be omitted.
[0039] Those skilled in the art know that the embodiments of the present disclosure can be a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0040] The features, structures, or characteristics described in the present disclosure can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0041] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0042] The accompanying drawings are only schematic illustrations of the present disclosure. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings do not necessarily have to correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0043] The flowcharts shown in the accompanying drawings are only exemplary illustrations, and do not necessarily include all the contents and steps, nor do they necessarily have to be executed in the described order. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0044] In the description of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B can represent A or B. The "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality of" means two or more. The terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit that they are different; the terms "comprising", "including", and "having" are used to indicate an open-ended inclusion meaning and mean that there can be additional elements / components / etc. in addition to the listed elements / components / etc.
[0045] In order to be able to more clearly understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0046] It should be noted that in the technical solution of the present disclosure, in terms of the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of user personal information, it complies with the provisions of relevant laws and regulations, is used for legal purposes, and does not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and safeguard the security of user personal information and network security.
[0047] First, some terms related to the embodiments of the present disclosure will be explained below to facilitate understanding by those skilled in the art.
[0048] Source domain: A dataset in a specific domain used to train an object detection model.
[0049] Target domain: A dataset in a specific domain used to evaluate an object detection model (or an object detection model).
[0050] Domain drift: Due to different data acquisition methods, lighting, resolution, scale, region, etc., there are differences in the domain distributions between the target domain and the source domain.
[0051] Domain adaptation: A transfer learning method aimed at solving the problem of the decline in model generalization performance caused by the data distribution differences between different domains; its basic idea is to align the data distributions of the source domain and the target domain to reduce the differences between different domains, so that the model trained on the source domain can be applied to the target domain.
[0052] Faster R-CNN: A classic object detection model that integrates feature extraction, region of interest candidate, detection box regression, and classification modules into one network, and has high object detection performance.
[0053] AdvGRL: The proposed adversarial gradient reversal layer used to achieve adversarial mining of hard examples; multiplying the gradient passed into AdvGRL by a negative number makes the training objectives of the networks before and after AdvGRL opposite. AdvGRL is a form of adversarial training that achieves domain adaptation by adding a reversal layer to the gradient. Specifically, when backpropagating, AdvGRL reverses the gradient, causing the downstream parameters to learn in the opposite direction. This mechanism helps the model find commonalities between the source domain and the target domain while reducing domain differences.
[0054] Image-level features: Refer to the output of the feature map sharing the backbone network; image-level features usually refer to the descriptive attributes extracted from the entire image.
[0055] Instance-level features: Feature vectors obtained from RoI (region of interest) pooling; instance-level features focus more on identifying specific individuals or instances in the image.
[0056] RPN: Region Proposal Network, an algorithm for generating candidate regions on the feature map.
[0057] ROI pooling: Region of Interest pooling.
[0058] Some noun concepts related to the embodiments of the present disclosure were introduced above. Next, the technical features involved in the embodiments of the present disclosure will be introduced.
[0059] Currently, cross-domain object detection technology is mainly implemented through deep learning methods, and there are some technical problems in related technologies.
[0060] 1. It is difficult to completely eliminate the inter-domain differences: Although related technologies can reduce the inter-domain differences through adversarial learning and transfer learning, it is still difficult to achieve complete alignment in complex scenarios, resulting in a decrease in detection accuracy.
[0061] 2. High computational complexity: The multi-scale feature fusion method has a large amount of computation and long processing time when dealing with large-scale data, making it difficult to meet the requirements of real-time detection. 3. Insufficient generalization ability: Existing models perform well in a specific domain, but have limited generalization ability in cross-domain detection and are difficult to adapt to changes between different domains.
[0062] The exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0063] Figure 1 A scenario diagram of a method for cross-domain object detection or a device for cross-domain object detection that can be applied to the embodiments of the present disclosure is shown.
[0064] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present disclosure.
[0065] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0066] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Among them, the terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, wearable devices, virtual reality devices, smart homes, etc.
[0067] The server 105 can be a server that provides various services, such as a background management server that supports the operations performed by the user using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as requests received, and feedback the processing results to the terminal devices.
[0068] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The present disclosure does not limit this.
[0069] The server 105 can, for example, obtain a training sample image and a domain label of the training sample image, where the domain label of the training sample image is the source domain or the target domain; the server 105 can, for example, perform feature extraction processing on the training sample image through a feature extraction network in the cross-domain target object detection model to obtain the image-level features of the training sample image; wherein, the cross-domain target object detection model further includes an image-level domain adaptation network, a target object detection network, and an instance-level domain adaptation network; the server 105 can, for example, perform image-level domain classification on the image-level features through the image-level domain adaptation network to determine the image-based domain classification result of the training sample image, and determine the first loss based on the domain label and the image-based domain classification result; the server 105 can, for example, perform object feature extraction on the image-level features through the target object detection network to obtain the instance-level features of the training sample image; the server 105 can, for example, perform feature extraction processing on the instance-level features through the instance-level domain adaptation network to obtain the object-based domain classification result of the training sample image, and determine the second loss based on the domain label and the object-based domain classification result; the server 105 can, for example, perform consistency loss calculation on the image-based domain classification result and the object-based domain classification result to determine the third loss; the server 105 can, for example, update the parameters of the cross-domain target object detection model based on the first loss, the second loss, and the third loss, so as to perform cross-domain target detection according to the cross-domain target object detection model with updated parameters.
[0070] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0071] Under the above system architecture, an embodiment of the present disclosure provides a method for cross-domain object detection, which can be executed by any electronic device with computing and processing capabilities.
[0072] This application aims to solve the deficiencies of existing cross-domain object detection methods in terms of inter-domain difference elimination, computational complexity, and generalization ability, and proposes a cross-domain object detection method based on multi-scale adversarial mining to improve the accuracy and efficiency of cross-domain detection.
[0073] Figure 2 is a flowchart of a method for cross-domain object detection shown according to an exemplary embodiment. The method provided by the embodiment of the present disclosure can be executed by any electronic device with computing and processing capabilities. For example, this method can be executed by the server or terminal device in the above Figure 1 embodiment, or can be jointly executed by the server and the terminal device. In the following embodiments, the server is taken as an example of the execution entity for illustration, but the present disclosure is not limited thereto.
[0074] Refer to Figure 2 , the method for cross-domain object detection provided by the embodiment of the present disclosure may include the following steps.
[0075] Step S202, obtain a training sample image and the domain label of the training sample image, where the domain label of the training sample image is the source domain or the target domain.
[0076] Among them, the training sample image may include an image of the source domain or an image of the target domain.
[0077] In some embodiments, during one training process, an image of the source domain or an image of the target domain can be input alone, or an image of the source domain and an image of the target domain can be input in pairs (as Figure 3 shown), and this application does not limit this.
[0078] During the training process, the training sample image may correspond to a domain label, and this domain label can be used to identify whether the training sample image comes from the source domain or the target domain.
[0079] Step S204, perform feature extraction processing on the training sample image through the feature extraction network in the cross-domain object detection model to obtain the image-level features of the training sample image. Among them, the cross-domain object detection model further includes an image-level domain adaptation network, an object detection network, and an instance-level domain adaptation network.
[0080] Instance-level features refer to features related to a specific instance or object. These features are typically used to describe the unique attributes or characteristics of that instance. In the context of image processing, instance-level features may include features such as shape, color, texture, etc. related to specific image objects (such as people, animals, buildings, etc.). These features help to distinguish different image objects and play an important role in tasks such as image retrieval, image classification, and object detection. Instance-level features are usually extracted from the original image data through feature extraction algorithms. These algorithms can identify and extract key information in the image, such as edges, corners, texture, etc., and convert this information into feature vectors that can be used for subsequent processing and analysis.
[0081] Image-level features refer to features related to the entire image. These features are typically used to describe the overall attributes or characteristics of the image, such as the brightness, contrast, color distribution, etc. of the image. In the context of image processing, image-level features may include global color histograms, global texture features, global shape features, etc. These features help to understand the overall content and structure of the image and play an important role in tasks such as image classification and scene recognition. Image-level features are usually extracted through global analysis of the entire image. Different from instance-level features, image-level features do not depend on specific objects or regions in the image, but rather perform an overall evaluation of the entire image.
[0082] In some embodiments, the above cross-domain object detection model may include a feature extraction network (such as Figure 3 the feature extraction network 301 shown), an image-level domain adaptation network (such as Figure 3 the image-level domain adaptation network 303 shown), an object detection network (such as Figure 3 the object detection network 302 shown), and an instance-level domain adaptation network (such as Figure 3 the instance-level domain adaptation network 304 shown).
[0083] Among them, the feature extraction network can be used to extract features from images. For example, it can be VGG, AlexNet, ResNet, etc., or it can also be a stack of ordinary convolutional layers. This application does not limit this.
[0084] In some embodiments, the image-level domain adaptation network can obtain an image-based domain classification result through the processing of image-level features.
[0085] In some embodiments, the instance-level domain adaptation network can obtain an instance (or object-based) domain classification result through the processing of instance-level features.
[0086] In some embodiments, the target object detection network can be used to identify the targets (or objects or target objects) in the training sample images, obtain the positions (e.g., the positions of the targets are boxed by bounding boxes) and / or categories of each target in the training sample images. The target object detection network can be, for example, the Faster R-CNN network or a part of the Faster R-CNN network, and can also be other networks capable of target recognition. This application does not limit this.
[0087] Step S206: Perform image-level domain classification on the image-level features through the image-level domain adaptation network, determine the image-based domain classification result of the training sample image, and determine the first loss based on the domain label and the image-based domain classification result.
[0088] The formula for the image-level domain adaptation loss can be expressed as the following formula: where D i represents the true domain label of the i-th image. When D i is equal to 0, the image comes from the source domain. When D i is equal to 1, the image comes from the target domain. p i u,v represents the predicted value at the u, v position on the feature map of the i-th image extracted by the extraction network. Among them, i, u, and v are all integers greater than or equal to 0.
[0089] Step S208: Extract object features from the image-level features through the target object detection network to obtain the instance-level features of the training sample image.
[0090] Step S210: Perform feature extraction processing on the instance-level features through the instance-level domain adaptation network, obtain the object-based domain classification result of the training sample image, and determine the second loss based on the domain label and the object-based domain classification result.
[0091] The instance-level features are feature vectors obtained from RoI (region of interest) pooling. Due to RoI pooling, the instance-level features are scale-independent because the features of all instances are mapped to the same size. Therefore, for instance-level domain adaptation, the loss calculation formula can be as follows: where p i,j represents the probability that the j-th region proposal in the i-th image comes from the target domain. Among them, i and j are integers greater than or equal to 1.
[0092] Step S212: Calculate the consistency loss between the image-based domain classification result and the object-based domain classification result to determine the third loss.
[0093] To further enhance consistency, this embodiment introduces a consistency regularizer. Since the multi-scale image-level domain classifier generates an output for each activation (each pixel point) of the image-level representation, the average value of all activations in the image is used as its image-level probability, and it is implemented as a global average pooling operation. The consistency loss corresponding to the consistency regularizer (i.e., the third loss) can be expressed as: where |I| represents the total number of activations in the feature map, and ‖·‖ 2 is the L2 norm.
[0094] Step S214: Based on the first loss, the second loss, and the third loss, update the parameters of the cross-domain object detection model so as to perform cross-domain object detection according to the cross-domain object detection model after parameter update.
[0095] In the above embodiment, the domain difference between the source domain and the target domain can be eliminated through the first loss between the domain label and the image-based domain classification result, the second loss between the domain label and the object-based domain classification result, and the third loss (consistency loss) between the image-based domain classification result and the object-based domain classification result, thereby improving the object detection accuracy of the object detection network.
[0096] Figure 4 is a flowchart of an image-level domain classification method shown according to an exemplary embodiment.
[0097] In some embodiments, the image-level domain adaptation network includes a multi-scale information aggregation sub-network, a first adversarial gradient reversal layer, and an image-level domain classifier.
[0098] Figure 5 The image-level domain adaptation network based on multi-scale adversarial mining shown can be the image-level domain adaptation network in this application.
[0099] Such as Figure 5 shown, the above image-level domain adaptation network 503 based on multi-scale adversarial mining can include a multi-scale information aggregation sub-network, a first adversarial gradient reversal layer (such as Figure 5 AdvGRL in 503) and an image-level domain classifier.
[0100] Among them, the specific network structure of the image-level domain adaptation network 503 based on multi-scale adversarial mining can refer to the network structure shown in Figure 6 shown.
[0101] Referring to Figure 4 , the above image-level domain classification method may include the following steps.
[0102] Step S402: Extract multi-scale information from the image-level features through the multi-scale information aggregation sub-network to obtain multi-scale information aggregation features.
[0103] In some embodiments, the image-level features can be subjected to feature extraction at multiple scales through a multi-scale information aggregation sub-network to obtain features at multiple different scales, and then the features at the multiple different scales are aggregated (such as concatenation, summation, or averaging processing) to obtain the multi-scale information aggregation feature.
[0104] As Figure 6 shown, the part enclosed by the left dashed box can be the multi-scale information aggregation sub-network, and the multi-scale information aggregation sub-network can include multiple convolutional layers, and the dilation rate of each convolutional layer can be different.
[0105] In some embodiments, the image-level features can be subjected to feature extraction through the above-mentioned multiple convolutional layers respectively, and then the features output by each convolutional layer can be of different scales.
[0106] In some embodiments, operations such as connection, convolution, activation, and gradient reversal can be performed on the features at multiple different scales to predict the domain classification result based on the image.
[0107] Step S404: Perform feature extraction processing on the multi-scale information aggregation feature through the first adversarial gradient reversal layer to obtain the multi-scale information aggregation depth feature.
[0108] Specifically, the features at multiple different scales can be subjected to a connection (Concat) operation, and then the connected features are subjected to feature extraction processing through the first adversarial gradient reversal layer to obtain the multi-scale information aggregation depth feature.
[0109] In some embodiments, an adversarial gradient reversal layer can be connected after the multi-scale information aggregation sub-network. Then, during forward propagation, the adversarial gradient reversal layer has no impact on the model parameters, but during backward parameter propagation, the gradient can be reversed after the adversarial gradient reversal layer, enabling the downstream parameters to learn in the opposite direction. This mechanism helps the model find commonalities between the source domain and the target domain while reducing the domain difference.
[0110] In some embodiments, during forward propagation, by performing feature extraction on the multi-scale information aggregation feature through the first adversarial gradient reversal layer, the depth information of the multi-scale information aggregation feature - the multi-scale information aggregation depth feature - can be obtained.
[0111] In some embodiments, when performing backward parameter update on the cross-domain target object detection model, the backward propagation directions of the parameters before and after the first adversarial gradient reversal layer are opposite.
[0112] Step S406: Perform image-level domain classification on the image-level features through an image-level domain classifier to determine the image-based domain classification result of the training sample image.
[0113] The above image-level domain classification enables classifiers such as sofmax to obtain domain classification results by classifying features.
[0114] The above method can obtain information at multiple scales in the image-level features through a multi-scale information aggregation sub-network, so as to improve the accuracy of the domain classification results.
[0115] Figure 7 It is a flowchart of a multi-scale information extraction method shown according to an exemplary embodiment.
[0116] In some embodiments, the multi-scale information aggregation sub-network may include multiple convolutional layers with different dilation rates (as shown in the left part of Figure 6 , the multi-scale information aggregation sub-network may include convolutional layers with dilation rates of 1, 2, and 4, etc.). It should be noted that to realize the flow of representative features, the dashed line is used to indicate the scale corresponding to the current convolutional layer.
[0117] Refer to Figure 7 , the above multi-scale information extraction method may include the following steps.
[0118] Step S702, perform feature extraction on the image-level features through multiple convolutional layers with different dilation rates to obtain feature maps with multiple different receptive fields.
[0119] In some embodiments, it can be through Figure 6 The convolutional layers with dilation rates of 1, 2, and 4 shown in perform feature extraction on the image-level features respectively to obtain features at multiple different scales, where different scales of features correspond to different receptive fields.
[0120] Step S704, perform information aggregation processing on the feature maps with multiple different receptive fields to obtain multi-scale information aggregation features.
[0121] As shown in Figure 6 , it is possible to perform information aggregation processing on the feature maps with multiple different receptive fields through a connection layer (Concat) to obtain multi-scale information aggregation features.
[0122] Through the above method, feature extraction can be performed on the image-level features from multiple scales to obtain image information of different receptive fields of the training sample images.
[0123] Figure 8 It is a flowchart of a gradient reversal method of an adversarial gradient reversal layer shown according to an exemplary embodiment.
[0124] In some embodiments, a Gradient Reversal Layer (GRL) is used for unsupervised domain adaptation in image classification tasks. Specifically, the input is kept unchanged during forward propagation, while the gradient is reversed by multiplying a negative scalar when propagating forward to the base network during training. A domain classifier is trained to maximize the probability of identifying the domain, while the base network is optimized to confuse the domain classifier. In this way, domain-invariant features are obtained to achieve domain adaptation. The forward propagation of GRL is defined as follows: R λ (v) = v, where v is the feature vector, and R λ is the forward propagation function executed by GRL. The backpropagation function of GRL is: where I is the identity matrix and -λ is a negative scalar.
[0125] Reference Figure 8 , the gradient reversal method of the adversarial gradient reversal layer provided in this application may include the following steps.
[0126] In some embodiments, the first adversarial gradient reversal layer may include a reversal degree parameter (such as λ in the above formula), and this reversal degree parameter can be used to weight the gradient of the propagation function in the first adversarial gradient reversal layer.
[0127] However, in the related art, λ of GRL is a fixed constant or a varying value, without considering the differences in the challenge levels of different training samples (i.e., the differences in domain classification difficulty). Therefore, this embodiment proposes a new adversarial gradient reversal layer (AdvGRL), which enhances the model's transfer learning ability under challenging samples by adversarially mining difficult example samples.
[0128] Step S802, when the first loss is greater than the preset threshold, the reversal degree parameter is equal to the preset constant.
[0129] Step S804, when the first loss is less than or equal to the preset threshold, the reversal degree parameter is inversely proportional to the first loss.
[0130] Specifically, as shown in formula (1), the above λ can be replaced with a new λadv, which is used to confuse the domain classifier during backpropagation. The calculation method of λadv is as follows:
[0131]
[0132] where L d is the loss of the domain classifier, for example, it can be the first loss in this application. α is a difficulty threshold used to determine whether the training sample is challenging, β is an overflow threshold to avoid generating too many gradients during backpropagation, and λ 0 = 1 is set as a fixed parameter in the experiment. In other words, if the loss L of the domain classifierd If it is smaller, the domain of the training samples can be more easily identified, and its features are not the domain-invariant features required by this application. Therefore, such training samples are more difficult examples for domain adaptation, and it is necessary to increase λadv to increase the parameter antagonism before and after the adversarial gradient reversal layer, so as to better find the domain-invariant features. Domain-Invariant Features is an important concept in the fields of deep learning and machine learning. It refers to the features that exist in different domains or datasets and have stability and consistency. These features are crucial for the generalization ability of the model and cross-domain transfer learning.
[0133] λ adv and L d The relationship between them is as Figure 9 shown. Here, an example of the proposed AdvGRL for adversarial mining of difficult training samples is given. In this example, λ 0 is set to 1, β = 30, and it is illustrated that more difficult training examples with a lower domain classifier loss L d will have a greater response.
[0134] Figure 10 is a flowchart of a method for obtaining a domain classification result based on an object according to an exemplary embodiment.
[0135] In some embodiments, the instance-level domain adaptation network may include a second adversarial gradient reversal layer and an instance-level domain classifier.
[0136] Referring to Figure 10 above, the above-mentioned domain classification result based on an object may include the following steps.
[0137] Step S1002, perform feature extraction processing on the instance-level features through the second adversarial gradient reversal layer to obtain instance-level depth features; wherein, when updating the parameters of the cross-domain object detection model, the backpropagation directions of the parameters before and after the second adversarial gradient reversal layer are opposite.
[0138] In some embodiments, both the above-mentioned first adversarial gradient reversal layer and the second adversarial gradient reversal layer may be the improved gradient reversal layer in this application, or an ordinary gradient reversal layer, and this application does not limit this.
[0139] Step S1004, perform classification processing on the instance-level depth features through the instance-level domain classifier to obtain a domain classification result based on an object.
[0140] Through the above method, the domain invariance between the source domain and the target domain can be better mined.
[0141] Figure 11It is a flowchart of an instance-level feature extraction method shown according to an exemplary embodiment.
[0142] In some embodiments, the above cross-domain object detection model may include a feature extraction network (such as Figure 5 the feature extraction network 501 shown), an image-level domain adaptation network (such as Figure 5 the image-level domain adaptation network 503 based on multi-scale adversarial mining shown), an object detection network (such as Figure 5 the object detection network 502 shown), and an instance-level domain adaptation network (such as Figure 5 the instance-level domain adaptation network 504 based on adversarial mining shown).
[0143] In some embodiments, the above object detection network may be a Fast R-CNN network, or a part of the Fast R-CNN network (such as Figure 5 the part of the Fast R-CNN network shown except for the feature extraction network 501), or other network structures capable of object recognition, and the present application does not limit this.
[0144] In some embodiments, the above object detection network may include a region proposal sub-network (such as Figure 5 the region proposal sub-network RPN 5021 shown) and a region of interest pooling sub-network (such as Figure 5 the region of interest pooling sub-network ROI Pooling 5022 shown).
[0145] Among them, the region proposal sub-network can be used to generate region bounding boxes in the training sample image, and each region bounding box may include the object to be recognized.
[0146] Referring to Figure 11 , the above instance-level feature extraction method may include the following steps.
[0147] Step S1102, perform region generation processing on the image-level features through the region proposal sub-network to obtain a plurality of region bounding boxes; wherein, the region bounding boxes are used to frame the position of the target object.
[0148] In some embodiments, the region proposal sub-network can be used to process the image-level features of the training sample image to generate a plurality of region bounding boxes, and the region bounding boxes can be used to frame the position of the target object (or the target) in the training sample image.
[0149] Step S1104, perform pooling processing on the plurality of region bounding boxes through the region of interest pooling sub-network to obtain instance-level features.
[0150] In some embodiments, the above-mentioned multiple region editing boxes may be pooled through a region of interest pooling sub-network to obtain instance-level features corresponding to each object.
[0151] Through the above method, the instance-level features corresponding to the targets in the training sample images can be obtained efficiently and accurately.
[0152] Figure 12 FIG. is a flowchart of a method for cross-domain object detection according to an exemplary embodiment.
[0153] In some embodiments, if the training sample image has object bounding boxes and object labels in each object bounding box, then the above-mentioned region bounding box acquisition method may include the following steps.
[0154] Step S1202: Generate regions from the image-level features through a region generation sub-network to obtain a plurality of region bounding boxes and the object existence probabilities in each region bounding box.
[0155] Among them, the object existence probability can be used to determine the probability that each bounding box accurately encloses the target object. The higher the object existence probability, the greater the possibility that the corresponding bounding box encloses the target object; the lower the object existence probability, the lower the possibility that the corresponding bounding box encloses the target object.
[0156] Step S1204: Input the instance-level features into the class classifier and bounding box regressor of the target object detection network to obtain the bounding box regression values of each region bounding box and the target recognition results in each region bounding box, where the bounding box regression value is used to measure the error between the region bounding box and the object bounding box.
[0157] In some embodiments, the above-mentioned bounding box regression value can be used to confirm the offset between the region bounding box generated by the region generation sub-network and the object bounding box in the training sample image.
[0158] Step S1206: Determine the fourth loss according to the object labels in each object bounding box and the object existence probabilities in each region bounding box.
[0159] In some embodiments, the fourth loss between the object labels in each object bounding box of each training sample image and the object existence probabilities in each region bounding box can be calculated.
[0160] Step S1208: Determine the fifth loss according to the object labels in each object bounding box and the target object recognition results in each region bounding box.
[0161] Step S1211: Determine the sixth loss according to the object bounding box, each region bounding box, and the bounding box regression value.
[0162] Step S1213, update the parameters of the cross-domain object detection model based on the first loss, the second loss, the third loss, the fourth loss, the fifth loss, and the sixth loss.
[0163] The above method can not only mine the domain-invariant features between the source domain images and the target domain images through the first loss, the second loss, and the third loss, but also accurately extract the target object and its corresponding category from the training sample images through the fourth loss, the fifth loss, and the sixth loss.
[0164] Figure 13 It is a flowchart of a cross-domain object detection method shown according to an exemplary embodiment.
[0165] Refer to Figure 13 , the above cross-domain object detection method may include the following steps.
[0166] Step S1302, obtain the image to be predicted.
[0167] The above image to be predicted may be a source domain image or a target domain image, and the present application does not limit this.
[0168] Step S1304, perform feature extraction processing on the image to be predicted through the feature extraction network in the cross-domain object detection model to obtain the image-level features of the image to be predicted.
[0169] Step S1306, perform region feature extraction on the image-level features of the image to be predicted through the object detection network to obtain the instance-level features of the image to be predicted and multiple region bounding boxes.
[0170] Step S1308, input the instance-level features of the image to be predicted into the class classifier and the bounding box regressor of the object detection network to obtain the bounding box regression values of each region bounding box in the image to be predicted and the target recognition results in each region bounding box.
[0171] In some embodiments, the object bounding box can be accurately determined according to the above multiple region bounding boxes and the bounding box regression values. Through the object bounding box, the position of the target object can be framed in the image to be predicted. Through the above target recognition results, it can be accurately known the recognition result of the target object framed by the object bounding box.
[0172] The present application also proposes a method for cross-domain object recognition.
[0173] This embodiment proposes a cross-domain object detection system and method based on multi-scale adversarial mining. As Figure 5 shown, the algorithm may include: an image-level domain adaptation module based on multi-scale adversarial mining and an instance-level domain adaptation module based on adversarial mining.
[0174] Figure 5 The overall framework of the algorithm is shown. The image-level domain adaptation module based on multi-scale adversarial mining obtains information at different scales through multi-scale information aggregation, and uses the adversarial gradient reversal layer (AdvGRL) to mine samples that are difficult to domain-confuse at the image level. The instance-level domain adaptation module based on adversarial mining obtains features with scale information through RoI pooling (RoIPooling), and uses AdvGRL to mine samples that are difficult to domain-confuse at the instance level. The entire network performs domain adaptation adversarial feature training at the image level and instance level respectively during the training phase, learns domain-invariant features, and makes the RPN (Region Proposal Network) have domain-invariant features through consistency regularization, thereby improving the generalization of object detection.
[0175] The following separately introduces the image-level domain adaptation network based on multi-scale adversarial mining, the adversarial gradient reversal layer, and the instance-level domain adaptation network based on adversarial mining proposed in this application.
[0176] Figure 6 is an image-level domain adaptation network based on multi-scale adversarial mining shown according to an exemplary embodiment.
[0177] As Figure 6 shown, the image-level domain adaptation network of multi-scale adversarial mining first performs dilated convolutions with multiple different dilation rates on the feature map of each input image to express multiple scale information. Then, channel communication and prediction of whether the feature comes from the source domain or the target domain are respectively performed through channel aggregation and convolution and the adversarial gradient reversal layer (AdvGRL).
[0178] As Figure 5 shown, the image-level domain adaptation network based on multi-scale adversarial mining may include a multi-scale information aggregation module and an adversarial gradient reversal layer AdvGRL.
[0179] (1) Multi-scale information aggregation module.
[0180] The multi-scale information aggregation module realizes multi-scale information aggregation through multiple dilated convolutions with different dilation rates. As Figure 6As shown on the left side, this module receives C feature maps as input, generates C feature maps as output, and does not define a loss function. Each layer in this module has C channels, which can be used to obtain multi-scale information of the feature maps without normalization. By learning context information, this module can easily fuse information at different scales, obtain feature representations with different receptive field sizes, and thus improve accuracy. This module consists of three layers, and each layer applies a 3×3 convolutional kernel with different dilation rates, which are 1, 2, and 4 respectively. The C feature maps obtained by the backbone network go through three different dilated convolutions respectively, and then follow a ReLU activation function to truncate negative values. Aggregation is performed in the channel dimension to obtain 3×C feature maps. The output of the dilated convolution is the same size as its input, so information at these different scales can be easily fused through the concatenation operator. Such feature representations have different scale information and can better capture the context information and global information of the target object. In addition, this module can also improve the robustness of the model because information at different scales can complement each other, thus reducing the dependence of the model on specific scale information.
[0181] (2) Adversarial Gradient Reversal Layer (AdvGRL).
[0182] The original Gradient Reversal Layer (GRL) is used for unsupervised domain adaptation in image classification tasks. Specifically, during forward propagation, the input remains unchanged, while during training, when propagating forward to the base network, the gradient is reversed by multiplying a negative scalar. A domain classifier is trained to maximize the probability of identifying the domain, while the base network is optimized to confuse the domain classifier. In this way, domain-invariant features are obtained to achieve domain adaptation. The forward propagation of GRL is defined as follows: R λ (v) = v, where v is the feature vector, and R λ is the forward propagation function executed by GRL. The backpropagation function of GRL is: where I is the identity matrix and -λ is a negative scalar.
[0183] (dR_λ) / dv = -λI in the backpropagation function means that during backpropagation, the gradient is multiplied by -λ. Here, I usually represents the identity matrix, meaning that the gradient is multiplied by -λ in each dimension, thus achieving the reversal and scaling of the gradient.
[0184] The specific value of λ usually needs to be determined through experiments, and it depends on the specific task and dataset. In practical applications, methods such as grid search, random search, or Bayesian optimization can be used to find the optimal value of λ. It should be noted that the choice of λ has a great impact on the performance of the model, so it should be carefully adjusted during model training.
[0185] In addition, λ is not a constant and can be adjusted during training to better adapt to data changes and the training requirements of the model. This adjustment can be achieved through methods such as learning rate decay and dynamic adjustment strategies.
[0186] In summary, λ is a key hyperparameter in GRL, which determines the degree of gradient inversion during backpropagation and has a great impact on the performance of the model. In practical applications, the value of λ needs to be carefully selected and adjusted to optimize the model performance.
[0187] However, the λ of the original GRL is a fixed constant or a varying value, without considering the differences in the challenge levels of different training samples. Therefore, this embodiment proposes a new adversarial gradient reversal layer (AdvGRL) to enhance the transfer learning ability of the model under challenging samples by adversarially mining hard example samples. Specifically, λ is replaced with a new λadv for confusing the domain classifier during backpropagation. The calculation method of λadv is as follows: where L d is the loss of the domain classifier. It is easier to handle when the loss is relatively large. Different samples need to be considered, as the difficulty of distinguishing the source domain and the target domain is different. If it is greater than a threshold, a constant is taken. If it is relatively small, taking a fixed constant will be ignored. Setting a slightly larger number like this is equivalent to amplification. α is a difficulty threshold used to determine whether a training sample is challenging, and β is an overflow threshold to avoid generating too many gradients during backpropagation, while λ 0 = 1 is set as a fixed parameter in the experiment. In other words, if the loss L d of the domain classifier is smaller, the domain of the training sample can be more easily identified, and its features are not the desired domain-invariant features. Therefore, such training samples are more difficult examples for domain adaptation. Domain-Invariant Features is an important concept in the fields of deep learning and machine learning. It refers to the features that exist in different domains or datasets and have stability and consistency. These features are crucial for the generalization ability of the model and cross-domain transfer learning.
[0188] The relationship between λ adv and L d is as Figure 9 shown. Here, an example of the proposed AdvGRL for adversarially mining hard training samples is given. In this example, λ 0 = 1 and β = 30 are set, and it is shown that harder training examples with a lower domain classifier loss L d will have a greater response.
[0189] (3) Loss function of the image-level domain adaptation network based on multi-scale adversarial mining.
[0190] Align the domain distributions through a min-max game: the domain discriminator minimizes the adversarial loss to distinguish the domain origin of the features, while the backbone network maximizes the loss to confuse the domain discriminator.
[0191] The designs of the multi-scale information aggregation module and the adversarial gradient reversal layer (AdvGRL) have been introduced successively before. In this embodiment, these two modules are used on the image-level domain classifier. For images, such as drone images, the scale changes of the objects captured from the top-down perspective are quite different. Although instance-level domain adaptation can solve this problem through RoI pooling, it is also meaningful to overcome this problem at the image level.
[0192] Image-level features refer to the feature map outputs of the shared backbone network, that is, Figure 5 the feature map represented by scale one in Figure 6 As shown in where D i represents the true domain label of the i-th image. When D i is equal to 0, the image comes from the source domain. When D i is equal to 1, the image comes from the target domain. p i u,v represents the predicted value at the u, v position on the feature map of the i-th image extracted by the feature extraction network of the image-level domain classification network.
[0193] In this embodiment, a min-max game is used to align the domain distributions: the domain discriminator distinguishes which domain the features belong to by minimizing the above-mentioned adversarial loss, while the backbone network confuses the domain discriminator by maximizing the loss. To conduct this min-max game, an adversarial gradient reversal layer is used to reverse the sign of the gradient when passing through the adversarial gradient reversal layer during the backpropagation process and simultaneously mine samples that are difficult to be domain-confused.
[0194] Next, the instance-level domain adaptation network based on adversarial mining will be introduced.
[0195] The instance-level features are the feature vectors obtained from RoI pooling. Due to RoI pooling, the instance-level features are scale-invariant because the features of all instances are mapped to the same size. Therefore, for instance-level domain adaptation, when reversing the gradients, an adversarial gradient reversal layer (AdvGRL) is used, and the loss calculation formula is as follows: where p i,j represents the probability that the j-th region proposal in the i-th image comes from the target domain. Similar to image-level adaptation, an AdvGRL is added before the domain classifier to apply the adversarial mining training strategy.
[0196] In addition, the present application also uses a consistency regularizer between the image-level domain adaptation network and the instance-level domain adaptation network.
[0197] In this embodiment, by enhancing the consistency between domain classifiers at different levels, it can help to learn a bounding box predictor with cross-domain robustness (i.e., the RPN in the Faster R-CNN model). Therefore, to further enhance this consistency, a consistency regularizer is introduced. Since the multi-scale image-level domain classifier generates an output for each activation (each pixel point) of the image-level representation, the average value of all activations in the image is used as its image-level probability, and it is implemented as a global average pooling operation. The loss of the consistency regularizer can be expressed as: where |I| represents the total number of activations in the feature map, and ‖·‖ 2 is the L2 norm.
[0198] Next, the optimization objectives of the cross-domain object detection model proposed in this embodiment are introduced.
[0199] (1) Faster R-CNN object detection network
[0200] In this embodiment, the widely used and powerful Faster R-CNN model can be adopted as the basic detector. Faster R-CNN is a two-stage detector, mainly composed of three parts: (1) a feature extractor that uses a convolutional neural network to extract features from the input image; (2) a region proposal network (RPN) that is responsible for providing candidate region proposal boxes; (3) a region of interest (RoI) head that is used to classify and regressively fine-tune the regions proposed by the RPN. The total loss function of Faster R-CNN can be defined as: where and are the loss functions of the RPN, the RoI-based regressor, and the classifier, respectively.
[0201] (2) Total training objective
[0202] The algorithm network framework of this embodiment is as Figure 5 shown. It can be based on the Faster R-CNN architecture, with domain adaptation components added, thus achieving domain-adaptive object detection. In Figure 5 , the original Faster R-CNN model is shown at the top. This model includes a shared feature extraction convolutional layer, an RPN layer, a RoI pooling layer, and two fully connected layers for extracting instance-level features.
[0203] This embodiment's algorithm introduces three novel components to achieve domain adaptation. First, a multi-scale adversarial mining image-level domain classifier is introduced, which is added after the last convolutional layer of the backbone network. Second, an adversarial mining instance-level domain classifier is introduced, which is added at the end of the RoI pooling features. These two classifiers are connected by consistency regularization to make the RPN domain-invariant. Finally, the training loss of this embodiment's algorithm network can be expressed as the weighted sum of the losses of each part, that is: where η is a parameter that balances Faster R-CNN and the newly added domain adaptation components.
[0204] This embodiment can use the standard SGD algorithm for end-to-end training. For the domain adaptation components, adversarial training is implemented using AdvGRL, and this layer can automatically reverse the gradient and automatically mine hard samples. During the training phase, the entire network is used for training. During the inference phase, the domain adaptation components can be directly removed, and only the original Faster R-CNN architecture with domain adaptation weights (what kind of weights are these) is used for object detection.
[0205] Next, this application will introduce the above method for cross-domain object detection in combination with specific application scenarios.
[0206] This embodiment details the specific implementation steps of a cross-domain object detection algorithm based on multi-scale adversarial mining, including backpropagation during the training process and the inference process.
[0207] Training process.
[0208] Step 1: Input image pairs.
[0209] Input dataset image pairs from the source domain and the target domain and
[0210] Step 2: Feature extraction.
[0211] Input and Extract image features through a feature extraction network (such as VGG, AlexNet, ResNet, etc.) to generate a feature map and
[0212] Step 3: Image-level domain adaptation.
[0213] For the feature maps generated in Step 2 and Calculate through an image-level domain adaptation network based on multi-scale adversarial mining.
[0214] Step 4: Generate candidate regions.
[0215] Input the feature maps generated in Step 2 and into the Region Proposal Network (RPN) to generate candidate regions of interest and calculate the RPN loss
[0216] Step 5: RoI pooling.
[0217] Input the results of the candidate regions of interest generated in Step 4 and the feature maps generated in Step 2 and into the RoI pooling module to generate feature maps of regions of interest with a fixed size and
[0218] Step 6: Process through the fully connected layer.
[0219] Input the feature maps of the regions of interest in Step 5 and into the fully connected layer and connect them to the class classifier and the bounding box regressor to calculate the regression loss and the classification loss.
[0220] Step 7: Instance-level domain adaptation.
[0221] Input the feature maps of the regions of interest in Step 5 and into the instance-level domain adaptation network based on adversarial mining to calculate the instance-level domain adaptation loss
[0222] Step 8: Consistency regularization.
[0223] Calculate the consistency regularization loss of the image-level domain adaptation network based on multi-scale adversarial mining and the instance-level domain adaptation network based on adversarial mining
[0224] Step 10: Calculate the total loss function.
[0225] Calculate the total loss function
[0226] Step 11: Backpropagation.
[0227] During backpropagation, the adversarial gradient reversal layer (AdvGRL) is used to automatically reverse the gradient for adversarial mining of indistinguishable samples.
[0228] Step 12: End-to-end training.
[0229] Use the standard stochastic gradient descent (SGD) algorithm for end-to-end training to update the network weights.
[0230] Inference process.
[0231] Step 13: Remove the domain adaptation component.
[0232] In the inference stage, remove the domain adaptation component and only retain the original Faster R-CNN architecture with domain adaptation weights for object detection.
[0233] Step 14: Feature extraction.
[0234] For each input image x, generate a feature map F through a pre-trained feature extraction network.
[0235] Step 15: Generate candidate regions.
[0236] Input the generated feature map F into the region proposal network (RPN) to generate candidate regions of interest.
[0237] Step 15: RoI pooling.
[0238] Perform RoI pooling on the candidate regions to generate a fixed-size feature map R of regions of interest.
[0239] Step 16: Process the fully connected layer.
[0240] Input the feature map R of regions of interest into the fully connected layer for class classification and bounding box regression, and calculate the classification loss and the classification loss
[0241] Step 17: Output the detection results.
[0242] Output the final object detection results, including the object class and the bounding box location.
[0243] The above-mentioned Embodiment 1 proposes a cross-domain object detection algorithm based on multi-scale adversarial mining, including image-level and instance-level domain adaptation; 2. A new adversarial gradient reversal layer (AdvGRL) is proposed for adversarial mining, which solves the problem of enhancing the transfer learning ability of the model under challenging samples. Specifically, AdvGRL can use the negative gradient in backpropagation to confuse the domain classifier, thereby generating domain-invariant features. In addition, AdvGRL can perform adversarial mining on difficult examples to further enhance the generalization ability of the model under challenging examples; 3. A multi-scale information aggregation module is proposed, which adopts multi-channel dilated convolution, and captures context information from different scales by using multiple dilated convolution layers with different dilation rates, so as to perform image-level domain adaptation to align the source domain and the target domain at different scales and reduce the image-level domain difference; 4. In this embodiment, without the need to label any target domain data, only the source domain data with label information and the target domain data without label information can be used to optimize the object detection performance of the target domain image with obvious domain differences.
[0244] This embodiment proposes a cross-domain object detection algorithm based on multi-scale adversarial mining. By combining image-level and instance-level domain adaptation, and introducing multi-scale information aggregation and adversarial gradient reversal layer (AdvGRL), the object detection performance of the target domain image is optimized. Without the need to label any target domain data, this embodiment realizes the effective optimization of the target domain data by the source domain data.
[0245] Compared with the prior art, the main advantages of the above-mentioned embodiment are as follows.
[0246] 1. Enhanced domain adaptation ability: The multi-scale adversarial mining method proposed by the present invention effectively reduces the domain difference between the source domain and the target domain by combining image-level and instance-level domain adaptation. This multi-level domain adaptation method significantly improves the performance of cross-domain object detection. Especially when the target domain data has large variations, it can still maintain high-precision detection results.
[0247] 2. Application of the adversarial gradient reversal layer (AdvGRL): AdvGRL is introduced for adversarial mining, and the domain classifier is confused by the negative gradient in backpropagation to generate domain-invariant features. AdvGRL not only improves the effect of domain adaptation, but also can perform adversarial mining on difficult examples to enhance the generalization ability of the model under challenging samples. This method significantly improves the transfer learning ability of the model, enabling it to maintain stable detection performance in a wider range of application scenarios.
[0248] 3. Innovation of the multi-scale information aggregation module: By adopting multi-channel dilated convolutions, context information from different scales is captured, further enhancing the effect of image-level domain adaptation. This method can align the source domain and the target domain at different scales, reducing the image-level domain difference, thereby improving the accuracy of cross-domain object detection.
[0249] 4. No need to label target domain data: In this embodiment, without the need to label any target domain data, only the source domain data with label information and the target domain data without label information are used to optimize the object detection performance of the target domain images with obvious domain differences. This greatly reduces the cost and complexity of data annotation, making this method more operable and economical in practical applications.
[0250] Compared with the prior art, the main creativity of the above embodiments lies in the following aspects.
[0251] 1. Proposal and application of the Adversarial Gradient Reversal Layer (AdvGRL): Compared with the prior art, the Adversarial Gradient Reversal Layer (AdvGRL) is proposed and applied for the first time to conduct adversarial mining. AdvGRL uses the negative gradient in backpropagation to confuse the domain classifier and generate domain-invariant features, thereby improving the domain adaptation ability of the model. In addition, AdvGRL can conduct adversarial mining for difficult-to-distinguish samples, further enhancing the generalization ability of the model under challenging examples. This innovative method significantly improves the effect and application scope of cross-domain object detection.
[0252] 2. Innovative design of the multi-scale information aggregation module: A new multi-scale information aggregation module is proposed, which captures context information from different scales by using multiple dilated convolutional layers with different dilation rates. Compared with the single-scale information processing method in the prior art, the multi-scale information aggregation module can more effectively align the image-level domain features of the source domain and the target domain, thereby reducing the domain difference and improving the accuracy and robustness of cross-domain object detection.
[0253] 3. Comprehensive consideration of image-level and instance-level domain adaptation: The present invention comprehensively considers image-level and instance-level domain adaptation in the algorithm design and proposes a complete multi-scale adversarial mining framework. This comprehensive domain adaptation method can more comprehensively reduce the domain difference between the source domain and the target domain compared with the single-level domain adaptation in the prior art, thereby significantly improving the overall performance of cross-domain object detection.
[0254] 4. Optimize the detection performance of unlabeled target domain data: Compared with the prior art, the present invention can optimize the target detection performance of target domain images with obvious domain differences only relying on the source domain data with label information and the target domain data without label information without the need to label any target domain data. This innovative design significantly reduces the cost and difficulty of data annotation, making this method more advantageous and operable in practical applications.
[0255] The embodiments proposed in this application can optimize the target detection performance of target domain images with obvious domain differences and lack of sample labels using source domain data without the need to label any target domain data. Considering the problems of scale changes and different sample distribution differences alignment difficulties in cross-domain target detection of images, this embodiment proposes a domain adaptive target detection method based on multi-scale adversarial mining. This method creatively realizes multi-scale information aggregation and difficult sample mining for features that are difficult to align on the basis of the domain adaptive target detection method based on adversarial feature learning. In this regard, the present invention respectively adopts dilated convolutions with different dilation rates and adversarial gradient reversal layers. Specifically, multi-scale information aggregation fuses feature maps of different scales by using dilated convolutions with different dilation rates. Difficult example sample mining uses an adversarial gradient reversal layer to improve the robustness to difficult samples that are difficult to align. Finally, multi-scale information aggregation and the adversarial gradient reversal layer are applied to the image-level domain adaptive network and the instance-level domain adaptive network, and consistency regularization is added to improve the performance of target detection.
[0256] It should be particularly noted that each step in each of the above embodiments of the method for cross-domain target detection can be cross, replaced, added, or deleted from each other. Therefore, these reasonable permutation and combination transformations for the method for cross-domain target detection should also belong to the protection scope of the present disclosure, and the protection scope of the present disclosure should not be limited to the embodiments.
[0257] Based on the same inventive concept, an apparatus for cross-domain target detection is also provided in the embodiments of the present disclosure, as shown in the following embodiments. Since the principle of solving problems in this apparatus embodiment is similar to that of the above method embodiment, the implementation of this apparatus embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be described again.
[0258] Figure 14 is a block diagram of an apparatus for cross-domain target detection shown according to an exemplary embodiment. Referring to Figure 14 , the apparatus 1400 for cross-domain target detection provided in the embodiments of the present disclosure may include: a sample acquisition module 1401, a feature extraction module 1402, a first loss determination module 1403, an instance-level feature acquisition module 1404, a second loss determination module 1405, a third loss determination module 1406, and a parameter update module 1407.
[0259] Among them, the sample acquisition module 1401 can be used to acquire training sample images and domain labels of the training sample images, where the domain labels of the training sample images are the source domain or the target domain; the feature extraction module 1402 can be used to perform feature extraction processing on the training sample images through the feature extraction network in the cross-domain object detection model to obtain the image-level features of the training sample images; among them, the cross-domain object detection model further includes an image-level domain adaptation network, an object detection network, and an instance-level domain adaptation network; the first loss determination module 1403 can be used to perform image-level domain classification on the image-level features through the image-level domain adaptation network, determine the image-based domain classification result of the training sample images, and determine the first loss based on the domain label and the image-based domain classification result; the instance-level feature acquisition module 1404 can be used to perform object feature extraction on the image-level features through the object detection network to obtain the instance-level features of the training sample images; the second loss determination module 1405 can be used to perform feature extraction processing on the instance-level features through the instance-level domain adaptation network to obtain the object-based domain classification result of the training sample images, and determine the second loss based on the domain label and the object-based domain classification result; the third loss determination module 1406 can be used to calculate the consistency loss between the image-based domain classification result and the object-based domain classification result to determine the third loss; the parameter update module 1407 can be used to update the parameters of the cross-domain object detection model based on the first loss, the second loss, and the third loss, so as to perform cross-domain object detection according to the cross-domain object detection model with updated parameters.
[0260] It should be noted here that the above sample acquisition module 1401, feature extraction module 1402, first loss determination module 1403, instance-level feature acquisition module 1404, second loss determination module 1405, third loss determination module 1406, and parameter update module 1407 correspond to S202 - S215 in the method embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above method embodiment. It should be noted that the above modules, as part of the device, can be executed in a computer system such as a set of computer-executable instructions.
[0261] In some embodiments, the image-level domain adaptation network includes a multi-scale information aggregation sub-network, a first adversarial gradient reversal layer, and an image-level domain classifier; among them, the first loss determination module 1403 can include: a multi-scale information extraction sub-module, a first gradient reversal sub-module, and a first domain classification sub-module.
[0262] Among them, the multi-scale information extraction sub-module can be used to perform multi-scale information extraction on the image-level features through the multi-scale information aggregation sub-network to obtain multi-scale information aggregation features; the first gradient reversal sub-module can be used to perform feature extraction processing on the multi-scale information aggregation features through the first adversarial gradient reversal layer to obtain multi-scale information aggregation depth features; the first domain classification sub-module can be used to perform image-level domain classification on the image-level features through the image-level domain classifier to determine the image-based domain classification result of the training sample image.
[0263] In some embodiments, when updating the parameters of the cross-domain target object detection model, the backpropagation directions of the parameters before and after the first adversarial gradient reversal layer are opposite.
[0264] In some embodiments, the first adversarial gradient reversal layer includes a reversal degree parameter, and the reversal degree parameter is used to weight the propagation function gradient in the first adversarial gradient reversal layer; when the first loss is greater than a preset threshold, the reversal degree parameter is equal to a preset constant; when the first loss is less than or equal to the preset threshold, the reversal degree parameter is inversely proportional to the first loss.
[0265] In some embodiments, the multi-scale information aggregation sub-network includes a plurality of convolutional layers with different dilation rates; among them, the multi-scale information extraction sub-module may include: a feature map extraction unit and an aggregation unit.
[0266] Among them, the feature map extraction unit can be used to perform feature extraction on the image-level features through a plurality of convolutional layers with different dilation rates to obtain feature maps with multiple different receptive fields; the aggregation unit can be used to perform information aggregation processing on the feature maps with multiple different receptive fields to obtain multi-scale information aggregation features.
[0267] In some embodiments, the instance-level domain adaptation network includes a second adversarial gradient reversal layer and an instance-level domain classifier, where the second loss determination module 1405 may include: an instance-level depth feature acquisition sub-module and a second domain classification result acquisition sub-module.
[0268] Among them, the instance-level depth feature acquisition sub-module can be used to perform feature extraction processing on the instance-level features through the second adversarial gradient reversal layer to obtain instance-level depth features; among them, when updating the parameters of the cross-domain target object detection model, the backpropagation directions of the parameters before and after the second adversarial gradient reversal layer are opposite; the second domain classification result acquisition sub-module can be used to perform classification processing on the instance-level depth features through the instance-level domain classifier to obtain the object-based domain classification result.
[0269] In some embodiments, the target object detection network includes a region generation sub-network and a region of interest pooling sub-network; among them, the instance-level feature acquisition module 1404 may include: a region bounding box acquisition sub-module and a pooling sub-module.
[0270] Among them, the region bounding box obtaining sub-module can be used to perform region generation processing on the image-level features through the region generation sub-network to obtain a plurality of region bounding boxes; wherein, the region bounding boxes are used to box the positions of the target objects; the pooling sub-module can be used to perform pooling processing on the plurality of region bounding boxes through the region of interest pooling sub-network to obtain instance-level features.
[0271] In some embodiments, if the training sample image has object bounding boxes and object labels in each object bounding box; among them, the region bounding box obtaining sub-module may include: an object existence probability determining unit.
[0272] Among them, the object existence probability determining unit can be used to perform region generation on the image-level features through the region generation sub-network to obtain a plurality of region bounding boxes and the object existence probability in each region bounding box.
[0273] The device 1400 for cross-domain object detection may further include: a target recognition result determining module, a fourth loss determining module, a fifth loss determining module, and a sixth loss determining module.
[0274] Among them, the target recognition result determining module can be used to input the instance-level features into the class classifier and the bounding box regressor of the target object detection network to obtain the bounding box regression values of each region bounding box and the target recognition results in each region bounding box, where the bounding box regression values are used to measure the error between the region bounding box and the object bounding box; the fourth loss determining module can be used to determine the fourth loss according to the object labels in each object bounding box and the object existence probability in each region bounding box; the fifth loss determining module can be used to determine the fifth loss according to the object labels in each object bounding box and the target object recognition results in each region bounding box; the sixth loss determining module can be used to determine the sixth loss according to the object bounding box, each region bounding box, and the bounding box regression values.
[0275] In some embodiments, the parameter update module includes: a parameter update sub-module.
[0276] Among them, the parameter update sub-module can be used to update the parameters of the cross-domain target object detection model based on the first loss, the second loss, the third loss, the fourth loss, the fifth loss, and the sixth loss.
[0277] In some embodiments, the device 1400 for cross-domain object detection may further include: a to-be-predicted image obtaining module, an image-level feature obtaining module of the to-be-predicted image, a bounding box determining module of the to-be-predicted image, and a target recognition result determining module.
[0278] Among them, the to-be-predicted image acquisition module can be used to acquire the to-be-predicted image; the image-level feature acquisition module of the to-be-predicted image can be used to perform feature extraction processing on the to-be-predicted image through the feature extraction network in the cross-domain object detection model to acquire the image-level features of the to-be-predicted image; the bounding box determination module of the to-be-predicted image can be used to perform regional feature extraction on the image-level features of the to-be-predicted image through the object detection network to acquire the instance-level features of the to-be-predicted image and multiple regional bounding boxes; the target recognition result determination module can be used to input the instance-level features of the to-be-predicted image into the class classifier and bounding box regressor of the object detection network to obtain the bounding box regression values of each regional bounding box in the to-be-predicted image and the target recognition results in each regional bounding box.
[0279] Since the functions of apparatus 1400 have been described in detail in their corresponding method embodiments, the present disclosure will not repeat them here.
[0280] The modules and / or sub-modules and / or units involved in the embodiments of the present disclosure can be implemented in software or in hardware. The described modules and / or sub-modules and / or units can also be provided in a processor. Among them, the names of these modules and / or sub-modules and / or units do not constitute a limitation to the modules and / or sub-modules and / or units themselves in some cases.
[0281] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a part of a module or program segment, and the part of the above module or program segment includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer program instructions.
[0282] In addition, the above accompanying drawings are only schematic illustrations of the processes included in the methods according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above accompanying drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.
[0283] Figure 15The figure shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. It should be noted that Figure 15 The illustrated electronic device 1500 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0284] As Figure 15 shown, the electronic device 1500 includes a central processing unit (CPU) 1501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1502 or the program loaded from the storage section 1508 into the random access memory (RAM) 1503. In the RAM 1503, various programs and data required for the operation of the electronic device 1500 are also stored. The CPU 1501, ROM 1502, and RAM 1503 are connected to each other via a bus 1504. The input / output (I / O) interface 1505 is also connected to the bus 1504.
[0285] The following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, etc.; an output section 1507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 15010 including a network interface card such as a LAN card, a modem, etc. The communication section 15010 performs communication processing via a network such as the Internet. A drive 1511 is also connected to the I / O interface 1505 as required. A removable medium 1512, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1511 as required so that a computer program read from it can be installed into the storage section 1508 as required.
[0286] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes computer program instructions for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 15010, and / or installed from the removable medium 1512. When the computer program is executed by the central processing unit (CPU) 1501, the above functions defined in the system of the present disclosure are executed.
[0287] It should be noted that the computer-readable storage medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in this disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable computer program instructions. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and this computer-readable storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program instructions contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0288] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be included in the device described in the above embodiments; or may exist alone without being assembled into the device. The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the device, the functions that the device can implement include: obtaining a training sample image and a domain label of the training sample image, where the domain label of the training sample image is the source domain or the target domain; performing feature extraction processing on the training sample image through a feature extraction network in a cross-domain object detection model to obtain image-level features of the training sample image; wherein, the cross-domain object detection model further includes an image-level domain adaptation network, an object detection network, and an instance-level domain adaptation network; performing image-level domain classification on the image-level features through the image-level domain adaptation network to determine an image-based domain classification result of the training sample image, and determining a first loss based on the domain label and the image-based domain classification result; performing object feature extraction on the image-level features through the object detection network to obtain instance-level features of the training sample image; performing feature extraction processing on the instance-level features through the instance-level domain adaptation network to obtain an object-based domain classification result of the training sample image, and determining a second loss based on the domain label and the object-based domain classification result; calculating a consistency loss between the image-based domain classification result and the object-based domain classification result to determine a third loss; updating the parameters of the cross-domain object detection model based on the first loss, the second loss, and the third loss, so as to perform cross-domain object detection according to the cross-domain object detection model with updated parameters.
[0289] According to one aspect of the present disclosure, there is provided a computer program product or a computer program, which includes computer program instructions stored in a computer-readable storage medium. Reading the computer program instructions from the computer-readable storage medium, and a processor executes the computer program instructions to implement the methods provided in the various alternative implementations of the above embodiments.
[0290] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several computer program instructions for causing an electronic device (such as a server or a terminal device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0291] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0292] It should be understood that the present disclosure is not limited to the detailed structures, drawing methods, or implementation methods shown herein. On the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method for cross-domain target detection, characterized in that: include: Acquire a training sample image and a domain label of the training sample image, wherein the domain label of the training sample image is a source domain or a target domain; Performing feature extraction processing on the training sample image through a feature extraction network in a cross-domain target object detection model to obtain image-level features of the training sample image; wherein the cross-domain target object detection model also includes an image-level domain adaptation network, a target object detection network, and an instance-level domain adaptation network; performing image-level domain classification on the image-level features through the image-level domain adaptation network, determining an image-based domain classification result of the training sample image, and determining a first loss based on the domain label and the image-based domain classification result; Extracting object features from the image-level features through the target object detection network to obtain instance-level features of the training sample image; performing feature extraction processing on the instance-level features through the instance-level domain adaptation network to obtain an object-based domain classification result of the training sample image, and determining a second loss based on the domain label and the object-based domain classification result; Performing consistency loss calculation on the image-based domain classification result and the object-based domain classification result to determine a third loss; Based on the first loss, the second loss and the third loss, parameters of the cross-domain target object detection model are updated, so as to perform cross-domain target detection according to the cross-domain target object detection model after parameter update.
2. The method according to claim 1, characterized in that: The image-level domain adaptation network includes a multi-scale information aggregation subnetwork, a first adversarial gradient reversal layer, and an image-level domain classifier; wherein, performing image-level domain classification on the image-level features through the image-level domain adaptation network to determine the image-based domain classification result of the training sample image includes: Extracting multi-scale information from the image-level features through the multi-scale information aggregation subnetwork to obtain multi-scale information aggregation features; Performing feature extraction processing on the multi-scale information aggregation feature through the first adversarial gradient inversion layer to obtain a multi-scale information aggregation deep feature; The image-level features are subjected to image-level domain classification by the image-level domain classifier to determine an image-based domain classification result of the training sample image.
3. The method according to claim 2, characterized in that: When updating parameters of the cross-domain target object detection model, the back propagation directions of the parameters before and after the first adversarial gradient reversal layer are opposite.
4. The method according to claim 2, characterized in that: The first adversarial gradient reversal layer includes a reversal degree parameter, and the reversal degree parameter is used to weight the propagation function gradient in the first adversarial gradient reversal layer; the method also includes: When the first loss is greater than a preset threshold, the reversal degree parameter is equal to a preset constant; When the first loss is less than or equal to the preset threshold, the reversal degree parameter is inversely proportional to the first loss.
5. The method according to claim 2, characterized in that: The multi-scale information aggregation subnetwork includes a plurality of convolutional layers with different expansion rates; wherein, the multi-scale information aggregation subnetwork is used to extract multi-scale information from the image-level features to obtain multi-scale information aggregation features, including: Extracting features of the image level features respectively through the multiple convolutional layers with different expansion rates to obtain feature maps of multiple different receptive fields; Information aggregation processing is performed on the feature maps of the multiple different receptive fields to obtain the multi-scale information aggregation feature.
6. The method according to claim 1, characterized in that: The instance-level domain adaptation network includes a second adversarial gradient reversal layer and an instance-level domain classifier, wherein the instance-level features are subjected to feature extraction processing by the instance-level domain adaptation network to obtain an object-based domain classification result of the training sample image, and a second loss is determined based on the domain label and the object-based domain classification result, including: The instance-level features are subjected to feature extraction processing by the second adversarial gradient reversal layer to obtain instance-level deep features; wherein, when the parameters of the cross-domain target object detection model are updated, the back propagation directions of the parameters before and after the second adversarial gradient reversal layer are opposite; The instance-level deep features are classified by the instance-level domain classifier to obtain the object-based domain classification result.
7. A device for cross-domain target detection, characterized in that: include: A sample acquisition module, used to acquire a training sample image and a domain label of the training sample image, wherein the domain label of the training sample image is a source domain or a target domain; A feature extraction module, used to perform feature extraction processing on the training sample image through a feature extraction network in a cross-domain target object detection model to obtain image-level features of the training sample image; wherein the cross-domain target object detection model also includes an image-level domain adaptation network, a target object detection network, and an instance-level domain adaptation network; a first loss determination module, configured to perform image-level domain classification on the image-level features through the image-level domain adaptation network, determine an image-based domain classification result of the training sample image, and determine a first loss based on the domain label and the image-based domain classification result; An instance-level feature acquisition module, used to extract object features from the image-level features through the target object detection network to obtain instance-level features of the training sample image; a second loss determination module, configured to perform feature extraction processing on the instance-level features through the instance-level domain adaptation network to obtain an object-based domain classification result of the training sample image, and determine a second loss based on the domain label and the object-based domain classification result; a third loss determination module, configured to perform consistency loss calculation on the image-based domain classification result and the object-based domain classification result to determine a third loss; A parameter updating module is used to update the parameters of the cross-domain target object detection model based on the first loss, the second loss and the third loss, so as to perform cross-domain target detection according to the cross-domain target object detection model after the parameter update.
8. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer program instructions; the processor calls the computer program instructions stored in the memory to implement the method for cross-domain target detection as described in any one of claims 1-6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method for cross-domain target detection according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising computer program instructions, wherein the computer program instructions are stored in a computer-readable storage medium, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Model training method and device and image detection method and device
CN122244427A