Trusted remote sensing target detection method and device and storage medium
By adopting a method based on uncertain perception dynamic context evidence fusion in remote sensing object detection, the problem of performance degradation in the prior art caused by excessive confidence and context introduction is solved, and high-precision and reliable remote sensing object detection in the risk-sensitive field is achieved.
Patent Information
- Application Number
- CN202510244087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing remote sensing object detection methods are difficult to apply to risk-sensitive fields due to overconfidence, and the introduction of context leads to performance degradation in some scenarios.
A trusted remote sensing object detection method based on uncertain perception dynamic context evidence fusion is adopted. By constructing an Oriented-RCNN-based model, a backbone network, a region generation network and a context-enhanced network are combined with feature extraction backbone network, a region generation network, and a context-enhanced network, and an adaptive semantic evidence-keeping strategy and D-S evidence fusion rules are adopted.
Overcoming the problem of overconfidence makes the object detection model suitable for risk-sensitive areas, while improving detection accuracy and reliability, avoiding performance degradation caused by context introduction.
Smart Images

Figure CN119942088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and in particular to a reliable remote sensing target detection method, device and storage medium. Background Art
[0002] In recent years, many remote sensing target detection methods have been widely used in various application scenarios using deep learning. However, since existing methods usually rely on imprecise hard annotations as supervision information, there is inevitably a problem of overconfidence, which is manifested in that the model sometimes makes unexpected, incorrect but overconfident predictions, thus limiting the application of such methods in many risk-sensitive fields (such as national defense security, disaster relief, etc.).
[0003] To this end, many methods alleviate this problem by modeling the uncertainty of the model. Common methods include MC-Dropout, deep integration networks, and evidence deep learning. The former two often have high computational overhead (MC-Dropout requires multiple forward propagations, and deep integration networks require reasoning on multiple models), while evidence deep learning only requires one forward propagation to model the uncertainty of the model, which has an advantage in computational overhead. However, evidence deep learning usually underfits difficult samples, resulting in a decrease in detection performance.
[0004] In addition, many methods recognize that referencing context in remote sensing images is necessary and effective. However, in some special scenes with dense targets, high noise and cluttered background, the introduction of context can be counterproductive.
[0005] Therefore, this application specifically proposes a reliable remote sensing target detection method to solve the above technical problems. Summary of the invention
[0006] The main purpose of the present invention is to provide an implementable trusted remote sensing target detection method based on uncertainty-aware dynamic contextual evidence fusion, so as to solve the problem that remote sensing target detection proposed in the background technology is difficult to apply in risk-sensitive fields due to overconfidence, and the problem of performance degradation caused by inappropriate introduction of context.
[0007] The present invention adopts the following technical solutions to solve the above technical problems: A reliable remote sensing target detection method comprises the following steps: S1. Obtain remote sensing target detection training data set; S2. Taking Oriented-RCNN as the baseline model, a credible remote sensing target detection model based on uncertainty-aware dynamic contextual evidence fusion is constructed. The feature extraction backbone network is constructed, and after the feature extraction backbone network is constructed, a region generation network and a context enhancement network are further constructed. The context enhancement network is used to extract feature maps after context enhancement. In addition, the original classification head is transformed into an evidence classification head based on an adaptive semantics-preserving evidence reduction strategy, while the regression head remains unchanged. S3. Training the credible remote sensing target detection model by training the data set; S4. Obtain a test data set for remote sensing target detection; S5. Input the test image in the test data set into the trained credible remote sensing target detection model, obtain the RoI region after context enhancement and the RoI region without context enhancement, and then obtain the classification result and the predicted box position through the evidence classification head and regression head of the credible remote sensing target detection model; S6. Perform DS evidence fusion on the classification results to obtain the fused classification results, and output them together with the prediction box position to obtain the final prediction result.
[0008] Preferably, the process of building a reliable remote sensing target detection model in step S2 includes: S21. Construct a feature extraction backbone network, assuming that the image input to the network is , is the resolution The RGB image is input into the feature extraction backbone network, and the output feature map is , Denoted as the feature map size, where Indicates the number of channels, represents the channel width, Indicates the channel height; S22. Construct a region generation network, which is used to generate the region from the feature map Generate candidate regions that may contain targets, input the output feature map into the region generation network, and get the output ,in Indicates the number of candidate regions that may contain targets, usually for different input images In terms of The number of is generally different, RoI uses a 5-dimensional vector represents the candidate region, where represents the center point coordinates, represents the width and height of the candidate region, Indicates the rotation angle; S23. Construct a context enhancement network, which is used to introduce more context information into the feature map. Input into the context enhancement network to generate a feature map with context information ,In the context enhancement network, the output shape of the feature map is consistent with the input shape; S24. Feature map With feature map Perform RoI-Align operation and use RoI-Align to obtain The feature maps in the region are obtained respectively. and ; S25. Build a regression head, which is responsible for the candidate The bounding box of the region is fine-tuned to make it more accurately match the position of the real target. Since the positioning of the object does not require the assistance of context information, it is directly Input regression header and get regression offset ; S26. Build a classification evidence header, which is responsible for each candidate The region extracts evidence and generates various types of evidence for classification and uncertainty modeling. The classification evidence head outputs two parts of total evidence and distribution of evidence categories ,in is the total number of categories; S27. The total amount of evidence and distribution of evidence categories Combination, can obtain evidence vector , for classification and uncertainty modeling. Preferably, the total training loss RPN is also calculated during the training process of step S3, and the output of RPN is , using a 5-dimensional vector represents the candidate region, denoted as ,in represents the center point coordinates, represents the width and height of the candidate region, Represents the rotation angle. The RPN loss consists of regression loss and classification loss, which are optimized using mean square error and cross entropy loss respectively. The RPN loss formula is as follows:
[0009] in, and are the weight hyperparameters for RPN regression and classification, respectively. is the mean square error loss, is the cross entropy loss.
[0010] Preferably, the regression head loss is also calculated during the training process of step S3, and the output of the regression head is the offset from the RoI to the actual prediction box Therefore, when calculating the regression head loss, it is necessary to first calculate the offset of each RoI to the GT box, including: Calculate the offset of each RoI region to the GT box ,have:
[0011] in, Represents the frame parameters of the GT frame area, Represents the box parameters of the bounding box area. The box parameters include , is the horizontal coordinate of the center point of the frame area, is the ordinate of the center point of the frame area, is the width of the frame area, is the height of the frame area, Indicates the rotation angle of the frame area; Calculate the mean square error for the offset .
[0012] Preferably, the training process of step S3 also calculates the loss of the evidence classification head, and the output of the evidence classification head is the evidence vector ,in is the total amount of evidence of all categories collected from the sample, is the evidence category distribution. According to the theory of evidence deep learning, the output of the model is a Dirichlet distribution with the distribution parameter ; At this time, the second type of maximum likelihood estimation in the theory of evidence deep learning is used for loss modeling, and the Divergence and The two regularization terms are calculated, and the loss formula is:
[0013] in The loss is shown in the following formula
[0014] in is the data category label One-hot encoding of is the Dirichlet intensity, that is , since the Dirichlet distribution is used to model the second-order distribution of the model, it is a conjugate prior for the categorical distribution, so the integral , represents the parameter of the classification distribution, which means the probability that the target belongs to each category.
[0015] The loss is shown in the following formula:
[0016] in , , represents the element-by-element multiplication of two vectors, for function, for function. Similarly, the Dirichlet distribution is the conjugate prior for the categorical distribution, so , represents the parameter of the classification distribution, which means the probability that the target belongs to each category, is the evidence vector output by the model, is the distribution parameter of the Dirichlet distribution, then The loss is shown in the following formula:
[0017] in is the data category label One-hot encoding of Distribution of evidence categories.
[0018] Preferably, the training process of step S3 also adopts an adaptive semantics-preserving evidence reduction strategy, and when updating parameters, In (distribution of evidence categories, indicating the percentage of each type of evidence in the total amount of evidence) is set to a constant, and its gradient is not back-propagated during the gradient back-propagation process. Therefore, the original semantics is maintained during the parameter update process and the difficult and easy samples are adaptively weighted.
[0019] Preferably, As a constant, it can also be regarded as A weight is used to adaptively weight difficult and easy samples, so that difficult samples will reduce the total amount of evidence, ensuring the model's ability to distinguish difficult and easy samples.
[0020] Preferably, the specific operation process in step S5 includes: S51. Through the feature extraction backbone network, the output is obtained ; S52. Through the region generation network, the output is obtained ; S53. Through the context enhancement network, the feature map output in step S51 is Input into the context enhancement network and get the output , its output shape is consistent with the input shape; S54. The feature map output in step S2 and Perform RoI-Align and get and ; S55. Send it to the regression head and output the position offset , and then and Send it to the evidence classification head and output the evidence vector without context enhancement and the context-enhanced evidence vector .
[0021] Preferably, the specific operation process of step S6 includes: S61. Based on evidence deep learning and subjective logic theory, the evidence vector Convert to belief value , there is uncertainty ; Defining the set of belief qualities after context enhancement and , and combined according to DS evidence theory, there is a belief quality set ,in:
[0022]
[0023]
[0024]
[0025] in represents the degree of conflict between two sets of belief qualities, and is the quality and uncertainty of the integrated belief, and are the belief values in the two belief quality sets before fusion. Similarly, and They are the uncertainty before fusion; S62. Perform NMS using the fused result as the threshold of NMS, and then output the final detection result.
[0026] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0027] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0028] It can be seen from the above technical solution that the present invention provides a reliable remote sensing target detection method. Compared with the prior art, the present invention has the following advantages: 1. The present invention can be used for two-stage remote sensing target detection by applying high-order probability modeling methods to deep learning of evidence, giving the target detection model the ability to output uncertainty, so as to overcome the problem of overconfidence in existing remote sensing target detection, making it suitable for risk-sensitive fields. Compared with traditional MC-Dropout and deep integration methods, it can greatly save data processing operation time.
[0029] 2. The present invention overcomes the problem of underfitting of difficult samples caused by indiscriminate simplification of non-GT evidence in traditional evidence deep learning by decoupling the total amount of evidence and the distribution of evidence categories and adopting an adaptive semantic-preserving evidence reduction strategy. This improves detection accuracy while maintaining the ability to evaluate uncertainty.
[0030] 3. Based on deep learning of evidence, the present invention dynamically fuses contextual evidence through DS evidence combination rules, which can achieve dynamic and credible fusion of prediction accuracy and reliability metrics while maintaining compatibility with existing context fusion methods.
[0031] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become easy to understand through the following description. Of course, it is not necessary to achieve all of the advantages described above simultaneously for any product implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 It is a schematic diagram of the overall operation flow of the present invention; Figure 2 This is a structural block diagram of the credible remote sensing target detection model system of the present invention. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0034] In the embodiment, see Figure 1 to Figure 2 .
[0035] like Figure 1 As shown, the trusted remote sensing target detection method proposed in the embodiment of the present invention is based on an adaptive semantic-preserving evidence reduction strategy and dynamic contextual evidence fusion, and proposes a trusted remote sensing target detection method, which solves the problem that remote sensing target detection is difficult to apply in risk-sensitive fields due to overconfidence, and at the same time improves the detection accuracy in contextual scenarios.
[0036] The specific steps include: Step 1: Obtain remote sensing target detection training data set. Define the training data set as , Representative For training data, they are remote sensing images , and a callout box .
[0037] The training images are three-channel RGB images. Since the number of targets in each image is uncertain, The number of annotations in a collection usually varies from image to image. Represents the center point coordinates and width and height of one of the annotation boxes. Indicates the label corresponding to this target.
[0038] Step 2, build Figure 2 The credible remote sensing target detection model based on uncertainty-aware dynamic contextual evidence fusion is shown; Specifically, it is divided into the following steps: Step 2.1, construction of feature extraction backbone network. Assume that the image input to the network is , is the resolution The RGB image is input into the feature extraction backbone network, and the output features are The traditional target detection model only uses the output of the last convolutional layer for subsequent detection, which leads to poor detection performance in small target scenes. Therefore, it is necessary and effective to use multi-scale fusion feature maps for prediction. The present invention uses FPN as a multi-scale fusion method. For the convenience of description, only one scale is used in the following text. To elaborate.
[0039] Step 2.2: Construction of the region generation network. The region generation network is used to generate candidate regions that may contain targets. In order to ensure that forward propagation can still be performed when the input scale changes, and to ensure that the input and output shapes remain consistent, the present invention adopts multiple The fully convolutional network composed of convolutional layers (filled with 1) is used as the region generation network. The region generation network first generates anchor boxes of different sizes, scales and directions at each pixel point. , each anchor box uses a 5-dimensional vector Then, the feature map output from step 2.1 is input into the region generation network to obtain the output for each anchor box. , which means the regression offset for each anchor box and the probability of whether it contains an object Then, according to the probability of whether the object is contained Perform NMS to obtain ,in Indicates the number of candidate regions that may contain targets, usually for different input images In terms of The number of is generally different, RoI also uses a 5-dimensional vector represents the candidate region, where represents the center point coordinates, represents the width and height of the candidate region, Indicates the rotation angle.
[0040] Step 2.3, construction of context enhancement network. The context enhancement network is used to introduce more context information into the feature map. To this end, the present invention uses dilated convolution to expand the perception of the feature map and can introduce more context information into the RoI after RoI-Align. The feature map output from step 2.1 is input into the context enhancement network, and then four feature maps are obtained through four dilated convolutions with dilated values of 0, 1, 2, and 3, respectively. , and then concatenate them to get , Its output shape is the same as its input shape.
[0041] Step 2.4, use RoI-Align to obtain the feature map in the RoI area. Perform RoI-Align on the output of step 2.1 and step 2.3, and get and .
[0042] Step 2.5, build the regression head. The regression head is responsible for fine-tuning the bounding box of the candidate region to make it more accurately match the position of the real target. Since the feature map after RoI-Align ensures that the shape remains unchanged, the regression head can be designed using a fully connected layer. Specifically, the regression head consists of two The convolutional layer and two fully connected layers are used. Since the location of objects does not require the assistance of context information, Input regression head to get output . Step 2.6, construction of classification evidence head. The classification evidence head is responsible for extracting evidence from each candidate region and generating various types of evidence for classification and uncertainty modeling. The classification evidence head in this invention is consistent with the regression head in step 2.5 in terms of structure, except that it has two output parts: total evidence amount and distribution of evidence categories , generated by two fully connected layers, where K is the total number of categories, and then combined into an evidence vector ,in is the total amount of evidence of all categories collected from the sample, Distribution of evidence categories.
[0043] Step 3: training the credible remote sensing target detection model based on uncertainty-aware dynamic context evidence fusion.
[0044] Step 3.1, calculate the RPN loss of the region generation network. Due to the existence of anchor boxes, the RPN loss of the region generation network takes an anchor box as a sample. The RPN loss is divided into two parts, regression loss and classification loss. The model needs to calculate the IoU for each GT and all anchor boxes, and assign it to the anchor box with the largest IoU. Each GT has only one anchor box assigned to it. Send the feature map to the RPN network to get the output for each anchor box. . For an anchor box, its regression loss is expressed as follows:
[0045] in is the mean square error, Indicates the anchor box assigned to GT, and the actual offset of GT box. For an anchor box, its classification loss is expressed as follows:
[0046] in is the cross entropy loss, so the overall loss of RPN is as follows:
[0047] in and are the weight hyperparameters for RPN regression and classification, respectively.
[0048] Step 3.2, calculate the regression head loss. The regression head loss refers to the regression loss part of the RPN loss in step 3.1, except that here the offset between RoI and GT is calculated.
[0049] Step 3.3, calculate the evidence classification head loss. The output of the evidence classification head is the evidence vector ,in is the total amount of evidence of all categories collected from the sample, is the evidence category distribution. According to the theory of evidence deep learning, the output of the model is a Dirichlet distribution, and its distribution parameter is The loss is given by the following formula:
[0050] in The loss is the second-type maximum likelihood estimate of the Dirichlet distribution, as shown in the following formula:
[0051] in is the data category label One-hot encoding of is the Dirichlet intensity, that is , since the Dirichlet distribution is used to model the second-order distribution of the model, it is a conjugate prior for the categorical distribution, so the integral , represents the parameter of the classification distribution, which means the probability that the target belongs to each category.
[0052] The loss is shown in the following formula:
[0053] in , , represents the element-by-element multiplication of two vectors, for function, for function. Similarly, , represents the parameter of the classification distribution, which means the probability that the target belongs to each category. is the evidence vector output by the model, is the distribution parameter of the Dirichlet distribution. The loss is shown in the following formula:
[0054] in is the data category label One-hot encoding of Distribution of evidence categories.
[0055] Step 3.4: Adaptive semantics-preserving evidence reduction strategy. and distribution of evidence categories Decoupled into two parts, so when updating parameters, In (evidence category distribution, indicating the percentage of each type of evidence in the total amount of evidence) is set as a constant, and its gradient is not back-propagated during the gradient back-propagation process, so the original semantics is maintained during the parameter update process. In addition, this design can make The optimization goal of this item is Convert to Therefore, After being treated as a constant, it can also be regarded as A weight is used to adaptively weight difficult and easy samples, so that difficult samples will reduce the total amount of evidence, ensuring the model's ability to distinguish difficult and easy samples.
[0056] Step 4, obtaining a remote sensing target detection test data set; Step 5: Input the test image into the trained remote sensing target detection model to obtain the context-enhanced RoI and the non-context-enhanced RoI respectively, and then obtain the classification result and the predicted box position through the evidence classification head and regression head.
[0057] Specifically, it is divided into the following steps: Step 5.1, through the feature extraction backbone network, get the output ; Step 5.2, through the region generation network, get the output ; Step 5.3: Input the feature map output from step 5.1 into the context enhancement network through the context enhancement network to obtain the output , its output shape is consistent with the input shape; Step 5.4, perform RoI-Align on the outputs of step 2.1 and step 2.3, and obtain and .
[0058] Step 5.5, Send it to the regression head and output the position offset , and then and Send it to the evidence classification head and output the evidence vector without context enhancement and the context-enhanced evidence vector .
[0059] Step 6: Perform DS evidence fusion on the classification results to obtain the fused classification results, and output them together with the prediction box position to obtain the final prediction results. Step 6.1, dynamic context evidence fusion. According to evidence deep learning and subjective logic theory, the evidence vector Can be converted into belief value , uncertainty The definition of the belief quality set after context enhancement is and the belief quality set without context enhancement is , according to DS evidence theory, the combined belief quality set is , as shown by the following formula:
[0060] A more specific calculation is shown in the following formula:
[0061] in represents the degree of conflict between two sets of belief qualities, , is the quality and uncertainty of the belief after integration. and is the belief value in the two belief quality sets before fusion. Similarly, and The uncertainty before fusion.
[0062] In step 6.2, the fused result is used as the threshold of NMS to perform NMS, and then the final detection result is output. The detection result includes the location of the prediction box, the classification result, and the uncertainty assessment.
[0063] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0064] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0065] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any of the trusted remote sensing target detection methods in the above embodiments.
[0066] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.
[0067] The embodiment of the present application also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. Memory, used to store computer programs; The processor is used to implement the above-mentioned reliable remote sensing target detection method when executing the program stored in the memory.
[0068] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industrial Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.
[0069] The communication interface is used for communication between the above electronic device and other devices.
[0070] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0071] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0072] It should also be noted that electronic devices also include terminal devices, which can also be called terminals, user equipment, mobile stations, mobile terminals, etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablet computers, computers with wireless transceiver functions, virtual reality terminal devices, augmented reality terminal devices, wireless terminals in industrial control, wireless terminals in unmanned driving, wireless terminals in remote surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technology and specific device form used by the terminal devices.
[0073] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk), etc.
[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
[0075] In addition, it should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back...), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components in a certain specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0076] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing in the full text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or schemes that A and B meet at the same time. In addition, in the embodiments of the present invention, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
Claims
1. A reliable remote sensing target detection method, characterized in that: The following steps are involved: S1. Obtain training data set; S2. Using Oriented-RCNN as the baseline model, a reliable remote sensing target detection model is constructed; S3. Train a credible remote sensing target detection model using a training data set; S4. Obtain a test data set for remote sensing target detection; S5. Input the test image in the test data set into the trained credible remote sensing target detection model, obtain the RoI region after context enhancement and the RoI region without context enhancement, and then use the credible remote sensing target detection model to obtain the classification result and the predicted box position; S6. Perform DS evidence fusion on the classification results to obtain the fused classification results, and output them together with the prediction box position to obtain the final prediction result.
2. The credible remote sensing target detection method according to claim 1, characterized in that: The construction process of the credible remote sensing target detection model in step S2 includes: S21. Build a feature extraction backbone network, input the RGB image into the feature extraction backbone network, and output the feature map , Denoted as the feature map size, where Indicates the number of channels, represents the channel width, Indicates the channel height; S22. Construct a region generation network to generate Generate candidate regions containing targets ,in Indicates the number of candidate regions containing the target; S23. Construct a context enhancement network to extract Generate feature maps with context information ,In the context enhancement network, the output shape of the feature map is consistent with the input shape; S24. Feature map With feature map Perform RoI-Align operation to obtain The feature maps in the region are obtained respectively. and ; S25. Build a regression head for The bounding box of the region is fine-tuned and extracted from the feature map Get the regression offset in ; S26. Construct a classification evidence header for use in Evidence extraction and total evidence generation within the region and distribution of evidence categories ,in is the total number of categories; S27. Combination to obtain evidence vector , for classification and uncertainty modeling, where represents the total amount of evidence of all categories collected from the sample, It indicates the percentage of each type of evidence in the total amount of evidence, that is, the distribution of evidence categories.
3. The credible remote sensing target detection method according to claim 2, characterized in that: The total training loss RPN is also calculated during the training process of step S3. The RPN loss consists of regression loss and classification loss, which are optimized using mean square error and cross entropy loss respectively. The RPN loss formula is as follows: in, and are the weight hyperparameters for RPN regression and classification, respectively. is the mean square error loss, is the cross entropy loss.
4. The credible remote sensing target detection method according to claim 3, characterized in that: The regression head loss is also calculated during the training process of the S3 step, including: Calculate the offset of each RoI region to the GT box ,have: in, Represents the frame parameters of the GT frame area, Represents the box parameters of the bounding box area. The box parameters include , is the horizontal coordinate of the center point of the frame area, is the ordinate of the center point of the frame area, is the width of the frame area, is the height of the frame area, Indicates the rotation angle of the frame area; Calculate the mean square error for the offset .
5. The credible remote sensing target detection method according to claim 2, characterized in that: The training process of step S3 also calculates the evidence classification head loss, uses the second-class maximum likelihood estimation in the evidence deep learning theory to model the loss, and adds Divergence and The two regularization terms are calculated, and the loss formula is: in, is the data category label One-hot encoding of and They are loss functions and The weight hyperparameters are represents the element-by-element multiplication of two vectors, for function, for function, represents the parameters of the categorical distribution, is the evidence vector output by the model, is the distribution parameter of the Dirichlet distribution.
6. The reliable remote sensing target detection method according to claim 5, characterized in that: The training process of step S3 also adopts an adaptive semantics-preserving evidence reduction strategy. When updating parameters, Set it as a constant, and do not back-propagate its gradient during the gradient back-propagation process to maintain the original semantics and adaptively weight the difficult and easy samples.
7. The credible remote sensing target detection method according to claim 2, characterized in that: In the specific operation flow of step S5: Will and Send it to the evidence classification head and output the evidence vector without context enhancement and the context-enhanced evidence vector .
8. The credible remote sensing target detection method according to claim 2, characterized in that: The specific operation process of step S6 includes: S61. Based on evidence deep learning and subjective logic theory, the evidence vector Convert to belief value , there is uncertainty ; Defining the set of belief qualities after context enhancement and , and combined according to DS evidence theory, there is a belief quality set ,in: in represents the degree of conflict between two sets of belief qualities, and is the quality and uncertainty of the integrated belief, and are the belief values in the two belief quality sets before fusion, and They are the uncertainty before fusion; S62. Perform NMS using the fused result as the threshold of NMS, and then output the final detection result.
9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.
10. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.