Remote sensing image target identification method
Through the methods of multi-scale feature extraction and fusion, dynamic weight allocation and adaptive heterogeneous support tensor machine, the problems of scale feature capture and multi-source information integration in traditional remote sensing image processing are solved, high-precision target detection and recognition are achieved, and the adaptability and accuracy of remote sensing image processing are improved.
Patent Information
- Application Number
- CN202510856156.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional convolutional neural networks have difficulty in simultaneously capturing the features of targets of different scales when processing remote sensing images, resulting in limited target detection accuracy. Existing methods are also unable to effectively integrate multi-source information and dynamically adjust feature importance, affecting the accuracy and adaptability of remote sensing image target recognition.
The method adopts multi-scale feature extraction and fusion, construction of heterogeneous feature tensors, feature processing based on dynamic weight allocation, target detection and positioning, and adaptive heterogeneous support tensor machine. Through multi-branch convolution and void convolution collaborative feature extraction, combined with state variable weight theory and tensor generalized Hough transform, multi-scale fusion of features and dynamic weight adjustment are realized, thereby improving the accuracy of target detection and recognition.
It significantly improves the accuracy and recall rate of remote sensing image target detection, enhances the ability to characterize targets of different sizes and shapes, improves the adaptability and accuracy of the model in complex environments, reduces the number of false alarms, and provides a reliable basis for target positioning and recognition.
Smart Images

Figure CN120708088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image target recognition method. Background Art
[0002] Remote sensing images contain a wealth of information, ranging from macroscopic geographical features to microscopic specific targets. These targets vary greatly in size, shape, and distribution. Traditional convolutional neural networks, due to their fixed kernel size and receptive field, struggle to simultaneously and effectively capture the features of targets of varying scales. Small targets are easily overlooked or their features are insufficiently extracted, while large targets can lose local details, limiting target detection accuracy.
[0003] Remote sensing images contain heterogeneous information from multiple sources, including spatial location, spectral information, and time series (multi-temporal remote sensing imagery). Efficiently integrating these different types of information to construct a data structure that comprehensively and accurately describes target characteristics has always been a challenge in remote sensing image processing. Previous simple data fusion methods failed to fully exploit the inherent connections between these multiple sources of information, hindering the representation and understanding of the target's complex characteristics.
[0004] Existing object recognition models typically use fixed weights to process features, failing to dynamically adjust feature importance based on image scenes and changes in target features. In complex remote sensing imagery, this fixed weighting approach results in insufficient focus on key features and a weaker ability to suppress interfering features, reducing the model's adaptability and accuracy.
[0005] When it comes to identifying target types in remote sensing images, traditional classification methods are prone to overfitting and blurred classification boundaries when processing high-dimensional, heterogeneous remote sensing data. Classic algorithms such as support vector machines have limitations when processing tensor data and cannot fully exploit the structural characteristics of multi-source remote sensing imagery data. As a result, target type recognition accuracy fails to meet the requirements of practical applications such as military target identification and transportation facility classification.
[0006] Therefore, the present invention proposes a remote sensing image target recognition method. Summary of the Invention
[0007] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0008] In view of the above problems and / or the problems existing in the prior art, the present invention is proposed.
[0009] Therefore, the purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for remote sensing image target recognition, which comprises the following steps: S100, multi-scale feature extraction and fusion; S200, constructing heterogeneous feature tensors; S300, feature processing based on dynamic weight allocation; S400, target detection and positioning; S500, target model identification; Preferably, the multi-scale feature extraction and fusion in S100 includes the following steps: S110, multi-branch convolution and dilated convolution collaborative feature extraction; S120, feature fusion strategy strengthens feature expression; In the S110, multi-branch convolution and dilated convolution collaborative feature extraction, multiple convolution branches with different convolution kernel sizes and expansion rates are constructed. ;in: The first branch adopts The group convolution of the convolution kernel is calculated as follows: ; in, Represents the number of groups. Through group convolution, the receptive field of the convolution kernel is expanded while the amount of convolution operations is reduced, thereby capturing a wider range of image information. It is a nonlinear activation function responsible for introducing nonlinear factors and enhancing the expressive power of the model; Represents the batch normalization operation, which is used to normalize the convolutional features, accelerate model convergence and improve training stability; Represents a convolution operation.
[0010] The second branch uses the expansion rate of The formula for dilated convolution is:
[0011] Continue to build other branches, such as the third branch adopts Group convolution of convolution kernels , and the expansion rate is of Dilated convolution as the fourth branch wait.
[0012] In the above S120, in the feature fusion strategy to enhance the feature expression, each branch feature is first spliced using the following formula: , connect the features of different branches along the channel dimension, so that the fused features cover multi-scale information and enrich the dimension of the features. Then, the features are further fused by element-by-element summation. Let the aggregation operation stands for element-wise sum, through the formula. Get the final fusion features .
[0013] Preferably, the step S200, constructing a heterogeneous feature tensor, includes the following steps: S210, determining tensor dimension based on multi-scale features; S220, tensor element assignment and structure construction; In the S210, in determining the tensor dimension based on multi-scale features, it is assumed that the fusion feature In the spatial dimension OK, Column, in the spectral dimension channels, taking into account the time series information (if any), with time steps (if it is a static image, ). Then, the constructed heterogeneous feature tensor For one rank tensor( or , depending on whether the time dimension is included), its dimension is expressed as T∈RHxWxCxT.
[0014] In the S220, tensor element assignment and structure construction, the multi-scale features extracted in S100 are filled into the heterogeneous feature tensor according to certain rules. For spatial location , in the spectral channel and time step According to the fusion features The corresponding information is assigned. If the fusion feature Central position ,aisle , time step The characteristic value of , then assign it to the heterogeneous feature tensor The corresponding element .
[0015] Preferably, the feature processing based on dynamic weight allocation in S300 includes the following steps: S310, determining a constant weight vector; S320, introduce state variable weight theory to calculate dynamic weight; S330, calculating dynamic weights and applying them to feature processing; In the step S310, determining the constant weight vector, after constructing the heterogeneous feature tensor Then, let the heterogeneous feature tensor The feature dimension is , constant weight vector ,satisfy and .
[0016] In step S320, the state variable weight theory is introduced to calculate the dynamic weight, and the state variable weight vector is defined. ,in is a vector containing the states of each characteristic index. The state variable weight vector must meet certain conditions, such as for any ,have or ; For each variable Continuous and strictly monotonically decreasing or monotonically increasing; for any combination about The following state-variable weight function is used to calculate the state-variable weight vector:
[0017] in, , Indicates the negation level, Indicates passing level, Indicates the motivation level, Incentive regulation level is a heterogeneous feature tensor Middle The standardized state value of a characteristic indicator.
[0018] The S330, calculates the dynamic weight and applies it to feature processing, through the constant weight vector and state variable weight vector The normalized product of gets the dynamic weight vector :
[0019] After obtaining the dynamic weight vector, apply it to the heterogeneous feature tensor For heterogeneous feature tensors Each feature dimension is weighted according to the dynamic weight. , the processed eigenvalues , is the original eigenvalue.
[0020] Preferably, the target detection and positioning in S400 includes the following steps: S410, Fundamentals of target detection based on tensor generalized Hough transform; S420, Voting mechanism and target positioning of remote sensing images; S430, false alarm removal and target determination; In the target detection basis based on tensor generalized Hough transform in S410, the template image is assumed to be , use edge detection operators (such as Canny operator) to get edge images ,in Indicates the location of the contour point in the template image .
[0021] Set the angle and scale interval to and , traverse the possible rotation angles and scale For the current and , calculate the new position of the contour point after rotation and scale change ,in is the position of the reference point, is the rotation matrix; at the same time, the new gradient angle of the contour point , and quantize the gradient angle as , is the gradient angle quantization interval.
[0022] Constructing a fifth-order tensor reference table , the first, second, third, fourth and fifth orders represent the horizontal space domain, vertical space domain, gradient angle domain, rotation angle domain and scale domain respectively. and scale , the contour point attributes obtained under the conditions , let the tensor reference the corresponding element in the table , complete the construction of the tensor reference table.
[0023] In the above S420, voting mechanism and target positioning of remote sensing images, for the remote sensing images to be detected , using a window of the same size as the template image to scan each position in the image , generating a series of slices For each slice, extract the gradient angle of its contour points And record the relevant attributes to construct a third-order tensor used to describe the slice outline information , where the first, second and third orders represent the horizontal spatial domain, vertical spatial domain and gradient angle domain respectively.
[0024] According to the tensor reference table constructed , perform rotation and scale-invariant contour matching, and calculate candidate reference points of remote sensing images If the number of votes and are all equal to 1, that is, there is a contour point with the same attributes in the template image and the slice, then the contour point (,,y,) in the slice will generate a vote for a specific rotation angle and scale. Candidate reference point The total number of votes received at a specific rotation angle and scale is expressed as To prevent the scale change from affecting the vote count, the scale corresponding to the potential object is used. Normalized votes, the final vote count of the candidate reference point is given by the formula Calculate and record the corresponding entry in 2D-HCS (Hough accumulation space) .
[0025] In the above S430, false alarm removal and target determination, since the contour extraction result is easily interfered with and produces false alarms, the target detection problem is transformed into an analytical optimization problem, the causes of false alarms are analyzed, and a false alarm removal strategy is proposed. The target detection problem is equivalent to:
[0026] Defining the Match Ratio (MR) score , which is used to distinguish targets from false alarms caused by complex contour interference (ICC). For interference similar to the target part (IPS), the matching sparsity score (MS) is introduced: , The final decision criteria are formulated by combining the MR score, MS score and vote count constructed by distinguishing the target and IPS. , Represent the threshold of vote count, MR score and MS score respectively. If it is equal to 1, the target is considered to exist in the corresponding reference point; otherwise, the target is considered not to exist. Non-maximum suppression is used to delete reference points whose votes are not the maximum value of the HCS neighborhood. , the final reference point The position of the detected target is regarded as the target, and a bounding box of appropriate size is used to mark the detection result to complete the detection and positioning of the target in the remote sensing image.
[0027] Preferably, in the step S500, identifying the target model: S510, Classification Model Construction Based on Adaptive Heterogeneous Support Tensor Machine; S520, solving the model and determining the classification hyperplane; S530, target model identification and result output; In the above S510, in the construction of the classification model based on the adaptive heterogeneous support tensor machine, the optimization problem of the adaptive heterogeneous support tensor machine is constructed. First, the tensor mode is defined. Product, Tensor With the matrix of The pattern product is expressed as .
[0028] For the adaptive heterogeneous support tensor machine, the optimization problem is:
[0029] in, To control the parameters of different mode weights, are projection vectors of different modes of the tensor, is the relaxation factor, which is used to measure the classification error and ensure the feasibility of the optimization problem constraints. It is a trade-off coefficient used to balance the impact of the maximum classification interval and the number of violation points on classification.
[0030] In the above S520, model solution and classification hyperplane determination, the above optimization problem is solved by alternating iterative optimization technology. In each iteration, other parameters are fixed and one parameter is optimized. When fixed , transform the optimization problem into a standard quadratic programming problem. Construct the Lagrangian function:
[0031] in, is the Lagrangian multiplier, Apply KKT conditions:
[0032] Substituting the above results into the Lagrangian function, we get the Lagrangian dual problem. By solving the dual problem, we get , and then update . Continue to iterate until the algorithm converges, thereby determining the parameters of the classification hyperplane.
[0033] In the above S530, target model identification and result output, after obtaining the classification hyperplane parameters, the heterogeneous feature tensor of each detected target area is obtained. , and input it into the trained adaptive heterogeneous support tensor machine model for calculation. , the model will output a predicted label , the label corresponds to different target model categories. , representing the three target types of "aircraft", "ship" and "vehicle" respectively. When the calculation results When , the target is determined to be an "aircraft"; when When , it is determined to be a “vehicle”.
[0034] The present invention provides a remote sensing image target recognition method, which has the following beneficial effects: 1. By designing a feature extraction structure that collaborates with multi-branch convolution and dilated convolution, with different branches employing varying kernel sizes and dilation rates, we are able to comprehensively capture the feature information of remote sensing images at multiple scales. This multi-scale feature fusion strategy not only enriches the feature dimensionality but also enhances the ability to represent objects of varying sizes and shapes, significantly improving the precision and recall of object detection. Whether it's tiny buildings or large bodies of water, these objects can be more accurately detected.
[0035] 2. Constructing a heterogeneous feature tensor based on multi-scale fusion features organically integrates multiple sources of information, including spatial, spectral, and temporal (if available) information from remote sensing images. This tensor structure preserves the inherent relationships between features, providing a rich and organized data foundation for subsequent feature processing and analysis. This facilitates a deeper understanding of target characteristics, improves the ability to represent complex targets, and ultimately enhances the accuracy and reliability of target recognition and positioning.
[0036] 3. By combining state-dependent weight theory to calculate dynamic weights and dynamically adjusting weights based on the actual state of feature indicators, the model can rationally highlight the role of important features and suppress the influence of secondary features in different image scenes and target characteristics. This significantly improves the model's focus and adaptability to different features, enabling accurate target identification even in complex backgrounds, effectively enhancing the model's performance in various complex environments.
[0037] 4. Utilizing the tensor generalized Hough transform to construct a reference table, combined with a voting mechanism and false alarm removal strategy, we can achieve accurate detection and localization of targets in remote sensing images. Through precise contour matching and strict decision-making criteria, we effectively reduce the number of false alarms and improve target positioning accuracy, providing a reliable foundation for subsequent target analysis and applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0039] Figure 1A diagram of a remote sensing image target recognition method according to the present invention; Figure 2 This is a diagram of the multi-scale feature extraction and fusion framework of the present invention; The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0040] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions of the present invention. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0041] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the embodiments of the specification.
[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0043] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0044] The present invention provides a remote sensing image target recognition method, which comprises the following steps: S100, multi-scale feature extraction and fusion; S200, constructing heterogeneous feature tensors; S300, feature processing based on dynamic weight allocation; S400, target detection and positioning; S500, target model identification; As a preferred embodiment of the present invention, the multi-scale feature extraction and fusion step S100 includes the following steps: S110, multi-branch convolution and dilated convolution collaborative feature extraction; S120, feature fusion strategy strengthens feature expression; In the S110, multi-branch convolution and dilated convolution collaborative feature extraction, a multi-branch convolution structure is designed for the multi-scale characteristics of remote sensing images and combined with dilated convolution technology for feature extraction. Construct multiple convolution branches with different convolution kernel sizes and expansion rates. ;in: The first branch adopts The group convolution of the convolution kernel is calculated as follows: .in, Represents the number of groups. Through group convolution, the receptive field of the convolution kernel is expanded while the amount of convolution operations is reduced, thereby capturing a wider range of image information. It is a nonlinear activation function responsible for introducing nonlinear factors and enhancing the expressive power of the model; Represents the batch normalization operation, which is used to normalize the convolutional features, accelerate model convergence and improve training stability; Represents a convolution operation.
[0045] The second branch uses the expansion rate of Dilated convolution, the formula is By increasing the dilation rate, dilated convolution effectively expands the receptive field, capturing multi-scale contextual information, and reducing data dimensionality while reducing computational complexity. Although dilated convolution may cause some feature information loss, subsequent branches will compensate for this loss.
[0046] Continue to build other branches, such as the third branch adopts Group convolution of convolution kernels , and the expansion rate is of Dilated convolution as the fourth branch Etc. Different branches extract features from images at different scales to obtain rich feature information.
[0047] In the above-mentioned S120, the feature fusion strategy strengthens the feature expression, and the features extracted by each branch are fused to generate a feature vector with stronger semantic expression ability.
[0048] First, perform splicing operation on each branch feature, and use the following formula for splicing; , the features of different branches are connected along the channel dimension, so that the fused features cover multi-scale information and enrich the dimension of the features.
[0049] Then, the feature is further fused by element-by-element summation. Assuming the aggregation operation stands for element-wise sum, through the formula. Get the final fusion features .
[0050] As a preferred embodiment of the present invention, the step S200 of constructing a heterogeneous feature tensor includes the following steps: S210, determining tensor dimension based on multi-scale features; S220, tensor element assignment and structure construction; In the above S210, determining the tensor dimension based on multi-scale features, after completing the multi-scale feature extraction and fusion in S100, the fused feature F is obtained. Based on the different attributes and dimensional information of the fused features, the dimensional parameters of the heterogeneous feature tensor are determined. Since remote sensing images have many features, such as spatial location and spectral information, these different types of feature information are integrated. Assuming that the fused feature In the spatial dimension OK, Column, in the spectral dimension channels, taking into account the time series information (if any), assuming that time steps (if it is a static image, ). Then, the constructed heterogeneous feature tensor For one rank tensor( or , depending on whether the time dimension is included), its dimensions are represented as T∈ RHxWxCxT.
[0051] In the S220, tensor element assignment and structure construction, the multi-scale features extracted in S100 are filled into the heterogeneous feature tensor according to certain rules. For spatial location , in the spectral channel and time step According to the fusion features For example, if the fusion feature Central position ,aisle , time step The characteristic value of , then assign it to the heterogeneous feature tensor The corresponding element This approach fully integrates multi-scale features into heterogeneous feature tensors, preserves the spatial, spectral, and temporal (if any) relationships between features, fully utilizes the multi-source information of remote sensing images, and provides a rich data structure foundation for subsequent tensor-based processing and analysis. This enables tensors to better represent the complex features of targets in remote sensing images and improve the accuracy and reliability of target recognition.
[0052] As a preferred embodiment of the present invention, the feature processing based on dynamic weight allocation in S300 includes the following steps: S310, determining a constant weight vector; S320, introduce state variable weight theory to calculate dynamic weight; S330, calculating dynamic weights and applying them to feature processing; In the step S310, determining the constant weight vector, after constructing the heterogeneous feature tensor Finally, in order to measure the importance of each feature index in target recognition, it is necessary to determine the constant weight vector. Assuming that the heterogeneous feature tensor The feature dimension is , constant weight vector ,satisfy and The constant weight vector can be determined by using a ranking-based weighting algorithm, which is calculated based on the degree of correlation between features. Based on the contribution of each feature dimension to target recognition, a matrix reflecting the feature correlation relationship is constructed. For example, if there is feature dimensions, construct The correlation matrix ,in Indicates the The feature dimension and By processing the correlation matrix, such as calculating the probability transfer matrix, we can finally get the constant weight vector , so that the constant weight vector can reflect the relative importance of each feature dimension.
[0053] In step S320, the state variable weight theory is introduced to calculate the dynamic weight. In combination with the state variable weight theory, the constant weight is adjusted according to the actual state of the characteristic index to obtain the dynamic weight. Define the state variable weight vector ,in is a vector containing the states of each characteristic index. The state variable weight vector must meet certain conditions, such as for any ,have or ; For each variable Continuous and strictly monotonically decreasing or monotonically increasing; for any combination about The following state-variable weight function is used to calculate the state-variable weight vector:
[0054] in, , Indicates the negation level, Indicates passing level, Indicates the motivation level, Incentive regulation level is a heterogeneous feature tensor Middle The standardized state value of a characteristic index is obtained by mapping the characteristic index value to the interval [0,1]. For example, for a characteristic index , its standardized value , is the maximum and minimum value of the feature index in the data set.
[0055] The S330, calculates the dynamic weight and applies it to feature processing, through the constant weight vector and state variable weight vector The normalized product of gets the dynamic weight vector :
[0056] After obtaining the dynamic weight vector, apply it to the heterogeneous feature tensor For heterogeneous feature tensors Each feature dimension is weighted according to the dynamic weight. For example, for the feature dimension , the processed eigenvalues , is the original feature value. In this way, under different image scenes and target features, the weights can be dynamically adjusted according to the actual status of the feature indicators, more reasonably highlighting the role of important features, suppressing the influence of minor features, improving the model's attention and adaptability to different features, and providing more representative features for subsequent target detection and recognition.
[0057] As a preferred embodiment of the present invention, the target detection and positioning in S400 includes the following steps: S410, Fundamentals of target detection based on tensor generalized Hough transform; S420, Voting mechanism and target positioning of remote sensing images; S430, false alarm removal and target determination; In the target detection basis based on tensor generalized Hough transform (S410), after completing the feature processing based on dynamic weight allocation (S300), a weighted heterogeneous feature tensor is obtained. Using the tensor generalized Hough transform for target detection, a reference table represented by a tensor is first constructed to record the contour characteristics of the target at different orientations and scales. For a given template image, its wheel point information is extracted using an edge detection operator. Suppose the template image is , use edge detection operators (such as Canny operator) to get edge images ,in Indicates the location of the contour point in the template image .
[0058] Set the angle and scale interval to and , traverse the possible rotation angles and scale For the current and , calculate the new position of the contour point after rotation and scale change ,in is the position of the reference point, is the rotation matrix; at the same time, the new gradient angle of the contour point , and quantize the gradient angle as , is the gradient angle quantization interval.
[0059] Constructing a fifth-order tensor reference table , the first, second, third, fourth and fifth orders represent the horizontal space domain, vertical space domain, gradient angle domain, rotation angle domain and scale domain respectively. and scale , the contour point attributes obtained under the conditions , let the tensor reference the corresponding element in the table , complete the construction of the tensor reference table.
[0060] In the above S420, voting mechanism and target positioning of remote sensing images, for the remote sensing images to be detected , using a window of the same size as the template image to scan each position in the image , generating a series of slices For each slice, extract the gradient angle of its contour points And record the relevant attributes to construct a third-order tensor used to describe the slice outline information , where the first, second and third orders represent the horizontal spatial domain, vertical spatial domain and gradient angle domain respectively.
[0061] According to the tensor reference table constructed , perform rotation and scale-invariant contour matching, and calculate candidate reference points of remote sensing images If the number of votes and are all equal to 1, that is, there is a contour point with the same attributes in the template image and the slice, then the contour point (,,y,) in the slice will generate a vote for a specific rotation angle and scale. Candidate reference point The total number of votes received at a specific rotation angle and scale is expressed as To prevent the scale change from affecting the vote count, the scale corresponding to the potential object is used. Normalized votes, the final vote count of the candidate reference point is given by the formula Calculate and record the corresponding entry in 2D-HCS (Hough accumulation space) .
[0062] In the above S430, false alarm removal and target determination, since the contour extraction result is easily interfered with and produces false alarms, the target detection problem is transformed into an analytical optimization problem, the causes of false alarms are analyzed, and a false alarm removal strategy is proposed. The target detection problem is equivalent to:
[0063] Defining the Match Ratio (MR) score , which is used to distinguish targets from false alarms caused by complex contour interference (ICC). For interference similar to the target part (IPS), the matching sparsity score (MS) is introduced: , The final decision criteria are formulated by combining the MR score, MS score and vote count constructed by distinguishing the target and IPS. , They represent the thresholds of vote count, MR score and MS score respectively. If it is equal to 1, the target is considered to exist in the corresponding reference point; otherwise, the target is considered not to exist. Using non-maximum suppression, delete those reference points whose votes are not the maximum value of the HCS neighborhood. , the final reference point The position of the detected target is regarded as the target, and a bounding box of appropriate size is used to mark the detection result to complete the detection and positioning of the target in the remote sensing image.
[0064] As a preferred embodiment of the present invention, the S500, target model identification S510, Classification Model Construction Based on Adaptive Heterogeneous Support Tensor Machine; S520, solving the model and determining the classification hyperplane; S530, target model identification and result output; In the above S510, the classification model based on the adaptive heterogeneous support tensor machine is constructed. After completing the target detection and positioning in S400, the detected target area and related feature information are obtained. Based on the heterogeneous feature tensors of these target areas, the adaptive heterogeneous support tensor machine is used to identify the target model. Assume that the heterogeneous feature tensor sample set obtained from the detection result is , the corresponding sample label is ( is the number of categories of the target model).
[0065] Drawing on the idea of supporting tensor machines, we construct an optimization problem for adaptive heterogeneous supporting tensor machines. First, we define the tensor model Product, Tensor With the matrix of The pattern product is expressed as .
[0066] For the adaptive heterogeneous support tensor machine, the optimization problem is:
[0067] in, To control the parameters of different mode weights, are projection vectors of different modes of the tensor, is the relaxation factor, which is used to measure the classification error and ensure the feasibility of the optimization problem constraints. It is a trade-off coefficient used to balance the impact of the maximum classification interval and the number of violation points on classification.
[0068] In the above S520, model solution and classification hyperplane determination, the above optimization problem is solved by alternating iterative optimization technology. In each iteration, other parameters are fixed and one parameter is optimized. For example, in the optimization When fixed , transform the optimization problem into a standard quadratic programming problem. Construct the Lagrangian function:
[0069] in, is the Lagrangian multiplier, Apply KKT conditions:
[0070] Substituting the above results into the Lagrangian function, we get the Lagrangian dual problem. By solving the dual problem, we get , and then update . Continue to iterate until the algorithm converges, thereby determining the parameters of the classification hyperplane.
[0071] In the above S530, target model identification and result output, after obtaining the classification hyperplane parameters, the heterogeneous feature tensor of each detected target area is obtained. , and input it into the trained adaptive heterogeneous support tensor machine model for calculation. , the model will output a predicted label , the label corresponds to different target model categories.
[0072] For example, if , representing the three target types of "aircraft", "ship" and "vehicle" respectively. When the calculation results When , the target is determined to be an "aircraft"; when When , it is determined to be a “vehicle”.
[0073] For multi-target detection, the feature tensor of each detected target region is processed one by one to obtain the corresponding target model identification results. These results are then presented visually, such as by drawing a bounding box around each detected target on the remote sensing image and annotating the identified target model next to the bounding box. A text report can also be generated, detailing each target's location coordinates, detection score, and corresponding target model for easy user review and analysis. Furthermore, the recognition results can be stored for further processing and analysis, such as performing target model statistical analysis on remote sensing images over a period of time.
[0074] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for remote sensing image target recognition, comprising the following steps: S100, multi-scale feature extraction and fusion; S200, constructing heterogeneous feature tensors; S300, feature processing based on dynamic weight allocation; S400, target detection and positioning; S500, target model identification.
2. The remote sensing image target recognition method according to claim 1, characterized in that: The multi-scale feature extraction and fusion step S100 includes the following steps: S110, multi-branch convolution and dilated convolution collaborative feature extraction; S120, feature fusion strategy strengthens feature expression; In the S110, multi-branch convolution and dilated convolution collaborative feature extraction, multiple convolution branches with different convolution kernel sizes and expansion rates are constructed. ;in: The first branch adopts The group convolution of the convolution kernel is calculated as follows: in, Represents the number of groups. Through group convolution, the convolution kernel receptive field is expanded while the amount of convolution operations is reduced, thereby capturing a wider range of image information. It is a nonlinear activation function responsible for introducing nonlinear factors and enhancing the expressive power of the model; Represents the batch normalization operation, which is used to normalize the convolutional features, accelerate model convergence and improve training stability; Represents the convolution operation; The second branch uses the expansion rate of The formula for dilated convolution is: Continue to build other branches, such as the third branch adopts Group convolution of convolution kernels , and the expansion rate is of Dilated convolution as the fourth branch ; In the above S120, in the feature fusion strategy to enhance the feature expression, each branch feature is first spliced using the following formula: , connect the features of different branches along the channel dimension, so that the fused features cover multi-scale information and enrich the dimension of the features. Then, further fuse the features by element-by-element summation. Set the aggregation operation Represents element-by-element summation, through the formula, Get the final fusion features .
3. The remote sensing image target recognition method according to claim 1, characterized in that: The step S200, constructing a heterogeneous feature tensor, includes the following steps: S210, determining tensor dimension based on multi-scale features; S220, tensor element assignment and structure construction; In the S210, in determining the tensor dimension based on multi-scale features, it is assumed that the fusion feature In the spatial dimension OK, Column, in the spectral dimension channels, taking into account the time series information (if any), with time steps (if it is a static image, ), then the constructed heterogeneous feature tensor For one rank tensor( or , depending on whether the time dimension is included), its dimension is expressed as T∈RHxWxCxT; In the S220, tensor element assignment and structure construction, the multi-scale features extracted in S100 are filled into the heterogeneous feature tensor according to certain rules. For spatial position , in the spectral channel and time step According to the fusion features The corresponding information is assigned, if the fusion feature Central position ,aisle , time step The characteristic value of , then assign it to the heterogeneous feature tensor The corresponding element .
4. The remote sensing image target recognition method according to claim 1, characterized in that: The feature processing based on dynamic weight allocation in S300 includes the following steps: S310, determining a constant weight vector; S320, introduce state variable weight theory to calculate dynamic weight; S330, calculating dynamic weights and applying them to feature processing; In the step S310, determining the constant weight vector, after constructing the heterogeneous feature tensor Then, let the heterogeneous feature tensor The feature dimension is , constant weight vector ,satisfy and ; In step S320, the state variable weight theory is introduced to calculate the dynamic weight, and the state variable weight vector is defined. ,in is a vector containing the states of each characteristic index. The state variable weight vector must meet certain conditions, such as for any ,have or ; For each variable Continuous and strictly monotonically decreasing or monotonically increasing; for any combination about Non-decreasing or non-increasing, the following state-variable weight function is used to calculate the state-variable weight vector: in, , Indicates the negation level, Indicates passing level, Indicates the motivation level, Incentive regulation level is a heterogeneous feature tensor Middle The standardized state value of each characteristic indicator; The S330, calculates the dynamic weight and applies it to feature processing, through the constant weight vector and state variable weight vector The normalized product of gets the dynamic weight vector : After obtaining the dynamic weight vector, apply it to the heterogeneous feature tensor In the feature processing of Each feature dimension is weighted according to the dynamic weight. , the processed eigenvalues , is the original eigenvalue.
5. The remote sensing image target recognition method according to claim 1, characterized in that: The target detection and positioning step S400 includes the following steps: S410, Fundamentals of target detection based on tensor generalized Hough transform; S420, Voting mechanism and target positioning of remote sensing images; S430, false alarm removal and target determination; In the target detection basis based on tensor generalized Hough transform in S410, the template image is assumed to be , use edge detection operators (such as Canny operator) to get edge images ,in Indicates the location of the contour point in the template image ; Set the angle and scale interval to and , traverse the possible rotation angles and scale , for the current and , calculate the new position of the contour point after rotation and scale change ,in is the position of the reference point, is the rotation matrix; at the same time, the new gradient angle of the contour point , and quantize the gradient angle as , is the gradient angle quantization interval; Constructing a fifth-order tensor reference table , the first, second, third, fourth and fifth orders represent the horizontal space domain, vertical space domain, gradient angle domain, rotation angle domain and scale domain respectively. and scale , the contour point attributes obtained under the conditions , let the tensor reference the corresponding element in the table , complete the construction of the tensor reference table; In the above S420, voting mechanism and target positioning of remote sensing images, for the remote sensing images to be detected , using a window of the same size as the template image to scan each position in the image , generating a series of slices , for each slice, extract the gradient angle of its contour points And record the relevant attributes to construct a third-order tensor used to describe the slice outline information , where the first, second and third orders represent the horizontal spatial domain, vertical spatial domain and gradient angle domain respectively; According to the tensor reference table constructed , perform rotation and scale-invariant contour matching, and calculate candidate reference points of remote sensing images If the number of votes and are all equal to 1, that is, there is a contour point with the same attributes in the template image and the slice, then the contour point (,,y,) in the slice will generate a vote for a specific rotation angle and scale, the candidate reference point The total number of votes received at a specific rotation angle and scale is expressed as To prevent the scale change from affecting the vote count, the scale corresponding to the potential object is used. Normalized votes, the final vote count of the candidate reference point is given by the formula Calculate and record the corresponding entry in 2D-HCS (Hough accumulation space) ; In the above S430, false alarm removal and target determination, since the contour extraction result is susceptible to interference and false alarms are generated, the target detection problem is converted into an analytical optimization problem, the causes of false alarms are analyzed and a false alarm removal strategy is proposed. The target detection problem is equivalent to: Defining the Match Ratio (MR) score , which is used to distinguish targets from false alarms caused by complex contour interference (ICC). For interference similar to the target part (IPS), the matching sparsity score (MS) is introduced: The final decision criteria are formulated by combining the MR score, MS score and vote count constructed by distinguishing the target and IPS. , Represent the threshold of vote number, MR score and MS score respectively. If If it is equal to 1, the target is considered to exist in the corresponding reference point; otherwise, the target is considered not to exist; Use non-maximum suppression to delete reference points whose votes are not the maximum value of the HCS neighborhood , the final reference point The position of the detected target is regarded as the target, and a bounding box of appropriate size is used to mark the detection result to complete the detection and positioning of the target in the remote sensing image.
6. The remote sensing image target recognition method according to claim 1, characterized in that: In the step S500, target model identification: S510, Classification Model Construction Based on Adaptive Heterogeneous Support Tensor Machine; S520, solving the model and determining the classification hyperplane; S530, target model identification and result output; In the S510, in the construction of the classification model based on the adaptive heterogeneous support tensor machine, an optimization problem of the adaptive heterogeneous support tensor machine is constructed; First define the tensor mode Product, Tensor With the matrix of The pattern product is expressed as ; For the adaptive heterogeneous support tensor machine, the optimization problem is: in, To control the parameters of different mode weights, are projection vectors of different modes of the tensor, is the relaxation factor, which is used to measure the classification error and ensure the feasibility of the optimization problem constraints. It is a trade-off coefficient used to balance the impact of the maximum classification interval and the number of illegal points on classification; In said S520, model solving and classification hyperplane determination, the above optimization problem is solved by an alternating iterative optimization technique; In each iteration, one of the parameters is optimized while the other parameters are fixed. When fixed , transform the optimization problem into a standard quadratic programming problem and construct the Lagrangian function: in, is the Lagrangian multiplier, Apply KKT conditions: Substituting the above results into the Lagrangian function, we get the Lagrangian dual problem. By solving the dual problem, we get , and then update , iterate continuously until the algorithm converges, thereby determining the parameters of the classification hyperplane; In the above S530, target model identification and result output, after obtaining the classification hyperplane parameters, the heterogeneous feature tensor of each detected target area is obtained. , input it into the trained adaptive heterogeneous support tensor machine model for calculation; By formula , the model will output a predicted label , the label corresponds to different target model categories, if , representing three target types: "aircraft", "ship" and "vehicle". When the calculation results When , the target is determined to be "aircraft"; when When , it is determined to be a "vehicle".