Image processing method and device, electronic equipment, readable storage medium and chip
By extracting and fusing target features and global local features in three-dimensional situational images, and generating feature descriptors in combination with orthogonal components, the problems of long image processing cycles and inaccurate extraction of key information are solved, and fast and accurate analysis and identification support is achieved.
Patent Information
- Application Number
- CN202411950780.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems in image processing with long image processing cycles and insufficient extraction of key information, especially when processing a large number of images, it is difficult to complete analysis and recognition in a short time.
By acquiring the detection area image and the target image, the first feature structure extraction model and the second feature structure extraction model respectively extract the target feature and the local global feature, and combine the orthogonal component and the feature descriptor to generate the feature descriptor for analysis and recognition.
It realizes rapid analysis and identification of key emergency areas in three-dimensional situation images. The generated feature descriptors are highly specific and robust, and can process a large number of images in a short time, supporting command decision-making.
Smart Images

Figure CN120047654A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image processing method, device, electronic device, readable storage medium and chip. Background Art
[0002] With the rapid development of artificial intelligence and image processing technology, it is particularly important to analyze and identify key emergency areas in the reconnaissance three-dimensional situation image data, which is of great value and significance for command decision-making. In the existing technical solutions, by obtaining the global image and local image of the reconnaissance area, the deep learning network is used to extract features from the global image and local image, and the extracted features are fused and processed to complete the key recognition and analysis of the image. However, in the existing technical solutions, there are problems such as the extraction of key information is not accurate enough, and when a large number of images need to be processed, it cannot be completed in a short time. Summary of the invention
[0003] In view of this, the present invention aims to solve the problems of long image processing cycle and inaccurate extraction of key information.
[0004] Specifically, the present invention is achieved through the following technical solutions:
[0005] A first aspect of the present invention provides an image processing method.
[0006] A second aspect of the present invention provides a device.
[0007] A third aspect of the present invention provides an electronic device.
[0008] A fourth aspect of the present invention provides a readable storage medium.
[0009] A fifth aspect of the present invention provides a chip.
[0010] The present invention provides an image processing method, comprising: acquiring an image of a detection area and a target image in the detection area image, wherein the target image includes at least one target feature; determining a first parameter of the target feature based on a first feature structure extraction model; determining a local feature of the detection area image and a second parameter corresponding to the local feature based on a second feature structure extraction model, wherein the local feature includes the target feature; determining a global feature of the detection area image and a third parameter corresponding to the global feature based on the second feature structure extraction model, wherein the global feature includes the local feature; determining a first orthogonal component based on the first parameter and the third parameter; determining a second orthogonal component based on the second parameter and the third parameter; determining a feature descriptor corresponding to the target feature based on the first orthogonal component, the second orthogonal component and the third parameter; marking the detection area image according to the feature descriptor to generate an image to be analyzed.
[0011] In some technical solutions, optionally, determining the first parameter of the target feature based on the first feature structure extraction model specifically includes: determining the target feature map according to the first feature extraction network corresponding to the first feature structure extraction model; determining the first spatial attention map according to the target feature map; determining the first parameter of the target feature according to the first spatial attention map.
[0012] In some technical schemes, optionally, a first spatial attention map is determined based on the target feature map, specifically including: processing the target feature map based on a convolution operation to determine the first feature map; processing the target feature map based on average pooling to determine the second feature map; performing a feature dimensionality reduction operation on the first feature map and the second feature map to determine a reduced dimensionality image; and determining the first spatial attention map based on the reduced dimensionality image.
[0013] In some technical solutions, optionally, local features of the detection area image and second parameters corresponding to the local features are determined based on a second feature structure extraction model, and the local features include target features. Specifically, the method includes: determining a local feature map according to a second feature extraction network corresponding to the second feature structure extraction model; determining a second spatial attention map according to the local feature map; and determining a second parameter corresponding to the local feature according to the second spatial attention map.
[0014] In some technical solutions, optionally, the global features of the detection area image and the third parameters corresponding to the global features are determined based on the second feature structure extraction model, and the global features include local features, specifically including: determining a global feature map according to a third feature extraction network corresponding to the second feature structure extraction model; determining a third spatial attention map according to the global feature map; determining a third parameter corresponding to the global feature according to the third spatial attention map.
[0015] In some technical solutions, optionally, determining the first orthogonal component according to the first parameter and the third parameter specifically includes: determining the first projection parameter according to the first parameter and the third parameter; determining the first orthogonal component according to the first parameter and the first projection parameter.
[0016] In some technical solutions, optionally, determining the second orthogonal component according to the second parameter and the third parameter specifically includes: determining the second projection parameter according to the second parameter and the third parameter; determining the second orthogonal component according to the second parameter and the second projection parameter.
[0017] In some technical schemes, optionally, a feature descriptor corresponding to the target feature is determined based on the first orthogonal component, the second orthogonal component and the third parameter, specifically including: determining a first orthogonal component vector based on the first orthogonal component; determining a second orthogonal component vector based on the second orthogonal component; determining a feature descriptor corresponding to the target feature based on the first orthogonal component vector, the second orthogonal component vector and the third parameter.
[0018] In some technical solutions, optionally, obtaining a reconnaissance area image and a target image within the reconnaissance area image specifically includes: obtaining image information to be identified; determining the reconnaissance area image in the image information to be identified; determining at least one target feature within the reconnaissance area image; and determining a target image corresponding to the at least one target feature.
[0019] The second aspect of the present invention provides a device, which includes: an acquisition module: used to acquire a detection area image and a target image in the detection area image, the target image including at least one target feature; a first extraction module: used to determine a first parameter of the target feature based on a first feature structure extraction model; a second extraction module: used to determine a local feature of the detection area image and a second parameter corresponding to the local feature based on a second feature structure extraction model, the local feature including the target feature; a third extraction module: used to determine a global feature of the detection area image and a third parameter corresponding to the global feature based on the second feature structure extraction model, the global feature including the local feature; a first calculation module: used to determine a first orthogonal component according to the first parameter and a third parameter; a second calculation module: used to determine a second orthogonal component according to the second parameter and the third parameter; a processing module: used to determine a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter; a marking module: used to mark the detection area image according to the feature descriptor to generate an image to be analyzed; the device implements the steps of the image processing method of the first aspect of the present invention.
[0020] The third aspect of the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the image processing method of the first aspect of the present invention are implemented.
[0021] A fourth aspect of the present invention provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the image processing method of the first aspect of the present invention are implemented.
[0022] A fifth aspect of the present invention provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the image processing method of the first aspect of the present invention.
[0023] The technical solution provided by the present invention brings at least the following beneficial effects:
[0024] The present invention proposes an image processing method, which extracts and learns features of the detected three-dimensional situation area and key targets therein based on fusion contrast learning, fully mines the key high- and low-dimensional feature information therein, and adopts the feature orthogonality idea to enhance the discriminability of the features learned by the model, removes the redundant components between the two, focuses on key information, and forms a feature descriptor with high specificity and strong robustness for analysis and identification of key emergency areas. It can complete important emergency analysis and identification of a large number of reconnaissance areas in a short time, and plays an important role in command decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0027] Figure 1 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0029] Figure 3 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0031] Figure 5 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0032] Figure 6 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0033] Figure 7 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0034] Figure 8 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0035] Fig. 9 A schematic diagram of a flow chart of an image processing method provided by an embodiment of the present invention;
[0036] Fig.10 A schematic diagram of the structure of a device provided by an embodiment of the present invention;
[0037] Fig.11 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0038] Fig.12 A schematic diagram of the structure of an image processing model provided by an embodiment of the present invention;
[0039] Fig.13 A schematic diagram of the structure of a first characteristic structure extraction model provided by an embodiment of the present invention;
[0040] Fig.14 A schematic diagram of the structure of a self-attention module provided in an embodiment of the present invention;
[0041] Fig.15 A schematic diagram of the structure of a second feature structure extraction model provided by an embodiment of the present invention;
[0042] Fig.16 A schematic diagram of the structure of a feature descriptor calculation module provided in an embodiment of the present invention.
[0043] in, Fig.10 and Fig.11 The corresponding relationship between the component names and numbers in is as follows:
[0044] 200 device, 201 acquisition module, 202 first extraction module, 203 second extraction module, 204 third extraction module, 205 first calculation module, 206 second calculation module, 207 processing module, 208 marking module, 300 electronic device, 302 processor, 304 memory. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] See also Figure 1 The first aspect of the present invention provides an image processing method, the image processing method comprising:
[0047] S102: Acquire a detection area image and a target image within the detection area image, wherein the target image includes at least one target feature;
[0048] S104: Determine a first parameter of the target feature based on the first feature structure extraction model;
[0049] S106: determining local features of the detection area image and second parameters corresponding to the local features based on the second feature structure extraction model, where the local features include target features;
[0050] S108: determining the global features of the detection area image and a third parameter corresponding to the global features based on the second feature structure extraction model, where the global features include local features;
[0051] S110: determining a first orthogonal component according to the first parameter and the third parameter;
[0052] S112: determining a second orthogonal component according to the second parameter and the third parameter;
[0053] S114: determining a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter;
[0054] S116: Marking the detection area image according to the feature descriptor to generate an image to be analyzed.
[0055] The image processing method provided in this embodiment first needs to obtain the detection area image and the target image in the detection area image through a device before processing the image. For example, the image can be obtained through a camera, a still camera and other devices. This is the first step of image processing, and the image information to be processed is obtained, wherein the target image includes one or more target features, such as the size, shape, and form of the target image, which is helpful for the subsequent processing of the target image and the identification and judgment of whether the target image is a key area in the regional image. The image processing mainly extracts features from the image through two models, including a first feature structure extraction model and a second feature structure extraction model. The target features of the target image are extracted according to the first feature structure extraction model, and the detailed information in the target image is captured using a relatively shallow network structure to obtain a first parameter corresponding to the target feature. The first parameter includes a target feature point, which is used to supplement and enhance the global features of the detection area later. According to the second feature structure extraction model, local features in the detection area image are extracted. Similarly, a relatively shallow network is used to capture local image information in the detection area image, and local features in the local image are extracted to obtain a second parameter corresponding to the local feature. The second parameter includes local feature points. Here, the local feature includes the target feature, which facilitates the enhanced comparison of the target feature with the global feature in the detection area image, and helps to analyze and identify the key emergency area of the detection area image through the orthogonal fusion strategy. According to the second feature structure extraction model, the global feature in the detection area image is extracted. A relatively deep network is used to capture the global image information in the detection area image, and the global feature in the global image is extracted to obtain a third parameter corresponding to the global feature. The third parameter includes global feature points. Here, the global feature includes the local feature, that is, the global feature also includes the target feature, which facilitates the feature fusion of the target feature and the local feature with the global feature, thereby generating a feature descriptor for determining whether the target image is a key area of the detection area image. It can be understood that the premise of generating the feature descriptor is to compare and learn the target feature with the global feature, and compare and learn the local feature with the global feature, and to obtain it by fusing the results of the above two comparison learning. Specifically, the projection parameter of the first parameter on the third parameter is first calculated through the projection process expression, that is, the first projection vector is obtained according to the target feature point and the global feature point through the projection process expression, and then the first projection vector is subtracted from the target feature point, and the first orthogonal component is obtained after calculation. Similarly, the projection parameter of the second parameter on the third parameter is calculated through the projection process expression, that is, the second projection vector is obtained according to the local feature point and the global feature point through the projection process expression, and then the second projection vector is subtracted from the local feature point, and the second orthogonal component is obtained after calculation. After the above process, two orthogonal components can be obtained, and each point on them is orthogonal to the global feature point of the reconnaissance area.Furthermore, the first and second orthogonal components are compressed and simplified through the pooling layer to obtain two orthogonal component vectors, which are then fused and spliced with the third parameter to generate feature descriptors for key emergency area analysis and identification. Finally, the target image in the detection area is marked according to the feature descriptor to generate an image to be analyzed for key identification, so that the system can determine whether the target area is an emergency key area.
[0056] It can be understood that the image processing method of the present application is to first obtain a detection area image and a target image in the detection area image, perform target feature extraction on the target image through a first feature structure extraction model, and determine a first parameter, the first parameter includes a target feature point, perform local feature extraction on the detection area image through a second feature structure extraction model, and determine a second parameter, the second parameter includes local feature points, and the local feature includes the target feature, perform global feature extraction on the detection area image through the second feature structure extraction model, and determine a third parameter, the third parameter includes global feature points, and the global feature includes local features, and then orthogonally fuse the first parameter, the second parameter, and the third parameter to generate feature description points to determine whether the target image is key area information in the regional image.
[0057] In simple terms, the present invention designs a deep learning model around the perspective of effectively orthogonally fusing multiple features of the reconnaissance area image and the target image therein to improve the feature discrimination ability. A fusion contrast learning network based on the orthogonal fusion of the global three-dimensional reconnaissance situation image and the typical target features therein is proposed. The model captures the information of typical targets and the information of the reconnaissance area, fully extracts the key high and low dimensional feature information therein, and removes the redundant components between the two through the proposed feature orthogonal method, focusing on the key information, and finally generating a feature descriptor for analyzing and identifying whether the area is a key emergency area. The feature descriptor generated by this design scheme has the characteristics of high specificity and strong robustness, which improves the pertinence of image processing and the accuracy of target selection, and can complete important emergency analysis and identification of a large number of reconnaissance areas in a short time, playing an important role in command decision-making.
[0058] In some embodiments, optionally, Figure 2 As shown, determining the first parameter of the target feature based on the first feature structure extraction model specifically includes:
[0059] S1042: determining a target feature map according to a first feature extraction network corresponding to the first feature structure extraction model;
[0060] S1044: Determine a first spatial attention map according to the target feature map;
[0061] S1046: Determine a first parameter of the target feature according to the first spatial attention map.
[0062] In this embodiment, the first feature structure extraction model includes a first feature extraction network and a first local enhancement network, wherein the first local enhancement network includes a multi-scale convolution module and a self-attention module. First, the target image is input into the first feature structure extraction model, and the target image is feature extracted through the first feature extraction network to generate a target feature map. The generated target feature map is then input into the first local enhancement network, and the multi-scale convolution module is used to obtain feature maps of different spatial receptive fields, that is, multi-scale feature maps, and downsampled feature maps, that is, feature maps after reducing the image size and retaining important information. Next, these feature maps are connected and feature dimension reduction operations are performed through the convolution layer. The dimension reduction operation can reduce the data dimension and reduce the algorithm calculation complexity. Then, the feature map output by the dimension reduction operation is passed to the self-attention module to further model the importance of each typical target feature point. Finally, the first spatial attention map is obtained after processing by the self-attention module. The first parameter of the target feature is obtained by subsequent modulation of the first spatial attention map.
[0063] The first feature extraction network is a deep learning network consisting of three dense blocks, each of which consists of a convolutional module consisting of multiple batch normalization layers, activation layers, and convolutional layers. The input of each convolutional module is the union of the outputs of all previous convolutional layers, and the features learned in this layer will be directly passed to all subsequent layers as input. Each dense block is connected to the subsequent dense blocks via a transition block, which consists of a batch normalization layer, a ReLU activation layer, a convolutional layer, and a pooling layer.
[0064] In a specific embodiment, in order to enhance the effectiveness of the features extracted by the first feature extraction network, the present invention also uses a supervised contrast learning strategy to constrain the feature learning process. t (·) to get the feature map r t , through the projection module H t (·) The typical target feature r t Mapped to the feature vector z t . A nonlinear multilayer perceptron is used to implement the projection module to improve the feature quality. Next, the feature vector z t Regularization is performed so that the inner product distance measurement is used. Finally, supervised contrastive learning loss is used on the regularized feature vector.
[0065] like Figure 3As shown, in some embodiments, optionally, determining a first spatial attention map according to the target feature map specifically includes:
[0066] S10442: Processing the target feature map based on a convolution operation to determine a first feature map;
[0067] S10444: Processing the target feature map based on average pooling to determine a second feature map;
[0068] S10446: Performing a feature dimension reduction operation on the first feature map and the second feature map to determine a dimension reduction image;
[0069] S10448: Determine a first spatial attention map based on the dimensionality reduction image.
[0070] In this embodiment, the target feature map output by the first feature extraction network is input into the first local enhancement network for processing. The first local enhancement network includes a multi-scale convolution module and a self-attention module, wherein the multi-scale convolution module includes three hole convolution layers and an average pooling. When the target feature map is input into the first local enhancement network, the target feature map is firstly processed by three hole convolution layers through convolution operation to obtain a first feature map. Specifically, different expansion rates are set by hole convolution to obtain feature maps of different receptive fields, i.e., the first feature map. The target feature map is processed by the average pooling branch to obtain a second feature map. Specifically, the target feature map is processed by average pooling to obtain a downsampled feature map, i.e., the second feature map. Next, the first feature map and the second feature map are connected and combined, and a feature dimension reduction operation is performed through a 1×1 convolution layer to obtain a reduced dimension image. The dimension reduction operation can reduce the data dimension and reduce the algorithm calculation complexity. Then, the reduced dimension image output by the dimension reduction operation is passed to the self-attention module to further model the importance of each typical target feature point. In the self-attention module, the input feature map is first processed using a 1×1 convolution layer and a batch normalization layer, and then a spatial attention map is generated by passing through a 1×1 convolution layer.
[0071] In summary, by adopting the above method and utilizing the first feature structure extraction model, the present invention can extract target features in the target image, thus preparing for the subsequent determination of feature descriptors through orthogonal fusion.
[0072] like Figure 4 As shown, in some embodiments, optionally, local features of the detection area image and second parameters corresponding to the local features are determined based on the second feature structure extraction model, and the local features include target features, specifically including:
[0073] S1062: Determine a local feature map according to a second feature extraction network corresponding to the second feature structure extraction model;
[0074] S1064: Determine a second spatial attention map according to the local feature map;
[0075] S1066: Determine a second parameter corresponding to the local feature according to the second spatial attention map.
[0076] In this embodiment, the second feature structure extraction model includes a second feature extraction network and a second local enhancement network. First, the detection area image is subjected to local feature extraction by the second feature extraction network in the second feature structure extraction model to obtain a local feature map, which is used to enhance and compare the global features of the local features in the detection area image. Here, the second feature extraction network is composed of a structure of three dense blocks. Then, the local feature map is input into the second local enhancement network to further extract local features. Here, the second local enhancement network is also composed of a multi-scale convolution module and a self-attention module. The local feature map is subjected to the second local enhancement module and the multi-scale convolution module is used to enhance the local features of the local features. Block, obtain feature maps of different spatial receptive fields, that is, multi-scale feature maps, and downsampled feature maps, that is, feature maps after reducing the image size and retaining important information. Next, connect these feature maps and perform feature dimension reduction operation through the convolution layer. The dimension reduction operation can reduce the data dimension and reduce the algorithm calculation complexity. Then, the feature map output by the dimension reduction operation is passed to the self-attention module to further model the importance of each typical target feature point. Finally, the second spatial attention map is obtained after processing by the self-attention module. The second parameter of the target feature is obtained by normalizing and modulating the subsequent features of the second spatial attention map.
[0077] In a specific embodiment, in the process of learning local features in the reconnaissance area, in order to enhance the effectiveness of the features extracted by the second feature extraction network, the present invention also uses a supervised contrast learning strategy to constrain the feature learning process. l Feed through projection module H lp (·) to get the eigenvector z lp Next, we regularize the feature vector so that the inner product can be used for distance measurement. Then, we apply supervised contrastive learning loss on the regularized feature vector.
[0078] In summary, by adopting the above method and utilizing the second feature structure extraction model, the present invention can extract local features in the detection area image, and prepare for the subsequent determination of feature descriptors through orthogonal fusion.
[0079] like Figure 5 As shown, in some embodiments, optionally, the global features of the detection area image and the third parameter corresponding to the global features are determined based on the second feature structure extraction model, and the global features include local features, specifically including:
[0080] S1082: Determine a global feature map according to a third feature extraction network corresponding to the second feature structure extraction model;
[0081] S1084: Determine a third spatial attention map according to the global feature map;
[0082] S1086: Determine a third parameter corresponding to the global feature according to the third spatial attention map.
[0083] In this embodiment, the second feature structure extraction model includes a third feature extraction network and a global module network. First, the detection area image is subjected to global feature extraction through the third feature extraction network in the second feature structure extraction model to obtain a global feature map for orthogonal fusion with local features and target features. Here, the third feature extraction network is composed of a dense block structure. Then, the global feature map is input into the global module network to further extract global features. Here, the global module network is composed of four dense blocks and a self-attention module. The input global feature map is processed using four dense blocks. Next, these processed feature maps are connected and subjected to feature dimension reduction operation through a convolution layer. The dimension reduction operation can reduce the data dimension and reduce the algorithm calculation complexity. Then, the feature map output by the dimension reduction operation is passed to the self-attention module to further model the importance of each typical target feature point. Unlike the self-attention module in the second local enhancement network, the self-attention module in the global module network replaces the last global average pool of the network with a gem pooling layer. Finally, the third spatial attention map is obtained after processing by the self-attention module. The third parameter of the target feature is obtained by normalizing and modulating the subsequent features of the third spatial attention map.
[0084] In a specific embodiment, during the global feature learning process of the reconnaissance area, supervised contrastive learning loss is also used to constrain it, and the above feature r lg Feed through projection module H lg (·) to get the eigenvector z lg ,Next, we regularize the feature vector and apply supervised ,contrastive learning loss on the regularized feature vector.
[0085] like Figure 6 As shown, in some embodiments, optionally, determining the first orthogonal component according to the first parameter and the third parameter specifically includes:
[0086] S1102: Determine a first projection parameter according to the first parameter and the third parameter;
[0087] S1104: Determine a first orthogonal component according to the first parameter and the first projection parameter.
[0088] In this embodiment, the first parameter includes the target feature point, that is, The third parameter includes the global feature point, which is r lg , we set the target feature point and the global feature point r lg Substitute the projection process expression into the projection process to obtain the first projection parameter, which is The expression of the projection process is as follows:
[0089]
[0090] Among them, ab, c are the indexes of feature points.
[0091] In the above formula, the symbol represents the dot product operation, |r lg | 2 Yes lg The l2 norm of :
[0092]
[0093] Then, the first parameter is subtracted from the first projection parameter, and the orthogonal component of the target feature point on the global feature point is obtained by the calculation method shown in the following formula, which is the first orthogonal component
[0094]
[0095] like Figure 7 As shown, in some embodiments, optionally, determining the second orthogonal component according to the second parameter and the third parameter specifically includes:
[0096] S1122: Determine a second projection parameter according to the second parameter and the third parameter;
[0097] S1124: Determine a second orthogonal component according to the second parameter and the second projection parameter.
[0098] In this embodiment, the second parameter includes local feature points, namely The third parameter includes the global feature point, which is r lg , we will local feature points and the global feature point r lg Substitute the projection process expression to obtain the second projection parameter,
[0099] That is The expression of the projection process is as follows:
[0100]
[0101] Among them, i, j, k are the indices of feature points.
[0102] In the above formula, the symbol represents the dot product operation, |r lg | 2 Yes lg The l2 norm of :
[0103]
[0104] Then the second parameter is subtracted from the second projection parameter, and the orthogonal component of the local feature point on the global feature point is obtained by the calculation method shown in the following formula, which is the second orthogonal component
[0105]
[0106] like Figure 8 As shown, in some embodiments, optionally, determining a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter specifically includes:
[0107] S1142: Determine a first orthogonal component vector according to the first orthogonal component;
[0108] S1144: Determine a second orthogonal component vector according to the second orthogonal component;
[0109] S1146: Determine a feature descriptor corresponding to the target feature according to the first orthogonal component vector, the second orthogonal component vector and the third parameter.
[0110] In this embodiment, each point on the first orthogonal component is orthogonal to the global feature point of the detection area, and each point on the second orthogonal component is orthogonal to the global feature point of the detection area. The first orthogonal component is then processed and aggregated using a pooling layer to generate a first orthogonal component vector. Similarly, the second orthogonal component is processed and aggregated using a pooling layer to generate a second orthogonal component vector. The first orthogonal component vector and the second orthogonal component vector generated by the above process are then feature fused and spliced with the third parameter to generate a feature descriptor that is ultimately used for key emergency area analysis and identification.
[0111] like Fig. 9 As shown, in some embodiments, optionally, acquiring the detection area image and the target image in the detection area image specifically includes:
[0112] S1022: Obtaining image information to be identified;
[0113] S1024: Determine the detection area image in the image information to be identified;
[0114] S1026: Determine at least one target feature in the detection area image;
[0115] S1028: Determine a target image corresponding to at least one target feature.
[0116] In this embodiment, it is first necessary to obtain the information of the image to be identified, that is, the original image, which is used to mark the identified image to determine whether it is a key area. For example, the acquisition method can be through a device such as a camera or a camera, and then the image to be identified is processed by cropping, locking, etc. to determine the detection area image, wherein the detection area image includes a target image, and the target image includes one or more target features. The target features can be features such as size and shape. This is the premise of the image processing method of the present invention. Feature extraction is performed through the target image and the detection area image, and finally a feature descriptor is generated for the system to determine whether the target image is a key emergency area.
[0117] like Fig.10 As shown, the second aspect of the present invention provides a device 200, which includes: an acquisition module 201: used to acquire an image of a detection area and a target image in the detection area image, the target image including at least one target feature; a first extraction module 202: used to determine a first parameter of the target feature based on a first feature structure extraction model; a second extraction module 203: used to determine a local feature of the detection area image and a second parameter corresponding to the local feature based on a second feature structure extraction model, the local feature including the target feature; a third extraction module 204: used to determine a global feature of the detection area image and a third parameter corresponding to the global feature based on the second feature structure extraction model, the global feature including the local feature; a first calculation module 205: used to determine a first orthogonal component according to the first parameter and the third parameter; a second calculation module 206: used to determine a second orthogonal component according to the second parameter and the third parameter; a processing module 207: used to determine a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter; a marking module 208: used to mark the detection area image according to the feature descriptor to generate an image to be analyzed. The device implements the steps of the image processing method of the first aspect of the present invention.
[0118] In this embodiment, the second aspect of the present invention provides a device 200, which includes: an acquisition module 201, a first extraction module 202, a second extraction module 203, a third extraction module 204, a first calculation module 205, a second calculation module 206, a processing module 207 and a marking module 208. Specifically, the acquisition module 201 is used to acquire the detection area image and the target image in the detection area image, wherein the target image includes at least one target feature; the first extraction module 202 is used to determine the first parameter of the target feature based on the first feature structure extraction model; the second extraction module 203 is used to determine the local feature of the detection area image and the second parameter corresponding to the local feature based on the second feature structure extraction model, wherein the local feature includes the target feature; the third extraction module 204 is used to determine the global feature of the detection area image and the third parameter corresponding to the global feature based on the second feature structure extraction model, wherein the global feature includes the local feature; the first calculation module 205 is used to determine the first orthogonal component according to the first parameter and the third parameter; the second calculation module 206 is used to determine the second orthogonal component according to the second parameter and the third parameter; the processing module 207 is used to determine the feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter; the marking module 208 is used to mark the detection area image according to the feature descriptor to generate an image to be analyzed.
[0119] In some embodiments, optionally, the device is also used to: determine a target feature map based on a first feature extraction network corresponding to the first feature structure extraction model; determine a first spatial attention map based on the target feature map; and determine a first parameter of the target feature based on the first spatial attention map.
[0120] In some embodiments, optionally, the device is also used to: process the target feature map based on a convolution operation to determine a first feature map; process the target feature map based on average pooling to determine a second feature map; perform feature dimensionality reduction operations on the first feature map and the second feature map to determine a reduced dimensionality image; and determine a first spatial attention map based on the reduced dimensionality image.
[0121] In some embodiments, optionally, the device is also used to: determine a local feature map based on a second feature extraction network corresponding to the second feature structure extraction model; determine a second spatial attention map based on the local feature map; and determine a second parameter corresponding to the local feature based on the second spatial attention map.
[0122] In some embodiments, optionally, the device is also used to: determine a global feature map based on a third feature extraction network corresponding to the second feature structure extraction model; determine a third spatial attention map based on the global feature map; and determine a third parameter corresponding to the global feature based on the third spatial attention map.
[0123] In some embodiments, optionally, the device is further used to: determine a first projection parameter according to the first parameter and the third parameter; and determine a first orthogonal component according to the first parameter and the first projection parameter.
[0124] In some embodiments, optionally, the device is also used to: determine a second projection parameter based on the second parameter and the third parameter; determine a second orthogonal component based on the second parameter and the second projection parameter.
[0125] In some embodiments, optionally, the device is also used to: determine a first orthogonal component vector based on the first orthogonal component; determine a second orthogonal component vector based on the second orthogonal component; and determine a feature descriptor corresponding to the target feature based on the first orthogonal component vector, the second orthogonal component vector and a third parameter.
[0126] In some embodiments, optionally, the device is also used to: obtain image information to be identified; determine a detection area image in the image information to be identified; determine at least one target feature in the detection area image; and determine a target image corresponding to the at least one target feature.
[0127] By extracting and learning the features of the detected three-dimensional situation area and the key targets therein based on fusion contrast learning, the key high- and low-dimensional feature information is fully mined, and the feature orthogonality idea is used to enhance the discriminability of the features learned by the model, remove the redundant components between the two, focus on key information, and form a highly specific and robust feature descriptor for analysis and identification of key emergency areas. This can complete important emergency analysis and identification of a large number of reconnaissance areas in a short period of time, which plays an important role in command decision-making.
[0128] like Fig.11 As shown, the third aspect of the present invention provides an electronic device 300, including a processor 302, a memory 304, and a program or instruction stored in the memory 304 and executable on the processor 302. When the program or instruction is executed by the processor 302, the steps of the image processing method of the first aspect of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0129] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned electronic devices and non-electronic devices.
[0130] The fourth aspect of the present invention provides a readable storage medium, on which programs or instructions are stored. When the programs or instructions are executed by a processor, the steps of the image processing method of the first aspect of the present invention are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.
[0131] The methods may be implemented in a variety of different ways depending on the specific features and / or example applications. For example, the methods may be implemented by a combination of hardware, firmware, and / or software. For example, in a hardware implementation, the processor may be implemented in one or more application-specific integrated circuits, digital signal processors, digital signal processing devices, programmable logic devices, field programmable gate arrays, controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the above functions, and / or combinations thereof.
[0132] A computer readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer readable storage medium may be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above devices, but is not limited thereto. A non-exhaustive list of more specific examples of computer readable storage media includes: a portable computer floppy disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, a static random access memory, a portable optical disk read-only memory, a digital universal disk, a memory card, a floppy disk, an encoding mechanical device (such as a punched card or a groove with a raised structure having instructions recorded thereon), and any suitable combination of the above devices. The computer readable storage medium used herein should not be understood as a transmission signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium, or an electrical signal transmitted through a wire, etc.
[0133] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0134] The fifth aspect of the present invention provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the image processing method of the first aspect of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0135] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0136] In a specific embodiment, the present invention provides an image processing method, and the related technical background includes feature projection technology and contrast learning technology. Among them:
[0137] Feature learning is an important step and one of the advantages of deep neural networks. Learning efficient representations or features from raw data is the key to the excellent performance of deep learning networks in various tasks. In order to further improve feature learning and make the features learned by the model more discriminative for classification tasks, the feature projection method was first proposed in natural language classification tasks. Its core idea is to map the features obtained by traditional feature extractors into a space orthogonal to common features, thereby obtaining orthogonal features that are more discriminative for classification.
[0138] Contrastive Learning (CL) has made great progress in learning effective features. The basic idea is to learn a hidden feature space and maximize the consistency between different enhanced views of the same image by comparing the consistency between different images in it. Self-supervised contrastive learning and supervised contrastive learning usually adopt a two-stage learning method. The first stage uses contrastive learning loss to learn features, and the second stage uses cross entropy loss to learn classifiers.
[0139] The specific process of the image processing method includes: a deep learning model is designed from the perspective of effectively orthogonally fusing multiple features of the three-dimensional reconnaissance situation image and the typical target image to improve the feature discrimination ability. A fusion contrast learning network based on the orthogonal fusion of the global three-dimensional reconnaissance situation image and the typical target features is proposed. The model captures the information of the typical target and the information of the reconnaissance area, and removes the redundant components between the two through the proposed feature orthogonal method, focusing on the key information, and finally generating a feature descriptor for analyzing and identifying whether the area is a key emergency area.
[0140] The deep learning network model proposed in the present invention is mainly composed of three parts: a typical target branch, a reconnaissance area branch and an orthogonal fusion module. The first two respectively perform effective feature extraction and learning on typical targets and three-dimensional reconnaissance areas, and the orthogonal fusion module effectively fuses these features to generate a feature identifier that is ultimately used for analysis and identification.
[0141] The typical target branch is mainly used to extract local features of typical targets, and a supervised contrastive learning strategy is used to effectively constrain the feature learning stage.
[0142] The reconnaissance region branch mainly extracts and learns the local features and global features of the three-dimensional reconnaissance region. In order to enhance the learning of local features, the proposed model is equipped with a local enhancement module composed of multi-channel convolution and self-attention mechanism to extract representative local features.
[0143] The above feature extraction processes all adopt a supervised contrastive learning strategy, and the learned features are further enhanced under the constraint of service contract loss (SCLoss) to be more discriminative.
[0144] Next, the two local features generated based on the above process are orthogonally fused with the global features respectively. Subsequently, the two orthogonal components and the global features are connected as complementary parts to form a feature descriptor for analysis and identification of key emergency areas.
[0145] In a specific embodiment, Fig.12 As shown in the figure, the image processing method process is as follows: first obtain the detection area image (i.e., ROI2), then obtain the target image (i.e., ROI1) from the detection area image, use the first feature extraction backbone network (i.e., Encoder1) to extract features of the target image to obtain a target feature map (i.e., Feature), then perform feature learning (i.e., SC Learning) on the target feature map, and at the same time input the target feature map into the local enhancement module (i.e., LocalEnhancementModule) for further feature extraction to determine the final target feature parameters (i.e., ROI1 Feature). The network structure of reconnaissance region feature extraction consists of a local branch and a global branch. The reconnaissance region image (i.e., ROI2) is input into the second feature extraction backbone network (i.e., Encoder2) to perform local feature extraction and global feature extraction on the reconnaissance region image, and generate local feature maps and global feature maps. The local feature map enters the local branch network, and then performs feature learning (i.e., SC Learning) on the local feature map. At the same time, the local feature map is input into the local enhancement module (i.e., Local Enhancement Module) for further feature extraction to determine the final local feature parameters (i.e., ROI2Local Feature). The global feature map enters the global module (i.e., Global Module), and then performs feature learning (i.e., SC Learning) on the global feature map. At the same time, the global feature map is further extracted to determine the final global feature parameters (i.e., ROI2Global Feature). Then, the target feature parameters (i.e., ROI1 Feature), local feature parameters (i.e., ROI2 Local Feature) and global feature parameters (i.e., ROI2 Global Feature) are input into the orthogonal fusion module (i.e., Orthogonal Fusion Module) for orthogonal fusion to generate feature descriptors (i.e., FeatureDescriptors), which are used to determine whether the target image is an emergency key area (i.e., Predictionof region Status) on the regional image.
[0146] In a specific embodiment, the typical target feature extraction and learning process is as follows:
[0147] The present invention designs a relatively shallow network structure to extract typical target features and capture detailed information in typical targets, which is used for subsequent supplementary enhancement of the global features of the reconnaissance situation and the construction of feature descriptors for key emergency area analysis and identification through an orthogonal fusion strategy.
[0148] The feature extraction and learning part of typical targets consists of two parts: the feature extraction backbone network and the local enhancement module. For the input ROI1, the feature extraction backbone network is first used for feature extraction, and then the multi-scale feature analysis and importance analysis are performed through the multi-scale convolution module and the self-attention module in the local enhancement module. The detailed structure of the network used for typical target feature extraction is shown in the figure below.
[0149] The feature extraction backbone network is a deep learning network consisting of three dense blocks, each of which consists of a convolutional module consisting of multiple batch normalization layers, activation layers, and convolutional layers. The input of each convolutional module is the union of the outputs of all previous convolutional layers, and the features learned in this layer will be directly passed to all subsequent layers as input. Each dense block is connected to the subsequent dense blocks via a transition block, which consists of a batch normalization layer, a ReLU activation layer, a convolutional layer, and a pooling layer.
[0150] In order to enhance the efficiency of the features extracted by the feature extraction backbone network, this chapter also uses the supervised contrast learning strategy to constrain the feature learning process. t (·) to get the feature map r t , through the projection module H t (·) The typical target feature r t Mapped to the feature vector z t . A nonlinear multilayer perceptron is used to implement the projection module to improve the feature quality. Next, the feature vector z t L2 regularization is performed so that the inner product distance measurement can be used. Finally, supervised contrastive learning loss is used on the regularized feature vector. The supervised contrastive learning loss used for local feature learning in a typical target region is shown in the following formula, where the temperature coefficient T is set to 0.1.
[0151]
[0152] Where L SCt (z i ) is the supervised contrast learning loss for local feature learning in the typical target region ROI1, z i is the eigenvector.ti is the anchor sample feature vector, Yes ti The set of all positive sample feature vectors, Yes ti The number of all positive samples. Indicates z tj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z ti The feature vector of the positive sample; z tk is the number of batches except z ti All sample feature vectors except itself. i, j, k are subscripts.
[0153] The local enhancement module consists of a multi-scale convolution module and a self-attention module. The former simulates the feature pyramid of different image scale changes based on dilated convolution, and the latter is used for importance modeling. Dilated convolution is a special convolution operation that expands the convolution kernel by adding some spaces between the convolution kernel elements. Downsampling is usually used to expand the receptive field and reduce the amount of calculation in deep neural networks, but the downsampling process will lose some input information and sacrifice spatial resolution to a certain extent; while dilated convolution can maintain a higher resolution while increasing the receptive field. The difference between dilated convolution and general convolution lies in the dilation rate. When different dilation rates are set for dilated convolution, the receptive field of the resulting feature map will be different. Based on this feature of dilated convolution, it can be used to capture multi-scale feature information.
[0154] In a specific embodiment, Fig.13 As shown in the figure, the detection area image is first obtained, and then the target image (i.e., ROI1) is obtained from the detection area image. The first feature extraction backbone network (i.e., Encoderl) is used to extract features of the target image to obtain a target feature map (i.e., Feature). Then, feature learning (i.e., SC Learning) is performed on the target feature map. At the same time, the target feature map is further extracted through a multiscale convolution module (i.e., Multiscale Convolution) and a self-attention module (i.e., Attention Module) to determine the final target feature parameters (i.e., ROI1 Feature).
[0155] The multi-scale convolution module designed in the present invention includes three hole convolution layers and one average pooling branch, and its structural schematic diagram is shown in the figure below.
[0156] Three of the dilated convolutional layers are used to obtain feature maps with different spatial receptive fields, and the average pooling branch is used to obtain the feature map after average pooling downsampling. Next, these features are connected and passed through a 1×1 convolutional layer for feature dimension reduction. Then, the feature map output by the dimension reduction operation is passed to the self-attention module to further model the importance of each typical target feature point. In the self-attention module, the input feature map is first processed using a 1×1 convolutional layer and a batch normalization layer, and then the subsequent features are normalized and modulated by generating a spatial attention map through a 1×1 convolutional layer, followed by a SoftPlus operation. The calculation formula is shown below.
[0157] Softplus(x)=log(1+e x );
[0158] Here, x is the input and e is a natural constant.
[0159] In a specific embodiment, Fig.14 As shown in the figure, specifically, the target feature map is input into the multiscale convolution module (Multiscale Convolution), and the feature maps of different spatial receptive fields are obtained through the three dilated convolution layers (i.e., DilatedConvolution) in the multiscale convolution module. The downsampled feature map is obtained through the average pooling branch. The average pooling branch includes the pooling layer (i.e., Pooling Layer), the convolution layer (i.e., Conv Layer), and the activation layer (i.e., ReLU). Then, the feature maps of different receptive fields and the downsampled feature maps are sent to the connection module (Concatenate) for connection, and feature dimension reduction operation is performed through a 1×1 convolution layer (i.e., Conv Layer) and an activation layer (i.e., ReLU). Then, the feature map output by the dimension reduction operation is passed to the self-attention module (i.e., Attention Module) to further model the importance of each typical target feature point. In the self-attention module, the input feature map is first processed using a 1×1 convolution layer (i.e., ConvLayer), a batch normalization layer (i.e., BN Layer), and an activation layer (i.e., ReLU). Then, the target feature parameters are finally determined by multiplication (i.e., Multiplication) through a 1×1 convolution layer (i.e., Conv Layer), a soft plus activation layer (i.e., SoftPlus), and an L2 norm (i.e., L2 Norm).
[0160] In another specific embodiment, the lung feature extraction and learning process is as follows:
[0161] For the entire reconnaissance area, the analysis area is larger and contains richer information, but also more noise information. Therefore, for the reconnaissance area, the present invention extracts two features, local and global, which complement and enhance each other to achieve the purpose of more comprehensive analysis of the reconnaissance area. The local features come from the shallower network, and the global features come from the deeper network.
[0162] The network structure of reconnaissance area feature extraction consists of a local branch and a global branch. The input reconnaissance area 3D R0I2 is processed by the feature extraction backbone network G l (·) Get the feature map r l , where the network G l (·) is composed of three dense blocks. In the process of local feature extraction, this feature map r l It is transmitted to the local enhancement module to obtain the local features of the reconnaissance area. The local enhancement module here is also composed of a multi-scale convolution module based on dilated convolution and a self-attention module for importance modeling. l After the above processing, the local features r of the reconnaissance area are generated. lp In the process of learning local features in the reconnaissance area, the present invention uses supervised contrastive learning loss to constrain it, and transforms the above feature map r l Feed through projection module H lp (·) to get the eigenvector z lp Next, we perform L2 regularization on the feature vector so that the inner product can be used for distance measurement. Then, supervised contrastive learning loss is applied on the regularized feature vector as shown below, where the temperature coefficient τ is set to 0.1.
[0163]
[0164] Where L sc_lp (z i ) is the supervised contrastive learning loss for local feature learning in the reconnaissance area, z i is the eigenvector. lpi is the anchor sample feature vector, Yes lpi The set of all positive sample feature vectors, Yes lpi The number of all positive samples. Indicates z lPj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z lpi The feature vector of the positive sample; z lpk is the number of batches except z lpi All sample feature vectors except itself. i, j, k are subscripts.
[0165] For the global branch, we will pass the feature extraction backbone network G l (·) Get the feature map r l After being sent to the fourth dense block and processed, the global feature map r of the reconnaissance area is finally obtained. lg , the whole global feature process is similar to the structure of densely connected network (DenseNet), where four dense blocks process the input data. The difference is that the global average pooling at the end of the network is replaced by GeM pooling. Specifically, let us represent the output feature map of dense block (dense block4) as
[0166] In the process of learning the global features of the reconnaissance area, supervised contrast learning loss is also used to constrain it. lg Feed through projection module H lg (·) to get the eigenvector z lg Next, we perform L2 regularization on the feature vector and apply supervised contrastive learning loss on the regularized feature vector, as shown in the following formula:
[0167]
[0168] Where L sc_lg (z i ) is the supervised contrastive learning loss for global feature learning in the reconnaissance region, z i is the eigenvector. lgi is the anchor sample feature vector, Yes lgi The set of all positive sample feature vectors, Yes lgi The number of all positive samples. Indicates z lgj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z lgi The feature vector of the positive sample; z lgk is the number of batches except z lgi All sample feature vectors except itself. i, j, k are subscripts.
[0169] In a specific embodiment, Fig.15As shown in the figure, the network structure of reconnaissance area feature extraction consists of a local branch and a global branch. The reconnaissance area image (i.e., ROI2) is input into the second feature extraction backbone network (i.e., Encoder2) to perform local feature extraction and global feature extraction on the reconnaissance area image to generate local feature maps and global feature maps. The local feature map enters the local branch network, and then feature learning (i.e., SCLearning) is performed on the local feature map. At the same time, the local feature map is input into the multiscale convolution module (i.e., Multiscale Convolution) and the self-attention module (i.e., Attention Module) for further feature extraction to determine the final local feature parameters (i.e., ROI2 LocalFeature). The global feature map enters the global module (i.e., Global Module) for further feature extraction. The global module (i.e., Global Module) includes four dense blocks (i.e., Dense Block4), batch normalization layer (i.e., BN Layer), activation layer (i.e., ReLU) and gem pooling layer (i.e., GeM Pooling Layer). Then, feature learning (i.e., SC Learning) is performed on the global feature map processed by the global module, and at the same time, the final global feature parameters (i.e., ROI2 Global Feature) are determined.
[0170] like Fig.16 As shown, this module focuses on the target feature r t , local features of the reconnaissance area t lp and the global feature r of the reconnaissance area lg As input, then calculate r t 、r lp Each local feature point To the lung global feature r lg Projection characteristics The expression of the projection process is as follows:
[0171]
[0172] Among them, a, b, c, i, j, k are the indices of feature points.
[0173] In the above formula, the symbol represents the dot product operation, |r lg | 2 Yes lg The l2 norm of :
[0174]
[0175] Here, c is the index and C is the number of dimensions in that dimension.
[0176] The orthogonal component is the key target feature r t , local features of the reconnaissance area r lP Its projection vector The difference between the key target feature r is obtained by the calculation method shown in the following formula t , local features of the reconnaissance area r lp Orthogonal components, here the orthogonal components of the key target features and the global features of the reconnaissance area are recorded as The orthogonal components of the local features of the reconnaissance area and the global features of the reconnaissance area are recorded as
[0177]
[0178] After the above process, two tensors can be obtained, and each point on them is related to the global feature r of the reconnaissance area. lg Orthogonal. Then the tensor is aggregated into a vector through the pooling layer, and finally two orthogonal component vectors are generated. The two orthogonal vectors generated above are combined with r lg The splicing is performed to generate the final feature descriptor used for analysis and identification of key emergency areas.
[0179] The present invention analyzes and identifies key emergency areas in the detected three-dimensional situation image data, which has important value and significance for command decision-making. The present invention extracts and learns the detected three-dimensional situation area and the key targets therein based on fusion contrast learning, fully mines the key high- and low-dimensional feature information therein, and adopts the feature orthogonal idea to enhance the discriminability of the features learned by the model, removes the redundant components between the two, focuses on key information, and forms a highly specific and robust feature descriptor for analysis and identification of key emergency areas. It can complete the important emergency analysis and identification of a large number of reconnaissance areas in a short time, which plays an important role in command decision-making.
[0180] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although the features may work as above in certain combinations and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.
[0181] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0182] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
[0183] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0184] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An image processing method, characterized in that: include: Acquire a detection area image and a target image within the detection area image, wherein the target image includes at least one target feature; Determine a first parameter of the target feature based on a first feature structure extraction model; Determine local features of the detection area image and second parameters corresponding to the local features based on a second feature structure extraction model, wherein the local features include the target features; Determine, based on the second feature structure extraction model, a global feature of the detection area image and a third parameter corresponding to the global feature, wherein the global feature includes the local feature; Determine a first orthogonal component according to the first parameter and the third parameter; Determine a second orthogonal component according to the second parameter and the third parameter; Determining a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter; The detection area image is marked according to the feature descriptor to generate an image to be analyzed.
2. The image processing method according to claim 1, characterized in that: The step of determining the first parameter of the target feature based on the first feature structure extraction model specifically includes: Determining a target feature map according to a first feature extraction network corresponding to the first feature structure extraction model; Determining a first spatial attention map according to the target feature map; According to the first spatial attention map, a first parameter of the target feature is determined.
3. The image processing method according to claim 2, characterized in that: Determining a first spatial attention map according to the target feature map specifically includes: Processing the target feature map based on a convolution operation to determine a first feature map; Processing the target feature map based on average pooling to determine a second feature map; Performing a feature dimensionality reduction operation on the first feature map and the second feature map to determine a dimensionality reduction image; According to the dimension-reduced image, a first spatial attention map is determined.
4. The image processing method according to claim 1, characterized in that: The determining of the local features of the detection area image and the second parameter corresponding to the local features based on the second feature structure extraction model, wherein the local features include the target features, specifically includes: Determining a local feature map according to a second feature extraction network corresponding to the second feature structure extraction model; Determining a second spatial attention map according to the local feature map; According to the second spatial attention map, a second parameter corresponding to the local feature is determined.
5. The image processing method according to claim 1, characterized in that: The determining of the global feature of the detection area image and the third parameter corresponding to the global feature based on the second feature structure extraction model, wherein the global feature includes the local feature, specifically includes: Determining a global feature map according to a third feature extraction network corresponding to the second feature structure extraction model; Determining a third spatial attention map according to the global feature map; According to the third spatial attention map, a third parameter corresponding to the global feature is determined.
6. The image processing method according to claim 1, characterized in that: The determining the first orthogonal component according to the first parameter and the third parameter specifically includes: Determine a first projection parameter according to the first parameter and the third parameter; The first orthogonal component is determined according to the first parameter and the first projection parameter.
7. The image processing method according to claim 1, characterized in that: The determining the second orthogonal component according to the second parameter and the third parameter specifically includes: Determine a second projection parameter according to the second parameter and the third parameter; The second orthogonal component is determined according to the second parameter and the second projection parameter.
8. The image processing method according to claim 1, characterized in that: The determining, according to the first orthogonal component, the second orthogonal component and the third parameter, a feature descriptor corresponding to the target feature specifically includes: Determine a first orthogonal component vector according to the first orthogonal component; Determine a second orthogonal component vector according to the second orthogonal component; A feature descriptor corresponding to the target feature is determined according to the first orthogonal component vector, the second orthogonal component vector and the third parameter.
9. The image processing method according to claim 1, characterized in that: The acquiring of the detection area image and the target image within the detection area image specifically includes: Obtaining image information to be identified; Determining the detection area image in the image information to be identified; Determining at least one target feature within the detection area image; A target image corresponding to at least one of the target features is determined.
10. A device, characterized in that: The device comprises: Acquisition module: used for acquiring an image of a detection area and an image of a target in the image of the detection area, wherein the image of the target includes at least one target feature; A first extraction module: used for determining a first parameter of the target feature based on a first feature structure extraction model; A second extraction module: used for determining local features of the detection area image and second parameters corresponding to the local features based on a second feature structure extraction model, wherein the local features include the target features; A third extraction module: used for determining the global feature of the detection area image and a third parameter corresponding to the global feature based on the second feature structure extraction model, wherein the global feature includes the local feature; A first calculation module: used for determining a first orthogonal component according to the first parameter and the third parameter; A second calculation module: used for determining a second orthogonal component according to the second parameter and the third parameter; A processing module: used for determining a feature descriptor corresponding to the target feature according to the first orthogonal component, the second orthogonal component and the third parameter; Marking module: used to mark the detection area image according to the feature descriptor to generate an image to be analyzed; The device implements the steps of the image processing method according to any one of claims 1 to 9.
11. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the image processing method as claimed in any one of claims 1 to 9.
12. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the image processing method according to any one of claims 1 to 9 are implemented.
13. A chip, characterized in that: The chip includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or an instruction to implement the steps of the image processing method according to any one of claims 1 to 9.