Image Processing Method, Apparatus, Computer Device, and Storage Medium
By introducing encoding modules, in-segment measurement modules and linear layers into the graph neural network, the image is extracted and classified, which solves the problem that traditional convolution cannot efficiently capture image features and improves the accuracy of image classification.
Patent Information
- Application Number
- CN202111248774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-10-26
AI Technical Summary
Traditional convolution filters images with single repetition, and cannot efficiently capture features with obvious differences, affecting the accuracy of image classification.
Graph neural network is adopted, including feature extraction network and feature classification network, and image features are classified encoding, iterative graph convolution processing and classification processing through encoding modules, in-segment measurement modules and linear layers.
The depth and aggregation ability of image feature extraction are improved, and the accuracy of image classification results is enhanced.
Smart Images

Figure CN113989523B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the continuous development of deep learning, it has strongly promoted the development of many fields such as image processing, computer vision, natural language processing, and machine translation, which fully demonstrates the powerful potential of deep learning.
[0003] For image classification problems, traditional convolutions are usually used to filter images in a single and repetitive manner to extract features. However, this method cannot efficiently capture significantly different features, which is not conducive to feature classification after feature extraction and is likely to affect the accuracy of classification results. Summary of the Invention
[0004] To solve the above technical problems, the present application provides an image processing method, apparatus, computer device, and storage medium.
[0005] In a first aspect, the present application provides an image processing method applied to a graph neural network including a feature extraction network and a feature classification network. The feature classification network includes an encoding module, a plurality of intra-segment metric modules, and a linear layer. The method includes:
[0006] Input a plurality of sample images into the feature extraction network, and perform feature extraction processing on the sample images through the feature extraction network to obtain image features corresponding to the sample images; wherein, the plurality of sample images include known-class images and unknown-class images;
[0007] Input the plurality of image features into the encoding module for class encoding to obtain a plurality of encoding features corresponding to the image features;
[0008] Input the plurality of encoding features into the plurality of intra-segment metric modules in sequence for iterative graph convolution processing to obtain a plurality of convolution features corresponding to the encoding features;
[0009] Input the plurality of convolution features into the linear layer for classification processing to obtain corresponding classification results; wherein, the classification results include the classes of the unknown-class images.
[0010] In a second aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0011] Input multiple sample images into the feature extraction network, and perform feature extraction processing on the sample images through the feature extraction network to obtain the image features corresponding to the sample images; wherein, the multiple sample images include known-class images and unknown-class images;
[0012] Input multiple of the image features into the encoding module for class encoding to obtain multiple encoding features corresponding to the image features;
[0013] Input the multiple encoding features into multiple of the intra-segment metric modules in sequence for iterative graph convolution processing to obtain multiple convolution features corresponding to the encoding features;
[0014] Input the multiple convolution features into the linear layer for classification processing to obtain corresponding classification results; wherein, the classification results include the classes of the unknown-class images.
[0015] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0016] Input multiple sample images into the feature extraction network, and perform feature extraction processing on the sample images through the feature extraction network to obtain the image features corresponding to the sample images; wherein, the multiple sample images include known-class images and unknown-class images;
[0017] Input multiple of the image features into the encoding module for class encoding to obtain multiple encoding features corresponding to the image features;
[0018] Input the multiple encoding features into multiple of the intra-segment metric modules in sequence for iterative graph convolution processing to obtain multiple convolution features corresponding to the encoding features;
[0019] Input the multiple convolution features into the linear layer for classification processing to obtain corresponding classification results; wherein, the classification results include the classes of the unknown-class images.
[0020] Based on the above image processing method, multiple sample images are input into the feature extraction network, and the feature extraction network performs feature extraction processing on the sample images to obtain the image features corresponding to the sample images. Among them, the multiple sample images include known-class images and unknown-class images, and the known-class images are used to identify the unknown-class images. The multiple image features are input into the encoding module for class encoding to obtain multiple encoding features corresponding to the image features. The multiple encoding features are sequentially input into multiple in-segment metric modules for iterative graph convolution processing to obtain multiple convolution features corresponding to the encoding features. By performing iterative graph convolution processing on the encoding features through multiple in-segment metric modules, the depth of the graph neural network is increased and the convolution aggregation ability is improved. The multiple convolution features are input into the linear layer for classification processing, and the classification result of the unknown-class image can be obtained. On the basis of increasing the depth and aggregation ability of the graph neural network, the accuracy of the classification result is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic structural diagram of a graph neural network in an embodiment;
[0024] Figure 2 It is a schematic flowchart of an image processing method in an embodiment;
[0025] Figure 3 It is a schematic structural diagram of a feature classification network in an embodiment;
[0026] Figure 4 It is a schematic structural diagram of an adjacency perception network in an embodiment;
[0027] Figure 5 It is a schematic structural diagram of an in-segment metric module in an embodiment;
[0028] Figure 6 It is a schematic structural diagram of a feature extraction network in an embodiment;
[0029] Figure 7 It is a structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0031] In one embodiment, Figure 1 is a schematic structural diagram of a graph neural network in one embodiment. Refer to Figure 1 , this image processing method is applied to a graph neural network including a feature extraction network and a feature classification network. Figure 1 PMGCN in Figure 2 is used to indicate the graph neural network, the feature extractor is used to indicate the feature extraction network, and the classifier is used to indicate the feature classification network. The feature classification network includes an encoding module, multiple intra-segment metric modules, and a linear layer. Figure 2 is a schematic flowchart of an image processing method in one embodiment. Refer to Figure 1 , an image processing method is provided. This embodiment mainly takes the application of this method to the graph neural network in the above
[0032] Step S210: Input multiple sample images into the feature extraction network, and perform feature extraction processing on the sample images through the feature extraction network to obtain the image features corresponding to the sample images.
[0033] Specifically, the multiple sample images include known-class images and unknown-class images. A known-class image refers to an image whose class has been recognized, and an unknown-class image refers to an image whose class has not been recognized. The known-class images are used to identify the class attributes of the unknown-class images. The corresponding known-class attributes are determined according to the known-class images. The class attribute of an unknown-class image is one of multiple known-class attributes. The class attribute can specifically be at least one or more of shape, symbol, animal, person, car, etc. As Figure 1 shown, the known-class attributes are respectively triangle, trapezoid, and circle. The class attribute of any unknown-class image is one of triangle, trapezoid, and circle. Feature extraction is performed on each sample image to obtain the image features, and each sample image corresponds to one image feature.
[0034] Step S220: Input the multiple image features into the encoding module for class encoding to obtain multiple encoded features corresponding to the image features.
[0035] Specifically, the encoding module in the feature classification network performs category encoding on each image feature, and splices each image feature with its corresponding category encoding to obtain encoded features corresponding to the image features. As Figure 3 shown, the image feature is represented by X f , the category encoding is represented by X l , and the encoded feature after splicing the image feature and the category encoding is represented by X f-l .
[0036] Step S230: Sequentially input the multiple encoded features into the multiple intra-segment metric modules for iterative graph convolution processing to obtain multiple convolution features corresponding to the encoded features.
[0037] Specifically, the number of intra-segment metric modules can perform iterative graph convolution processing of the encoded features through multiple intra-segment metric modules, increasing the learning depth of the graph neural network, improving the aggregation ability for the encoded features, and facilitating the improvement of the accuracy of subsequent feature classification.
[0038] Step S240: Input the multiple convolution features into the linear layer for classification processing to obtain corresponding classification results; wherein, the classification results include the category of the unknown category image.
[0039] Specifically, input the convolution features after multiple iterative graph convolution processes into the linear layer, calculate the probability that the convolution feature belongs to each known category attribute, and use the known category attribute with the highest probability as the classification result corresponding to the convolution feature. In this way, determine the category attribute of the unknown category image, Figure 3 where P q is used to indicate the classification result.
[0040] In one embodiment, the intra-segment metric module includes an adjacency perception network and a graph convolution network. The step of sequentially inputting the multiple encoded features into the multiple intra-segment metric modules for iterative graph convolution processing to obtain multiple convolution features corresponding to the encoded features includes: inputting the multiple encoded features into the adjacency perception network to determine the adjacency relationship between the encoded features and generate an adjacency matrix; inputting the multiple encoded features and the adjacency matrix into the graph convolution network for graph convolution processing to obtain multiple segmented features corresponding to the encoded features; and sequentially inputting the multiple segmented features into the multiple intra-segment metric modules for iterative graph convolution processing to obtain the convolution features.
[0041] Specifically, the adjacency matrix is used to represent the adjacency relationship between pairwise coding features among multiple coding features. Based on the adjacency relationship between each coding feature, the matching degree between the known-class image and the unknown-class image can be determined, thereby determining the class attribute of the unknown-class image. Each intra-segment metric module includes an adjacency perception network and a graph convolutional network. As Figure 3 shown, the input of each non-first intra-segment metric module is the output of the previous intra-segment metric module. The same operations are performed within each intra-segment metric module, that is, the adjacency perception module in each intra-segment metric module will generate a corresponding adjacency matrix according to the input, and the graph convolutional network in each intra-segment metric module will perform graph convolutional processing on the input and the corresponding adjacency matrix. That is to say, the adjacency matrix will be updated once for each passing of an intra-segment metric module. The number of intra-segment metric modules is n. The larger n is, the deeper the depth of the graph neural network and the stronger the aggregation ability, but the more likely the over-smoothing problem occurs. Therefore, for different graph convolutional tasks, custom adjustments are made according to the corresponding depth requirements and over-smoothing requirements.
[0042] In one embodiment, the adjacency perception network includes a learning network, a residual network, and an output layer. The step of inputting the multiple coding features into the adjacency perception network to determine the adjacency relationship between each coding feature and generate an adjacency matrix includes: inputting multiple absolute values of feature differences into the learning network to learn the adjacency relationship between the two coding features corresponding to the absolute values of feature differences, and obtaining multiple candidate matrices for indicating the adjacency relationship; inputting the multiple absolute values of feature differences into the residual network to obtain multiple feedback matrices corresponding to the absolute values of feature differences; inputting the sum of the multiple candidate matrices and the corresponding feedback matrices into the output layer to generate the adjacency matrix.
[0043] Specifically, within the adjacency perception network, first calculate the absolute value of the feature difference between pairwise coding features. The absolute value of the feature difference is used to indicate the absolute value of the difference between the corresponding numerical values of any two coding features. Then, input the absolute value of the feature difference into the input layer and the hidden layer in the learning network successively. As Figure 4 shown, X fin is used to indicate a coding feature, X fini and X finj respectively correspond to two different coding features, FC 1 is the input layer, FC 2 -FC K is the hidden layer. After multiple-layer learning in the learning network, a candidate matrix for indicating the adjacency relationship between pairwise coding features is obtained. As Figure 4 shown in, at FC KThe multiple vertically arranged points afterwards are the candidate matrix. To avoid overfitting of the learning network, the absolute value of the feature difference also needs to be input into the residual network for learning to obtain a feedback matrix that also indicates the adjacency relationship between pairwise encoded features, such as Figure 4 in FC r The multiple vertically arranged points afterwards are the feedback matrix. After combining the candidate matrix and the feedback matrix, they are input into the output layer. The output layer integrates the sum of the candidate matrix and the corresponding feedback matrix between pairwise encoded features into an adjacency value, that is Figure 4 the FC in , generating an adjacency matrix, which is a matrix composed of multiple adjacency values.
[0044] In one embodiment, the step of inputting the multiple encoded features and the adjacency matrix into the graph convolutional network for graph convolutional processing to obtain multiple piecewise features corresponding to the encoded features includes: using the multiple encoded features and the adjacency matrix as the input of the first graph convolutional layer in the multiple graph convolutional layers, and using the adjacency matrix and the output features of the previous graph convolutional layer as the input of the non-first graph convolutional layers in the multiple graph convolutional layers to perform graph convolutional processing, generating multiple piecewise features corresponding to the encoded features.
[0045] Specifically, the graph convolutional network includes multiple graph convolutional layers, as Figure 5 shown, within the first segment metric module, GC 1 -GC m respectively correspond to m graph convolutional layers. SRMLP is used to indicate the adjacency-aware network. One adjacency matrix is shared within one segment metric module, that is, each graph convolutional layer within the segment metric module uses the same adjacency matrix as the input. And except that the input of the first graph convolutional layer is the encoded features, the input of other non-first graph convolutional layers is the concatenated feature obtained by concatenating the output features of the previous graph convolutional layer and the input features of the previous graph convolutional layer. In this way, each graph convolutional layer will perform graph convolutional processing by combining the input and output of the previous convolutional layer. After passing through the last graph convolutional layer, the convolved piecewise features are output, and these piecewise features will be used as the input of the next segment metric module to perform graph convolutional processing according to the same process as above. Here, only the execution process within the first segment metric module is illustrated as an example, and the execution processes within other non-first segment metric modules are the same. Among them, the larger the number m of graph convolutional layers, the deeper the depth of the graph neural network, but at the same time, the more likely it is to have the over-smoothing problem. Therefore, for different graph convolutional tasks, custom adjustments are made according to the corresponding depth requirements and over-smoothing requirements. The total number of graph convolutional layers in the entire graph neural network is m*n. Exemplarily, m is less than or equal to 4, and n is less than or equal to 3. In this way, on the basis of ensuring sufficient depth of the graph neural network, the over-smoothing phenomenon is avoided. Among them, the number of graph convolutional layers within each segment metric module can be the same or different.
[0046] In one embodiment, the step of inputting a plurality of sample images into the feature extraction network and performing feature extraction processing on the sample images through the feature extraction network to obtain the image features corresponding to the sample images includes: inputting a plurality of the sample images into the feature extraction module to obtain sample features corresponding to the sample images; inputting the sample features into the high-level convolution module for convolution processing to obtain corresponding feedback features; and inputting the sample features and the feedback features into the low-level convolution module, and performing convolution processing after splicing the sample features and the guidance features to obtain the image features.
[0047] Specifically, the feature extraction network includes a feature extraction module, a high-level convolution module, and a low-level convolution module. The feature extraction module performs feature extraction on the sample images to obtain corresponding sample features. If the traditional convolution method is used to perform convolution processing on the sample features, features with obvious differences cannot be obtained. The high-level convolution module performs convolution on the sample features to obtain feedback features, and the feedback features after convolution are spliced with the sample features. The low-level convolution module performs convolution processing on the spliced features to extract more abstract features and magnify the distinctiveness of the features, that is, to obtain image features with obvious differences.
[0048] In one embodiment, the step of inputting the sample features into the high-level convolution module for convolution processing to obtain corresponding feedback features includes: inputting the sample features into the high-level convolution network for convolution processing to obtain corresponding guidance features; and inputting the guidance features into the splitting module for splitting to obtain at least two split features; wherein, the feedback feature is one of the at least two split features.
[0049] Specifically, the high-level convolution module includes a high-level convolution network and a splitting module. Inside the high-level convolution module, the high-level convolution network first performs convolution processing on the sample features to obtain guidance features, and the splitting module splits the guidance features after convolution. If the data volume for processing by the low-level convolution module using the complete guidance features after convolution is large, then partial features after splitting are used to be spliced with the sample features for convolution processing in the low-level convolution module, which is convenient for improving the distinctiveness of the features after convolution.
[0050] In one embodiment, inputting the guiding feature into the splitting module for splitting to obtain at least two splitting features includes: inputting the guiding feature into the splitting module for splitting to obtain a third feature and a fourth feature; the low-level convolutional module includes a low-level convolutional network and a splicing module, and inputting the sample feature and the feedback feature into the low-level convolutional module, performing convolutional processing on the spliced sample feature and guiding feature to obtain the image feature, including: inputting the third feature and the sample feature into the splicing module for splicing to obtain a combined feature; inputting the combined feature into the low-level convolutional network for convolutional processing to obtain the convolved combined feature; inputting the convolved combined feature and the sampled fourth feature into the splicing module for splicing to generate the image feature.
[0051] Specifically, as Figure 6 shown, RSGI-Conv is used to indicate the feature extraction network, RSGI is used to indicate the high-level convolutional network, the sample feature is f in , and the shape corresponding to the sample feature is (C in , H in , W in ). In the high-level convolutional network, the sample feature passing through Conv means performing convolutional processing on the sample feature to obtain a guiding feature with a shape of (C, H in , W in ). The guiding feature passing through Split means splitting the guiding feature to obtain a third feature f with a shape of g2 , and a fourth feature f with a shape of g1 . The larger the value of n, the larger the shape of the third feature compared to the fourth feature. When n is 2, the shapes of the third feature and the fourth feature are the same.
[0052] Splicing the third feature and the sample feature to obtain a combined feature, denoted as f in-g2 , with a shape of . After the combined feature passes through Conv for convolutional processing, the convolved combined feature is obtained, denoted as f out1 , with a shape of . At the same time, the sampled fourth feature is denoted as f out2 , with a shape of . Splicing f out1 and f out2 to generate an image feature, denoted as f out , with a shape of (C out , H out , W out ).
[0053] In one embodiment, the step of inputting the plurality of image features into the encoding module for category encoding to obtain a plurality of encoding features corresponding to the image features comprises: inputting the plurality of image features into the encoding module, and using known category labels and unknown category labels to respectively perform category labeling on the image features corresponding to the known category images and the image features corresponding to the unknown category images to obtain a plurality of encoding features.
[0054] Specifically, the corresponding known category attributes can be determined according to the known category image, and each known category attribute corresponds to a known category label. The known category label is used to classify the image features corresponding to the known category image. For example, the known category attributes include triangle, trapezoid and circle, and the known category label corresponding to the triangle is S1, and the label of S1 is encoded as (1,0,0), the known category label corresponding to the trapezoid is S2, and the label of S2 is encoded as (0,1,0), and the known category label corresponding to the circle is S3, and the label of S3 is encoded as (0,0,1). There are usually several category attributes, and the label encoding is composed of several characters. Each character in the label encoding corresponds to a known category, and the image features of the known category image belonging to the triangle category are marked as S1, and the image features of other categories are marked in the same way. The unknown category label is represented by a character different from the known category label. In this embodiment, the label encoding of the unknown category label is (0,0,0). The image features corresponding to all unknown category images are marked with the unknown category label. The coded features obtained after marking the image features are the features after the image features and their corresponding label encodings are spliced.
[0055] Figure 2 FIG. 1 is a flow chart of an image processing method in one embodiment. It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0056] Figure 7 The internal structure diagram of a computer device in one embodiment is shown. The computer device may specifically be Figure 1 The terminal or server equipped with the graph neural network in . Figure 7As shown, the computer device includes a processor and a memory connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement an image processing method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the image processing method.
[0057] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0058] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any of the above embodiments is implemented.
[0059] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the method described in any of the above embodiments is implemented.
[0060] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by a computer program instructing relevant hardware. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0061] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0062] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. An image processing method, characterized in that, it is applied to a graph neural network including a feature extraction network and a feature classification network, the feature classification network includes an encoding module, a plurality of intra-segment metric modules and a linear layer, and the method includes: Inputting a plurality of sample images into the feature extraction network, and performing feature extraction processing on the sample images through the feature extraction network to obtain image features corresponding to the sample images; wherein, the plurality of sample images include known-class images and unknown-class images; Inputting the plurality of image features into the encoding module for class encoding to obtain a plurality of encoding features corresponding to the image features; Sequentially inputting the plurality of encoding features into the plurality of intra-segment metric modules for iterative graph convolution processing to obtain a plurality of convolution features corresponding to the encoding features; Inputting the plurality of convolution features into the linear layer for classification processing to obtain corresponding classification results; wherein, the classification results include the classes of the unknown-class images; The intra-segment metric module includes an adjacency perception network and a graph convolution network. The step of sequentially inputting the plurality of encoding features into the plurality of intra-segment metric modules for iterative graph convolution processing to obtain a plurality of convolution features corresponding to the encoding features includes: inputting the plurality of encoding features into the adjacency perception network to determine the adjacency relationship between each of the encoding features and generate an adjacency matrix; inputting the plurality of encoding features and the adjacency matrix into the graph convolution network for graph convolution processing to obtain a plurality of segment features corresponding to the encoding features; sequentially inputting the plurality of segment features into the plurality of intra-segment metric modules for iterative graph convolution processing to obtain the convolution features; The adjacency perception network includes a learning network, a residual network and an output layer. The step of inputting the plurality of encoding features into the adjacency perception network to determine the adjacency relationship between each of the encoding features and generate an adjacency matrix includes: inputting a plurality of absolute values of feature differences into the learning network to learn the adjacency relationship between two of the encoding features corresponding to the absolute values of the feature differences to obtain a plurality of candidate matrices for indicating the adjacency relationship; wherein, the absolute value of the feature difference is used to indicate the absolute value of the difference between the corresponding values of any two of the encoding features; inputting the plurality of absolute values of the feature differences into the residual network to obtain a plurality of feedback matrices corresponding to the absolute values of the feature differences; inputting the sum of the plurality of candidate matrices and the corresponding feedback matrices into the output layer to generate the adjacency matrix.
2. The method according to claim 1, characterized in that, the graph convolution network includes a plurality of graph convolution layers. The step of inputting the plurality of encoding features and the adjacency matrix into the graph convolution network for graph convolution processing to obtain a plurality of segment features corresponding to the encoding features includes: Use the multiple encoded features and the adjacency matrix as the input of the first graph convolutional layer in the multiple graph convolutional layers, and use the adjacency matrix and the output features of the previous graph convolutional layer as the input of the non-first graph convolutional layers in the multiple graph convolutional layers to perform graph convolutional processing, generating multiple segmented features corresponding to the encoded features.
3. The method according to claim 1, wherein, the feature extraction network includes a feature extraction module, a high-level convolutional module, and a low-level convolutional module. The step of inputting multiple sample images into the feature extraction network and performing feature extraction processing on the sample images through the feature extraction network to obtain the image features corresponding to the sample images includes: Inputting multiple of the sample images into the feature extraction module to obtain sample features corresponding to the sample images; Inputting the sample features into the high-level convolutional module for convolutional processing to obtain corresponding feedback features; Inputting the sample features and the feedback features into the low-level convolutional module, and performing convolutional processing on the concatenated sample features and feedback features to obtain the image features.
4. The method according to claim 3, wherein, the high-level convolutional module includes a high-level convolutional network and a splitting module. The step of inputting the sample features into the high-level convolutional module for convolutional processing to obtain corresponding feedback features includes: Inputting the sample features into the high-level convolutional network for convolutional processing to obtain corresponding guiding features; Inputting the guiding features into the splitting module for splitting to obtain at least two split features; wherein, the feedback feature is one of the at least two split features.
5. The method according to claim 4, wherein, the step of inputting the guiding features into the splitting module for splitting to obtain at least two split features includes: Inputting the guiding features into the splitting module for splitting to obtain a third feature and a fourth feature; the low-level convolutional module includes a low-level convolutional network and a concatenation module. The step of inputting the sample features and the feedback features into the low-level convolutional module, and performing convolutional processing on the concatenated sample features and feedback features to obtain the image features includes: Inputting the third feature and the sample features into the concatenation module for concatenation to obtain a combined feature; Inputting the combined feature into the low-level convolutional network for convolutional processing to obtain the convolutional combined feature; Inputting the convolutional combined feature and the sampled fourth feature into the concatenation module for concatenation to generate the image features.
6. The method according to claim 1, wherein, the step of inputting the multiple image features into the encoding module for category encoding to obtain multiple encoded features corresponding to the image features includes: Inputting the multiple image features into the encoding module, and using known category labels and unknown category labels to respectively perform category marking on the image features corresponding to the known category images and the image features corresponding to the unknown category images to obtain the multiple encoded features.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Multi-label image classification method, device and equipment based on image convolution
CN109816009A
Graph data identification method and device, computer device and storage medium
CN110363086A