Training Method, Extraction Method and Device for Image Feature Extraction Network
Patent Information
- Application Number
- CN202111398839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-19
Smart Images

Figure CN114359650B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular, to a training method, an extraction method, and a device for an image feature extraction network. Background Art
[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0003] In the training method of the image feature extraction network in the related art, when training the image feature extraction network, the training effect is not good, resulting in the weak learning ability of the image feature extraction network for image feature representation learning. For how to enable the image feature extraction network to perform more effective image feature representation learning, there is no effective solution in the related art. Summary of the Invention
[0004] Embodiments of the present application provide a training method, an extraction method, a device, an electronic device, a computer-readable storage medium, and a computer program product for an image feature extraction network, which can enable the image feature extraction network to perform more effective image feature representation learning.
[0005] The technical solution of the embodiments of the present application is implemented as follows:
[0006] Embodiments of the present application provide a training method for an image feature extraction network. The image feature extraction network includes a first encoder and a second encoder. The method includes:
[0007] Invoking the first encoder to perform encoding processing based on the first image sample to obtain the first low-dimensional feature of the first image sample, and respectively invoking the first processing network and the third processing network based on the first low-dimensional feature to correspondingly obtain the first high-dimensional feature and the third high-dimensional feature;
[0008] Invoking the second encoder to perform encoding processing based on the second image sample to obtain the second low-dimensional feature of the second image sample, and respectively invoking the second processing network and the fourth processing network based on the second low-dimensional feature to correspondingly obtain the second high-dimensional feature and the fourth high-dimensional feature;
[0009] Determine a first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determine a second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, and perform gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value;
[0010] Perform momentum update on the parameters of the second encoder according to the parameters of the first encoder after gradient update.
[0011] An embodiment of the present application provides a training device for an image feature extraction network, including:
[0012] A first processing module, configured to call a first encoder to perform encoding processing on a first image sample to obtain a first low-dimensional feature of the first image sample, and respectively call a first processing network and a third processing network based on the first low-dimensional feature to correspondingly obtain a first high-dimensional feature and a third high-dimensional feature;
[0013] A second processing module, configured to call a second encoder to perform encoding processing on a second image sample to obtain a second low-dimensional feature of the second image sample, and respectively call a second processing network and a fourth processing network based on the second low-dimensional feature to correspondingly obtain a second high-dimensional feature and a fourth high-dimensional feature;
[0014] A first update module, configured to determine a first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determine a second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, and perform gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value;
[0015] A second update module, configured to perform momentum update on the parameters of the second encoder according to the parameters of the first encoder after gradient update.
[0016] An embodiment of the present application provides a feature extraction device based on an image feature extraction network, which is applied to the image feature extraction network. The feature extraction device based on the image feature extraction network includes: a first encoding module, configured to perform encoding processing on the trained first encoder based on a first preprocessed image corresponding to the image to be processed, so as to obtain an encoded feature of the first preprocessed image; a second encoding module, configured to perform encoding processing on the trained second encoder based on a second preprocessed image corresponding to the image to be processed, so as to obtain an encoded feature of the second preprocessed image; a fusion module, configured to perform fusion processing on the encoded feature of the first preprocessed image and the encoded feature of the second preprocessed image, so as to obtain a feature of the image to be processed.
[0017] An embodiment of the present application provides an electronic device, including:
[0018] A memory, configured to store executable instructions;
[0019] A processor, configured to implement the training method of the image feature extraction network provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0020] An embodiment of the present application provides an electronic device, including:
[0021] A memory, configured to store executable instructions;
[0022] A processor, configured to implement the feature extraction method based on the image feature extraction network provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0023] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the training method of the image feature extraction network provided by the embodiment of the present application when executed.
[0024] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the feature extraction method based on the image feature extraction network provided by the embodiment of the present application when executed.
[0025] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the training method or the extraction method of the image feature extraction network described above in the embodiment of the present application.
[0026] The embodiment of the present application has the following beneficial effects:
[0027] By separately calling the first encoder and the second encoder for encoding processing based on the first image sample and the second image sample, respectively processing the obtained low-dimensional features by calling different processing networks to obtain corresponding high-dimensional features, determining the cross-loss value through the obtained high-dimensional features, thereby using the cross-loss value to update the gradient of the first encoder, and performing momentum update on the second encoder according to the first encoder after gradient update. By enabling the first encoder and the second encoder to interact during training, it is possible to adopt completely different update methods for different encoders during training, and thus enable the image feature extraction network to perform more effective representation learning of image features. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figures 1A - 1B FIG. is a schematic structural diagram of the training system architecture of the image feature extraction network provided by an embodiment of the present application;
[0029] Figures 2A - 2B FIG. is a schematic structural diagram of the training device of the image feature extraction network provided by an embodiment of the present application;
[0030] Figures 3A - 3C FIG. is a schematic flowchart of the training method of the image feature extraction network provided by an embodiment of the present application;
[0031] Figure 3D FIG. is a schematic flowchart of the feature extraction method based on the image feature extraction network provided by an embodiment of the present application;
[0032] Figures 4A - 4F FIG. is a schematic diagram of the principle of the training method of the image feature extraction network provided by an embodiment of the present application;
[0033] Figure 4G FIG. is a schematic diagram of the principle of the feature extraction method based on the image feature extraction network provided by an embodiment of the present application;
[0034] Figures 5A - 5B FIG. is a schematic diagram of the principle of the training method of the image feature extraction network provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0036] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0037] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0039] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0040] 1) Encoder: It is a neural network structure that can learn an efficient representation of input data through unsupervised learning. The efficient representation of input data is called coding (Codings), and its dimension is generally much smaller than that of the input data, so that the encoder can be used for image dimensionality reduction processing. At the same time, the encoder can be used as a feature detector (Feature Detectors) and applied to the pre-training of deep neural networks. In the embodiments of this application, the first encoder and the second encoder can be two different types of encoders. Among them, the second encoder can be an encoder with a deep self-attention transformation function, and the image features extracted by the second encoder can have better spatial focus properties.
[0041] 2) Artificial Intelligence (AI): It is the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics.
[0042] 3) High-dimensional features / Low-dimensional features: Also known as high-dimensional image features / low-dimensional image features, the dimension of high-dimensional image features is higher than that of low-dimensional image features. Image features mainly include color features, texture features, shape features, and spatial relationship features of the image. Among them, color features are a type of global feature, which describe the surface properties of the scene corresponding to the image or image region; texture features are also a type of global feature, which also describe the surface properties of the scene corresponding to the image or image region; shape features have two types of representation methods, one is contour features, and the other is region features. The contour features of an image mainly target the outer boundary of an object, while the region features of an image are related to the entire shape region; spatial relationship features refer to the mutual position or relative direction relationship between multiple objects segmented from an image, and these relationships can also be divided into connection / adjacency relationships, overlap / overlap relationships, and inclusion / containment relationships, etc.
[0043] 4) Gradient update: Gradient update of the parameters of the encoder refers to the process of updating the parameter values along the direction of the negative gradient of the parameters, where the direction of the negative gradient is the direction of gradient descent.
[0044] 5) Momentum Update: It is a parameter update method of the encoder that accelerates gradient update by optimizing the update of relevant directions and weakening the update of irrelevant directions.
[0045] 6) Projector: It is a multi-layer perceptron network, which can be composed of fully connected layers, activation layers, and normalization layers, for example.
[0046] 7) Prodictor: It is a multi-layer perceptron network. The structures of the predictor and the projector can be the same, that is, it can be composed of fully connected layers, activation layers, and normalization layers; of course, different structures can also be adopted. For example, the projector can be composed of two fully connected layers, activation layers, and normalization layers, and the predictor can be composed of one fully connected layer, activation layers, and normalization layers.
[0047] 8) Multilayer Perceptron (MLP): It is a forward-structured artificial neural network, which is mainly responsible for mapping a set of input vectors to a set of output vectors.
[0048] 9) Multi-head Self-attention (MHSA): It is a special structure in the sub-encoder, which contains multiple processes of linear processing, weighted multiplication processing, and summation processing. The multi-head self-attention layer can enable the sub-encoder to output higher-quality image features.
[0049] 10) Convolutional Neural Networks (CNN): A type of feed-forward neural network (FNN) that contains convolutional computations and has a deep structure, and is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input images according to their hierarchical structure.
[0050] During the implementation of the embodiments of this application, the applicant found the following problems in the related art: In the related art, the self-supervised training framework mainly uses two first encoder branches to learn the features of the same image under two different preprocessings, and by minimizing the distance between the two features of the same image in the high-dimensional space or maximizing the distance between the features of different images in the high-dimensional space, so that the first encoder can learn effective features under different input images. However, in the related art, when training the image feature extraction network, the spatial attention characteristics of the encoder are not fully utilized, resulting in poor training effects. There is no effective solution in the related art for how to improve the training effect of the image feature extraction network.
[0051] In some embodiments, in the application scenario of medical image processing, it is usually difficult to obtain training samples with annotation information (i.e., medical image samples), resulting in poor training effects when the related technology trains the image feature extraction network. Through the training method of the image feature extraction network provided by the embodiments of this application, the image feature extraction network can be effectively trained in the application scenario of medical image processing, enabling the image feature extraction network to perform more effective representation learning of image features.
[0052] In other embodiments, in the application scenario of face recognition, since it is difficult to obtain effective training samples with annotation information in the application scenario of face recognition, the training effect is poor when the related technology trains the image feature extraction network. Through the training method of the image feature extraction network provided by the embodiments of this application, the image feature extraction network can be effectively trained in the application scenario of face recognition, enabling the image feature extraction network to perform more effective representation learning of image features.
[0053] The embodiments of the present application provide a training method, an extraction method, a device, an electronic device, a computer-readable storage medium, and a computer program product for an image feature extraction network, which can enable the image feature extraction network to perform more effective feature representation learning of images. The following describes an exemplary application of the training device for the image feature extraction network provided by the embodiments of the present application. The training device for the image feature extraction network provided by the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), etc., or can also be implemented as a server.
[0054] See Figure 1A , Figure 1A FIG. is an optional schematic architecture diagram of a training system 100 for an image feature extraction network provided by the embodiments of the present application. To implement the application scenarios for training the image feature extraction network (for example, the application scenario for training the image feature extraction network can be to train the image feature extraction network for medical images, and the application scenario for training the image feature extraction network can also be to train the image feature extraction network for face recognition images), the terminal (exemplarily shows the terminal 400) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0055] The terminal 400 is used for the user to use the client 410, which is displayed on the graphical interface 410-1 (exemplarily shows the graphical interface 410-1). The terminal 400 and the server 200 are connected to each other through a wired or wireless network.
[0056] In some embodiments, the server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0057] In some embodiments, the terminal 400 receives the image feature extraction network sent by the server 200, trains the image feature extraction network, obtains the trained image feature extraction network, and the terminal 400 sends the trained image feature extraction network to the server 200. The server 200 invokes the trained image feature extraction network based on the image to be processed to perform image feature extraction processing, and obtains the image features of the image to be processed.
[0058] In other embodiments, the server 200 trains the image feature extraction network, obtains the trained image feature extraction network, and the server 200 sends the trained image feature extraction network to the terminal 400. After receiving the trained image feature extraction network sent by the server 200, the terminal 400 displays a reminder message indicating that the image feature extraction network has been trained in the graphical interface 410-1. The terminal 400 invokes the trained image feature extraction network based on the image to be processed to perform image feature extraction processing, and obtains the image features of the image to be processed.
[0059] In some embodiments, the server 200 may be a server cluster or a distributed system composed of multiple servers. Taking the distributed system as a blockchain system as an example, multiple servers can form a blockchain network, and the server 200 is a node on the blockchain network.
[0060] Next, an exemplary application of the blockchain network will be described by taking multiple terminals accessing the blockchain network to implement the training of the image feature extraction network as an example.
[0061] In some embodiments, refer to Figure 1B , Figure 1B An optional architecture schematic diagram of the training system 100 of the image feature extraction network provided in the embodiments of the present application. Multiple terminals involved in the image feature extraction network participate in the training of the image feature extraction network, such as the terminal 500 and the terminal 600. After obtaining the authorization of the blockchain management platform 800, the client 510 of the terminal 500 and the client 610 of the terminal 600 can both access the blockchain network 700.
[0062] The terminal 500 sends an image feature extraction network acquisition request to the blockchain management platform 800 (the terminal 600 sends an image feature extraction network acquisition request to the blockchain management platform 800). The blockchain management platform 800 generates a corresponding update operation according to the image feature extraction network acquisition request. The update operation specifies the smart contract to be called to implement the update operation / query operation and the parameters to be passed to the smart contract. The transaction also carries the digital signature signed by the web page, and sends the update operation to the blockchain network 700.
[0063] When nodes 210-1, 210-2, and 210-3 in the blockchain network 700 receive an update operation, they verify the digital signature of the update operation. After the digital signature verification is successful, they confirm whether the client 510 has the access permission according to the identity of the client 510 carried in the update operation. Any verification failure in the digital signature and permission verification will result in a failed acquisition. After successful verification, the signing node 210 signs its own digital signature (for example, encrypts the digest of the transaction using the private key of node 210-1) and continues to broadcast in the blockchain network 700.
[0064] Nodes 210-1, 210-2, 210-3, etc. with sorting functions in the blockchain network 700, after receiving a successfully verified acquisition, fill the acquisition request into a new block and broadcast it to the nodes in the blockchain network 700 that provide consensus services.
[0065] The nodes in the blockchain network 700 that provide consensus services perform a consensus process on the new block to reach an agreement. The nodes that provide the ledger function append the new block to the tail of the blockchain and execute the acquisition request in the new block: for the submitted acquisition request of the image feature extraction network, update the key-value pair corresponding to the image feature extraction network in the status database; for the acquisition request to query the image feature extraction network, query the key-value pair corresponding to the image feature extraction network from the status database and return the image feature extraction network. After the terminals 500 and 600 receive the image feature extraction network returned by the blockchain network 700, the terminals 500 and 600 train the image feature extraction network, obtain the trained image feature extraction network, and display a training success prompt message in the graphical interfaces 510-1 and 610-1. The terminals 500 and 600 send the trained image feature extraction network to the blockchain network 700, and the blockchain network 700 invokes the trained image feature extraction network based on the image to be processed to perform image feature extraction processing and obtain the image features of the image to be processed.
[0066] See Figure 2A , Figure 2A is a schematic structural diagram of a server 200 for the training method of the image feature extraction network provided by an embodiment of the present application. Figure 2A The shown server 200 includes: at least one processor 460, a memory 450, and at least one network interface 420. Each component in the server 200 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2A all kinds of buses are labeled as the bus system 440.
[0067] The processor 460 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0068] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 460.
[0069] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0070] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which will be exemplarily described below.
[0071] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks.
[0072] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.
[0073] In some embodiments, the training device of the image feature extraction network provided by the embodiments of the present application can be implemented in software. Figure 2A Shown is the training device 455A of the image feature extraction network stored in the memory 450, which can be software in the form of programs and plugins, etc., and includes the following software modules: a first processing module 4551A, a second processing module 4552A, a first update module 4553A, and a second update module 4554A. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0074] In some other embodiments, the training device of the image feature extraction network provided by the embodiments of the present application may be implemented in a hardware manner. As an example, the training device of the image feature extraction network provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the image feature extraction network provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0075] See Figure 2B , Figure 2B is another schematic structural diagram of the feature extraction server 200 based on the image feature extraction network provided by the embodiments of the present application. Figure 2B The server 200 shown includes: at least one processor 660, a memory 650, and at least one network interface 620. Each component in the server 200 is coupled together through a bus system 640. It can be understood that the bus system 640 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 640 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2B all kinds of buses are labeled as the bus system 640.
[0076] The processor 660 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0077] The memory 650 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 650 optionally includes one or more storage devices that are physically located away from the processor 660.
[0078] The memory 650 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 650 described in the embodiments of the present application is intended to include any suitable type of memory.
[0079] In some embodiments, the memory 650 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which will be exemplarily described below.
[0080] The operating system 651 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks.
[0081] The network communication module 652 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 620. Exemplary network interfaces 620 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.
[0082] In some embodiments, the training device of the image feature extraction network provided by the embodiments of the present application can be implemented in software. Figure 2B Shown is a feature extraction device 655B based on an image feature extraction network stored in the memory 650, which may be software in the form of programs and plugins, etc., and includes the following software modules: a first encoding module 6551B, a second encoding module 6552B, and a fusion module 6553B. These modules are logical, and thus can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0083] In some other embodiments, the feature extraction device of the image feature extraction network provided by the embodiments of the present application can be implemented in a hardware manner. As an example, the feature extraction device of the image feature extraction network provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the feature extraction method of the image feature extraction network provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0084] The training method of the image feature extraction network provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server provided by the embodiments of the present application.
[0085] In some embodiments, Figure 4A is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiments of the present application. Refer to Figure 4A , the image feature extraction network includes a first encoder and a second encoder. By training the first encoder and the second encoder in the image feature extraction network through the training method of the image feature extraction network provided by the embodiments of the present application, the training effect of the image feature extraction network can be significantly improved, and the feature extraction ability of the image feature extraction network can be effectively enhanced.
[0086] Next, refer to Figure 3A , Figure 3A is a schematic flowchart of the training method of the image feature extraction network provided by the embodiments of the present application, which will be described in combination with Figure 3A the steps 101 to 104 shown. The execution subject of the following steps 101 to 104 can be the aforementioned server.
[0087] In step 101, based on the first image sample, the first encoder is called for encoding processing to obtain the first low-dimensional feature of the first image sample, and based on the first low-dimensional feature, the first processing network and the third processing network are respectively called to correspondingly obtain the first high-dimensional feature and the third high-dimensional feature.
[0088] In some embodiments, the first low-dimensional feature is a low-dimensional feature relative to the first high-dimensional feature and the third high-dimensional feature, that is, the feature dimension of the first low-dimensional feature is lower than the feature dimensions of the first high-dimensional feature and the third high-dimensional feature.
[0089] In some embodiments, the first encoder may be an encoder based on a convolutional neural network, where the encoder based on the convolutional neural network may be a ResNet50 encoder, a ResNet101 encoder, a ResNet152 encoder, etc. The first encoder is used to encode an image sample to obtain a low-dimensional feature of the image sample. The first processing network is used to process the low-dimensional feature of the image sample to obtain a corresponding high-dimensional feature, and the third processing network is used to process the low-dimensional feature of the image sample to obtain a corresponding high-dimensional feature, that is, both the first processing network and the third processing network can play a role in increasing the feature dimension of the image sample. Among them, in order to ensure the difference in the processing process, the structures of the first processing network and the third processing network are different, and the training update methods are different.
[0090] As an example, refer to Figure 4B , the first encoder is called to perform encoding processing based on the first image sample x1 to obtain the first low-dimensional feature Z1 of the first image sample x1, and the first processing network and the third processing network are respectively called based on the first low-dimensional feature Z1 to correspondingly obtain the first high-dimensional feature f1(x) and the third high-dimensional feature f3(x).
[0091] In some embodiments, refer to Figure 4D , Figure 4D is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiments of the present application. The first processing network includes a first mapper and a first predictor. The third processing network includes a third encoder with spatial attention characteristics, a third mapper, and a third predictor. Among them, the structures of the first mapper and the first predictor may be the same. The main structures of the first mapper and the first predictor may be composed of a multilayer perceptron (MLP). The main structures of the first mapper and the first predictor may include fully connected layers (FC), activation layers, and batch normalization (BN) layers. The structures of the third mapper and the third predictor may be the same as those of the first mapper and the first predictor. That is, the above-mentioned first mapper and third mapper are mappers with the same structure but different parameters, and the first predictor and third predictor are predictors with the same structure but different parameters.
[0092] In this way, by setting different numbers of mappers and predictors in the first processing network and the third processing network, the differential processing of the first processing network and the third processing network can be ensured, so that the high-dimensional features processed by the first processing network and the third processing network maintain a certain difference, so that the corresponding loss values have sufficient diversity, which is convenient for subsequent training of the image feature extraction network.
[0093] In some embodiments, before encoding the first image sample by invoking the first encoder in step 101 above, the first image sample and the second image sample can be obtained in the following manner: the image sample to be processed is subjected to a first preprocessing to obtain the first image sample; the image sample to be processed is subjected to a second preprocessing to obtain the second image sample; wherein the first preprocessing and the second preprocessing are inverse processing procedures to each other.
[0094] As an example, when the first preprocessing is image blurring processing, the second preprocessing can be image sharpening processing; when the first preprocessing is image sharpening processing, the second preprocessing can be image blurring processing. The image sharpening processing and the image blurring processing are inverse processing procedures to each other, that is, the first preprocessing and the second preprocessing are inverse processing procedures to each other.
[0095] As an example, referring to Figure 4B , the image sample X to be processed is subjected to a first preprocessing to obtain the first image sample x1; the image sample X to be processed is subjected to a second preprocessing to obtain the second image sample x2. That is, the first image sample x1 is the input of the first encoder, and the second image sample x2 is the input of the second encoder.
[0096] In this way, by differentially processing the image sample to be processed, the first image sample and the second image sample are obtained, such that the first image sample and the second image sample are not exactly the same. Since both the first image sample and the second image sample are derived from the same image sample to be processed, the first image sample and the second image sample have a certain similarity. Thus, the first image sample and the second image sample are similar but not exactly the same, which is convenient for subsequent training of the image feature extraction network.
[0097] In some embodiments, referring to Figure 3C , in step 101 above, based on the first low-dimensional feature, the first processing network and the third processing network are respectively invoked to obtain the first high-dimensional feature and the third high-dimensional feature, which can be achieved by executing step 1011 to step 1012.
[0098] In step 1011, based on the first low-dimensional feature, the first mapper in the first processing network is invoked for mapping processing to obtain a first mapping processing result; based on the first mapping processing result, the first predictor in the first processing network is invoked for prediction processing, and the obtained first prediction processing result is used as the first high-dimensional feature.
[0099] As an example, referring to Figure 4D, the first encoder is called based on the first image sample x1 for encoding processing to obtain the first low-dimensional feature Z1 of the first image sample x1. Based on the first low-dimensional feature Z1, the first mapper in the first processing network is called for mapping processing to obtain the first mapping processing result; based on the first mapping processing result, the first predictor in the first processing network is called for prediction processing, and the obtained first prediction processing result is used as the first high-dimensional feature f1(x).
[0100] In step 1012, the third encoder in the third processing network is called based on the first low-dimensional feature for encoding processing to obtain the third encoding result; based on the third encoding result, the third mapper in the third processing network is called for mapping processing to obtain the third mapping processing result; based on the third mapping processing result, the third predictor in the third processing network is called for prediction processing, and the obtained third prediction processing result is used as the third high-dimensional feature.
[0101] As an example, refer to Figure 4D , the first encoder is called based on the first image sample x1 for encoding processing to obtain the first low-dimensional feature Z1 of the first image sample x1. The third encoder in the third processing network is called based on the first low-dimensional feature Z1 for encoding processing to obtain the third encoding result; based on the third encoding result, the third mapper in the third processing network is called for mapping processing to obtain the third mapping processing result; based on the third mapping processing result, the third predictor in the third processing network is called for prediction processing, and the obtained third prediction processing result is used as the third high-dimensional feature f3(x).
[0102] In some embodiments, the step of calling the third encoder in the third processing network based on the first low-dimensional feature in step 1012 for encoding processing to obtain the third encoding result can be implemented in the following manner: calling the first convolutional layer based on the first low-dimensional feature for convolutional processing to obtain the first convolutional processing result; calling the self-attention layer based on the first convolutional processing result for attention processing to obtain the attention processing result; calling the second convolutional layer based on the attention processing result for convolutional processing to obtain the second convolutional processing result; adding the second convolutional processing result and the first low-dimensional feature to obtain the third encoding result.
[0103] As an example, refer to Figure 4F , Figure 4F is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiments of the present application, Figure 4F What is shown by reference numeral 41 in Figure 4FShown in reference numeral 41 is a schematic structural diagram of a sub-encoder provided in an embodiment of the present application. The third encoder includes a plurality of serial sub-encoders with the same structure.
[0104] As an example, refer to Figure 4F Reference numeral 41. Taking the processing process of a serial sub-encoder as an example, based on the first low-dimensional feature, a first convolutional layer is called for convolutional processing to obtain a first convolutional processing result R1. Based on the first convolutional processing result R1, a self-attention layer is called for attention processing to obtain an attention processing result R2. Based on the attention processing result R2, a second convolutional layer is called for convolutional processing to obtain a second convolutional processing result R3; the second convolutional processing result R3 and the first low-dimensional feature are added to obtain a third encoding result R4.
[0105] In some embodiments, the above-mentioned calling of the self-attention layer for attention processing based on the first convolutional processing result to obtain the attention processing result can be implemented in the following manner: the first convolutional processing result is respectively subjected to a first linear processing, a second linear processing, and a third linear processing to correspondingly obtain a first linear processing result, a second linear processing result, and a third linear processing result; the first linear processing result and the second linear processing result are respectively multiplied by corresponding weights to obtain a first multiplication processing result corresponding to the first linear processing result and a second multiplication processing result corresponding to the second linear processing result; the first multiplication processing result and the second multiplication processing result are multiplied to obtain a third multiplication processing result; the first linear processing result and the timing information are respectively multiplied by corresponding weights to obtain a fourth multiplication processing result corresponding to the first linear processing result and a fifth multiplication processing result corresponding to the timing information; the fourth multiplication processing result and the fifth multiplication processing result are multiplied to obtain a sixth multiplication processing result; the third multiplication processing result and the sixth multiplication processing result are summed, and the obtained summation processing result is normalized to obtain a normalized processing result; the normalized processing result and the third linear processing result are multiplied to obtain the attention processing result.
[0106] In some embodiments, refer to Figure 4F Reference numeral 42 Figure 4F Reference numeral 42 is a schematic structural diagram of the self-attention layer provided in an embodiment of the present application. The attention processing of the first convolutional processing result can be performed through the structure of the self-attention layer shown in Figure 4F Reference numeral 42 to obtain the attention processing result.
[0107] As an example, refer to Figure 4F, the first convolution processing result R1 is respectively subjected to a first linear processing, a second linear processing, and a third linear processing to correspondingly obtain a first linear processing result, a second linear processing result, and a third linear processing result. The first linear processing result and the second linear processing result are respectively multiplied by corresponding weights, that is, the first linear processing result is multiplied by the weight value q corresponding to the first linear processing result, and the second linear processing result is multiplied by the weight value k corresponding to the second linear processing result, to respectively obtain a first multiplication processing result w1 corresponding to the first linear processing result and a second multiplication processing result w2 corresponding to the second linear processing result; the first multiplication processing result w1 and the second multiplication processing result w2 are multiplied to obtain a third multiplication processing result w3; the first linear processing result and the timing information are respectively multiplied by corresponding weights, that is, the first linear processing result is multiplied by the corresponding weight q, and the timing information is multiplied by the corresponding weight p, to respectively obtain a fourth multiplication processing result w4 corresponding to the first linear processing result and a fifth multiplication processing result w5 corresponding to the timing information; the fourth multiplication processing result w4 and the fifth multiplication processing result w5 are multiplied to obtain a sixth multiplication processing result w6; the third multiplication processing result w3 and the sixth multiplication processing result w6 are summed, and the obtained sum processing result is normalized to obtain a normalized processing result w7; the normalized processing result w7 and the third linear processing result w8 are multiplied (that is, the third linear processing result is multiplied by the corresponding weight v to obtain the third linear processing result w8 for multiplying with the normalized processing result w7), to obtain an attention processing result R2.
[0108] In step 102, based on the second image sample, a second encoder is called for encoding processing to obtain a second low-dimensional feature of the second image sample, and based on the second low-dimensional feature, a second processing network and a fourth processing network are respectively called to correspondingly obtain a second high-dimensional feature and a fourth high-dimensional feature.
[0109] As an example, refer to Figure 4B , based on the second image sample x2, a second encoder is called for encoding processing to obtain a second low-dimensional feature Z2 of the second image sample, and based on the second low-dimensional feature Z2, a second processing network and a fourth processing network are respectively called to correspondingly obtain a second high-dimensional feature f2(x) and a fourth high-dimensional feature f4(x).
[0110] In some embodiments, the second processing network includes a second mapper, and the fourth processing network includes a fourth encoder and a fourth mapper with spatial focus characteristics. Among them, the structures of the second mapper and the fourth mapper may be the same. The main structures of the second mapper and the fourth mapper may be composed of a Multilayer Perceptron (MLP), and the main structures of the second mapper and the fourth mapper may include fully connected layers (FC), an activation layer, and a Batch Normalization (BN) layer. That is, the above-mentioned second mapper and fourth mapper may be mappers with the same structure but different parameters.
[0111] In some embodiments, referring to Figure 3C , in step 102 above, the second processing network and the fourth processing network are respectively called based on the second low-dimensional feature, and the corresponding second high-dimensional feature and fourth high-dimensional feature can be obtained by executing steps 1021 to 1022.
[0112] In step 1021, the second mapper in the second processing network is called based on the second low-dimensional feature for mapping processing, and the obtained second mapping processing result is used as the second high-dimensional feature.
[0113] As an example, referring to Figure 4D , the second encoder is called based on the second image sample x2 for encoding processing to obtain the second low-dimensional feature Z2 of the second image sample x2. The second mapper in the second processing network is called based on the second low-dimensional feature Z2 for mapping processing, and the obtained second mapping processing result is used as the second high-dimensional feature f2(x).
[0114] In step 1022, the fourth encoder in the fourth processing network is called based on the second low-dimensional feature for encoding processing to obtain a fourth encoding result; the fourth mapper in the fourth processing network is called based on the fourth encoding result for mapping processing, and the obtained fourth mapping processing result is used as the fourth high-dimensional feature.
[0115] As an example, referring to Figure 4D , the second encoder is called based on the second image sample x2 for encoding processing to obtain the second low-dimensional feature Z2 of the second image sample x2. The fourth encoder in the fourth processing network is called based on the second low-dimensional feature Z2 for encoding processing to obtain a fourth encoding result; the fourth mapper in the fourth processing network is called based on the fourth encoding result for mapping processing, and the obtained fourth mapping processing result is used as the fourth high-dimensional feature f4(x).
[0116] In some embodiments, the above-mentioned fourth encoder includes a plurality of serial sub-encoders with the same structure, and each serial sub-encoder includes a first convolutional layer, a second convolutional layer, and a self-attention layer. That is, the third encoder and the fourth encoder described above can be encoders with the same structure.
[0117] As an example, refer to Figure 4F , Figure 4F In the schematic structural diagram of the fourth encoder provided by the embodiment of the present application shown by reference numeral 41, the fourth encoder includes a plurality of serial sub-encoders with the same structure, and each serial sub-encoder includes a first convolutional layer, a second convolutional layer, and a self-attention layer. The third encoder and the fourth encoder provided in the embodiment of the present application can be two encoders with the same structure and different parameters.
[0118] In some embodiments, the encoding process of calling the fourth encoder in the fourth processing network based on the second low-dimensional feature in step 1022 to obtain the fourth encoding result can be achieved in the following manner: calling the first convolutional layer based on the second low-dimensional feature for convolutional processing to obtain a first convolutional processing result; calling the self-attention layer based on the first convolutional processing result for attention processing to obtain an attention processing result; calling the second convolutional layer based on the attention processing result for convolutional processing to obtain a second convolutional processing result; adding the second convolutional processing result and the second low-dimensional feature to obtain the fourth encoding result.
[0119] As an example, refer to Figure 4F In reference numeral 42, taking the processing process of one serial sub-encoder as an example, calling the first convolutional layer based on the second low-dimensional feature for convolutional processing to obtain a first convolutional processing result R1; calling the self-attention layer based on the first convolutional processing result R1 for attention processing to obtain an attention processing result R2; calling the second convolutional layer based on the attention processing result R2 for convolutional processing to obtain a second convolutional processing result R3; adding the second convolutional processing result R3 and the second low-dimensional feature to obtain the fourth encoding result R4.
[0120] In some embodiments, the above-mentioned attention processing includes linear processing, multiplication processing, summation processing, and normalization processing. That is, the process of attention processing can be achieved through linear processing, multiplication processing, summation processing, and normalization processing. Among them, linear processing can be a process of linear matrix transformation, multiplication processing can be a process of performing multiplication operations, summation processing can be a process of performing summation operations, and normalization processing is a process of enhancing attention. The processed result after normalization can significantly improve the attention of the data to be processed compared with the data to be processed before normalization.
[0121] In some embodiments, the above-mentioned self-attention layer is called based on the first convolution processing result for focus processing to obtain a focus processing result, which can be implemented in the following manner: the first convolution processing result is respectively subjected to a first linear processing, a second linear processing, and a third linear processing to correspondingly obtain a first linear processing result, a second linear processing result, and a third linear processing result; the first linear processing result and the second linear processing result are respectively multiplied by corresponding weights to obtain a first multiplication processing result corresponding to the first linear processing result and a second multiplication processing result corresponding to the second linear processing result; the first multiplication processing result and the second multiplication processing result are multiplied to obtain a third multiplication processing result; the first linear processing result and the timing information are respectively multiplied by corresponding weights to obtain a fourth multiplication processing result corresponding to the first linear processing result and a fifth multiplication processing result corresponding to the timing information; the fourth multiplication processing result and the fifth multiplication processing result are multiplied to obtain a sixth multiplication processing result; the third multiplication processing result and the sixth multiplication processing result are summed, and the obtained sum processing result is normalized to obtain a normalization processing result; the normalization processing result and the third linear processing result are multiplied to obtain a focus processing result.
[0122] As an example, refer to Figure 4FAt label 42, the first convolution processing result R1 is respectively subjected to first linear processing, second linear processing, and third linear processing, corresponding to obtaining a first linear processing result, a second linear processing result, and a third linear processing result. The first linear processing result and the second linear processing result are respectively multiplied by corresponding weights, that is, the first linear processing result is multiplied by the weight value q corresponding to the first linear processing result, and the second linear processing result is multiplied by the weight value k corresponding to the second linear processing result, respectively obtaining a first multiplication processing result w1 corresponding to the first linear processing result and a second multiplication processing result w2 corresponding to the second linear processing result; the first multiplication processing result w1 and the second multiplication processing result w2 are multiplied to obtain a third multiplication processing result w3; the first linear processing result and the timing information are respectively multiplied by corresponding weights, that is, the first linear processing result is multiplied by the corresponding weight q, and the timing information is multiplied by the corresponding weight p, respectively obtaining a fourth multiplication processing result w4 corresponding to the first linear processing result and a fifth multiplication processing result w5 corresponding to the timing information; the fourth multiplication processing result w4 and the fifth multiplication processing result w5 are multiplied to obtain a sixth multiplication processing result w6; the third multiplication processing result w3 and the sixth multiplication processing result w6 are summed, and the obtained summation processing result is normalized to obtain a normalized processing result w7; the normalized processing result w7 and the third linear processing result w8 are multiplied (that is, the third linear processing result is multiplied by the corresponding weight v to obtain the third linear processing result w8 for multiplying with the normalized processing result w7), obtaining a focus processing result.
[0123] In step 103, a first cross-loss value between the first encoder and the second encoder is determined according to the first high-dimensional feature and the third high-dimensional feature, and a second cross-loss value between the first encoder and the second encoder is determined according to the second high-dimensional feature and the fourth high-dimensional feature. The parameters of the first encoder are gradient-updated based on the first cross-loss value and the second cross-loss value.
[0124] In some embodiments, the way of gradient update can be achieved by means of batch gradient update, stochastic gradient update, etc. Among them, the implementation method of batch gradient update is to use all parameters to update each parameter when updating each parameter, that is, to update each parameter according to the negative direction of the gradient of each parameter. The implementation method of stochastic gradient update is to use the partial derivative of the loss function of each parameter to obtain the corresponding gradient, and use the obtained gradient to update.
[0125] As an example, refer to Figure 4B , a first cross-loss value L between the first encoder and the second encoder is determined according to the first high-dimensional feature f1(x) and the third high-dimensional feature f3(x) att1, and determine the second cross-loss value \(L\) between the first encoder and the second encoder according to the second highest-dimensional feature \(f2(x)\) and the fourth highest-dimensional feature \(f4(x)\) att2 , based on the first cross-loss value \(L\) att1 and the second cross-loss value \(L\) att2 perform gradient update on the parameters of the first encoder.
[0126] As an example, the first cross-loss value \(L\) att1 and the second cross-loss value \(L\) att2 can be expressed as:
[0127] \(L\) att1 = f3(x) - f1(x) (1)
[0128] \(L\) att2 = f4(x) - f2(x) (2)
[0129] In some embodiments, before performing gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value in step 103 above, the parameters of the first processing network and the third processing network can also be updated by the following method: based on the first cross-loss value and the second cross-loss value, perform gradient update on the parameters of the first processing network and the third processing network in parallel.
[0130] As an example, referring to Figure 4B , based on the first cross-loss value \(L\) att1 and the second cross-loss value \(L\) att2 , perform gradient update on the parameters of the first processing network and the third processing network in parallel.
[0131] As an example, referring to Figure 4C , based on the first cross-loss value \(L\) att1 and the second cross-loss value \(L\) att2 , perform gradient update on the parameters of the first predictor in the first processing network, and the third encoder and the third predictor in the third processing network in parallel.
[0132] As an example, referring to Figure 4D , based on the first cross-loss value \(L\) att1 and the second cross-loss value \(L\) att2 , perform gradient update on the parameters of the first mapper and the first predictor in the first processing network, and the third encoder, the third mapper and the third predictor in the third processing network in parallel.
[0133] In some embodiments, referring to Figure 3B , Figure 3BIt is a schematic flowchart of a method for training an image feature extraction network provided by an embodiment of the present application. Before determining the second cross-loss value between the first encoder and the second encoder in step 103 above, the update of the first encoder is achieved by executing steps 105 to 106.
[0134] In step 105, according to the difference between the first high-dimensional feature and the second high-dimensional feature, determine the first loss value of the first encoder, and according to the difference between the third high-dimensional feature and the fourth high-dimensional feature, determine the second loss value of the first encoder.
[0135] As an example, see Figure 4B , according to the difference between the first high-dimensional feature f1(x) and the second high-dimensional feature f2(x), determine the first loss value L C of the first encoder, and according to the difference between the third high-dimensional feature f3(x) and the fourth high-dimensional feature f4(x), determine the second loss value L t of the first encoder. Among them, the expressions of the first loss value L C and the second loss value L t can be:
[0136] L t = f4(x) - f3(x) (3)
[0137] L c = f4(x) - f3(x) (4)
[0138] In step 106, based on the first loss value and the second loss value, perform gradient update on the parameters of the first encoder.
[0139] As an example, see Figure 4B , based on the first loss value L C and the second loss value L t perform gradient update on the parameters of the first encoder.
[0140] In some embodiments, before performing gradient update on the parameters of the first encoder based on the first loss value and the second loss value in step 106 above, the parameters of the first processing network and the third processing network can also be updated in the following manner: based on the first loss value and the second loss value, perform gradient update on the parameters of the first processing network and the third processing network in parallel.
[0141] In step 104, according to the parameters of the first encoder after gradient update, perform momentum update on the parameters of the second encoder.
[0142] As an example, see Figure 4B , according to the parameters of the first encoder after gradient update, perform momentum update on the parameters of the second encoder.
[0143] In some embodiments, the process of momentum update is an extension of gradient update, and the way of momentum update is more efficient than that of gradient update. Momentum update, also known as gradient update based on momentum, is an update method that accelerates the change of the gradient vector in the relevant direction and finally realizes accelerated convergence.
[0144] In some embodiments, referring to Figure 3B , the momentum update of the parameters of the second encoder in step 104 above can be achieved by executing steps 1041 to 1043.
[0145] In step 1041, multiply the parameters of the first encoder after gradient update by the weight value corresponding to the first encoder to obtain the first momentum update parameter.
[0146] As an example, when the parameters of the first encoder before gradient update are , the parameters of the first encoder after gradient update are The weight value corresponding to the first encoder is α, and the first momentum update parameter is Then the first momentum update parameter is The expression of can be:
[0147]
[0148] In step 1042, multiply the parameters of the second encoder before momentum update by the weight value corresponding to the second encoder to obtain the second momentum update parameter.
[0149] As an example, when the parameters of the second encoder before momentum update are The weight value corresponding to the second encoder is β, and the second momentum update parameter is ε. Then the expression of the second momentum update parameter ε can be:
[0150]
[0151] Among them, the sum of the weight value corresponding to the first encoder and the weight value corresponding to the second encoder is 1, that is, α + β = 1.
[0152] In step 1043, sum the first momentum update parameter and the second momentum update parameter, and determine the result of the summation process as the parameters of the second encoder after momentum update.
[0153] As an example, when the parameters of the second encoder after momentum update are , the parameters of the second encoder after momentum update are The expression of can be:
[0154]
[0155] In some embodiments, before performing momentum update on the parameters of the second encoder in step 104 above, the second processing network and the fourth processing network are updated by performing the following operations: according to the parameters of the updated first processing network, perform momentum update on the parameters of the second processing network; according to the parameters of the updated third processing network, perform momentum update on the parameters of the fourth processing network.
[0156] As an example, refer to Figure 4B , perform momentum update on the parameters of the second processing network according to the parameters of the first processing network after gradient update, and perform momentum update on the parameters of the fourth processing network according to the parameters of the third processing network after gradient update.
[0157] As an example, refer to Figure 4C , since the second processing network is an empty network, there is no need to perform momentum update on the parameters of the second processing network according to the parameters of the first processing network after gradient update. And since the fourth processing network includes a fourth encoder, then momentum update can be performed on the parameters of the fourth encoder in the fourth processing network according to the parameters of the third encoder in the third processing network after gradient update.
[0158] As an example, refer to Figure 4D , perform momentum update on the parameters of the second mapper in the second processing network according to the parameters of the first mapper in the first processing network after gradient update, perform momentum update on the parameters of the fourth encoder in the fourth processing network according to the parameters of the third encoder in the third processing network after gradient update, and perform momentum update on the parameters of the fourth mapper in the fourth processing network according to the parameters of the third mapper in the third processing network after gradient update.
[0159] In this way, since the efficiency of momentum update is significantly higher than that of gradient update, by performing momentum update on the parameters of the second processing network and the fourth processing network, the update efficiency is significantly improved.
[0160] In some embodiments, since the inputs of the first processing network and the third processing network are both the first low-dimensional features, and the inputs of the second processing network and the fourth processing network are both the second low-dimensional features, the first processing network and the third processing network need to adopt different structures to complete the conversion from low-dimensional features to high-dimensional features, and the second processing network and the fourth processing network need to adopt different structures to complete the conversion from low-dimensional features to high-dimensional features, so as to ensure that there is a certain degree of difference between the obtained first high-dimensional features and third high-dimensional features, and there is a certain degree of difference between the second high-dimensional features and fourth high-dimensional features, ensuring that the first cross-loss value and the second cross-loss value calculated subsequently are non-zero values, because when the first cross-loss value and the second cross-loss value are zero values, it will cause the first encoder and the second encoder not to be updated, thus affecting the update effect.
[0161] For example, see Figure 4D , Figure 4D which is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiment of the present application. When the third processing network includes a third encoder, that is, the third processing network includes a third encoder, a third mapper, and a third predictor. At this time, since the third processing network has an additional third encoder compared with the first processing network, it is ensured that there is a certain degree of difference between the obtained first high-dimensional feature and the third high-dimensional feature. At this time, the first mapper and the third mapper can be mappers with the same structure, and the number of layers of the first mapper and the third mapper can be the same. The first predictor and the third predictor can be predictors with the same structure, and the number of layers of the first predictor and the third predictor can be the same.
[0162] For example, see Figure 4E , Figure 4E which is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiment of the present application. When the third processing network does not include a third encoder, that is, the third processing network includes a third mapper and a third predictor. At this time, in order to ensure that there is a certain degree of difference between the obtained first high-dimensional feature and the third high-dimensional feature, the first mapper and the third mapper can be mappers with the same structure, but it is necessary to ensure that the number of layers of the first mapper and the third mapper is different, so that there is a certain degree of difference in the processing results between the first mapper and the third mapper. Similarly, the first predictor and the third predictor can be predictors with the same structure, but it is necessary to ensure that the number of layers of the first predictor and the third predictor is different, so that there is a certain degree of difference in the processing results between the first predictor and the third predictor, so that the corresponding first cross-loss value has sufficient diversity to ensure the update effect of the image feature extraction network.
[0163] The feature extraction method based on the image feature extraction network provided by the embodiment of the present application will be described in combination with the exemplary applications and implementations of the server provided by the embodiment of the present application.
[0164] Next, see Figure 3D , Figure 3D which is a schematic flowchart of the feature extraction method based on the image feature extraction network provided by the embodiment of the present application, and will be described in combination with Figure 3D the steps 201 to 203 shown.
[0165] In step 201, the trained first encoder is called based on the first preprocessed image corresponding to the image to be processed for encoding to obtain the encoded feature of the first preprocessed image.
[0166] For example, see Figure 4G , Figure 4GIt is a schematic diagram of the principle of the feature extraction method based on the image feature extraction network provided by the embodiment of the present application. Based on the first preprocessed image x1 corresponding to the image X to be processed, the trained first encoder is called for encoding processing to obtain the encoded feature y1 of the first preprocessed image.
[0167] In step 202, based on the second preprocessed image corresponding to the image to be processed, the trained second encoder is called for encoding processing to obtain the encoded feature of the second preprocessed image.
[0168] As an example, refer to Figure 4G , based on the second preprocessed image x2 corresponding to the image X to be processed, the trained second encoder is called for encoding processing to obtain the encoded feature y2 of the second preprocessed image.
[0169] In step 203, the encoded feature of the first preprocessed image and the encoded feature of the second preprocessed image are fused to obtain the feature of the image to be processed.
[0170] As an example, refer to Figure 4G , the encoded feature y1 of the first preprocessed image and the encoded feature y2 of the second preprocessed image are fused to obtain the feature y of the image to be processed. Among them, the fusion process can be a process of adding each corresponding feature in the encoded feature y1 of the first preprocessed image and the encoded feature y2 of the second preprocessed image.
[0171] In this way, by training the first encoder and the second encoder, the trained first encoder and the trained second encoder are obtained respectively. By using the trained first encoder and the trained second encoder to extract image features, the obtained encoded features can achieve more perfect performance in downstream image classification, image detection, image segmentation and other processing processes.
[0172] Next, the exemplary application of the embodiment of the present application in an actual application scenario will be described.
[0173] Embodiments of this application can have the following application scenarios. For example, the application of artificial intelligence in the field of medical health (i.e., medical AI (Artificial Intelligence)) has been very extensive. From the perspective of application scenarios, it mainly includes medical imaging, drug discovery, health management, wearable devices and other main scenarios. For example, in the medical imaging application scenario, it is usually difficult to obtain training samples with annotation information (i.e., medical imaging samples), which makes it impossible to effectively train the image feature extraction network, resulting in an unsatisfactory feature extraction effect for medical images. The training method of the image feature extraction network provided by the embodiments of this application can be widely applied to application scenarios where training samples lack annotation information or where it is difficult to obtain annotation information of training samples, effectively improving the training effect of the image feature extraction network.
[0174] The training method of the image feature extraction network provided by the embodiments of this application uses parallel CNN encoders (i.e., the first encoder and the second encoder above) and Transformer encoders (i.e., the third encoder and the fourth encoder above), and introduces a training strategy for the image feature extraction network, enabling the CNN encoder to achieve better performance in downstream tasks (such as image classification, object detection, semantic segmentation, etc.), so that the CNN encoder can more effectively learn higher-quality, more spatially focused high-dimensional features.
[0175] The training method of the image feature extraction network provided by the embodiments of this application introduces a Transformer encoder (i.e., the third encoder and the fourth encoder above) for auxiliary training in the design. The embodiments of this application use the Transformer encoder (i.e., the third encoder and the fourth encoder above) to specifically guide and supervise the features of the CNN encoder (i.e., the first encoder and the second encoder above) during the training process, so that the CNN encoder (i.e., the first encoder and the second encoder above) can more effectively learn higher-quality, more spatially focused high-dimensional features. The image feature extraction network provided by the embodiments of this application introduces 4 forward branches (i.e., the first processing network, the second processing network, the third processing network, and the fourth processing network above), mainly including 2 CNN branches (i.e., the first processing network and the second processing network above) and two Transformer branches (i.e., the third processing network and the fourth processing network above). See Figure 5A , Figure 5AIt is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiments of the present application. The CNN branch includes two CNN encoders, two projectors, and a predictor. The Transformer branch includes two Transformer encoders, two projectors, and a predictor. In the specific implementation of the embodiments of the present application, first, the picture x is preprocessed in two different ways. Subsequently, the two pictures (x1 and x2) after the two different preprocessings are respectively sent into two paths of the CNN branch to obtain two types of high-dimensional features f1(x) and f2(x). At the same time, the two pictures with the two different preprocessings also need to be input into two paths of the Transformer branch to obtain another two types of high-dimensional features f3(x) and f4(x). In one training, the embodiments of the present application initially train the CNN branch by minimizing the distance between f1(x) and f2(x), and initially train the Transformer branch by minimizing the distance between f3(x) and f4(x). At the same time, the embodiments of the present application make the CNN branch and the Transformer branch interact by minimizing the distance between f1(x) and f3(x), and minimizing the distance between f2(x) and f4(x), that is, using the more spatially focused features output by the Transformer branch to supervise the output features of the CNN branch, so as to gradually strengthen the effective feature extraction ability of the CNN encoder during the training process. The embodiments of the present application also propose three loss functions, which can help the CNN encoder perform more effective representation learning.
[0176] The embodiments of the present application help the CNN encoder (i.e., the first encoder and the second encoder above) to perform more effective learning by introducing two paths of Transformer encoders (i.e., the third encoder and the fourth encoder above). Through the supervision of the CNN encoder by the Transformer encoder on high-level features, the CNN encoder can more effectively extract higher-quality and more spatially focused features on different input images. The CNN encoder (i.e., the first encoder and the second encoder above) trained by the embodiments of the present application can achieve better performance in downstream image classification, image detection, and image segmentation tasks.
[0177] In the embodiments of the present application, by comparing positive and negative samples (i.e., the first image sample and the second image sample above), the additional computational overhead and memory overhead are reduced. At the same time, the embodiments of the present application combine the advantages of the Transformer encoder (i.e., the third encoder and the fourth encoder above) with relatively high spatial focus and consideration of global features, and propose a self-supervised learning framework for the parallel CNN encoder (i.e., the first encoder and the second encoder above) and the Transformer encoder (i.e., the third encoder and the fourth encoder above). Through the targeted design of the loss function, the image feature extraction network trained in the embodiments of the present application can more effectively learn a higher-quality CNN encoder (i.e., the first encoder and the second encoder above) with higher spatial attention.
[0178] In the embodiments of the present application, referring to Figure 5A , first, the input picture x (i.e., image x) is preprocessed twice in different ways to obtain two positive samples x1 and x2. Two sets of features of x1 and x2 are respectively extracted by two CNN encoders of two CNN branches. Among them, the CNN encoder of one CNN branch inputs the extracted features into a mapper (projector1) and a predictor (predictor1) respectively to obtain high-dimensional features f1(x). At the same time, the features extracted by the other CNN are only input into the mapper with momentum update (momentum projector1) to obtain high-dimensional features f2(x). In addition, the two sets of features extracted by the two CNN branches are also input into the Transformer branch at the same time. Among them, one Transformer1 branch extracts features with higher spatial focus, and inputs these features into the mapper (projector1) and the predictor (predictor1) in sequence to obtain high-dimensional features f3(x). The other Transformer branch (momentum Transformer1) also extracts corresponding features and inputs them into the mapper with momentum update (momentum projector1) to obtain high-dimensional features f4(x). Thus, the forward propagation process is completed. Subsequently, through the targeted loss function (L c 、L att 、L t)Perform backpropagation learning on the design. During the backpropagation learning process, only update one CNN encoder and one Transformer encoder in the CNN branch and the Transformer branch (i.e., CNN encoder1 and Transformer1), while the other CNN encoder and Transformer encoder in the CNN branch and the Transformer branch (i.e., momentum CNN encoder1 and momentum Transformer1) will be updated using the means of momentum update.
[0179] The CNN encoder obtained by training the embodiments of the present application can effectively extract image features. The CNN encoders provided by the embodiments of the present application (i.e., the first encoder and the second encoder above) have certain applicability. Therefore, there is good flexibility in the selection of the CNN encoder. For example, ResNet50 encoder, ResNet101 encoder, ResNet152 encoder, etc. can all be used as the CNN encoders provided by the embodiments of the present application (i.e., the first encoder and the second encoder above).
[0180] The Transformer encoder provided by the embodiments of the present application (i.e., the third encoder and the fourth encoder above) takes the output of the CNN encoder (i.e., the first encoder and the second encoder above) as input and outputs features with more spatial focus properties. The Transformer provided by the embodiments of the present application mainly includes 4 serial modules, and a single module is shown in Figure 5B , Figure 5B is a schematic diagram of the principle of the training method of the image feature extraction network provided by the embodiments of the present application. A single module mainly includes a 1x1 convolutional layer (i.e., conv layer(1×1)), a multi-head self-attention layer (Multi-head Self-attention, MHSA), and an additional 1x1 convolutional layer (i.e., conv layer(1×1)). Among them, the multi-head self-attention layer can well learn features with spatial focus properties. The structure of the multi-head self-attention layer can include a first linear processing structure (W q ), a second linear processing structure (W k ), a third linear processing structure (W v ), disordered information (pos), a normalization processing structure (softmax), a product structure (i.e., the multiplication processing described above), and a summation structure (i.e., the summation processing described above).
[0181] The main structures of the projector and predictor provided by the embodiments of this application are Multilayer Perceptrons (MLPs). Both the projector and predictor include two fully connected layers (FCs), an activation layer, and a Batch Normalization (BN) layer.
[0182] The training method of the image feature extraction network provided by the embodiments of this application includes three types of loss functions. First, by comparing the distances between f1(x) and f2(x), the first type of loss function is used to narrow the distance between f1(x) and f2(x) in the high-level space to better train the CNN branch. Second, in the embodiments of this application, by comparing the distances between f3(x) and f4(x), the loss function of the Transformer branch is designed to narrow the distance between f3(x) and f4(x) in the high-level space to better train the Transformer branch. Third, in the embodiments of this application, through the design of the interaction loss function, the distances between f1(x) and f3(x) and between f2(x) and f4(x) are compared, and f1(x) is supervised by f3(x), while f2(x) is supervised by f4(x). The above three parts of the loss function are mainly used to update the CNN branch and the Transformer branch.
[0183] In the embodiments of this application, after the upper branches of the CNN branch and the Transformer branch are updated, their lower branches mainly update the parameters by momentum update. The so-called momentum update mainly refers to using the parameters of the upper branches of the current CNN branch and Transformer branch, as well as the parameter information at the previous moment, to update the network parameters of their lower branches by momentum.
[0184] The embodiments of this application creatively introduce a Transformer branch. By utilizing the unique spatial attention characteristics of the Transformer encoder to guide and supervise the training of the CNN encoder, the trained CNN encoder can more effectively extract higher-quality and more spatially focused features. The actual tests of downstream tasks show that the CNN encoder trained by the embodiments of this application has higher performance. At the same time, the embodiments of this application can play a positive role in improving the performance of different CNN encoders. Therefore, the training method of the image feature extraction network provided by the embodiments of this application has a certain degree of universality.
[0185] Next, an exemplary structure of the training device 455A of the image feature extraction network provided in the embodiments of the present application as a software module will be further described. In some embodiments, as Figure 2A shown, the software module stored in the training device 455A of the image feature extraction network in the memory 450 may include: a first processing module 4551A, configured to call a first encoder to perform encoding processing on a first image sample to obtain a first low-dimensional feature of the first image sample, and based on the first low-dimensional feature, respectively call a first processing network and a third processing network to correspondingly obtain a first high-dimensional feature and a third high-dimensional feature; a second processing module 4552A, configured to call a second encoder to perform encoding processing on a second image sample to obtain a second low-dimensional feature of the second image sample, and based on the second low-dimensional feature, respectively call a second processing network and a fourth processing network to correspondingly obtain a second high-dimensional feature and a fourth high-dimensional feature; a first update module 4553A, configured to determine a first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determine a second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, and perform gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value; a second update module 4554A, configured to perform momentum update on the parameters of the second encoder according to the updated parameters of the first encoder.
[0186] In some embodiments, the above-mentioned training device 455A of the image feature extraction network further includes: a first determination module, configured to determine a first loss value of the first encoder according to the difference between the first high-dimensional feature and the second high-dimensional feature, and determine a second loss value of the first encoder according to the difference between the third high-dimensional feature and the fourth high-dimensional feature; a third update module, configured to perform gradient update on the parameters of the first encoder based on the first loss value and the second loss value.
[0187] In some embodiments, the above-mentioned second update module 4554A is further configured to multiply the parameters of the first encoder after gradient update by the weight value corresponding to the first encoder to obtain a first momentum update parameter; multiply the parameters of the second encoder before momentum update by the weight value corresponding to the second encoder to obtain a second momentum update parameter; wherein the sum of the weight value corresponding to the first encoder and the weight value corresponding to the second encoder is 1; sum the first momentum update parameter and the second momentum update parameter, and determine the sum result as the parameters of the second encoder after momentum update.
[0188] In some embodiments, the above-mentioned first processing module 4551A is further configured to perform mapping processing by invoking a first mapper in the first processing network based on the first low-dimensional feature to obtain a first mapping processing result; perform prediction processing by invoking a first predictor in the first processing network based on the first mapping processing result, and use the obtained first prediction processing result as the first high-dimensional feature; perform encoding processing by invoking a third encoder in the third processing network based on the first low-dimensional feature to obtain a third encoding result; perform mapping processing by invoking a third mapper in the third processing network based on the third encoding result to obtain a third mapping processing result; perform prediction processing by invoking a third predictor in the third processing network based on the third mapping processing result, and use the obtained third prediction processing result as the third high-dimensional feature.
[0189] In some embodiments, the training device 455A of the above-mentioned image feature extraction network further includes: a fourth update module, configured to perform gradient update on the parameters of the first processing network and the third processing network in a parallel manner based on the first cross-loss value and the second cross-loss value.
[0190] In some embodiments, the above-mentioned first processing module 4551A is further configured to perform convolution processing by invoking a first convolutional layer based on the first low-dimensional feature to obtain a first convolution processing result; perform attention processing by invoking a self-attention layer based on the first convolution processing result to obtain an attention processing result; perform convolution processing by invoking a second convolutional layer based on the attention processing result to obtain a second convolution processing result; perform addition processing on the second convolution processing result and the first low-dimensional feature to obtain a third encoding result.
[0191] In some embodiments, the above-mentioned second processing module 4552A is further configured to perform mapping processing by invoking a second mapper in the second processing network based on the second low-dimensional feature, and use the obtained second mapping processing result as the second high-dimensional feature; perform encoding processing by invoking a fourth encoder in the fourth processing network based on the second low-dimensional feature to obtain a fourth encoding result; perform mapping processing by invoking a fourth mapper in the fourth processing network based on the fourth encoding result, and use the obtained fourth mapping processing result as the fourth high-dimensional feature.
[0192] In some embodiments, the training device 455A of the above-mentioned image feature extraction network further includes: a fifth update module, configured to perform momentum update on the parameters of the second processing network according to the updated parameters of the first processing network; perform momentum update on the parameters of the fourth processing network according to the updated parameters of the third processing network.
[0193] In some embodiments, the above-mentioned second processing module 4552A is further configured to perform convolution processing on the first convolutional layer based on the second low-dimensional feature to obtain a first convolution processing result; perform attention processing on the self-attention layer based on the first convolution processing result to obtain an attention processing result; perform convolution processing on the second convolutional layer based on the attention processing result to obtain a second convolution processing result; perform an addition process on the second convolution processing result and the second low-dimensional feature to obtain a fourth encoding result.
[0194] In some embodiments, the above-mentioned second processing module 4552A is further configured to perform first linear processing, second linear processing, and third linear processing on the first convolution processing result respectively to obtain a first linear processing result, a second linear processing result, and a third linear processing result; perform a multiplication process on the first linear processing result and the second linear processing result with corresponding weights respectively to obtain a first multiplication processing result corresponding to the first linear processing result and a second multiplication processing result corresponding to the second linear processing result; perform a multiplication process on the first multiplication processing result and the second multiplication processing result to obtain a third multiplication processing result; perform a multiplication process on the first linear processing result and the timing information with corresponding weights respectively to obtain a fourth multiplication processing result corresponding to the first linear processing result and a fifth multiplication processing result corresponding to the timing information; perform a multiplication process on the fourth multiplication processing result and the fifth multiplication processing result to obtain a sixth multiplication processing result; perform a summation process on the third multiplication processing result and the sixth multiplication processing result, perform a normalization process on the obtained summation processing result to obtain a normalization processing result; perform a multiplication process on the normalization processing result and the third linear processing result to obtain an attention processing result.
[0195] In some embodiments, the training device 455A of the above-mentioned image feature extraction network further includes: a first preprocessing module, configured to perform first preprocessing on the to-be-processed image sample to obtain a first image sample; a second preprocessing module, configured to perform second preprocessing on the to-be-processed image sample to obtain a second image sample; wherein, the first preprocessing and the second preprocessing are reverse processing processes to each other.
[0196] In some embodiments, such as Figure 2BAs shown in the figure, the software module stored in the feature extraction device 655B of the image feature extraction network in the memory 650 may include: a first encoding module 6551B, configured to call the trained first encoder to perform encoding processing on the first preprocessed image corresponding to the image to be processed, so as to obtain the encoded features of the first preprocessed image; a second encoding module 6552B, configured to call the trained second encoder to perform encoding processing on the second preprocessed image corresponding to the image to be processed, so as to obtain the encoded features of the second preprocessed image; and a fusion module 6553B, configured to perform fusion processing on the encoded features of the first preprocessed image and the encoded features of the second preprocessed image, so as to obtain the features of the image to be processed.
[0197] An embodiment of the present application provides an electronic device, including: a memory, configured to store executable instructions; and a processor, configured to implement the training method of the image feature extraction network provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0198] An embodiment of the present application provides an electronic device, including: a memory, configured to store executable instructions; and a processor, configured to implement the feature extraction method based on the image feature extraction network provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0199] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the training method of the image feature extraction network provided by the embodiment of the present application when executed.
[0200] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the feature extraction method based on the image feature extraction network provided by the embodiment of the present application when executed.
[0201] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the training method of the image feature extraction network described above in the embodiment of the present application.
[0202] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the feature extraction method based on the image feature extraction network described above in the embodiment of the present application.
[0203] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0204] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0205] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program in question, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).
[0206] As an example, the executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0207] In summary, the embodiments of the present application have the following beneficial effects:
[0208] (1) By setting different numbers of mappers and predictors in the first processing network and the third processing network, the differential processing of the first processing network and the third processing network can be ensured, so that the high-dimensional features obtained by the first processing network and the third processing network maintain a certain difference, thereby making the corresponding loss values have sufficient diversity, which is convenient for subsequent training of the image feature extraction network.
[0209] (2) By differentially processing the image samples to be processed, the first image sample and the second image sample are obtained, so that the first image sample and the second image sample are not completely the same. Since the first image sample and the second image sample both originate from the same image sample to be processed, the first image sample and the second image sample have a certain similarity. Thus, the first image sample and the second image sample are similar but not completely the same, so that it is convenient for subsequent training of the image feature extraction network.
[0210] (3) By training the first encoder and the second encoder respectively, the trained first encoder and the trained second encoder are obtained. Image features are extracted by the trained first encoder and the trained second encoder, so that the obtained encoded features achieve more perfect performance in downstream processes such as image classification, image detection, and image segmentation.
[0211] (4) By respectively calling the first encoder and the second encoder based on the first image sample and the second image sample for encoding processing, different processing networks are respectively called for processing the obtained low-dimensional features to obtain corresponding high-dimensional features. The cross-loss value is determined by the obtained high-dimensional features, so as to use the cross-loss value to perform gradient update on the first encoder, and perform momentum update on the second encoder according to the first encoder after gradient update. Since completely different update methods are adopted for different encoders during training, the training efficiency and training effect of the image feature extraction network are effectively improved.
[0212] (5) Since the efficiency of momentum update is significantly higher than that of gradient update, by performing momentum update on the parameters of the second processing network and the fourth processing network, the update efficiency is significantly improved.
[0213] (6) Since the first predictor and the third predictor can be predictors with the same structure, but it is necessary to ensure that the number of layers of the first predictor and the third predictor is different, so that there is a certain degree of difference in the processing results of the first predictor and the third predictor, so that the corresponding first cross-loss value has sufficient diversity to ensure the update effect of the image feature extraction network.
[0214] (7) By enabling the first encoder and the second encoder to interact during training, completely different update methods can be adopted for different encoders during training, and thus more effective representation learning of image features can be performed by the image feature extraction network.
[0215] (8) Normalization processing is a processing process to improve concentration. The processing result after normalization can significantly improve the concentration of the data to be processed compared with the data to be processed before normalization.
[0216] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A training method for an image feature extraction network, characterized in that The image feature extraction network includes a first encoder and a second encoder, and the method includes: Invoking the first encoder based on a first image sample for encoding processing to obtain a first low-dimensional feature of the first image sample, and respectively invoking a first processing network and a third processing network based on the first low-dimensional feature to correspondingly obtain a first high-dimensional feature and a third high-dimensional feature; Invoking the second encoder based on a second image sample for encoding processing to obtain a second low-dimensional feature of the second image sample, and respectively invoking a second processing network and a fourth processing network based on the second low-dimensional feature to correspondingly obtain a second high-dimensional feature and a fourth high-dimensional feature; Determining a first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determining a second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, and performing gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value; Performing a multiplication process on the parameters of the first encoder after gradient update and the weight value corresponding to the first encoder to obtain a first momentum update parameter; Performing a multiplication process on the parameters of the second encoder before momentum update and the weight value corresponding to the second encoder to obtain a second momentum update parameter; wherein, the sum of the weight value corresponding to the first encoder and the weight value corresponding to the second encoder is 1; Performing a summation process on the first momentum update parameter and the second momentum update parameter, and determining the result of the summation process as the parameters of the second encoder after momentum update.
2. The method according to claim 1, wherein Before determining the first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determining the second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, the method further includes: Determining a first loss value of the first encoder according to the difference between the first high-dimensional feature and the second high-dimensional feature, and determining a second loss value of the first encoder according to the difference between the third high-dimensional feature and the fourth high-dimensional feature; Performing gradient update on the parameters of the first encoder based on the first loss value and the second loss value.
3. The method according to claim 1, wherein The first processing network includes a first mapper and a first predictor, and the third processing network includes a third encoder with spatial focus characteristics, a third mapper, and a third predictor; The respectively invoking the first processing network and the third processing network based on the first low-dimensional feature to correspondingly obtain a first high-dimensional feature and a third high-dimensional feature includes: Invoking the first mapper in the first processing network based on the first low-dimensional feature for mapping processing to obtain a first mapping processing result; invoking the first predictor in the first processing network based on the first mapping processing result for prediction processing, and taking the obtained first prediction processing result as the first high-dimensional feature; Based on the first low-dimensional feature, call the third encoder in the third processing network for encoding processing to obtain a third encoding result; based on the third encoding result, call the third mapper in the third processing network for mapping processing to obtain a third mapping processing result; based on the third mapping processing result, call the third predictor in the third processing network for prediction processing, and use the obtained third prediction processing result as the third high-dimensional feature; Before performing gradient update on the parameters of the first encoder based on the first cross-entropy loss value and the second cross-entropy loss value, the method further includes: Based on the first cross-entropy loss value and the second cross-entropy loss value, perform gradient update on the parameters of the first processing network and the third processing network in parallel.
4. The method according to claim 3, wherein The third encoder includes a plurality of serially-connected sub-encoders with the same structure, and each serially-connected sub-encoder includes a first convolutional layer, a second convolutional layer, and a self-attention layer; The step of calling the third encoder in the third processing network for encoding processing based on the first low-dimensional feature to obtain a third encoding result includes: Based on the first low-dimensional feature, call the first convolutional layer for convolutional processing to obtain a first convolutional processing result; based on the first convolutional processing result, call the self-attention layer for attention processing to obtain an attention processing result; Based on the attention processing result, call the second convolutional layer for convolutional processing to obtain a second convolutional processing result; perform an addition operation on the second convolutional processing result and the first low-dimensional feature to obtain a third encoding result.
5. The method according to claim 1, wherein The second processing network includes a second mapper, and the fourth processing network includes a fourth encoder and a fourth mapper with spatial attention characteristics; The step of calling the second processing network and the fourth processing network respectively based on the second low-dimensional feature to obtain a second high-dimensional feature and a fourth high-dimensional feature includes: Based on the second low-dimensional feature, call the second mapper in the second processing network for mapping processing, and use the obtained second mapping processing result as the second high-dimensional feature; Based on the second low-dimensional feature, call the fourth encoder in the fourth processing network for encoding processing to obtain a fourth encoding result; based on the fourth encoding result, call the fourth mapper in the fourth processing network for mapping processing, and use the obtained fourth mapping processing result as the fourth high-dimensional feature; Before multiplying the parameters of the first encoder after gradient update by the weight value corresponding to the first encoder to obtain a first momentum update parameter, the method further includes: According to the updated parameters of the first processing network, perform momentum update on the parameters of the second processing network; According to the updated parameters of the third processing network, perform momentum update on the parameters of the fourth processing network.
6. The method according to claim 5, wherein The fourth encoder includes a plurality of serially-connected sub-encoders with the same structure, and each serially-connected sub-encoder includes a first convolutional layer, a second convolutional layer, and a self-attention layer; Invoking a fourth encoder in the fourth processing network for encoding processing based on the second low-dimensional feature to obtain a fourth encoding result, including: Invoking a first convolutional layer for convolutional processing based on the second low-dimensional feature to obtain a first convolutional processing result; invoking a self-attention layer for attentiveness processing based on the first convolutional processing result to obtain an attentiveness processing result; Invoking a second convolutional layer for convolutional processing based on the attentiveness processing result to obtain a second convolutional processing result; adding the second convolutional processing result and the second low-dimensional feature to obtain a fourth encoding result.
7. The method according to claim 6, wherein: The attentiveness processing includes linear processing, multiplication processing, summation processing, and normalization processing; Invoking the self-attention layer for attentiveness processing based on the first convolutional processing result to obtain an attentiveness processing result, including: Performing first linear processing, second linear processing, and third linear processing on the first convolutional processing result respectively to obtain a first linear processing result, a second linear processing result, and a third linear processing result; multiplying the first linear processing result and the second linear processing result with corresponding weights respectively to obtain a first multiplication processing result corresponding to the first linear processing result and a second multiplication processing result corresponding to the second linear processing result; multiplying the first multiplication processing result and the second multiplication processing result to obtain a third multiplication processing result; multiplying the first linear processing result and the timing information with corresponding weights respectively to obtain a fourth multiplication processing result corresponding to the first linear processing result and a fifth multiplication processing result corresponding to the timing information; multiplying the fourth multiplication processing result and the fifth multiplication processing result to obtain a sixth multiplication processing result; Summing the third multiplication processing result and the sixth multiplication processing result, normalizing the obtained summation processing result to obtain a normalization processing result; multiplying the normalization processing result and the third linear processing result to obtain an attentiveness processing result.
8. The method according to claim 1, characterized in that Before invoking the first encoder for encoding processing based on the first image sample, the method further includes: Performing a first preprocessing on the image sample to be processed to obtain the first image sample; Performing a second preprocessing on the image sample to be processed to obtain the second image sample; wherein the first preprocessing and the second preprocessing are inverse processing processes to each other.
9. A feature extraction method based on an image feature extraction network, applied to the image feature extraction network according to any one of claims 1 to 8, characterized in that, The method includes: Invoking the trained first encoder for encoding processing based on the first preprocessed image corresponding to the image sample to be processed to obtain the encoding feature of the first preprocessed image; Invoking the trained second encoder for encoding processing based on the second preprocessed image corresponding to the image sample to be processed to obtain the encoding feature of the second preprocessed image; Fusing the encoding feature of the first preprocessed image and the encoding feature of the second preprocessed image to obtain the feature of the image sample to be processed.
10. A training device for an image feature extraction network, characterized in that, The image feature extraction network includes a first encoder and a second encoder, and the apparatus includes: A first processing module, configured to call the first encoder for encoding processing based on a first image sample, obtain a first low-dimensional feature of the first image sample, and respectively call a first processing network and a third processing network based on the first low-dimensional feature to correspondingly obtain a first high-dimensional feature and a third high-dimensional feature; A second processing module, configured to call the second encoder for encoding processing based on a second image sample, obtain a second low-dimensional feature of the second image sample, and respectively call a second processing network and a fourth processing network based on the second low-dimensional feature to correspondingly obtain a second high-dimensional feature and a fourth high-dimensional feature; A first updating module, configured to determine a first cross-loss value between the first encoder and the second encoder according to the first high-dimensional feature and the third high-dimensional feature, and determine a second cross-loss value between the first encoder and the second encoder according to the second high-dimensional feature and the fourth high-dimensional feature, and perform gradient update on the parameters of the first encoder based on the first cross-loss value and the second cross-loss value; A second updating module, configured to multiply the parameters of the first encoder after gradient update by the weight value corresponding to the first encoder to obtain a first momentum update parameter; multiply the parameters of the second encoder before momentum update by the weight value corresponding to the second encoder to obtain a second momentum update parameter; wherein, the sum of the weight value corresponding to the first encoder and the weight value corresponding to the second encoder is 1; sum the first momentum update parameter and the second momentum update parameter, and determine the sum result as the parameters of the second encoder after momentum update.
11. The apparatus according to claim 10, wherein A first determining module, configured to determine a first loss value of the first encoder according to the difference between the first high-dimensional feature and the second high-dimensional feature, and determine a second loss value of the first encoder according to the difference between the third high-dimensional feature and the fourth high-dimensional feature; A third updating module, configured to perform gradient update on the parameters of the first encoder based on the first loss value and the second loss value.
12. The device according to claim 10, characterized in that, The first processing network includes a first mapper and a first predictor, and the third processing network includes a third encoder with spatial focus characteristics, a third mapper, and a third predictor; The first processing module is further configured to call the first mapper in the first processing network for mapping processing based on the first low-dimensional feature to obtain a first mapping processing result; call the first predictor in the first processing network for prediction processing based on the first mapping processing result, and use the obtained first prediction processing result as the first high-dimensional feature; Call the third encoder in the third processing network for encoding processing based on the first low-dimensional feature to obtain a third encoding result; call the third mapper in the third processing network for mapping processing based on the third encoding result to obtain a third mapping processing result; Based on the third mapping processing result, call the third predictor in the third processing network to perform prediction processing, and use the obtained third prediction processing result as the third high-dimensional feature; A fourth update module, configured to perform gradient update on the parameters of the first processing network and the third processing network in parallel based on the first cross-loss value and the second cross-loss value.
13. The device according to claim 12, characterized in that, The third encoder includes a plurality of serially-connected sub-encoders with the same structure, and each serially-connected sub-encoder includes a first convolutional layer, a second convolutional layer, and a self-attention layer; The first processing module is further configured to perform convolutional processing on the first low-dimensional feature by calling the first convolutional layer to obtain a first convolutional processing result; perform attention processing on the first convolutional processing result by calling the self-attention layer to obtain an attention processing result; Perform convolutional processing on the attention processing result by calling the second convolutional layer to obtain a second convolutional processing result; perform an addition process on the second convolutional processing result and the first low-dimensional feature to obtain a third encoding result.
14. A feature extraction device based on an image feature extraction network, applied to the image feature extraction network according to any one of claims 1 to 9, characterized in that, The feature extraction device based on the image feature extraction network includes: A first encoding module, configured to perform encoding processing on the first preprocessed image corresponding to the image to be processed by calling the trained first encoder to obtain the encoded feature of the first preprocessed image; A second encoding module, configured to perform encoding processing on the second preprocessed image corresponding to the image to be processed by calling the trained second encoder to obtain the encoded feature of the second preprocessed image; A fusion module, configured to perform fusion processing on the encoded feature of the first preprocessed image and the encoded feature of the second preprocessed image to obtain the feature of the image to be processed.
15. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the training method of the image feature extraction network according to any one of claims 1 to 8, or the feature extraction method based on the image feature extraction network according to claim 9 when executing the executable instructions stored in the memory.
16. A computer-readable storage medium stores executable instructions, characterized in that, The executable instructions, when executed by the processor, implement the training method of the image feature extraction network according to any one of claims 1 to 8, or the feature extraction method based on the image feature extraction network according to claim 9.
17. A computer program product, comprising computer instructions, characterized in that, The computer instructions, when executed by the processor, implement the training method of the image feature extraction network according to any one of claims 1 to 8, or the feature extraction method based on the image feature extraction network according to claim 9.
Citation Information
Patent Citations
Image semantic segmentation method and device based on codec
CN111292330A