Image data processing method and device, computer device and storage medium
Patent Information
- Application Number
- CN202211070487.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-02
AI Technical Summary
[0003]传统的图像数据处理方法,基于卷积神经网络进行编码处理,得到编码特征,受限于卷积神经网络复杂的网络结构,训练时间长且对计算资源的要求高,极大的限制了图像数据处理方法在实际落地中的应用
[0020]上述图像数据处理方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,获取图像样本,对该图像样本进行图像分割,得到该图像样本对应的多个图像子块,然后,基于多层神经网络中每一层各自对应的分组信息,采用分组信息相同的各层共用同一个注意力层的方式,使用该多层神经网络分别对多个图像子块中的各目标子块进行注意力池化处理和特征编码处理,得到每一目标子块各自对应的编码特征。上述图像数据处理过程中,通过对多层神经网络中各层分别配置对应的分组信息,且分组信息相同的各层共用同一个注意力层的方式,可以实现灵活的参数共享,能匹配不同场景下的实际计算资源,有利于扩展图像数据处理方法的应用场景。
Smart Images

Figure CN117036368B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to an image data processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of deep learning technology, image data processing methods based on deep learning technology have emerged. These methods encode image samples to obtain their coded features, enabling subsequent tasks such as image recognition, image classification, or image reconstruction.
[0003] Traditional image data processing methods rely on convolutional neural networks for encoding to obtain coded features. However, due to the complex network structure of convolutional neural networks, the training time is long and the requirements for computing resources are high, which greatly limits the application of image data processing methods in practical applications. Summary of the Invention
[0004] Therefore, it is necessary to provide an image data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can expand application scenarios to address the above-mentioned technical problems.
[0005] Firstly, this application provides an image data processing method. The method includes:
[0006] Obtain image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples;
[0007] Based on the grouping information corresponding to each layer in the multi-layer neural network, attention pooling and feature encoding are performed on each target sub-block in the plurality of image sub-blocks to obtain the encoded features corresponding to each target sub-block; layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0008] Secondly, this application provides an image data processing apparatus. The apparatus includes:
[0009] An image sub-block acquisition module is used to acquire image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples.
[0010] The encoding module is used to perform attention pooling and feature encoding on each target sub-block in the plurality of image sub-blocks based on the grouping information corresponding to each layer in the multi-layer neural network, so as to obtain the encoding features corresponding to each target sub-block; the layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0011] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0012] Obtain image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples;
[0013] Based on the grouping information corresponding to each layer in the multi-layer neural network, attention pooling and feature encoding are performed on each target sub-block in the plurality of image sub-blocks to obtain the encoded features corresponding to each target sub-block; layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0014] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0015] Obtain image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples;
[0016] Based on the grouping information corresponding to each layer in the multi-layer neural network, attention pooling and feature encoding are performed on each target sub-block in the plurality of image sub-blocks to obtain the encoded features corresponding to each target sub-block; layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0017] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0018] Obtain image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples;
[0019] Based on the grouping information corresponding to each layer in the multi-layer neural network, attention pooling and feature encoding are performed on each target sub-block in the plurality of image sub-blocks to obtain the encoded features corresponding to each target sub-block; layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0020] The aforementioned image data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire an image sample, segment the image sample to obtain multiple image sub-blocks corresponding to the image sample, and then, based on the grouping information corresponding to each layer in a multi-layer neural network, use a method where layers with the same grouping information share the same attention layer. The multi-layer neural network is then used to perform attention pooling and feature encoding processing on each target sub-block within the multiple image sub-blocks to obtain the encoded features corresponding to each target sub-block. In this image data processing process, by configuring corresponding grouping information for each layer in the multi-layer neural network, and by having layers with the same grouping information share the same attention layer, flexible parameter sharing can be achieved. This allows for matching with actual computing resources in different scenarios, which is beneficial for expanding the application scenarios of the image data processing method. Attached Figure Description
[0021] Figure 1 This is an application environment diagram of the image data processing method in some embodiments;
[0022] Figure 2 This is a flowchart illustrating the image data processing method in some embodiments;
[0023] Figure 3 This is a schematic diagram illustrating the key areas of focus for each layer of a multi-layer neural network on the same image sub-block in some embodiments;
[0024] Figure 4 This is a schematic diagram illustrating the image restoration effect during the pre-training stage in some embodiments;
[0025] Figure 5 This is a flowchart illustrating the image data processing method in some other embodiments;
[0026] Figure 6 This is a comparative diagram of the embodiments of this application and conventional solutions;
[0027] Figure 7 This is a schematic diagram of the image data processing process in some embodiments;
[0028] Figure 8 This is a structural block diagram of the image data processing apparatus in some embodiments;
[0029] Figure 9 This is a diagram showing the internal structure of a computer device in some embodiments. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0032] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing, tracking, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR (optical character recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D (3D) technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0033] The solutions provided in this application involve artificial intelligence technologies such as computer vision and deep learning, and are specifically illustrated through the following embodiments:
[0034] The image data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on the cloud or other servers. Specifically, during image data processing, server 104: acquires an image sample, performs image segmentation on the image sample to obtain multiple image sub-blocks corresponding to the image sample, and, based on the grouping information corresponding to each layer in a multi-layer neural network, uses a method where layers with the same grouping information share the same attention layer. This multi-layer neural network is then used to perform attention pooling and feature encoding processing on each target sub-block within the multiple image sub-blocks to obtain the encoded features corresponding to each target sub-block.
[0035] In some embodiments, when the data processing capability of terminal 102 meets the data processing requirements, the image data processing method provided in this application embodiment may be applied only to terminal 102. Specifically, terminal 102 acquires image samples and determines the encoding features of each target sub-block in the image samples based on a multi-layer neural network.
[0036] The terminal 102 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. This invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0037] In some embodiments, such as Figure 2 As shown, an image data processing method is provided. This embodiment illustrates the method applied to server 104. It is understood that this method can also be applied to terminal 102, and further to a system including terminal 102 and server 104, and implemented through the interaction between terminal 102 and server 104. In this embodiment, the method includes the following steps:
[0038] Step S202: Obtain an image sample, perform image segmentation on the image sample, and obtain multiple image sub-blocks corresponding to the image sample.
[0039] The image samples can be training samples used during the image model training phase, or image samples to be processed during the image model application phase. The image model can refer to a pre-trained image model unrelated to a specific task, or a target image model matched to specific tasks such as image classification or image recognition. In the image model training phase, if the image model is a pre-trained model, the image samples can be obtained from an open-source model pre-training dataset; this open-source model pre-training dataset could be, for example, ImageNet, CIFAR100, or iNat19, etc.; if the image model is an image classification model, the image samples can be obtained from a training sample set carrying classification labels. In the image model application phase, if the image model is an image classification model, the image samples can be obtained from a set of unlabeled samples to be classified.
[0040] Furthermore, image segmentation refers to the process of dividing an image into multiple image sub-blocks, each of which may or may not be the same size. In a specific application, each image sub-block is a rectangular sub-block of the same size to facilitate subsequent feature encoding processing.
[0041] Specifically, the server can acquire an image sample and perform image segmentation on it to obtain multiple image sub-blocks corresponding to the image sample. The specific method by which the server performs image segmentation to obtain these multiple image sub-blocks is not unique. For example, the server can use the Unet image segmentation network or traditional image segmentation methods to directly segment the image sample into multiple image sub-blocks. Alternatively, the server can first extract features from the image sample to obtain a feature image corresponding to the image sample, and then segment the feature image into multiple image sub-blocks, which are the multiple image sub-blocks corresponding to the image sample.
[0042] Step S204: Based on the grouping information corresponding to each layer in the multi-layer neural network, the multi-layer neural network is used to perform attention pooling and feature encoding on each target sub-block in multiple image sub-blocks to obtain the encoded features corresponding to each target sub-block.
[0043] In this context, a multilayer neural network refers to an artificial neural network containing multiple intermediate layers. Specifically, in this application, each layer of the multilayer neural network includes an attention layer and an encoding layer, and layers with the same grouping information share the same attention layer. That is, each layer in the multilayer neural network needs to perform attention pooling and feature encoding on the target sub-block, and layers with the same grouping information share parameters during the attention pooling process. The grouping information corresponding to a certain layer in the multilayer neural network refers to the information used to characterize the group of that layer. This grouping information can include at least one of numbers and letters. Taking a multilayer neural network with six layers as an example, the grouping information of each layer can be (1,2,3,4,5,6), meaning each layer has different grouping information; it can also be (1,1,1,1,1,1), meaning each layer has the same grouping information; it can also be (1,2,2,3,4,5), meaning the second and third layers have the same grouping information, and so on.
[0044] A target sub-block refers to the processing object of a multi-layer neural network, and this target sub-block is one of multiple image sub-blocks. It can be understood that each target sub-block input to the multi-layer neural network can be at least a portion of multiple image sub-blocks. In a specific application, the server performs masking processing on image samples to filter out the unmasked target sub-blocks from multiple image sub-blocks.
[0045] Masking refers to the process of masking at least a portion of an image region, or at least a portion of multiple image sub-blocks. The masking position is usually random, while the masking ratio can be random or fixed. The fixed masking ratio could be, for example, 10%, 20%, or 30%, etc. A target sub-block refers to an unmasked image sub-block among multiple image sub-blocks corresponding to an image sample. Correspondingly, image sub-blocks other than the target sub-block are masked image sub-blocks, which contain missing image information.
[0046] Specifically, the server performs masking processing on image samples, and the method for selecting the target sub-block from multiple image sub-blocks is not unique. For example, the server can first perform image segmentation and masking processing on the same image sample to obtain a mask image, and then determine the masked sub-blocks and the unmasked target sub-blocks from the multiple image sub-blocks based on the position information of each image sub-block corresponding to the image sample in the mask image. Alternatively, the server can first perform image segmentation on the image sample to obtain multiple image sub-blocks, then mask at least a portion of these multiple image sub-blocks, and determine the target sub-block and masked sub-blocks from these multiple image sub-blocks.
[0047] Furthermore, pooling is essentially dimensionality reduction sampling, which can both preserve the image features of the target sub-block and reduce image information redundancy. Attention pooling is a general pooling method with input allocation preferences, allowing neural networks to ignore unimportant features and focus on calculating useful features. This improves computational speed while discarding useless features that interfere with the fitting results. In neural networks, attention pooling is mainly implemented through attention scores. An attention score is a value between 0 and 1. Under the attention pooling mechanism, the sum of the attention scores of each term is 1, and each attention score represents the attention weight assigned to the current term. For different regions within the same target sub-block, regions with higher attention scores contain more useful features.
[0048] Feature coding refers to the process of converting image data into coded features based on image feature coding algorithms. These image feature coding algorithms can specifically be spatial coding algorithms, transform coding algorithms, etc. Among them, spatial coding algorithms can be coding algorithms based on convolutional neural networks.
[0049] Specifically, such as Figure 3 As shown, for multi-layer neural networks, each layer identifies a similar region of focus for the same sub-block of an image, such as... Figure 3The darker areas are represented by the color. In other words, using cross-layer parameter sharing for the attention layer does not significantly impact the performance of the pre-trained model. On the other hand, since the input to the attention layer is segmented image blocks, and each pixel in each block participates in the computation, attention pooling for each block is equivalent to pixel-level attention scoring of the image samples. That is, the attention layer is pixel-level, and cross-layer parameter sharing can significantly reduce the number of parameters in image data processing, thereby reducing the number of parameters in the pre-trained model determined by the image data processing method, which is beneficial for improving data processing efficiency. Based on this, users can group the layers in the multi-layer neural network according to the model accuracy and computational resource requirements of specific application scenarios, determining the corresponding grouping information for each layer. Then, the server, based on the grouping information of each layer in the multi-layer neural network, uses layers with the same grouping information to share the same attention layer, and uses this multi-layer neural network to perform attention pooling and feature encoding processing on each target block, obtaining the encoded features corresponding to each target block.
[0050] It is understandable that if the grouping information corresponding to each layer in a multi-layer neural network is different, then cross-layer parameter sharing will not occur during image data processing; however, if each layer in a multi-layer neural network corresponds to the same grouping information, then all layers will share the same attention layer during image data processing. In other words, by configuring corresponding grouping information for each layer in a multi-layer neural network, flexible parameter sharing can be achieved to match the needs of different application scenarios.
[0051] In a specific application, the server can achieve parameter sharing between attention layers through the Cross Layer Parameter Sharing (CLPS) module. Specifically, in a multi-layer neural network, each layer performs attention pooling and feature encoding on the target sub-blocks sequentially according to its corresponding stacking order. Assuming that the grouping information of the second and third layers in the multi-layer neural network is the same, then during the image data processing in the second layer, attention pooling is first performed using the attention layer to obtain an attention score. During the image data processing in the third layer, the same attention layer as the second layer is called through CLPS, and the parameters of that attention layer are shared to obtain the corresponding attention score.
[0052] Furthermore, in some embodiments, the grouping information includes pooling grouping information and encoding grouping information. In a multi-layer neural network, layers with the same pooling grouping information share a single attention layer, and layers with the same encoding grouping information share a single encoding layer. The server can achieve parameter sharing between attention layers through a pooling parameter sharing module and parameter sharing between encoding layers through an encoding parameter sharing module. For details on the specific implementation of encoding parameter sharing, please refer to the description of pooling parameter sharing above; it will not be repeated here. Specifically, by configuring corresponding pooling grouping information and encoding grouping information for each layer in a multi-layer neural network, cross-layer pooling parameter sharing and / or cross-layer encoding parameter sharing can be achieved, which is beneficial for further improving the flexibility of parameter sharing and expanding the application scenarios of image data processing methods.
[0053] The image data processing method described above acquires an image sample, segments the image sample to obtain multiple image sub-blocks, and then, based on the grouping information of each layer in a multi-layer neural network, uses a method where layers with the same grouping information share the same attention layer. This multi-layer neural network is then used to perform attention pooling and feature encoding on each target sub-block within the multiple image sub-blocks, obtaining the encoded features corresponding to each target sub-block. In this image data processing process, by configuring corresponding grouping information for each layer in the multi-layer neural network, and ensuring that layers with the same grouping information share the same attention layer, flexible inter-layer parameter sharing can be achieved. This allows for matching with actual computing resources in different scenarios, which is beneficial for expanding the application scenarios of image data processing methods.
[0054] As mentioned above, after obtaining the encoded features, tasks such as image recognition, image classification, or image reconstruction can be performed based on these features. Taking image reconstruction as an example, in some embodiments, the target sub-block is obtained by masking the image sample and filtering it from multiple image sub-blocks. In this embodiment, the image data processing method further includes: decoding the image sample based on the encoded features corresponding to each target sub-block and the learnable features corresponding to the masked sub-blocks other than the target sub-blocks in the multiple image sub-blocks to obtain the pre-trained reconstructed image corresponding to the image sample.
[0055] The specific limitations regarding masking image samples and selecting target sub-blocks from multiple image sub-blocks are detailed above and will not be repeated here. Learnable features refer to the feature information used to characterize the features of the mask sub-blocks. During the image masking autoencoding process, each mask sub-block is represented by a shared learnable feature. Specifically, the server can determine the order of each image sub-block based on its position information within the image sample. Then, the learnable features and each encoded feature are sorted according to the order of their corresponding image sub-blocks, input into the decoder for decoding, and the decoder output is linearly projected to obtain the pre-trained reconstructed image corresponding to the image sample.
[0056] like Figure 4 The diagram shows the original image, mask image, and pre-trained reconstructed image for each of the six image samples. For each image sample, the first image is the mask image, the second is the pre-trained reconstructed image, and the third is the original image. Figure 4 It can be seen that by performing mask autoencoder processing on image samples, the missing pixels can be reconstructed, and the pre-trained reconstructed image corresponding to the image sample can be obtained.
[0057] In the above embodiments, after obtaining the encoding features of the target sub-block, the missing image of the mask sub-block is reconstructed based on the encoding features. This can be applied to the image model pre-training process and is beneficial to expanding the application scenarios of the pre-training method determined by the image data processing method.
[0058] In some embodiments, the number of image samples is multiple. In this embodiment, the image data processing method further includes: obtaining a pre-trained image model when the pre-trained reconstructed image corresponding to each image sample and the original image meet the similarity condition; and using the pre-trained image model as the teacher model to supervise the training of the initial student model to obtain a target image model that matches the image processing task.
[0059] The similarity condition can be either that the similarity between the pre-trained reconstructed image and the original image corresponding to the same image sample is greater than a similarity threshold, or that the similarity between the pre-trained reconstructed image and the original image corresponding to the same image sample is greater than or equal to a similarity threshold. This similarity can be represented by at least one of image similarity and semantic similarity. Furthermore, the teacher model refers to the model used to provide guidance during the knowledge distillation process, and correspondingly, the student model refers to the model being guided during the knowledge distillation process.
[0060] Specifically, if the pre-trained reconstructed image corresponding to each image sample satisfies the similarity condition with the original image, it indicates that the learning objective of the pre-training process has been achieved, and a pre-trained image model with a certain learning ability can be obtained. Then, the server can use the pre-trained image model as the teacher model, inputting training samples into both the teacher model and the initial student model simultaneously. By referring to the prediction results of the training samples obtained from the teacher model and the target loss function between the teacher model and the initial student model, the initial student model is trained under supervision to obtain a target image model that matches the image processing task.
[0061] In the above embodiments, knowledge distillation is performed on the initial student model based on the trained pre-trained image model to obtain the target image model. This can effectively improve the initial student model's ability to acquire scene information and its data generalization ability, thereby improving the accuracy of the target image model's prediction results.
[0062] In some embodiments, the image processing task is an image classification task; the target image model is an image classification model. In this embodiment, using a pre-trained image model as the teacher model to supervise the training of the initial student model to obtain a target image model that matches the image processing task includes: using a pre-trained image model as the teacher model to supervise the training of the initial student model to obtain an image classification model that matches the image classification task.
[0063] The initial student model comprises multiple attention layers, and at least a portion of the training parameters for each attention layer are identical during supervised training. "At least a portion of the training parameters for each attention layer are identical" can mean that at least a portion of the training parameters within each attention layer are the same, or that each attention layer uses at least a portion of its training parameters during training, or that a portion of the training parameters used by a subset of attention layers are the same during training. In a specific application, during supervised training, the server obtains the grouping information corresponding to each attention layer in the initial student model and uses a method where layers with the same grouping information share a single attention layer to supervise the training of the initial model.
[0064] Specifically, considering that for multi-layer neural networks, each layer identifies a similar region of focus for the same image sub-block, using cross-layer parameter sharing for the attention layer does not significantly impact model performance. However, it can substantially reduce the number of model parameters and improve data processing efficiency. Based on this, the server uses a pre-trained image model as the teacher model and employs cross-layer attention parameter sharing to supervise the training of the initial student model, thereby obtaining an image classification model that matches the image classification task.
[0065] In the above embodiments, it is equivalent to using cross-layer attention sharing to fine-tune the model in the downstream task stage, which can further improve data processing efficiency while expanding the application scenarios of image data processing methods.
[0066] In some embodiments, image segmentation is performed on an image sample to obtain multiple image sub-blocks corresponding to the image sample, including: extracting features from the image sample to obtain image features of the image sample, and determining the feature image represented by the image features; and performing image segmentation on the feature image to obtain multiple image sub-blocks corresponding to the image sample.
[0067] Image features refer to the characteristic information used to characterize the features of an image sample. The specific form of image features can be a vector or a matrix, and the characteristics of the image sample represented by the image features can be at least one of color, line direction, or semantics. Specifically, for an image sample, the original image is a low-level, pixel-level representation containing a large amount of redundant information. Based on this, after the server obtains the image sample, it can extract features from the image sample to obtain the image features of the image sample, further determine the feature image represented by the image features, and then perform image segmentation on the feature image to obtain multiple feature image sub-blocks corresponding to the feature image, which are the multiple image sub-blocks corresponding to the image sample.
[0068] In the above embodiments, image samples are first converted into feature images through feature extraction, and then image segmentation is performed to obtain multiple image sub-blocks corresponding to the image samples. This is equivalent to the input of the multi-layer neural network changing from the original image to the feature image, which helps the multi-layer neural network model extract higher-level semantics from the image samples, thereby improving the model accuracy of the neural network model obtained based on image data processing.
[0069] Furthermore, the specific algorithm for extracting image features from image samples is not unique. For example, a server can use a feature extraction network to perform feature extraction and normalization on image samples to obtain their image features. This feature extraction process may include convolution and pooling processes, etc. The number of convolutions in the convolution process can be one, two, or more.
[0070] In some embodiments, feature extraction of an image sample to obtain the image features of the image sample includes: performing a first convolution process and a second convolution process on the image sample in sequence to obtain candidate image features of the image sample; and performing pooling processing on each candidate image feature to obtain the image features of the image sample.
[0071] In this process, the size of the first convolution kernel in the first convolution process is larger than the size of the second convolution kernel in the second convolution process. For example, if the first convolution kernel size is 7, the second convolution kernel size could be 3, 4, 5, or 6. Specifically, after the server acquires an image sample, it sequentially performs the first and second convolution processes to obtain candidate image features. Then, pooling is applied to each candidate image feature to obtain the final image feature. The specific pooling algorithm can be average pooling or max pooling.
[0072] Furthermore, the server can design the parameters in the first convolutional processing, the second convolutional processing, and the pooling processing to ensure that the size of the feature image remains consistent with the output size of image patch embedding modules in other ViT-B structures. This improves the compatibility of image data processing methods with other correlation methods and further expands the application scenarios of image data processing methods. In a specific application, the server maps the 3 channels of the original image to 768 dimensions and obtains the image features of the image samples through two convolutional layers and one max pooling layer. Each convolutional layer consists of a three-layer network structure: Conv2d→Flatten→BatchNorm. The parameters of the first convolutional layer are kernel size = 7 and stride = 2, the parameters of the second convolutional layer are kernel size = 4 and stride = 4, and the parameters of the max pooling layer are kernel size = 3 and stride = 2. After processing by two convolutional layers and the max pooling layer, the image size is reduced to 1 / 16 of the original image, maintaining the same output size as image patch embedding modules in other ViT-B (Vision Transformer) structures.
[0073] In the above embodiments, it is equivalent to using progressive convolution to extract image features from image samples. This can remove redundant information from image samples while ensuring the robustness of the image data processing method. It can also ensure the anti-noise interference performance of the image data processing method in various application scenarios, which is conducive to further expanding the application scenarios of the image data processing method.
[0074] In some embodiments, the process of determining grouping information includes: obtaining the number of layers and the number of groups in a multi-layer neural network; determining the number of layers corresponding to each group based on the number of layers and the number of groups; and determining the grouping information corresponding to each layer in the multi-layer neural network based on the number of layers.
[0075] In this multi-layer neural network, the number of groups is less than the number of layers. The difference in the number of layers between groups is less than a set difference in the number of layers, and the layers within the same group are consecutive layers in the multi-layer neural network. This set difference in the number of layers can be, for example, 1, 2, or 3. Taking a multi-layer neural network with 3 groups and 12 layers as an example: If the set difference in the number of layers is 1, then layers 1 to 4 form one group, layers 5 to 8 form another group, and layers 9 to 12 form another group; if the set difference in the number of layers is 3, then layers 1 to 3 can form one group, layers 4 to 8 can form another group, and layers 9 to 12 can form another group, or layers 1 to 4 can form one group, layers 5 to 7 can form another group, and layers 8 to 12 can form another group, or layers 1 to 4 can form one group, layers 5 to 8 can form another group, and layers 9 to 12 can form another group, and so on.
[0076] Specifically, the server can obtain the number of layers and the number of groups in the multi-layer neural network. Then, based on the number of layers and the number of groups, it determines the number of layers corresponding to each group. Finally, based on the number of layers, it determines the grouping information corresponding to each layer in the multi-layer neural network to ensure that the difference in the number of layers between groups is less than the set difference in the number of layers, and that the layers in the same group are consecutive layers in the multi-layer neural network.
[0077] In a specific application, the layer difference is set to 2. Since the layer difference between groups is less than 2, the number of layers in each group is equal, or the layer difference is 1. Based on this, the server can first divide the number of layers by the number of groups to obtain a quotient and a remainder. Then, the quotient is used as the initial layer number for each group. Next, a remainder of target groups are randomly selected from each group, and the initial layer number is increased by 1 to obtain the layer number for each target group. The layer numbers of the remaining non-target groups are equal to the initial layer numbers, thus obtaining the grouping information corresponding to each layer in the multilayer neural network. It can be understood that if the remainder is 0, the number of layers in each group is equal, and all are initial layer numbers equal to the quotient.
[0078] Since the attention scores between consecutive layers in a multi-layer neural network are more similar, in the above embodiment, during the process of determining grouping information, the difference in the number of layers between each group is less than the set difference in the number of layers, and each layer in the same group is a consecutive layer in the multi-layer neural network, which can further reduce the impact of cross-layer parameter sharing on the model performance of the neural network model obtained based on image data processing methods.
[0079] In some embodiments, the attention layer has at least two attention heads; the process of performing attention pooling on the target sub-block includes: extracting features from the target sub-block based on the target attention head to obtain first candidate sub-feature information of the target sub-block, and determining shared parameters in the feature extraction process; using the shared parameters, extracting features from the target sub-block based on each of the other attention heads (excluding the target attention head) among the at least two attention heads to obtain second candidate sub-feature information corresponding to each attention head; and fusing the first candidate sub-feature information and each of the second candidate sub-feature information to obtain the target sub-feature information of the target sub-block.
[0080] In this context, the target attention head can be any one of at least two attention heads. An attention layer with at least two attention heads is also called a multi-head attention layer. For example... Figure 3 In the diagram, "L6H5" represents the fifth attention head in the sixth attention layer, and "L6H6" represents the sixth attention head in the sixth attention layer. Figure 3 It can be seen that "L6H5" and "L6H6" focus on similar regions, meaning that within the same attention layer, different attention heads focus on similar regions. Based on this, the server can further improve data processing efficiency in image data processing by sharing parameters among different attention heads within the same attention layer.
[0081] Specifically, the server extracts features from the target sub-block based on the target attention head to obtain the first candidate sub-feature information of the target sub-block, and determines the shared parameters in the feature extraction process. Then, using the shared parameters, the server extracts features from the target sub-block based on each of the at least two attention heads other than the target attention head, obtaining multiple second candidate sub-feature information of the target sub-block. Finally, the server fuses the first candidate sub-feature information and each of the second candidate sub-feature information to obtain the target sub-feature information of the target sub-block. The shared parameters can be at least one of a query weight matrix, a key weight matrix, or a value weight matrix. The specific method for fusing the first and second candidate sub-feature information can be averaging or summing, etc.
[0082] Taking an attention layer comprising three attention heads, sharing parameters including a query weight matrix and a key weight matrix, as an example: Specifically, the server can perform feature extraction on the target sub-block based on any one of the three attention heads, obtaining the first candidate sub-feature information of the target sub-block, and determining the query weight matrix and key weight matrix in the feature extraction process. Then, the server shares this query weight matrix and key weight matrix with the other two attention heads, and uses this query weight matrix and key weight matrix to perform feature extraction on the target sub-block based on the other two attention heads, obtaining two second candidate sub-feature information of the target sub-block. Finally, the server fuses the first candidate sub-feature information and each of the second candidate sub-feature information to obtain the target sub-feature information of the target sub-block.
[0083] It should be noted that in other embodiments, attention heads can be grouped according to the actual needs of the application scenario, and the parameters of each attention head in the same group can be shared to ensure that the image data processing method can be applied to various different application scenarios, thereby further improving the flexibility of the image data processing method.
[0084] In the above embodiments, for multi-head attention layers, parameter sharing is also performed between different attention heads in the same attention layer, which can further improve the data processing efficiency in the image data processing process while ensuring model performance.
[0085] In some embodiments, the feature encoding process includes: obtaining the pooling result obtained by attention pooling of the target sub-block; performing depthwise separable convolution on the pooling result to obtain the initial encoded features of the target sub-block; and performing secondary encoding on the initial encoded features based on a multilayer perceptron to obtain the encoded features of the target sub-block.
[0086] The prototype of depthwise separable convolution can be considered to originate from the Inception module in convolutional neural networks. Its convolution computation is divided into two parts: first, spatial convolution (depthwise convolution) is performed on each channel (depth), and the outputs are concatenated; then, pointwise convolution is performed using a unit convolution kernel to obtain the feature map. Compared to conventional convolution operations, depthwise separable convolution has a lower parameter count and computational cost. Specifically, for a target sub-block, the server can obtain the pooling result obtained by attention pooling on the target sub-block, and then perform depthwise separable convolution on the pooling result to obtain the initial encoded features of the target sub-block. Then, a secondary encoding process is performed on the initial encoded features based on a multilayer perceptron to obtain the encoded features of the target sub-block.
[0087] In the above embodiments, the pooling results are first processed by depthwise separable convolution, and then a secondary encoding process is performed to obtain the encoding features of the target sub-block. This is equivalent to adding a lightweight convolutional layer to the feedforward module while keeping other operations in the feedforward module unchanged. This can further explore the spatial relationship between image sub-blocks while reducing the amount of increase in model parameters, thereby ensuring the data processing efficiency of the pre-trained model obtained based on the image data processing method.
[0088] In some embodiments, such as Figure 5 As shown, the image data processing method includes:
[0089] Step S501: Obtain image samples;
[0090] Step S502: Perform a first convolution process and a second convolution process on the image sample in sequence to obtain the candidate image features of the image sample; the size of the first convolution kernel in the first convolution process is larger than the size of the second convolution kernel in the second convolution process.
[0091] Step S503: Perform pooling processing on each candidate image feature to obtain the image features of the image sample;
[0092] Step S504: Determine the feature image represented by the image features;
[0093] Step S505: Perform image segmentation on the feature image to obtain multiple image sub-blocks corresponding to the image samples;
[0094] Step S506: Perform masking on each image sub-block and filter out the target sub-block from each image sub-block;
[0095] Step S507: Obtain the number of layers and the number of groups in the multi-layer neural network; the number of groups is less than the number of layers.
[0096] Step S508: Determine the number of layers corresponding to each group based on the number of layers and the number of groups; the difference in the number of layers between groups is less than the set difference in the number of layers.
[0097] Step S509: Based on the number of layers, determine the grouping information corresponding to each layer in the multi-layer neural network; layers in the same group are consecutive layers in the multi-layer neural network.
[0098] Step S510: Based on the grouping information corresponding to each layer in the multi-layer neural network, the multi-layer neural network is used to perform attention pooling processing on each target sub-block to obtain the pooling result corresponding to each target sub-block. In the multi-layer neural network, the layers with the same grouping information share the same attention layer, and the attention layer is a multi-head attention layer. The parameters used by each attention head in the same attention layer are the same.
[0099] Step S511: Perform depthwise separable convolution processing on the pooling result to obtain the initial encoding features of the target sub-block;
[0100] Step S512: Perform secondary encoding processing on the initial encoding features based on the multilayer perceptron to obtain the encoding features of the target sub-block;
[0101] Step S513: Obtain the learnable features corresponding to the mask sub-blocks other than the target sub-blocks in multiple image sub-blocks;
[0102] Step S514: Based on each encoded feature and each learnable feature, decode the image sample to obtain the pre-trained reconstructed image corresponding to the image sample.
[0103] Step S515: Under the condition that the pre-trained reconstructed image corresponding to each image sample and the original image meet the similarity condition, the pre-trained image model is obtained.
[0104] Step S516: Using the pre-trained image model as the teacher model, supervise the training of the initial student model to obtain an image classification model that matches the image classification task; the initial student model includes multiple attention layers, and at least some of the training parameters of each attention layer are the same during the supervised training process.
[0105] In some embodiments, the image data processing method provided in this application can be applied to image classification scenarios. In this scenario, the server acquires an image sample to be classified, performs image segmentation on the image sample to obtain multiple image sub-blocks corresponding to the image sample, and then, based on the grouping information corresponding to each layer in the multi-layer neural network, the server uses the method of having layers with the same grouping information share the same attention layer to perform attention pooling and feature encoding processing on each image sub-block respectively, to obtain the encoded features corresponding to each image sub-block, so as to classify the image sample subsequently.
[0106] In some embodiments, the image data processing method provided in this application can be applied to image model pre-training scenarios. The resulting pre-trained image model can be applied to various image-related downstream tasks such as image recognition and object image detection. Specifically, after obtaining a pre-trained image model using the ImageNet dataset, users can fine-tune the model according to their defined task requirements and dataset. This process is very lightweight and fast, and does not require a large amount of task-specific labeled data or long training time (compared to training a model from scratch). Furthermore, this application can also be applied to more complex and customized application scenarios. Users can change the pre-trained dataset, i.e., the model's prior knowledge, according to their applicable scenarios, thereby obtaining better accuracy in various refined downstream tasks. Figure 6As shown, compared with other solutions, the scheme in this application can flexibly adjust the parameter sharing method during the model pre-training process according to user needs, and can customize the pre-training dataset according to downstream tasks. It features fast training speed, flexible customization, and efficient iteration. Figure 6 As shown, compared with other solutions, the solution proposed in this application can save a lot of computing resources and significantly improve training speed.
[0107] Specifically, during image model pre-training, based on MAE (masked auto-encoder), a lighter-weight masked auto-encoder compared to traditional MAE is obtained through optimizations in the following aspects: First, instead of directly dividing the original image into small image patches, convolutional operations are added to extract features, using feature patches to replace the original image patches as the model input; second, parameter sharing technology and structural modifications to the model pipeline are used to accelerate the model's convergence speed and reduce the number of parameters. As shown in Table 1, comparing the proposed solution with traditional solutions, it is easy to see that the proposed solution has significant advantages over traditional solutions in terms of training time, number of iterations, and model accuracy. Furthermore, if the proposed solution is applied to the model pre-training and fine-tuning stages, the model accuracy is even higher, and the advantages are even more pronounced.
[0108] Table 1: Comparison of training time, number of iterations, and model accuracy between the proposed solution and traditional solutions.
[0109] ViT-B [Dosovitskiy et al., 2020] 300 100 1X 77.9 MoCo v3 [Chen et al., 2021] 300 100 1X 82.4 BEiT [Bao et al., 2021] 800 267 0.37X 83.2 DINO [Caron et al., 2021] 400 133 0.75X 82.9 MAE [He et al., 2021] 200 33 3X 82.0 MAE [He et al., 2021] 400 65 1.5X 82.3 MAE [He et al., 2021] 600 100 1X 82.8 MAE [He et al., 2021] 800 130 0.77 83 The method described in this application (applies only to the pre-training phase) 100 20 5X 82.1 The method described in this application (applied to the pre-training and fine-tuning stages) 100 20 5X 82.8
[0110] like Figure 7As shown, for any image sample, the original image of the sample is input into the Convolutional Progressive Patch Embedding (CPPE) module. Image features are extracted through progressive convolutional layers to obtain a feature map, which is then divided into image patches for vector mapping. It can be understood that the original image is a low-level, pixel-level representation containing a large amount of redundant information. This application addresses this problem by adding convolutional operations. Specifically, this application follows the original ViT-B design, mapping the 3 channels of the original image to 768 dimensions. The original module consists of a simple three-layer network: Conv2d→Flatten→BatchNorm. An additional convolutional sub-module is added on top of this. That is, the CPPE module in this application consists of two convolutional layers and one max-pooling layer. The parameters of the first convolutional layer are kernel size = 7 and stride = 2, the parameters of the second convolutional layer are kernel size = 4 and stride = 4, and the parameters of the max-pooling layer are kernel size = 3 and stride = 2. After processing through two convolutional layers, the image size is reduced to 1 / 16 of the original, maintaining the same output size as other ViT-B structure image patch embedding modules. This module's design has been meticulously crafted and extensively experimentally validated, overcoming the poor performance of single-layer convolutional modules in some scenarios due to noise interference, thus increasing the method's robustness. Through the CPPE module, the model's input changes from the original image to a feature map, which helps the model extract higher-level semantics from the image and improves the model's accuracy.
[0111] After obtaining the feature map, such as Figure 7 As shown, the server performs image segmentation on the feature map, obtaining multiple image sub-blocks corresponding to the image samples. A portion of these sub-blocks are then masked, and the unmasked target sub-blocks are selected and input into the improved encoder (CLAME encoder) of this application. The encoder is a multi-layer neural network architecture; for example, a 12-layer Transformer encoder. During the technical exploration process, as... Figure 3 As shown, by visualizing the attention parameter weights of each layer of the Transformer structure, two key conclusions can be drawn: the distribution of attention weights is singular, and most of them are concentrated in local locations. Figure 3 The darkest diagonal line; the attention distributions of different layers, and even different attention heads within the same layer, are very similar. This indicates that for the Transformer structure, different heads in each layer extract similar regions of focus for the same image sub-block, and parameter sharing has little impact on model performance.
[0112] To achieve flexible parameter sharing, we introduce the concept of "Group". Each layer is assigned to a designated Group, which contains multiple adjacent layers. Attention layers within the same Group share parameters. The number of groups is represented by the `num_hidden_groups` parameter, and the number of layers is represented by `num_hidden_layers`. If `num_hidden_groups` is 2 and `num_hidden_layers` is 12, then the layers are divided into two groups: layers 1-6 are in the first group, layers 7-12 are in the second group, and so on. If `num_hidden_groups` is 1, then all layers share the same Transformer's attention layers; if `num_hidden_groups` equals `num_hidden_layers`, then parameter sharing for attention layers does not occur. In other words, in practical applications, users can achieve flexible parameter sharing by modifying the values of the `num_hidden_groups` and `num_hidden_layers` parameters according to the accuracy and computational resource requirements of the specific task. Since the attention layer of the model is pixel-level, cross-layer parameter sharing can significantly reduce the actual number of parameters in the model. Furthermore, the cross-layer parameter sharing strategy can be implemented through the Cross Layer Parameter Sharing (CLPS) module.
[0113] For each block, after obtaining the pooling result of the target sub-block through attention pooling, the pooling result is input into the Convolutional Multi-Layer Perceptron (CMLP) in that block. For example... Figure 7 As shown, in this module, to further utilize the spatial relationships between image patches, a lightweight depth-wise convolutional layer is added to the MLP (Multilayer Perceptron) head of the feedforward module, while keeping other operations in the feedforward layer unchanged. That is, the convolutional multilayer perceptron in this application does not perform additional scaling of each layer; it remains compatible with the original MAE or VIT architecture. The computation flow in the depth-wise convolutional layer is as follows:
[0114] x1 = BN(Conv2d(Reshape(X)) input )))
[0115] In the formula, X input x1 represents the pooling result of the target sub-block; x2 represents the initial encoded features obtained after depthwise separable convolution processing.
[0116] The computation process in MLP is as follows:
[0117] x2 = Flatten(x1),
[0118] x3 = Linear1(x2),
[0119] x out =Linear2(x3)
[0120] In the formula, x2 is the dimensionality reduction result obtained after dimensionality reduction processing of the initial encoded feature x1; x3 is the initial transformation result obtained after performing a linear transformation on the dimensionality reduction result; and x is the initial transformation result obtained after dimensionality reduction processing. out The coding features of the target sub-block obtained by performing a quadratic linear transformation on the initial transformation result x3.
[0121] After encoding, the encoded features corresponding to each target sub-block can be obtained, while other mask sub-blocks are represented by shared learnable features. The encoded features and learnable features are input into the convolution decoder for decoding, and the pre-trained reconstructed image of the image sample can be obtained.
[0122] In a specific application, the ImageNet dataset is used as the training dataset. During the pre-training phase, a self-supervised approach is adopted, aiming to recover the original image after occluding 25% of the original image. In downstream image tasks, we fine-tune the trained model based on different targets and datasets. Specifically, the pre-trained image model consists of a Transformer architecture with a 12-layer encoder and a 4-layer decoder, with a default input size of 224*224. The progressive convolutional image patch linear embedding module has a kernel size of 7 and a stride of 2 for the first convolutional layer, and a kernel size of 4 and a stride of 4 for the second convolutional layer. Detailed model training parameters are shown in Table 2 below.
[0123] Table 2: Pre-training parameters
[0124] Optimization methods AdamW Initial learning rate 5.6e-4 Weight decay 0.05 Optimize momentum <![CDATA[β1,β2=0.9,0.95]]> Size of each batch of data 1024 Learning rate decay method Cosine decay Warm-up training times 20 Data Augmentation Random cropping
[0125] The feature matrix output by this module maintains the same size as the original VIT architecture. All experimental results are based on training on an NVIDIA Tesla V100 GPU for 150 epochs. We have demonstrated the effectiveness of our proposed module through a series of experiments, and the specific results are shown in Table 3.
[0126] Table 3: Validation Results of the Three Core Modules in this Application
[0127] 83.5 81.38 √ 84.1 82.03 √ √ 87.1 82.58 √ √ √ 61.1 82.09 √ √ Pre-training only 61.1 / 87.1 82.71
[0128] To demonstrate the good generalization performance of our proposed pre-trained model, we performed transfer learning on multiple datasets, including ImageNet, CIFAR100, and iNat19, achieving excellent results. Specific results are shown in Table 4.
[0129] Table 4: Accuracy Comparison of this Solution and Traditional Solution in Downstream Tasks
[0130] ViT [Dosovitskiy et al., 2020] 300 77.9 87.1 - Deit-B [Touvron et al., 2021] 300 81.8 90.8 77.7 CeiT-T [Yuan et al., 2021] 300 76.4 88.4 72.8 CeiT-S [Yuan et al., 2021] 300 82.0 90.8 78.9 CvT-13 [Yuan et al., 2021] 300 81.6 87.3 - CvT-21 [Yuan et al., 2021] 300 82.5 89.0 - MAE [Yuan et al., 2021] 1600 83.6 - 80.5 The method described in this application (applies only to the pre-training phase) 100 82.1 85.1 76.9 The method described in this application (applied to the pre-training and fine-tuning stages) 100 82.8 86.0 77.5
[0131] Therefore, it can be seen that in the pre-training stage, the image data processing method of this application can reduce the number of parameters in the pre-trained model and improve the training speed, while achieving good recovery results for images of various categories. Furthermore, the image data processing method of this application also achieves extremely high accuracy in different downstream tasks.
[0132] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0133] Based on the same inventive concept, this application also provides an image data processing apparatus for implementing the image data processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image data processing apparatus embodiments provided below can be found in the limitations of the image data processing method described above, and will not be repeated here.
[0134] In some embodiments, such as Figure 8 As shown, an image data processing apparatus 800 is provided, including: an image sub-block acquisition module 802 and an encoding module 804, wherein:
[0135] The image sub-block acquisition module 802 is used to acquire image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples.
[0136] The encoding module 804 is used to perform attention pooling and feature encoding on each target sub-block in multiple image sub-blocks based on the grouping information corresponding to each layer in the multi-layer neural network, so as to obtain the encoding features corresponding to each target sub-block; the layers with the same grouping information in the multi-layer neural network share the same attention layer.
[0137] In some embodiments, the image data processing apparatus 800 further includes a grouping information determination module, configured to: obtain the number of layers and the number of groups in the multi-layer neural network; the number of groups is less than the number of layers; determine the number of layers corresponding to each group based on the number of layers and the number of groups; the difference in the number of layers between groups is less than a set difference in the number of layers; determine the grouping information corresponding to each layer in the multi-layer neural network based on the number of layers; and the layers in the same group are consecutive layers in the multi-layer neural network.
[0138] In some embodiments, the attention layer has at least two attention heads. In this embodiment, the encoding module 806 includes an attention pooling unit, specifically configured to: extract features from a target sub-block based on a target attention head to obtain first candidate sub-feature information of the target sub-block, and determine shared parameters in the feature extraction process; the target attention head is any one of at least two attention heads; using the shared parameters, extract features from the target sub-block based on each of the other attention heads (excluding the target attention head) among the at least two attention heads to obtain second candidate sub-feature information corresponding to each attention head; and fuse the first candidate sub-feature information and each of the second candidate sub-feature information to obtain target sub-feature information of the target sub-block.
[0139] In some embodiments, the encoding module 806 includes an encoding unit, specifically used for: obtaining the pooling result obtained by performing attention pooling on the target sub-block; performing depthwise separable convolution on the pooling result to obtain the initial encoding features of the target sub-block; and performing secondary encoding on the initial encoding features based on a multilayer perceptron to obtain the encoding features of the target sub-block.
[0140] In some embodiments, the image sub-block acquisition module 802 includes: a feature extraction unit, used to extract features from the image sample to obtain the image features of the image sample and determine the feature image represented by the image features; and a segmentation unit, used to segment the feature image to obtain multiple image sub-blocks corresponding to the image sample.
[0141] In some embodiments, the feature extraction unit is specifically used to: perform a first convolution process and a second convolution process on the image sample in sequence to obtain candidate image features of the image sample; the size of the first convolution kernel in the first convolution process is larger than the size of the second convolution kernel in the second convolution process; and perform pooling processing on each candidate image feature to obtain the image features of the image sample.
[0142] In some embodiments, target sub-blocks are obtained by masking image samples and filtering them from multiple image sub-blocks. In this embodiment, the image data processing apparatus 800 further includes: a target sub-block filtering module, used to mask image samples and filter target sub-blocks from multiple image sub-blocks; and a decoding module, used to decode image samples based on the encoding features corresponding to each target sub-block and the learnable features corresponding to masked sub-blocks other than the target sub-blocks in the multiple image sub-blocks, to obtain a pre-trained reconstructed image corresponding to the image sample.
[0143] In some embodiments, the number of image samples is multiple. In this embodiment, the image data processing apparatus 800 further includes: a pre-trained image model determination module, used to obtain a pre-trained image model when the pre-trained reconstructed image corresponding to each image sample and the original image meet the similarity condition; and a target image model determination module, used to supervise the training of the initial student model using the pre-trained image model as the teacher model to obtain a target image model that matches the image processing task.
[0144] In some embodiments, the image processing task is an image classification task; the target image model is an image classification model. In this embodiment, the target image model determination module is specifically used to: use a pre-trained image model as a teacher model to supervise the training of an initial student model to obtain an image classification model that matches the image classification task; the initial student model includes multiple attention layers, and at least some of the training parameters of each attention layer are the same during the supervised training process.
[0145] Each module in the aforementioned image data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0146] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data involved in the image data processing method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an image data processing method.
[0147] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0148] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the image data processing method described above.
[0149] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the image data processing method described above.
[0150] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the image data processing method described above.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0152] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0153] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0154] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An image data processing method, characterized in that, The method includes: Obtain image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples; Based on the grouping information corresponding to each layer in the multi-layer neural network, attention pooling and feature encoding are performed on each target sub-block in the plurality of image sub-blocks to obtain the encoded features corresponding to each target sub-block; layers with the same grouping information in the multi-layer neural network share the same attention layer.
2. The method according to claim 1, characterized in that, The process of determining the grouping information includes: Obtain the number of layers and the number of groups in the multilayer neural network; the number of groups is less than the number of layers. The number of layers corresponding to each group is determined based on the number of layers and the number of groups; the difference in the number of layers between groups is less than a set difference in the number of layers. Based on the number of layers, the grouping information corresponding to each layer in the multi-layer neural network is determined; the layers in the same group are consecutive layers in the multi-layer neural network.
3. The method according to claim 2, characterized in that, The attention layer has at least two attention heads; The process of performing attention pooling on the target sub-block includes: Based on the target attention head, feature extraction is performed on the target sub-block to obtain the first candidate sub-feature information of the target sub-block, and the shared parameters in the feature extraction process are determined; the target attention head is any one of the at least two attention heads; Using the shared parameters, feature extraction is performed on the target sub-block based on each of the at least two attention heads other than the target attention head, to obtain the second candidate sub-feature information corresponding to each attention head. The target sub-feature information of the target sub-block is obtained by fusing the first candidate sub-feature information and each of the second candidate sub-feature information.
4. The method according to claim 1, characterized in that, The process of feature encoding includes: Obtain the pooling result obtained by performing attention pooling on the target sub-block, and perform depthwise separable convolution on the pooling result to obtain the initial encoding features of the target sub-block; The initial encoded features are processed by a multilayer perceptron to obtain the encoded features of the target sub-block.
5. The method according to claim 1, characterized in that, The step of segmenting the image sample to obtain multiple image sub-blocks corresponding to the image sample includes: Feature extraction is performed on the image sample to obtain the image features of the image sample, and the feature image represented by the image features is determined; The feature image is segmented to obtain multiple image sub-blocks corresponding to the image sample.
6. The method according to claim 5, characterized in that, The step of extracting features from the image samples to obtain the image features of the image samples includes: The image sample is subjected to a first convolution process and a second convolution process in sequence to obtain candidate image features of the image sample; the size of the first convolution kernel in the first convolution process is larger than the size of the second convolution kernel in the second convolution process. Pooling is performed on each of the candidate image features to obtain the image features of the image sample.
7. The method according to any one of claims 1 to 6, characterized in that: The target sub-block is obtained by masking the image samples and filtering them from the plurality of image sub-blocks; The method further includes: Based on the encoding features corresponding to each target sub-block and the learnable features corresponding to the mask sub-blocks other than the target sub-blocks in the plurality of image sub-blocks, the image samples are decoded to obtain the pre-trained reconstructed images corresponding to the image samples.
8. The method according to claim 7, characterized in that, The number of image samples is multiple; the method further includes: A pre-trained image model is obtained when the pre-trained reconstructed image corresponding to each of the image samples satisfies the similarity condition with the original image. Using the pre-trained image model as the teacher model, supervised training is performed on the initial student model to obtain a target image model that matches the image processing task.
9. The method according to claim 8, characterized in that, The image processing task is an image classification task; the target image model is an image classification model. The step of using the pre-trained image model as the teacher model to supervise the training of the initial student model to obtain a target image model that matches the image processing task includes: Using the pre-trained image model as the teacher model, supervised training is performed on the initial student model to obtain an image classification model that matches the image classification task; the initial student model includes multiple attention layers, and at least some of the training parameters of each attention layer are the same during the supervised training process.
10. An image data processing apparatus, characterized in that, The device includes: An image sub-block acquisition module is used to acquire image samples, perform image segmentation on the image samples, and obtain multiple image sub-blocks corresponding to the image samples. The encoding module is used to perform attention pooling and feature encoding on each target sub-block in the plurality of image sub-blocks based on the grouping information corresponding to each layer in the multi-layer neural network, so as to obtain the encoding features corresponding to each target sub-block; the layers with the same grouping information in the multi-layer neural network share the same attention layer.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-level image compression method using Transform
CN113709455A
Image detection method and device, electronic equipment and storage medium
CN114663670A