Data Processing Method, Apparatus, Storage Medium, and Computer Device

By transforming the dimensions of intermediate variables, the TPU architecture can make full use of the advantages of parallel computing, solving the problem of inefficiency of TPUs when processing visual data, and achieving more efficient data processing.

CN119851051BActive Publication Date: 2025-06-10PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510294247.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-10
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In the prior art, when processing visual data, the TPU architecture cannot fully utilize the advantages of parallel computing due to the dimension problems of intermediate variables, resulting in low data processing efficiency.

Method used

By transforming the dimensions of the intermediate variable, the last dimension of each intermediate variable is the number of images of the batch image, avoiding the problem that the TPU cannot process in parallel due to the last dimension of 1.

Benefits of technology

Give full play to the computing potential of high-performance TPUs, improve TPU usage, and thus improve data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851051B_ABST
    Figure CN119851051B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a data processing method, apparatus, storage medium, and computer device. First, the first image features of a batch of images are obtained and input into the current processing module of the visual state space model. Through four-way scanning, an initial global receptive field is obtained, and then feature adjustment and dimensional transformation are performed to make the final dimension of the intermediate variable the number of batch images. The total output features are determined by combining the model weight parameters of the current processing module. Then, the inverse transformation is performed on the total output features, and the second image features are obtained through four-way scanning and integration. When the current processing module is the last one, the image classification result is determined based on the second image features. This method avoids the inability of the TPU to perform parallel processing due to the final dimension being 1 through reasonable dimensional transformation, fully exploits the computing potential of the high-performance TPU, improves the TPU utilization rate, and effectively improves the data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a data processing method, apparatus, storage medium, and computer device. Background Art

[0002] The Visual State Space Model (VMamba) is a new type of visual backbone network model, which aims to reduce the computational complexity to linear time complexity while retaining the advantageous features of Vision Transformers, such as the global receptive field and dynamic weight parameters. The design inspiration of VMamba comes from the Mamba state space language model, which has demonstrated the ability to efficiently model long sequences in natural language processing (NLP) tasks. VMamba transfers this concept to the visual field and processes visual data with linear complexity by introducing a Visual State Space (VSS) module and a 2D Selective Scan (SS2D) module. Extensive experiments have shown the excellent performance of VMamba in various visual perception tasks, including image classification, object detection, and semantic segmentation, highlighting its advantage in input scaling efficiency compared with existing benchmark models.

[0003] In the related art, in the SS2D module, there is a process that needs to be executed cyclically. This process needs to be executed sequentially L times, where L is the number of image patches (Tokens) obtained by splitting the image being processed. In the shallow Cross-Scan module, the number of image patches L is as high as 56 * 56. The large value of L results in a very large number of loop iterations, and the loop needs to be executed sequentially for each image patch, making it difficult to utilize the parallel computing advantage of the TPU architecture. In addition, the parallelism of the Tensor Processing Unit (TPU) architecture for four-dimensional data occurs in the second and fourth dimensions. In each actual loop, the shape of the intermediate variables involved in the prior art is mostly long and narrow, and the dimension of the last dimension is only 1, which is not friendly to batch data transmission and TPU parallel computing, making it difficult to utilize the computing potential of the high-performance TPU and resulting in a low TPU utilization rate and low data processing efficiency. Therefore, the related art urgently needs to propose a data processing method to solve the above technical problems. Summary of the Invention

[0004] The main purpose of this application is to provide a data processing method, apparatus, storage medium, and computer device, which can utilize the computing potential of the high-performance TPU and improve the TPU utilization rate, thereby enhancing the data processing efficiency.

[0005] In a first aspect, an embodiment of this application provides a data processing method, including:

[0006] Obtain the first image feature of each image in the batch of images;

[0007] Input the first image feature into the current processing module in the visual state space model, perform four-way scanning on each first image feature to obtain an initial global receptive field;

[0008] Perform different feature adjustments on the initial global receptive field to obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable respectively;

[0009] Perform dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field. The last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images;

[0010] Determine the total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module;

[0011] Perform inverse dimensionality transformation on the total output feature to obtain an inverse transformation result;

[0012] Perform four-way scanning integration on the inverse transformation result to obtain a second image feature;

[0013] When the current processing module is the last module, determine the image classification result based on the second image feature.

[0014] In a second aspect, an embodiment of the present application provides a data processing device, including:

[0015] An acquisition unit, configured to acquire the first image feature of each image in the batch of images;

[0016] An input unit, configured to input the first image feature into the current processing module in the visual state space model, perform four-way scanning on each first image feature to obtain an initial global receptive field;

[0017] An adjustment unit, configured to perform different feature adjustments on the initial global receptive field to obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable respectively;

[0018] A transformation unit for performing dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field, where the last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images;

[0019] A first determination unit for determining a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module;

[0020] An inverse transformation unit for performing inverse dimensionality transformation on the total output feature to obtain an inverse transformation result;

[0021] An integration unit for performing four-way scanning integration on the inverse transformation result to obtain a second image feature;

[0022] A second determination unit for determining an image classification result based on the second image feature when the current processing module is the last module.

[0023] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the data processing method as described in any one of the above.

[0024] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data processing method as described in any one of the above.

[0025] In the embodiment of the present application, by obtaining the first image feature of each image in a batch of images; inputting the first image feature into the current processing module in the visual state space model, performing four-way scanning on each first image feature to obtain an initial global receptive field; performing different feature adjustments on the initial global receptive field to respectively obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable; performing dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field, where the last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images; determining a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module; performing inverse dimensionality transformation on the total output feature to obtain an inverse transformation result; performing four-way scanning integration on the inverse transformation result to obtain a second image feature; when the current processing module is the last module, determining an image classification result based on the second image feature corresponding to each module. Compared with the related art where the parallel computing power of the TPU cannot be fully utilized due to the dimensionality problem of intermediate variables, in the embodiment of the present application, by performing dimensionality transformation on the intermediate variables, the last dimension of each intermediate variable is the number of images in the batch of images, avoiding the problem that the TPU cannot be processed in parallel due to the last dimension being 1, thereby exerting the computing potential of the high-performance TPU and increasing the TPU utilization rate, and further improving the data processing efficiency.

[0026] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained by the structures specifically pointed out in the specification, the claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0028] Figure 1 It is a schematic flowchart of the data processing method provided by the embodiment of the present application.

[0029] Figure 2Schematic diagram of data processing during four-way scanning of the Cross-Scan module provided by an embodiment of the present application.

[0030] Figure 3 Schematic diagram of transforming the dimension of an intermediate variable provided by an embodiment of the present application.

[0031] Figure 4 Schematic diagram of calculating the total output feature provided by an embodiment of the present application.

[0032] Figure 5 Specific schematic diagram of the inverse transformation process provided by an embodiment of the present application.

[0033] Figure 6 Schematic diagram of the four-way scanning integration process provided by an embodiment of the present application.

[0034] Figure 7 Schematic diagram of the structure of the data processing device provided by an embodiment of the present application.

[0035] Figure 8 Schematic diagram of the structure of the computer device provided by an embodiment of the present application. Specific implementation manners

[0036] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0037] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution rules. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0038] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0039] Visual State Space Model (VMamba): It is a new type of visual backbone network model, aiming to reduce the computational complexity to linear time complexity while retaining the advantageous features of Vision Transformers, such as global receptive fields and dynamic weight parameters. Specifically, (1) In terms of the state space model, the design of VMamba is inspired by the state space language model (Mamba), which has demonstrated the ability to efficiently model long sequences in natural language processing (NLP) tasks. VMamba transplanted this concept into the visual field and processed visual data with linear complexity by introducing the Visual State Space (VSS) module and the 2D Selective Scan (SS2D) module. (2) In terms of global receptive fields and dynamic weights, VMamba achieved global receptive fields and dynamic weights through its core Visual State Space (VSS) module and Selective Scan (SS2D) module, which helps collect context information from different sources and perspectives and provides rich feature representations for visual tasks. (3) In terms of linear time complexity, a significant feature of VMamba is its linear time complexity, which means that as the input scale increases, VMamba has obvious advantages in computational efficiency compared with traditional methods. Based on the current VMamba model architecture, the present invention further improves the performance of VMamba through optimization in data transfer. (4) In terms of the Cross-Scan module (CSM), to solve the problem that visual signals do not have the natural orderliness like text sequences, VMamba designed the Cross-Scan module. The Cross-Scan module adopts a four-way scan strategy to ensure that each element in the feature integrates information from all other positions in different directions to form a global receptive field without increasing the linear computational complexity. (5) In terms of experimental verification, extensive experiments have demonstrated the excellent performance of VMamba in various visual perception tasks, including image classification, object detection, and semantic segmentation, highlighting its advantages in input scaling efficiency compared with existing benchmark models.

[0040] Tensor Processing Unit (TPU): Focuses on the research and development, promotion, and application of computing power products such as processors, adheres to the ecological concept of comprehensive open source and openness, and leads the innovation of intelligent computing technology. It has built a full-scenario product matrix covering "cloud, edge, and terminal", and has been widely applied and recognized by users in multiple scenarios such as urban operation, intelligent manufacturing, large model applications, and intelligent terminals. The TPU accelerator of SuiNeng has highly parallel computing capabilities, which can significantly improve the inference speed and efficiency of AI models. It provides a more stable and reliable computing power infrastructure for the global open source community and developers.

[0041] 2D Selective Scanning (SS2D) Module: It is a functional module used to perform specific scanning and processing on two-dimensional data in the field of image processing. Its scanning mechanism is as follows: It will scan two-dimensional data, such as the pixel matrix of an image, according to specific rules. This scanning is not a simple row-by-row or column-by-column scanning, but selective. It may selectively scan specific rows, columns, or regions based on certain features of the data or a pre-set algorithm to obtain more valuable information. During the scanning process, the module will perform specific calculation operations, such as numerical calculations and feature extraction on the scanned data. At the same time, data scheduling will also be carried out, that is, according to the requirements of calculations and subsequent processing, reasonably arrange the flow and storage of data to ensure that the data can be efficiently processed and utilized.

[0042] The role of using the SS2D module is as follows: It helps to extract specific features of an image. By selectively scanning different regions of the image, it can capture feature information such as edges, textures, and colors in the image, providing a basis for subsequent image analysis, recognition, and other tasks. The data can be processed according to the scanning results to achieve data dimensionality reduction, remove some redundant information, and at the same time retain the features important for subsequent tasks, thereby improving the efficiency of data processing and the performance of the model. Or in some cases, enhance the processing of data in specific regions to highlight the key information in the image. In complex image processing or deep learning models, the SS2D module usually works in cooperation with other modules. For example, in the VMamba model, it can cooperate with the Cross-Merge module, etc., to provide support for the image feature processing and inference process of the entire model, and provide more appropriate and valuable input data for other modules through selective scanning.

[0043] In related technologies, in the 2D Selective Scanning (SS2D) module, there is a process that needs to be executed in a loop. This process needs to be executed sequentially L times, where L is the number of image patches (Tokens) obtained by splitting the image being processed. In the shallow Cross-Scan module, the number of image patches L is as high as 56 * 56. The large value of L results in a very large number of loop iterations for this process, and the loop needs to be executed sequentially, making it difficult to take advantage of the parallel computing advantages of the TPU architecture. In addition, the parallelism of the TPU architecture for four-dimensional data occurs in the second dimension and the fourth dimension. In each actual loop, the shape of the intermediate variables involved in the calculation is mostly long and narrow (such as: [batch size, 1536, 1, 1]), and the dimension of the last dimension is only 1, which is not friendly to batch data transmission and TPU parallel computing, resulting in low IO efficiency and low TPU utilization rate, and it is difficult to take advantage of the computing potential of high-performance TPUs.

[0044] In the embodiments of the present application, to solve the above problems, the first image feature of each image in a batch of images is obtained; the first image feature is input into the current processing module in the visual state space model, and each first image feature is scanned in four directions to obtain an initial global receptive field; different feature adjustments are performed on the initial global receptive field to respectively obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable; dimensionality transformation is performed on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field. The last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images; based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module, a total output feature is determined; dimensionality inverse transformation is performed on the total output feature to obtain an inverse transformation result; four-way scan integration is performed on the inverse transformation result to obtain a second image feature; when the current processing module is the last module, an image classification result is determined based on the second image feature corresponding to each module. Compared with the related art where the parallel computing power of the TPU cannot be fully utilized due to the dimensionality problem of intermediate variables, in the embodiments of the present application, by performing dimensionality transformation on the intermediate variables, the last dimension of each intermediate variable is the number of images in the batch of images, avoiding the problem that the TPU cannot be processed in parallel caused by the last dimension being 1, thereby exerting the computing potential of the high-performance TPU and increasing the TPU utilization rate, and further improving the data processing efficiency. For details, please continue to refer to the following specific embodiments.

[0045] In this embodiment, a description will be made from the perspective of a data processing device, which may be specifically integrated in a client equipped with a storage unit and installed with a microprocessor and having computing capabilities.

[0046] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the data processing method provided by the embodiments of the present application. The data processing method includes:

[0047] In step 201, the first image feature of each image in a batch of images is obtained.

[0048] Among them, batch images refer to a group of images processed at one time. These images are usually organized together in a data structure to improve processing efficiency by using parallel computing. When the first image feature is the feature representation obtained after preprocessing each image (such as the preprocessing in the data preparation stage) and is used for subsequent input into the Block module of the visual state space model for further processing, the first image feature contains some basic feature information of the image, such as the initially extracted features like color, texture, and edges. When the first image feature is the image feature output by a certain Block module of the visual state space model, the first image feature is the result of a series of complex processes. It is the product of deep processing of the input features, which has undergone operations such as shape adjustment, dimension segmentation, element summation, and dimension swapping, integrating four-way scanning information, and is generally a summary and refinement of the previously processed multi-dimensional feature information.

[0049] Specifically, obtain the pre-trained model weights of the VMamba model, compile them into a model file that can be executed on the Synnax TPU (for example: a model file in Bmodel format), and transfer it to the storage space of the Synnax TPU device. Package and preprocess a batch (for example: 16 images) of image data to be classified, and transfer it to the storage space of the Synnax TPU device. Taking 16 images with three RGB channels as an example, first scale each image to a data block with a size of [3, 224, 224]. Among them, 3 is the number of RGB channels, and 224 is the length and width of the scaled image. Then, normalize each image dataset separately, with the specific formula being (x - min) / (max - min), so that its numerical distribution is [0, 1] to obtain a new data block. Among them, x is the image data block, min is the minimum value of the data block, and max is the maximum value of the data block. Then, normalize the above new data block, specifically by subtracting the three means (for example, [0.485, 0.456, 0.406]) statistically obtained on the training dataset of the model from the corresponding positions on the three channels, and dividing by the three standard deviations (for example, [0.229, 0.224, 0.225]) statistically obtained on the training dataset of the model on the three channels to obtain the normalized data block. Finally, package the normalized data blocks of 16 images into a data block with a size of [16, 3, 224, 224] to obtain the preprocessed image data block.

[0050] The inference process of the VMamba model is carried out on the Synnovation TPU device. The specific process is to input a batch of preprocessed image data blocks in the above data preparation stage into the VMamba model in the TPU device, and then use the Synnovation TPU device to sequentially execute the image block embedding layer, multiple Block modules, and the last few layers of the VMamba model. Finally, the inference result output by the model is obtained for the stage of parsing the model inference result. The first several layers of the VMamba model, specifically in order, are a 2D convolutional layer, a GELU activation layer, a 2D convolutional layer, and a LayerNorm normalization layer. The role of these four neural networks is to map the preprocessed image data blocks into a higher-dimensional data space to obtain the first image features, which facilitates subsequent feature extraction in the Block modules. The image feature size of the first image features obtained at this time is [batch size, C, H_input, W_input]. Where C is the depth of the image data block features, H_input is the height of the image, and W_input is the width of the image.

[0051] In step 202, the first image features are input into the current processing module of the visual state space model, and each of the first image features is scanned in four directions to obtain an initial global receptive field.

[0052] Among them, the visual state space model includes multiple Block modules. Each Block module includes a Cross-Scan module. The multiple Block modules are processed sequentially in order, and each Block module has its own independent model weight parameters. For the current Block module to be processed, each first image feature is scanned in four directions through the Cross-Scan module to obtain an initial global receptive field. The initial global receptive field is the feature representation obtained through the four-way scan. It contains the information integrated from the image features in different directions, enabling the model to capture more comprehensive context information in the image and providing a richer information basis for subsequent feature processing.

[0053] In some embodiments, the scanning of each of the first image features in four directions to obtain an initial global receptive field includes:

[0054] (1) For each of the first image features, the first image feature is copied to obtain two first image features. The first dimension of the first image feature is the depth of the image data block features, the second dimension is the longitudinal quantity of the image data block, and the third dimension is the transverse quantity of the image data block;

[0055] (2) Multiply the third dimension and the fourth dimension of one of the first image features, and use the result as the third dimension to obtain a third image feature;

[0056] (3) Swap the third dimension and the fourth dimension of the other first image feature, multiply the swapped third dimension and the swapped fourth dimension, and use the result as the third dimension to obtain a fourth image feature;

[0057] (4) Stack the third image feature and the fourth image feature in the dimension of the depth of the image data block feature to obtain a stacked feature;

[0058] (5) Copy the stacked feature to obtain two stacked features;

[0059] (6) Stack each stacked feature in the dimension of the depth of the image data block feature to obtain an initial global receptive field.

[0060] Among them, in the related technology, in the Cross-Scan module, a local flipping method of the current image feature map is used to implement the four-way scanning strategy of the technology. In the Cross-Merge module, a local flipping method of the current image feature map is also used to merge the image information obtained by the four-way scanning strategy. However, in the SOPHON TPU architecture, the execution efficiency of flipping the image data is very low, resulting in an operator that flips a large number of images without computational load, but occupies a large amount of inference time. Based on this, the embodiments of the present application provide a brand-new four-way scanning strategy.

[0061] Specifically, as Figure 2 shown, Figure 2This is a schematic diagram of data processing during the four-way scan of the Cross-Scan module provided in the embodiments of this application. During the inference process of each Block module in the VMamba model, when the Cross-Scan module is executed, the data flipping operation is not used. Instead, a specific calculation and data scheduling strategy is used in the subsequent 2D Selective Scan (SS2D) module to achieve the same effect. In the Cross-Scan module, the first image feature (with a feature dimension size of [the number of images in the batch of images, C, H, W], where C is the depth of the image data block feature, H is the number of vertical image data blocks, and W is the number of horizontal image data blocks) of the input to this module is copied to obtain a total of two identical first image features. Then, one of them is shape-adjusted, that is, H in the third dimension and W in the fourth dimension of the first image feature are multiplied, and the feature size of the obtained third image feature is [the number of images, C, H*W]; for the other one, first H in the third dimension and W in the fourth dimension are swapped, and then the shape is adjusted, that is, W in the third dimension after swapping and H in the fourth dimension after swapping are multiplied, and the feature size of the obtained fourth image feature is also [the number of images, C, H*W]; then, the above two first image features are stacked in the C dimension to obtain a stacked feature with a feature size of [the number of images, 2*C, H*W]; then, the stacked feature is copied and stacked in the C dimension to obtain an initial global receptive field with a data block size of [the number of images, 4*C, H*W]. The above steps complete the four-way scan process of the Cross-Scan module, ensuring that each element in the feature integrates information from all other positions in different directions to form an initial global receptive field.

[0062] In step 203, different feature adjustments are performed on the initial global receptive field to obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable respectively.

[0063] Among them, during the inference process of each Block module of the VMamba model, after the Cross-Scan module is normally executed, there may be other operators in front of the 2D Selective Scan (SS2D) module. Specifically, the initial global receptive field with a feature size of [number of images, 4*C, L] obtained in step 202, where L is H*W, needs to go through a 1D convolutional layer, and the convolutional result is sliced to obtain the first intermediate variable, the second intermediate variable, and the third intermediate variable. Then, the first intermediate variable is input into the next 1D convolutional layer to obtain a new first intermediate variable with a size of [batch size, 4*C, L, N] to replace the original first intermediate variable, obtaining the initial first intermediate variable. The second and third intermediate variables are repeated C times in the third dimension respectively to obtain new second and third intermediate variables with a size of [batch size, 4*C, N, L] and replace the original second and third intermediate variables, obtaining the initial third intermediate variable. Then, the second intermediate variable is further subjected to matrix multiplication and shape transformation with the weight parameters in the model to obtain an updated second intermediate variable with a size of [batch size, 4*C, L, N] and replace the original second intermediate variable, obtaining the initial second intermediate variable. Among them, L is H*W, representing the total number of image patches after image slicing; N is a preset hyperparameter that controls the number of parameters, which is set to 1 in the embodiments of this application.

[0064] In step 204, the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field are subjected to dimensionality transformation to obtain the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field. The last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images of the batch of images.

[0065] Among them, before executing the 2D Selective Scan (SS2D) module, the intermediate variables participating in the operation are subjected to dimensionality transformation, and the dimension representing the batch size (i.e., the number of images) is placed at the last dimension, which is more friendly to batch data transmission and TPU parallel computing, better leveraging the parallel computing ability of the TPU architecture for four-dimensional data and improving the IO efficiency and TPU utilization rate.

[0066] Specifically, since the size of the initial first intermediate variable is [number of images, 4*C, L, N], the size of the initial second intermediate variable is [number of images, 4*C, L, N], the size of the initial third intermediate variable is [number of images, 4*C, N, L], and the size of the initial global receptive field is [number of images, 4*C, L], it can be seen that the last dimension of the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field is not the number of images. Therefore, the dimension of the number of images is placed at the last dimension to facilitate the use of the TPU.

[0067] In some embodiments, the first dimension of the initial first intermediate variable and the initial second intermediate variable is the number of images in the batch of images, the second dimension is the depth of the image data block features at a specified multiple, the third dimension is the total number of image data blocks, and the fourth dimension is a preset hyperparameter. The first dimension of the initial third intermediate variable is the number of images, the second dimension is the depth of the image data block features at a specified multiple, the third dimension is the preset hyperparameter, and the fourth dimension is the total number of image data blocks. The first dimension of the initial global receptive field is the number of images, the second dimension is the depth of the image data block features at a specified multiple, and the third dimension is the total number of image data blocks;

[0068] The step of performing dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field includes:

[0069] (1) Swap the first dimension and the fourth dimension of the initial first intermediate variable, the initial second intermediate variable, and the initial third intermediate variable respectively to obtain the first intermediate variable, the second intermediate variable, and the third intermediate variable;

[0070] (2) Swap the first dimension and the third dimension of the initial global receptive field to obtain the global receptive field.

[0071] Among them, please refer to Figure 3 , Figure 3 is a schematic diagram of performing dimensionality transformation on the intermediate variables provided by the embodiments of the present application. For the first intermediate variable with a shape of [number of images, 4*C, L, N], swapping the first and fourth dimensions results in a first intermediate variable with a shape of [N, 4*C, L, number of images]; for the second intermediate variable with a shape of [number of images, 4*C, L, N], swapping the first and fourth dimensions results in a second intermediate variable with a shape of [N, 4*C, L, number of images]; for the third intermediate variable with a shape of [number of images, 4*C, N, L], swapping the first and fourth dimensions results in a third intermediate variable with a shape of [L, 4*C, N, number of images]; for the initial global receptive field with a shape of [number of images, 4*C, L], swapping the first and third dimensions results in a global receptive field with a shape of [L, 4*C, number of images].

[0072] In step 205, based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module, the total output feature is determined.

[0073] Among them, the model weight parameters are the parameters learned by the visual state space model during the training process, which are used to transform and process the input data and determine the behavior and performance of the model. Different modules have different weight parameters, and these parameters are continuously optimized through training to enable the model to accurately complete tasks, such as image classification. The total output feature is the final feature representation obtained by calculating and fusing the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters. It synthesizes the information of multiple intermediate variables and the knowledge carried by the model weights, providing a key feature basis for subsequent tasks (such as classification).

[0074] Specifically, specific operations are performed on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the module to finally obtain the total output feature.

[0075] In some embodiments, determining the total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module includes:

[0076] (1) Divide the current second dimension of the first intermediate variable, the second intermediate variable, and the third intermediate variable to obtain two divided first intermediate variables, two divided second intermediate variables, and two divided third intermediate variables. Divide the second dimension of the product of the global receptive field and the model weight parameters of the module to obtain two divided fourth intermediate variables;

[0077] (2) Extract the currently to-be-calculated first variable slice from one of the divided first intermediate variables, one of the divided second intermediate variables, one of the divided third intermediate variables, and one of the divided fourth intermediate variables in the first preset order, and extract the currently to-be-calculated second variable slice from the other divided first intermediate variable, the other divided second intermediate variable, the other divided third intermediate variable, and the other divided fourth intermediate variable in the second preset order. The first preset order is opposite to the second preset order;

[0078] (3) Calculate each first variable slice to obtain a first partial result, and calculate each second variable slice to obtain a second partial result;

[0079] (4) Store the first partial result in the first preset order, store the second partial result in the second preset order, and increment the loop execution count by one to obtain the current loop execution count;

[0080] (5) When the number of times the current loop is executed has not reached half of the total number of image data blocks, return to execute the step of extracting the currently to-be-calculated first variable slice from one of the divided first intermediate variables, one of the divided second intermediate variables, one of the divided third intermediate variables, and one of the divided fourth intermediate variables respectively according to the first preset order, and extracting the currently to-be-calculated second variable slice from another one of the divided first intermediate variables, another one of the divided second intermediate variables, another one of the divided third intermediate variables, and another one of the divided fourth intermediate variables respectively according to the second preset order, until the number of times the current loop is executed reaches half of the total number of image data blocks, obtaining a plurality of stored first partial results and a plurality of stored second partial results, and splicing each of the first partial results and each of the second partial results to obtain the total output feature.

[0081] Among them, the 2D selective scan (SS2D) module is executed using variables with the batch size dimension placed at the back, and a calculation and data scheduling strategy using local variable caching and an avoidable flipping operation is adopted to retain as many intermediate variables as possible in the cache of the TPU, reducing the number and total amount of data transfers. The size of the model weight parameter is [4*C, 1].

[0082] Specifically, as Figure 4 shown, Figure 4Schematic diagram for calculating the total output features provided by the embodiments of this application. First, the 4*C dimensions of the first intermediate variable, the second intermediate variable, and the third intermediate variable are respectively segmented to obtain two segmented first intermediate variables after segmentation, two segmented second intermediate variables after segmentation, and two segmented third intermediate variables after segmentation. Then, the 4*C dimension (i.e., the second dimension) of the product of the global receptive field and the model weight parameters of the module is segmented to obtain two segmented fourth intermediate variables. Then, L loops are executed. In each loop, the segmented first variable slices are sequentially extracted from the L dimension (i.e., the first preset order) and in reverse order (i.e., the second preset order) to participate in the operation, obtaining partial results of the module output feature one (with a shape of [1, 2*C, 5 number of images]) and partial results of the module output feature two (with a shape of [1, 2*C, number of images]). After L loops, the complete module output feature one (i.e., multiple first partial results) and the module output feature two (i.e., multiple second partial results, with a shape of [L, 2*C, number of images]) are obtained, and they are concatenated in the 2*C dimension to obtain the total output feature with a final shape of [L, 4*C, number of images]. That is, in each loop, different slices (the first variable slice and the second variable slice) are sequentially selected from each intermediate variable through the first preset order and the second preset order for calculation. Finally, the first partial results calculated from the first variable slices are stored in the first preset order, and the second partial results calculated from the second variable slices are stored in the second preset order until the number of times the current loop is executed reaches half of the total number of image data blocks, obtaining multiple first partial results and multiple second partial results for concatenation, and obtaining the total output feature with a size of [L, 4*C, number of images].

[0083] Thus, when executing the Cross-Scan module, instead of using the data flipping operation, a specific calculation and data scheduling strategy is used in the 2D Selective Scan (SS2D) module to play the same role, avoiding the problem that the execution efficiency of flipping the image data is very low, resulting in a large number of images being flipped by an operator without computational load, occupying a large amount of inference time, and improving the inference efficiency.

[0084] In step 206, an inverse dimension transformation is performed on the total output feature to obtain an inverse transformation result.

[0085] Among them, the inverse dimension transformation is an operation opposite to the previous dimension transformation, restoring the dimension of the total output feature to a structure suitable for subsequent processing. Since a dimension transformation was performed before to facilitate certain calculations, the dimension now needs to be restored to connect with other modules or operations. The inverse transformation result is the result obtained after the total output feature undergoes the inverse dimension transformation, and its dimension structure matches the input dimension structure required for the subsequent processing of the model at this stage.

[0086] In some embodiments, the first dimension of the total output feature is the total number of image data blocks, the second dimension is the depth of the image data block features at a specified multiple, and the third dimension is the number of images. Performing an inverse dimensional transformation on the total output feature to obtain an inverse transformation result includes:

[0087] (1) Swapping the first dimension and the third dimension of the total output feature to obtain a dimension-transformed total output feature;

[0088] (2) Reshaping the second dimension in the dimension-transformed total output feature to obtain an inverse transformation result. The first dimension of the inverse transformation result is the number of images, the second dimension is the specified multiple, the third dimension is the depth of the image data block features, the fourth dimension is the vertical quantity, and the fifth dimension is the horizontal quantity.

[0089] Among them, please refer to Figure 5 , Figure 5 which is a specific schematic diagram of the inverse transformation process provided by the embodiments of the present application. Performing an inverse dimensional transformation on the output result of the 2D selective scanning module (i.e., the total output feature), and bringing the dimension representing the number of images to the front as the first dimension. Specifically, for the module output result with a shape of [L, 4*C, number of images], swapping the first and third dimensions, the obtained output result has a shape of [number of images, 4*C, L], and adjusting the dimensions to obtain an inverse transformation result with a shape of [number of images, 4, C, H, W].

[0090] In step 207, performing four-way scanning integration on the inverse transformation result to obtain a second image feature.

[0091] Among them, the second image feature is the image feature obtained after four-way scanning integration by the Cross-Merge module. It is the result of re-integrating information after the total output feature undergoes an inverse dimensional transformation. Compared with the previous features, it may be more representative and discriminative, providing a better feature representation for final image classification and other tasks.

[0092] In some embodiments, the horizontal quantity is the same as the vertical quantity. Performing four-way scanning integration on the inverse transformation result to obtain a second image feature includes:

[0093] (1) Using the total number of image data blocks obtained by multiplying the fourth dimension and the fifth dimension of the inverse transformation result as the fourth dimension to obtain a first feature representation;

[0094] (2) Divide the first feature representation in the third dimension to obtain a first sub - feature representation and a second sub - feature representation. The first dimension of the first sub - feature representation and the second sub - feature representation is the number of images, the second dimension is half of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0095] (3) Perform a bit - by - bit summation of the first sub - feature representation and the second sub - feature representation to obtain a second feature representation. The first dimension of the second image feature representation is the number of images, the second dimension is half of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0096] (4) Divide the second feature representation in the third dimension to obtain a third sub - feature representation and a fourth sub - feature representation. The first dimension of the third sub - feature representation and the fourth sub - feature representation is the number of images, the second dimension is one - quarter of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0097] (5) Reshape the fourth dimension of the third sub - feature representation to obtain the vertical quantity as the fourth dimension and the horizontal quantity as the fifth dimension;

[0098] (6) Reshape the fourth dimension of the fourth sub - feature representation to obtain the horizontal quantity as the fourth dimension and the vertical quantity as the fifth dimension;

[0099] (7) Perform a bit - by - bit summation of the third sub - feature representation and the fourth sub - feature representation to obtain the second image feature.

[0100] Among them, as Figure 6 shown, Figure 6Schematic diagram of the four-way scan integration process provided by the embodiments of the present application. When the Cross-Merge module is executed, the data flipping operation is not used. The batch size dimension that is default in the first dimension is omitted. Specifically, in the Cross-Merge module, the inverse transformation result of the input of this module (the image feature dimension size is [number of images, 4, C, H, W]) is shape-adjusted, that is, the total number of image data blocks obtained by multiplying the fourth dimension and the fifth dimension of the inverse transformation result is used as the fourth dimension, and the obtained feature size is the first feature representation of [number of images, 4, C, H*W]; then, the first feature representation is segmented in the third dimension to obtain two first sub-feature representations and second sub-feature representations with a feature size of [number of images, 2, C, H*W]; then, the elements at the same positions of the current two features are summed, and the obtained feature size is the second feature representation of [number of images, 2, C, H*W]; then it is segmented in the third dimension to obtain two third sub-feature representations and fourth sub-feature representations with a feature size of [number of images, C, H*W]; then, the third sub-feature representation is shape-adjusted, that is, the fourth dimension is reshaped to obtain the vertical quantity as the third dimension and the horizontal quantity as the fourth dimension, and the obtained feature size is [number of images, C, H, W]; the fourth sub-feature representation is shape-adjusted and the dimensions H and W are swapped, that is, the fourth dimension of the fourth sub-feature representation is reshaped to obtain the horizontal quantity as the third dimension and the vertical quantity as the fourth dimension, and the obtained feature size is [number of images, C, W, H]. In the embodiments of the present application, H and W are equal. Finally, the above reshaped third sub-feature representation and fourth sub-feature representation are summed for the elements at the same positions to obtain the final output feature with a size of [number of images, C, H, W]. The above steps complete the process of the Cross-Merge module integrating the four-way scan image features to obtain a single image feature with a size of [number of images, C, H, W]. It greatly reduces the time consumption of the Cross-Scan module and the Cross-Merge module, and further improves the model inference speed.

[0101] In step 208, when the current processing module is the last module, based on the second image feature, determine the image classification result.

[0102] Among them, since there are multiple Block modules in the VMamba model, if the current processing module is the last module, it means that there are no other Block modules subsequently. Then, through the further processing of the last several layers (such as the layernrom normalization layer, the global average layer, and the linear classification layer) of the VMamba model, the second image feature block is mapped to the corresponding image category, which is convenient for subsequent parsing of the inference result.

[0103] Specifically, obtain the inference results of the VMamba model on the Synnax TPU. According to the type of task and the dataset label list, parse the inference results to obtain the classification results, detection results, semantic segmentation results, etc. of the model for the current batch of image data. Taking image classification as an example, calculate the maximum value of the classification prediction probabilities of each sample in different categories in the inference results, and take the label corresponding to the maximum value as the predicted image classification result.

[0104] In some embodiments, the method further includes:

[0105] When the current processing module is not the last module, determine the second image feature as the first image feature, determine the next module as the current processing module, and return to execute the step of inputting the first image feature into the current processing module of the visual state space model, and performing four-way scanning on each of the first image features to obtain the initial global receptive field until the current processing module is the last module.

[0106] Among them, since multiple Block modules are executed sequentially, if the current processing module is not the last module, it means that the second image feature needs to be used as the input of the next Block module. Then, determine the second image feature as the first image feature, determine the next module as the current processing module, and return to execute the step of inputting the first image feature into the current processing module of the visual state space model, and performing four-way scanning on each of the first image features to obtain the initial global receptive field until the current processing module is the last module.

[0107] As can be seen from the above, in the embodiment of the present application, the first image feature of each image in a batch of images is obtained; the first image feature is input into the current processing module in the visual state space model, and each first image feature is scanned in four directions to obtain an initial global receptive field; different feature adjustments are performed on the initial global receptive field to respectively obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable; dimensionality transformation is performed on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field, and the last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images; based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module, a total output feature is determined; dimensionality inverse transformation is performed on the total output feature to obtain an inverse transformation result; four-way scan integration is performed on the inverse transformation result to obtain a second image feature; when the current processing module is the last module, an image classification result is determined based on the second image feature corresponding to each module. Compared with the related art where the parallel computing power of the TPU cannot be fully utilized due to the dimensionality problem of intermediate variables, in the embodiment of the present application, by performing dimensionality transformation on the intermediate variables, the last dimension of each intermediate variable is the number of images in the batch of images, avoiding the problem that the TPU cannot be processed in parallel caused by the last dimension being 1, thereby exerting the computing potential of the high-performance TPU and improving the TPU utilization rate, and further improving the data processing efficiency.

[0108] For the specific implementation of each of the above steps, reference may be made to the previous embodiments, which will not be elaborated here.

[0109] To facilitate better implementation of the data processing method provided in the embodiment of the present application, the embodiment of the present application also provides an apparatus based on the above data processing method. The meanings of the nouns are the same as those in the above data processing method, and the specific implementation details can refer to the description in the method embodiment.

[0110] Please refer to Figure 7 , Figure 7 FIG. 6 is a schematic structural diagram of the data processing apparatus provided in the embodiment of the present application. The data processing apparatus is applied to a computer device. The data processing apparatus may include an acquisition unit 601, an input unit 602, an adjustment unit 603, a transformation unit 604, a first determination unit 605, an inverse transformation unit 606, an integration unit 607, a second determination unit 608, and the like.

[0111] The acquisition unit 601 is configured to acquire the first image feature of each image in a batch of images;

[0112] An input unit 602 for inputting the first image feature into the current processing module in the visual state space model, performing four-way scanning on each of the first image features to obtain an initial global receptive field;

[0113] An adjustment unit 603 for performing different feature adjustments on the initial global receptive field to respectively obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable;

[0114] A transformation unit 604 for performing dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field, where the last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images;

[0115] A first determination unit 605 for determining a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module;

[0116] An inverse transformation unit 606 for performing inverse dimensionality transformation on the total output feature to obtain an inverse transformation result;

[0117] An integration unit 607 for performing four-way scanning integration on the inverse transformation result to obtain a second image feature;

[0118] A second determination unit 608 for, when the current processing module is the last module, determining an image classification result based on the second image feature.

[0119] In some embodiments, the first dimension of the initial first intermediate variable and the initial second intermediate variable is the number of images in the batch of images, the second dimension is the feature depth of the image data blocks at a specified multiple, the third dimension is the total number of image data blocks, and the fourth dimension is a preset hyperparameter; the first dimension of the initial third intermediate variable is the number of images, the second dimension is the feature depth of the image data blocks at a specified multiple, the third dimension is the preset hyperparameter, and the fourth dimension is the total number of image data blocks; the first dimension of the initial global receptive field is the number of images, the second dimension is the feature depth of the image data blocks at a specified multiple, and the third dimension is the total number of image data blocks;

[0120] A first swapping subunit for respectively swapping the first dimension and the fourth dimension of the initial first intermediate variable, the initial second intermediate variable, and the initial third intermediate variable to obtain a first intermediate variable, a second intermediate variable, and a third intermediate variable;

[0121] A second swapping subunit, configured to swap the first dimension and the third dimension of the initial global receptive field to obtain a global receptive field.

[0122] In some embodiments, the first determination unit 605 includes:

[0123] A dividing subunit, configured to divide the current second dimension of the first intermediate variable, the second intermediate variable, and the third intermediate variable to obtain two divided first intermediate variables, two divided second intermediate variables, and two divided third intermediate variables, and divide the second dimension of the product of the global receptive field and the model weight parameters of the module to obtain two divided fourth intermediate variables;

[0124] An extracting subunit, configured to extract a currently to-be-calculated first variable slice from one of the divided first intermediate variables, one of the divided second intermediate variables, one of the divided third intermediate variables, and one of the divided fourth intermediate variables respectively according to a first preset order, and extract a currently to-be-calculated second variable slice from the other divided first intermediate variable, the other divided second intermediate variable, the other divided third intermediate variable, and the other divided fourth intermediate variable respectively according to a second preset order, where the first preset order is opposite to the second preset order;

[0125] A calculating subunit, configured to calculate each of the first variable slices to obtain a first partial result, and calculate each of the second variable slices to obtain a second partial result;

[0126] A storing subunit, configured to store the first partial result according to the first preset order, store the second partial result according to the second preset order, and increment the loop execution count by one to obtain the current loop execution count;

[0127] A splicing subunit, configured to, when the number of times of the current loop execution does not reach half of the total number of the image data blocks, return to execute the step of respectively extracting a currently to-be-calculated first variable slice from one of the divided first intermediate variables, one of the divided second intermediate variables, one of the divided third intermediate variables, and one of the divided fourth intermediate variables according to a first preset order, and respectively extracting a currently to-be-calculated second variable slice from another one of the divided first intermediate variables, another one of the divided second intermediate variables, another one of the divided third intermediate variables, and another one of the divided fourth intermediate variables according to a second preset order, until the number of times of the current loop execution reaches half of the total number of the image data blocks, obtaining a plurality of stored first partial results and a plurality of stored second partial results, splicing each of the first partial results and each of the second partial results, and obtaining a total output feature.

[0128] In some embodiments, the input unit 602 includes:

[0129] A first copying subunit, configured to copy each of the first image features to obtain two of the first image features, where a first dimension of the first image feature is the depth of the image data block feature, a second dimension is the longitudinal number of the image data blocks, and a third dimension is the lateral number of the image data blocks;

[0130] A first multiplication operation subunit, configured to multiply a third dimension and a fourth dimension of one of the first image features, and use the result as the third dimension to obtain a third image feature;

[0131] A third swapping subunit, configured to swap a third dimension and a fourth dimension of the other first image feature, multiply the swapped third dimension and the swapped fourth dimension, and use the result as the third dimension to obtain a fourth image feature;

[0132] A first stacking subunit, configured to stack the third image feature and the fourth image feature in a dimension of the depth of the image data block feature to obtain a stacked feature;

[0133] A second copying subunit, configured to copy the stacked feature to obtain two of the stacked features;

[0134] A second stacking subunit, configured to stack each of the stacked features in a dimension of the depth of the image data block feature to obtain an initial global receptive field.

[0135] In some embodiments, a first dimension of the total output feature is the total number of the image data blocks, a second dimension is a specified multiple of the depth of the image data block feature, and a third dimension is the number of the images. The inverse transformation unit 606 includes:

[0136] A fourth swapping subunit, configured to swap the first dimension and the third dimension of the total output feature to obtain a dimension-transformed total output feature;

[0137] A first reshaping subunit, configured to reshape the second dimension in the dimension-transformed total output feature to obtain an inverse transformation result, where the first dimension of the inverse transformation result is the number of images, the second dimension is the specified multiple, the third dimension is the depth of the image data block feature, the fourth dimension is the vertical quantity, and the fifth dimension is the horizontal quantity.

[0138] In some embodiments, the horizontal quantity is the same as the vertical quantity. The integration unit 607 includes:

[0139] A second multiplication operation subunit, configured to use the total number of image data blocks obtained by multiplying the fourth dimension and the fifth dimension of the inverse transformation result as the fourth dimension to obtain a first feature representation;

[0140] A first splitting subunit, configured to split the first feature representation in the third dimension to obtain a first sub-feature representation and a second sub-feature representation, where the first dimension of the first sub-feature representation and the second sub-feature representation is the number of images, the second dimension is half of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0141] A first bitwise summation subunit, configured to perform bitwise summation on the first sub-feature representation and the second sub-feature representation to obtain a second feature representation, where the first dimension of the second image feature representation is the number of images, the second dimension is half of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0142] A second splitting subunit, configured to split the second feature representation in the third dimension to obtain a third sub-feature representation and a fourth sub-feature representation, where the first dimension of the third sub-feature representation and the fourth sub-feature representation is the number of images, the second dimension is one-fourth of the specified multiple, the third dimension is the depth of the image data block feature, and the fourth dimension is the total number of image data blocks;

[0143] A second reshaping subunit, configured to reshape the third dimension of the third sub-feature representation to obtain the vertical quantity as the third dimension and the horizontal quantity as the fourth dimension;

[0144] A third reshaping subunit, configured to reshape the third dimension of the fourth sub-feature representation to obtain the horizontal quantity as the third dimension and the vertical quantity as the fourth dimension;

[0145] A second bitwise summing subunit, configured to perform bitwise summation on the third sub-feature representation and the fourth sub-feature representation to obtain a second image feature.

[0146] In some embodiments, the apparatus further includes:

[0147] An execution unit, configured to, when the current processing module is not the last module, determine the second image feature as the first image feature, determine the next module as the current processing module, and return to execute the step of inputting the first image feature into the current processing module in the visual state space model and performing four-way scanning on each of the first image features to obtain an initial global receptive field, until the current processing module is the last module.

[0148] For the specific implementation of each of the above units, reference may be made to the previous embodiments and will not be elaborated herein.

[0149] As can be seen from the above, in the embodiment of the present application, the acquisition unit 601 acquires the first image feature of each image in a batch of images; the input unit 602 inputs the first image feature into the current processing module in the visual state space model and performs four-way scanning on each of the first image features to obtain an initial global receptive field; the adjustment unit 603 performs different feature adjustments on the initial global receptive field to respectively obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable; the transformation unit 604 performs dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field, and the last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images; the first determination unit 605 determines a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module; the inverse transformation unit 606 performs inverse dimensionality transformation on the total output feature to obtain an inverse transformation result; the integration unit 607 performs four-way scanning integration on the inverse transformation result to obtain a second image feature; the second determination unit 608 determines an image classification result based on the second image feature when the current processing module is the last module. Compared with the related art, due to the dimensionality problem of intermediate variables, the parallel computing power of the TPU cannot be fully utilized. In the embodiment of the present application, by performing dimensionality transformation on the intermediate variables, the last dimension of each intermediate variable is the number of images in the batch of images, avoiding the problem that the TPU cannot be processed in parallel due to the last dimension being 1, thereby exerting the computing potential of the high-performance TPU and improving the TPU utilization rate, and further enhancing the data processing efficiency.

[0150] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated herein.

[0151] Refer to Figure 8 , Figure 8 FIG. is a block diagram of a part of a computer device 1000 acting as a server for implementing the embodiments of the present disclosure. The computer device 1000 may vary greatly due to different configurations or performances, and may include one or more tensor processing units (Tensor Processing Unit, abbreviated as TPU) 622 (for example, one or more tensor processing units) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 600. Further, the tensor processing unit 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the computer device 1000.

[0152] The computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0153] The tensor processing unit 622 in the computer device 1000 may be used to execute the data processing method of the embodiments of the present disclosure, for example:

[0154] Obtain the first image feature of each image in the batch of images;

[0155] Input the first image feature into the current processing module of the visual state space model, perform four-way scanning on each first image feature, and obtain an initial global receptive field;

[0156] Perform different feature adjustments on the initial global receptive field to obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable respectively;

[0157] Perform dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field. The last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images;

[0158] Based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module, determine the total output feature;

[0159] Perform inverse dimensionality transformation on the total output feature to obtain an inverse transformation result;

[0160] Perform four-way scan integration on the inverse transformation result to obtain a second image feature;

[0161] When the current processing module is the last module, determine the image classification result based on the second image feature.

[0162] The embodiments of the present disclosure also provide a computer-readable storage medium. The computer-readable storage medium is used to store program code, and the program code is used to execute the data processing methods in the foregoing various embodiments.

[0163] The embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the data processing method implemented above. For example:

[0164] Obtain the first image feature of each image in the batch of images;

[0165] Input the first image feature into the current processing module of the visual state space model, and perform four-way scanning on each first image feature to obtain an initial global receptive field;

[0166] Perform different feature adjustments on the initial global receptive field to obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable respectively;

[0167] Perform dimensionality transformation on the initial first intermediate variable, the initial second intermediate variable, the initial third intermediate variable, and the initial global receptive field to obtain a first intermediate variable, a second intermediate variable, a third intermediate variable, and a global receptive field. The last dimension of the first intermediate variable, the second intermediate variable, the third intermediate variable, and the global receptive field is the number of images in the batch of images;

[0168] Determine the total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameters of the current processing module;

[0169] Perform an inverse dimensional transformation on the total output feature to obtain an inverse transformation result;

[0170] Perform four-way scanning integration on the inverse transformation result to obtain a second image feature;

[0171] When the current processing module is the last module, determine the image classification result based on the second image feature.

[0172] In addition, the terms "including" and "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0173] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0174] It should be understood that in the description of the embodiments of this application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0175] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0176] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0178] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0179] It should also be understood that the various embodiments provided by the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0180] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0181] The above is a specific description of the implementation manners of the present application. However, the present application is not limited to the above implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A data processing method, characterized in that: include: Obtaining the first image feature of each image in the batch of images; Inputting the first image features into a current processing module in a visual state space model, performing four-directional scanning on each of the first image features to obtain an initial global receptive field; Performing different feature adjustments on the initial global receptive field, respectively obtaining an initial first intermediate variable, an initial second intermediate variable and an initial third intermediate variable, wherein the first dimension of the initial first intermediate variable and the initial second intermediate variable is the number of images of the batch images, the second dimension is the feature depth of the image data blocks of a specified multiple, the third dimension is the total number of image data blocks, and the fourth dimension is a preset hyperparameter; the first dimension of the initial third intermediate variable is the number of images, the second dimension is the feature depth of the image data blocks of a specified multiple, the third dimension is the preset hyperparameter, and the fourth dimension is the total number of image data blocks; the first dimension of the initial global receptive field is the number of images, the second dimension is the feature depth of the image data blocks of a specified multiple, and the third dimension is the total number of image data blocks; Respectively swapping the first dimension and the fourth dimension of the initial first intermediate variable, the initial second intermediate variable, and the initial third intermediate variable to obtain a first intermediate variable, a second intermediate variable, and a third intermediate variable; The first dimension and the third dimension of the initial global receptive field are swapped to obtain a global receptive field, wherein the first intermediate variable, the second intermediate variable, the third intermediate variable, and the last dimension of the global receptive field are the number of images in the batch of images; Determine a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and a model weight parameter of a current processing module; Performing an inverse dimensionality transformation on the total output feature to obtain an inverse transformation result; Performing four-way scanning integration on the inverse transformation result to obtain a second image feature; When the current processing module is the last module, an image classification result is determined based on the second image feature.

2. The data processing method according to claim 1, characterized in that: The determining of the total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and the model weight parameter of the current processing module includes: The current second dimension of the first intermediate variable, the second intermediate variable and the third intermediate variable is divided to obtain two divided first intermediate variables, two divided second intermediate variables and two divided third intermediate variables, and the second dimension of the product of the global receptive field and the model weight parameter of the module is divided to obtain two divided fourth intermediate variables; Extract the first variable slice to be calculated from one of the first intermediate variables after division, one of the second intermediate variables after division, one of the third intermediate variables after division, and one of the fourth intermediate variables after division in a first preset order, and extract the second variable slice to be calculated from another of the first intermediate variables after division, another of the second intermediate variables after division, another of the third intermediate variables after division, and another of the fourth intermediate variables after division in a second preset order, wherein the first preset order is opposite to the second preset order; Calculate each of the first variable slices to obtain a first partial result, and calculate each of the second variable slices to obtain a second partial result; The first part of the results is stored in the first preset order, the second part of the results is stored in the second preset order, and the number of loop executions is increased by one to obtain the current number of loop executions; When the number of executions of the current loop does not reach half of the total number of the image data blocks, the steps of extracting the first variable slice to be calculated from one of the first intermediate variables after division, one of the second intermediate variables after division, one of the third intermediate variables after division and one of the fourth intermediate variables after division in the first preset order, and extracting the second variable slice to be calculated from another of the first intermediate variables after division, another of the second intermediate variables after division, another of the third intermediate variables after division and another of the fourth intermediate variables after division in the second preset order are returned to be executed until the number of executions of the current loop reaches half of the total number of the image data blocks, and a plurality of first partial results and a plurality of second partial results are stored, and each of the first partial results and each of the second partial results are spliced ​​to obtain the total output feature.

3. The data processing method according to any one of claims 1 or 2, characterized in that: The four-directional scanning of each of the first image features to obtain an initial global receptive field includes: For each of the first image features, the first image feature is copied to obtain two first image features, wherein the first dimension of the first image feature is the feature depth of the image data block, the second dimension is the longitudinal number of the image data blocks, and the third dimension is the lateral number of the image data blocks; multiplying the third dimension of one of the first image features by the fourth dimension, and taking the result as the third dimension, to obtain a third image feature; Swapping the third dimension and the fourth dimension of another of the first image features, and multiplying the swapped third dimension by the swapped fourth dimension, using the result as the third dimension, to obtain a fourth image feature; stacking the third image feature and the fourth image feature in the dimension of feature depth of the image data block to obtain a stacked feature; Copying the stacking feature to obtain two stacking features; Each of the stacked features is stacked in the dimension of the feature depth of the image data block to obtain an initial global receptive field.

4. The data processing method according to claim 3, characterized in that: The first dimension of the total output feature is the total number of the image data blocks, the second dimension is the feature depth of the image data blocks of a specified multiple, and the third dimension is the number of images. The inverse dimension transformation of the total output feature to obtain the inverse transformation result includes: Swapping the first dimension and the third dimension of the total output feature to obtain a total output feature after dimension transformation; The second dimension of the total output feature after the dimensional transformation is reshaped to obtain an inverse transformation result, wherein the first dimension of the inverse transformation result is the number of images, the second dimension is the specified multiple, the third dimension is the feature depth of the image data block, the fourth dimension is the longitudinal number, and the fifth dimension is the lateral number.

5. The data processing method according to claim 4, characterized in that: The horizontal number is the same as the vertical number, and the inverse transformation result is subjected to four-way scanning integration to obtain a second image feature, including: The total number of image data blocks obtained by multiplying the fourth dimension and the fifth dimension of the inverse transformation result is used as the fourth dimension to obtain a first feature representation; The first feature representation is divided in a third dimension to obtain a first sub-feature representation and a second sub-feature representation, wherein the first dimension of the first sub-feature representation and the second sub-feature representation is the number of images, the second dimension is half of the specified multiple, the third dimension is the feature depth of the image data block, and the fourth dimension is the total number of the image data blocks; Performing a bitwise summation on the first sub-feature representation and the second sub-feature representation to obtain a second feature representation, wherein a first dimension of the second image feature representation is the number of images, a second dimension is half of the specified multiple, a third dimension is the feature depth of the image data block, and a fourth dimension is the total number of the image data blocks; The second feature representation is divided in a third dimension to obtain a third sub-feature representation and a fourth sub-feature representation, wherein the first dimension of the third sub-feature representation and the fourth sub-feature representation is the number of images, the second dimension is one-fourth of the specified multiple, the third dimension is the feature depth of the image data block, and the fourth dimension is the total number of the image data blocks; Reshape the third dimension represented by the third sub-feature to obtain the longitudinal quantity as the third dimension and the lateral quantity as the fourth dimension; Reshape the third dimension represented by the fourth sub-feature to obtain the horizontal quantity as the third dimension and the vertical quantity as the fourth dimension; The third sub-feature representation and the fourth sub-feature representation are bitwise summed to obtain a second image feature.

6. The data processing method according to claim 4, characterized in that: The method further comprises: When the current processing module is not the last module, the second image feature is determined as the first image feature, the next module is determined as the current processing module, and the step of inputting the first image feature into the current processing module in the visual state space model and performing a four-way scan on each of the first image features to obtain an initial global receptive field is returned to execution until the current processing module is the last module.

7. A data processing device, characterized in that: include: An acquisition unit, used for acquiring a first image feature of each image in a batch of images; An input unit, used for inputting the first image features into a current processing module in a visual state space model, performing four-directional scanning on each of the first image features to obtain an initial global receptive field; an adjustment unit, configured to perform different feature adjustments on the initial global receptive field, and obtain an initial first intermediate variable, an initial second intermediate variable, and an initial third intermediate variable, respectively; the first dimension of the initial first intermediate variable and the initial second intermediate variable is the number of images of the batch images, the second dimension is the feature depth of the image data blocks of a specified multiple, the third dimension is the total number of image data blocks, and the fourth dimension is a preset hyperparameter; the first dimension of the initial third intermediate variable is the number of images, the second dimension is the feature depth of the image data blocks of a specified multiple, the third dimension is the preset hyperparameter, and the fourth dimension is the total number of image data blocks; the first dimension of the initial global receptive field is the number of images, the second dimension is the feature depth of the image data blocks of a specified multiple, and the third dimension is the total number of image data blocks; Transformation unit, comprising: A first swapping subunit is used to swap the first dimension and the fourth dimension of the initial first intermediate variable, the initial second intermediate variable and the initial third intermediate variable, respectively, to obtain a first intermediate variable, a second intermediate variable and a third intermediate variable; A second swapping subunit is configured to swap the first dimension and the third dimension of the initial global receptive field to obtain a global receptive field, wherein the first intermediate variable, the second intermediate variable, the third intermediate variable, and the last dimension of the global receptive field are the number of images in the batch of images; A first determining unit, configured to determine a total output feature based on the first intermediate variable, the second intermediate variable, the third intermediate variable, the global receptive field, and a model weight parameter of a current processing module; An inverse transformation unit, used to perform a dimensional inverse transformation on the total output feature to obtain an inverse transformation result; An integration unit, used for performing four-way scanning integration on the inverse transformation result to obtain a second image feature; The second determining unit is used to determine the image classification result based on the second image feature when the current processing module is the last module.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the data processing method according to any one of claims 1 to 6.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the data processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image classification method and device, equipment and storage medium

    CN117636014A

  • Image semantic segmentation method and device fused with Fourier transform

    CN119206236A