Data processing method and device, electronic equipment and storage medium

By combining an autoregressive sequence generation sub-model with a convolutional network, a hybrid network structure is constructed, dynamically adjusting the number of channels processed. This solves the problem of high computational complexity in existing neural networks, achieving efficient image processing results suitable for both mobile and server applications.

CN117151164BActive Publication Date: 2026-06-02BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2022-05-18
Publication Date
2026-06-02

Smart Images

  • Figure CN117151164B_ABST
    Figure CN117151164B_ABST
Patent Text Reader

Abstract

The method comprises: inputting an obtained to-be-processed image into a convolution network to obtain to-be-processed features of the to-be-processed image; processing the to-be-processed features based on a target network model and a preset channel ratio to obtain target features, wherein the target network model comprises an autoregressive sequence generation submodel and the convolution network, and the preset channel ratio is a ratio between an autoregressive channel processing quantity corresponding to the autoregressive sequence generation submodel and a convolution channel processing quantity corresponding to the convolution network; and performing analysis processing on the to-be-processed image based on the target features. The technical scheme of the embodiment of the present disclosure combines the autoregressive sequence generation submodel and the convolution network to construct a hybrid network structure, thereby combining the respective advantages of the two submodels, improving the feature extraction efficiency of the model, and simplifying the computational complexity of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the outstanding performance of neural networks in the field of natural language processing, the field of computer vision has also begun to apply neural networks to process visual images.

[0003] While existing client-side neural networks can process the data, their performance is unsatisfactory. Therefore, a deep neural network with a self-attention mechanism is proposed. This network can significantly reduce computational complexity; however, it incurs high computational costs. Summary of the Invention

[0004] This disclosure provides a data processing method, apparatus, electronic device, and storage medium to reduce the computational requirements of the model, reduce the amount of computation during model operation, and improve the model's data processing efficiency.

[0005] In a first aspect, embodiments of this disclosure provide a data processing method, the method comprising:

[0006] The acquired image to be processed is input into a convolutional network to obtain the features to be processed in the image;

[0007] Based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network.

[0008] The image to be processed is analyzed and processed based on the target features.

[0009] Secondly, embodiments of this disclosure also provide a data processing apparatus, the apparatus comprising:

[0010] An image input module is used to input the acquired image to be processed into a convolutional network to obtain the features to be processed in the image;

[0011] The feature processing module is used to process the features to be processed based on the target network model and the preset channel ratio to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network.

[0012] The image analysis module is used to analyze and process the image to be processed based on the target features.

[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0014] One or more processors;

[0015] Storage device for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method as described in any of the embodiments of this disclosure.

[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a data processing method as described in any of the embodiments of this disclosure.

[0018] The technical solution of this disclosure first inputs the acquired image to be processed into a convolutional network to obtain the features to be processed in the image. Further, based on the autoregressive sequence generation sub-model, the convolutional network, and the corresponding number of channels in each target network model, the features to be processed are processed sequentially to obtain target features. Finally, the image to be processed is analyzed and processed based on the target features. By combining the autoregressive sequence generation sub-model and the convolutional network, and dynamically adjusting the number of channels corresponding to the autoregressive sequence generation sub-model and the convolutional network according to the image processing requirements of the image to be processed, the advantages of the two sub-models are combined, improving the model feature extraction efficiency and simplifying the model's computational complexity. Attached Figure Description

[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0020] Figure 1 This is a schematic flowchart of a data processing method provided in an embodiment of this disclosure;

[0021] Figure 2 This is a schematic flowchart of a data processing method provided in an embodiment of this disclosure;

[0022] Figure 3 This is a network architecture diagram obtained by combining convolutional network and target network models provided in the embodiments of this disclosure;

[0023] Figure 4This is a network architecture diagram of the target network model provided in the embodiments of this disclosure;

[0024] Figure 5 This is a network architecture diagram of the target network model provided in the embodiments of this disclosure;

[0025] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this disclosure;

[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0038] When processing data using traditional deep neural network models, the models typically consume significant computational resources. This indirectly necessitates that the devices deploying the models possess substantial computing power, making them unsuitable for resource-constrained mobile deployments. Therefore, the technical solution provided in this disclosure introduces a hybrid network structure that combines an autoregressive sequence generation sub-model with a convolutional network. The autoregressive sequence generation sub-model is a deep neural network with a self-attention mechanism. Based on this, after convolutional processing of the input data to obtain intermediate features, the autoregressive sequence generation sub-model is used to further process these intermediate features, yielding higher-order, more abstract features, thereby improving data processing accuracy. Simultaneously, convolutional networks have advantages in extracting local feature information, further simplifying the model's computational speed.

[0039] The technical solutions provided in this disclosure can be deployed on mobile devices or on servers, and can meet any computing power requirements of the server.

[0040] Figure 1This is a schematic diagram of a data processing method provided in an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to situations where two neural network models with different advantages are combined to construct a hybrid network structure, thereby improving the data processing accuracy and efficiency of the model. The method can be executed by a data processing device, which can be implemented in the form of software and / or hardware, or optionally by an electronic device, such as a mobile terminal, a PC, or a server.

[0041] like Figure 1 As shown, the method includes:

[0042] S110. Input the obtained image to be processed into the convolutional network to obtain the features to be processed of the image.

[0043] Before introducing the solutions of the embodiments of this disclosure, it should first be noted that the model built based on the embodiments of this disclosure can be deployed on a server or a client. The server can be a targeted service program that provides services and resources to the client; the device running the server is the server itself. Correspondingly, the client is a program that provides local services to the user, corresponding to a specific server. Furthermore, the client and server can communicate based on the Hypertext Transfer Protocol (HTTP). For example, the network model in the embodiments of this disclosure can be integrated into application software that supports various functions such as special effects video processing and natural language processing. This software can be installed on an electronic device, optionally a mobile terminal or a PC. The application software can be software that processes data such as images, videos, and audio; specific application software will not be described in detail here, as long as it can process data such as images, videos, and audio. It can also be a specially developed application program that adds and displays special effects, or it can be integrated into a corresponding page, allowing users to process relevant data through the integrated page on a PC.

[0044] In this embodiment, the image to be processed can be an image acquired in real time by the image acquisition device of the terminal device and automatically uploaded to the corresponding server by the application software, or it can be an image selected by the user from the image storage space and actively uploaded to the server corresponding to the application software. In specific application scenarios, the image to be processed can be acquired in real time or periodically, and this embodiment does not impose specific limitations on this.

[0045] It should be noted that, for the network model in this embodiment, the image to be processed can also be a type of data to be processed. The type of data to be processed is determined by the functions provided by the application software. Therefore, when the application software provides audio processing or text processing functions to the user, the data to be processed can also be audio data, text data, and video data, etc. This embodiment will not elaborate further on this.

[0046] In practical applications, before processing the image to be processed based on the optimized autoregressive sequence generation model, a convolutional network (CNN) can be used to perform preliminary processing on the image to extract corresponding local features in order to improve the model's computational speed. A CNN is a type of feedforward neural network that includes convolutional computation and has a deep structure; it is one of the representative algorithms of deep learning. CNNs also possess representation learning capabilities, enabling translation-invariant classification of input information according to their hierarchical structure. This will not be elaborated further in the embodiments disclosed herein. The image to be processed is input into the CNN, and the output features are local image features of the image to be processed. Compared with global image features, local image features are rich in quantity, have low correlation between features, and do not affect the detection and matching of other features due to the disappearance of some features under occlusion. Therefore, the features to be processed can also be understood as a local representation of image features, reflecting the local features present in the image to be processed. These features are applicable to applications such as image matching and processing. The advantage of this setup is that it can improve the efficiency of subsequent feature extraction, thereby accelerating the model's computational speed.

[0047] S120. Based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features.

[0048] In this embodiment, once the features to be processed of the image are obtained, these features can be input into each target network model for processing. The target network model is a deep learning network model comprising an autoregressive sequence generation sub-model and a convolutional network. The autoregressive sequence generation sub-model can be a lightweight optimized transformer model, which is a model based on a self-attention mechanism to accelerate deep learning algorithms. The autoregressive sequence generation sub-model includes two normalization layers (Layer Norm or Batch Norm), a self-attention layer, and a multi-layer perceptron (MLP) layer.

[0049] The normalization layer is an algorithm used in deep network models to accelerate neural network training, convergence speed, and stability. Specifically, during model training, the normalization layer normalizes the training data to stabilize the forward input distribution and accelerate model convergence.

[0050] Self-attention subnetworks are neural networks used to implement the self-attention mechanism of a model. The self-attention mechanism is a mechanism for filtering out less important information from a large amount of information. Attention types are divided into spatial attention and temporal attention. In practical applications, it can also be divided into Soft Attention and Hard Attention. For Soft Attention, all data is paid attention to, and corresponding attention weights are calculated without setting filtering conditions. For Hard Attention, after generating all attention values, some attention values ​​that do not meet the conditions are filtered out, even if their attention weights are 0. This disclosure will not elaborate further. A multilayer perceptron is a feedforward artificial neural network model used to map multiple input datasets to a single output dataset. Those skilled in the art should understand that a typical multilayer perceptron includes three layers: an input layer, a hidden layer, and an output layer. Furthermore, the different layers of a multilayer perceptron neural network are fully connected, meaning that any neuron in one layer is connected to all neurons in the next layer. This disclosure will not elaborate further.

[0051] The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network.

[0052] It should be noted that when each target network model receives the features corresponding to the image to be processed, the autoregressive sequence generation sub-model in the target network model has certain advantages in extracting global features, maintaining high accuracy during feature extraction. Meanwhile, the convolutional network in the target network model has certain advantages in extracting local features, effectively improving the feature extraction rate. Therefore, before processing the features based on the target network model, different allocation ratios can be set according to the user's image processing needs. This allows the number of channels processed by the autoregressive sequence generation sub-model and the convolutional network to be dynamically adjusted, ensuring the total number of channels processed is the same as the number of channels in the image to be processed. For example, when the image processing requirement is to obtain a high-precision target image, the proportion of channels processed by the autoregressive sequence generation sub-model can be higher than that of the convolutional network; when the image processing requirement is to speed up image processing, the proportion of channels processed by the convolutional network can be higher than that of the autoregressive sequence generation sub-model. For example, if the number of channels for the current feature to be processed is 1000, and the channel processing ratio of the autoregressive sequence generation sub-model is 0.25 according to image processing requirements, then the number of channels processed by the autoregressive sequence generation sub-model is 250, and correspondingly, the number of channels processed by the convolutional network is 750.

[0053] In practical applications, after obtaining the features of the image to be processed and inputting these features into the target network model, the allocation ratio of the number of channels processed can be determined according to the current image processing requirements. This determines the number of channels processed corresponding to the autoregressive sequence generation sub-model and the number of channels processed corresponding to the convolutional network in the target network model. Furthermore, based on the autoregressive sequence generation sub-model, convolutional network, and their corresponding number of channels processed in each target network model, the features to be processed are processed sequentially. The features obtained after processing by the autoregressive sequence generation sub-model and the convolutional network are then concatenated to obtain the target features. In this way, the advantages of the autoregressive sequence generation sub-model and the convolutional network can be combined, and the ratio of the number of channels processed can be dynamically adjusted according to user needs, greatly improving the efficiency of server deployment.

[0054] It should also be noted that, since the solution in this embodiment combines an autoregressive sequence generation sub-model and a convolutional network to obtain a novel network model structure, in order to ensure the stability of the target network model, after obtaining the target features, the method further includes: inputting the target features into a normalization layer for normalization processing. In practical applications, the normalization layer (Batch Norm) can normalize the target features output by the target network model, thereby training the model's stability.

[0055] S130. Analyze and process the image to be processed based on the target features.

[0056] In this embodiment, after the target network models process the features to be processed and obtain the target features, the target features can be input into the feature analysis network deployed on the server or client side to realize the analysis and processing of the image to be processed. Optionally, the analysis and processing includes one or more of the following: scene classification; object detection; instance segmentation; 2D / 3D pose estimation. Therefore, it can be understood that the feature analysis network can select a variety of functional models according to the above business requirements, such as models for performing scene classification tasks, object detection tasks, or various stylized image processing tasks. Correspondingly, the processing results of these models are the processing results of various business requirements. For example, when the feature analysis network includes scene classification tasks, the target processing result is a classification matrix of multiple scene types; when the feature analysis network includes stylized image processing tasks, the target processing result is an image of a specific style corresponding to the image to be processed.

[0057] The technical solution of this disclosure first inputs the acquired image to be processed into a convolutional network to obtain the features to be processed in the image. Further, based on the target network model and a preset channel ratio, the features to be processed are processed to obtain target features. Finally, the image to be processed is analyzed and processed based on the target features. By combining the autoregressive sequence generation sub-model and the convolutional network, and dynamically adjusting the number of channels processed by the autoregressive sequence generation sub-model and the convolutional network according to the image processing requirements of the image to be processed, the advantages of the two sub-models are combined, improving the model feature extraction efficiency and simplifying the model's computational complexity.

[0058] Figure 2 This is a schematic flowchart of a data processing method provided in an embodiment of this disclosure. Based on the above embodiment, after obtaining the features to be processed from the image to be processed, it further includes a technical feature of determining the number of convolutional channels processed by the autoregressive sequence generation sub-model and the convolutional network, respectively. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiment will not be repeated here.

[0059] like Figure 2 As shown, the method specifically includes the following steps:

[0060] S210. Input the obtained image to be processed into the convolutional network to obtain the features to be processed of the image.

[0061] S220. Perform feature downsampling on the feature to be processed to update the feature to be processed.

[0062] In this embodiment, after obtaining the features to be processed of the image, since the Layer Norm layer in the subsequent target network model does not have the function of data downsampling, it is necessary to input the features to be processed into the pooling layer for downsampling processing in order to update the features to be processed.

[0063] In practical applications, downsampling is a multi-rate digital signal processing technique, which is also a process of reducing the signal sampling rate. It is typically used to change image resolution or image data size. For example, after downsampling an image with an original resolution of H×W by a factor of 2, the resolution of the resulting feature is H / 2×W / 2. Generally, the downsampling coefficients can be preset manually or automatically according to the user's image processing needs. Furthermore, the downsampling coefficients for each layer of the network can be the same or different; this disclosure does not specifically limit this.

[0064] In this embodiment, after the image to be processed is processed by the last convolutional network and the features to be processed are obtained, downsampling can be used to reduce the dimensions of each feature while retaining most of the important information, thereby achieving the updating of the features to be processed.

[0065] It should be noted that, in order to further improve the feature extraction rate of the features to be processed, and since convolutional networks have certain advantages in extracting local features, before downsampling the features to be processed, the following steps are also included: convolutional processing of the features to be processed based on at least one convolutional network to update the features to be processed.

[0066] In practical applications, after processing the image to be processed using a convolutional network to obtain the features to be processed, the obtained features are only local features in the image to be processed. There may be problems such as incomplete feature information or failure to extract the required feature information. Therefore, after obtaining the features to be processed, the features to be processed are input into at least one convolutional network for processing, thereby extracting multiple local feature information based on multiple convolutional networks to update the features to be processed.

[0067] For example, see Figure 3 As shown, when the client receives the image to be processed as input, it can perform a convolution operation with a kernel size of 3×3 and a stride of 2 on the image. Furthermore, a local region can be treated as a block and stacked multiple times based on a CNN, such as... Figure 1As shown, layers N1 can be stacked 3 times, layers N2 can be stacked 5 times, and layers N3 can be stacked 12 times. It can be understood that for each convolutional network, the input data needs to be convolved in the first layer. At the same time, when the above processing steps are divided into multiple stages according to the specific stacking process, the local features output by each stage are different. After processing by multiple convolutional networks, multiple local features corresponding to the image to be processed are obtained, and the features to be processed are updated based on the obtained multiple local features.

[0068] S230. Based on the convolutional layer and the preset channel ratio, determine the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network.

[0069] In this embodiment, since the autoregressive sequence generation sub-model in the target network model does not have the function of changing the number of image channels, before processing the features to be processed based on the autoregressive sequence generation sub-model, it is necessary to predetermine the channel processing ratio of the target network model to process the image to be processed based on the convolutional layer, so that the autoregressive sequence generation sub-model and the convolutional network can process the features to be processed according to the pre-set number of channels.

[0070] The channel processing ratio can be set according to the image processing task of the current image to be processed. For example, if the current image processing task is to obtain a feature image with high accuracy, the channel processing ratio of the autoregressive sequence generation sub-model can be higher than that of the convolutional network. If the current image processing task is to improve the processing speed of the feature image, the channel processing ratio of the convolutional network can be higher than that of the autoregressive sequence generation sub-model.

[0071] It should be noted that the channel processing ratio is set as a parameter in the convolutional layer. After downsampling to obtain the features to be processed, when processing based on the convolutional layer, it can be based on a ratio parameter in the convolutional layer to determine the number of autoregressive channel processing corresponding to the autoregressive sequence generation sub-model, and the number of convolutional channel processing corresponding to the convolutional network.

[0072] For example, when determining the channel processing ratio of the target network model, one can refer to... Figure 3When the number of target network models is N4, the first target network model in N4 needs to determine the channel processing ratio. After determining the channel processing ratio, the other target network models in N4 need to process the features to be processed according to the same channel processing ratio. Similarly, when the features to be processed are input into the number of target network models N5 for processing, the first target network model in N5 also needs to determine the channel processing ratio, and the other target network models in N5 need to process the features to be processed according to the same channel processing ratio.

[0073] In this embodiment, after determining the channel processing ratio of the target network model for the image to be processed, the number of channels of the image to be processed can be allocated to the autoregressive sequence generation sub-model and the convolutional network according to the channel processing ratio. Specifically, the number of channels processed by the autoregressive sequence generation sub-model is the autoregressive channel processing number, and the number of channels processed by the convolutional network is the convolutional channel processing number. Thus, the features to be processed are processed according to the determined channel processing numbers to meet the processing requirements of the current image.

[0074] For example, when the total number of channels in the image to be processed is 1000, and the channel processing ratio determined according to the current image processing task is 1:3, then the number of channels processed by the autoregressive sequence generation sub-model can be determined to be 250, and the number of channels processed by the convolutional network can be determined to be 750.

[0075] S240. Based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features.

[0076] In this embodiment, the number of target network models includes at least one, and the at least one target network model is sequentially connected. After the features to be processed are input into the target network model, in order to further improve the model processing efficiency, the features to be processed can be processed separately according to the autoregressive sequence generation sub-model and the convolutional network, and the processed feature information can be concatenated together to finally obtain the target features. In this way, the advantages of the autoregressive sequence generation sub-model and the convolutional network can be combined to not only improve the model feature extraction rate, but also simplify the model computation complexity.

[0077] Optionally, based on the target network model and a preset channel ratio, the features to be processed are processed to obtain target features, including: processing the features to be processed based on the autoregressive sequence generation sub-model and the number of autoregressive channels to obtain a first feature to be spliced; processing the first feature to be spliced ​​based on the convolutional network and the number of convolutional channels to obtain a second feature to be spliced; and splicing the first feature to be spliced ​​and the second feature to be spliced ​​together to obtain the target features.

[0078] In practical applications, since there is at least one target network model, for each target network model, a sub-model, a convolutional network, and the corresponding number of channels can be generated to process the features to be processed.

[0079] Specifically, the features to be processed are input into the autoregressive sequence generation sub-model. Based on the pre-determined number of processing channels corresponding to the number of autoregressive channels, the features to be processed are processed, and the output features are the first features to be concatenated. Furthermore, since the number of processing channels corresponding to the autoregressive sequence generation sub-model and the convolutional network has been pre-determined before processing the features based on the target network model, when processing the features based on the convolutional network, it is only necessary to input the features processed by the autoregressive sequence generation sub-model into the convolutional network. Based on this, the first features to be concatenated are input into the convolutional network, and based on the number of processing channels corresponding to the number of convolutional channels, the first features to be concatenated are processed, and the processed features are output, which are the second features to be concatenated. In this embodiment, since the target network model processes the features to be processed based on the autoregressive sequence generation sub-model and the convolutional network separately, after obtaining the first and second features to be concatenated, it is necessary to concatenate the two features to obtain the target concatenated features.

[0080] It should be noted that the processing method for each target network model is the same. In this embodiment, the processing method for one target network model is used as an example. If there are multiple target network models, the above processing method can be executed in a loop.

[0081] In practical applications, when there are multiple target network models, after obtaining the target splicing features output by the first target network model, these can be used as the input to the next target network model, i.e., the features to be processed. The steps for determining the target splicing features described above are repeated, and the target splicing features output by the last target network model are taken as the target features. The following section combines... Figure 4 The process by which the target network model processes the features to be processed is explained in detail.

[0082] See Figure 4As shown, based on the autoregressive sequence generation sub-model and the number of processing channels for the autoregressive channels, the features to be processed are processed to obtain the first feature to be concatenated, including: for each channel in the number of processing channels for the autoregressive channels: inputting the feature to be processed into the first sub-model to obtain the first output feature; performing residual processing on the first output feature and the feature to be processed to obtain the first residual feature; inputting the first residual feature into the second sub-model to obtain the second output feature; and performing residual processing on the second output feature and the first residual feature to obtain the first feature to be concatenated.

[0083] Among them, the first and second sub-modules are in the autoregressive sequence generation sub-model.

[0084] It should be noted that since the autoregressive sequence generation sub-model does not have the function of changing the number of channels processed, the processing of the features to be processed based on the autoregressive sequence generation sub-model is based on the processing channels corresponding to the predetermined number of autoregressive channels. Specifically, the features to be processed are input into the first sub-module, and the features to be processed by the first sub-module are processed to obtain the first output feature. Then, residual processing is performed on the first output feature and the features to be processed to obtain the first residual feature. Here, the residual is the difference between the actual observed value and the estimated value. In this embodiment, by performing residual processing on the first output feature and the features to be processed, the reliability of the processing result of the self-attention layer can be examined and verified. This embodiment will not be elaborated further here. Further, the first residual feature is input into the second sub-module to obtain the second output feature, and residual processing is performed again on the second output feature and the first residual feature to finally obtain the first feature to be spliced. It should be noted that the process of performing residual processing on the second output feature and the first residual feature is similar to the process of performing residual processing on the first output feature and the feature to be processed, and will not be described again in this embodiment.

[0085] Optionally, the first submodule is used to process the features to be processed to obtain the first output features, including: normalizing the features to be processed based on the normalization layer to obtain the first normalized features; and inputting the first normalized features into the self-attention layer to obtain the first output features.

[0086] The first submodule includes a normalization layer and a self-attention layer. For example, as shown below... Figure 5 As shown, in practical applications, after the feature to be processed is input into the first sub-module, the feature to be processed is first normalized based on the normalization layer to obtain a higher-order feature, namely the first normalized feature. Further, the first normalized feature is processed based on the self-attention layer to determine the weight of each sub-feature contained in the first feature. After associating each weight with each sub-feature, the first output feature can be obtained.

[0087] Optionally, the first residual feature is input into the second submodule to obtain the second output feature, including: normalizing the first residual feature based on the normalization layer to obtain the second normalized feature; and processing the second normalized feature based on the multilayer perceptron layer to obtain the second output feature.

[0088] The second submodule includes a normalization layer and a multilayer perceptron layer. For example, as shown... Figure 5 As shown, in practical applications, after obtaining the first output feature and processing the first output feature and the feature to be processed, the first residual feature can be obtained. The first residual feature is input into the second sub-module. First, the first residual feature is normalized based on the normalization layer to obtain the second normalized feature. Further, the second normalized feature is processed based on the multilayer perceptron layer to finally obtain the second output feature.

[0089] Furthermore, after obtaining the first feature to be concatenated, the first feature to be concatenated is input into a convolutional network for processing. Optionally, based on the convolutional network and the processing channels corresponding to the number of convolutional channels, the first feature to be concatenated is processed to obtain a second feature to be concatenated, including: for each channel in the processing channels of the number of convolutional channels: processing the first feature to be concatenated sequentially based on at least three convolutional layers in the convolutional network to obtain the feature to be applied; determining the second feature to be concatenated by processing the residual between the feature to be applied and the first feature to be concatenated.

[0090] It should be noted that when processing the first feature to be concatenated using a convolutional network, the processing should be performed according to the predetermined number of processing channels corresponding to the number of convolutional channels. See [link / reference] Figure 5 As shown, in this embodiment, the convolutional network is composed of at least three convolutional layers with different kernels stacked together. For example, the kernel size of the first convolutional layer is 1×1, the kernel size of the second convolutional layer is 3×3, and the kernel size of the third convolutional layer is 1×1. The first feature to be concatenated is processed sequentially by each convolutional layer in the convolutional network, and the resulting feature is the feature to be applied. Then, residual processing is performed between the feature to be applied and the first feature to be concatenated to evaluate the processing result of the convolutional network, and finally, the second feature to be concatenated is obtained.

[0091] In practical applications, after obtaining the first and second features to be spliced, the obtained features need to be spliced ​​to finally obtain the target features.

[0092] It should also be noted that the model structure in the autoregressive sequence generation sub-model can be arbitrary, as long as its general structure is the same as that of the autoregressive sequence generation sub-model, it is within the protection scope of this technical solution. That is, changes to the individual layers included in the autoregressive generation sub-model will not affect this technical solution. The technical solution provided in this disclosure is mainly used to determine the channel processing ratio of the target network model, determine the number of autoregressive channel processing and the number of convolutional channel processing based on the channel processing ratio, and process the features based on the autoregressive sequence generation sub-model, the convolutional network, and the corresponding number of channel processing to obtain the target features.

[0093] S250. Analyze and process the image to be processed based on the target features.

[0094] The technical solution of this disclosure first inputs the acquired features to be processed into a convolutional network to obtain the features to be processed. Then, downsampling is performed on the features to be processed to update them. Based on the convolutional layer and the channel processing ratio set in the convolutional layer, the number of autoregressive channel processing corresponding to the autoregressive sequence generation sub-model and the number of convolutional channel processing corresponding to the convolutional network are determined. Further, based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features. Finally, the image to be processed is analyzed and processed based on the target features. By combining the autoregressive sequence generation sub-model and the convolutional network, and dynamically adjusting the number of channel processing corresponding to the autoregressive sequence generation sub-model and the convolutional network according to the image processing requirements of the image to be processed, the advantages of the two sub-models are combined, improving the model feature extraction efficiency and simplifying the model's computational complexity.

[0095] Figure 6 This is a schematic diagram of a data processing device structure provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the device includes: an image input module 610, a feature processing module 620, and an image analysis module 630.

[0096] The image input module 610 is used to input the acquired image to be processed into the convolutional network to obtain the features to be processed of the image;

[0097] The feature processing module 620 is used to process the features to be processed based on the target network model and the preset channel ratio to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network.

[0098] Image analysis module 630 is used for image analysis and processing based on target features.

[0099] Based on the above technical solutions, after obtaining the features to be processed of the image to be processed, the device further includes: a feature update module and a channel processing ratio determination module.

[0100] The feature update module is used to downsample the features to be processed in order to update the features to be processed.

[0101] The channel processing quantity determination module is used to determine the number of autoregressive channel processing corresponding to the autoregressive sequence generation sub-model and the number of convolutional channel processing corresponding to the convolutional network, based on the convolutional layer and the preset channel ratio.

[0102] Based on the above technical solutions, the feature processing module 620 includes a first feature processing unit, a second feature processing unit, a feature splicing unit, and a feature determination unit.

[0103] The first feature processing unit is used to process the features to be processed based on the autoregressive sequence generation sub-model and the processing channels corresponding to the number of autoregressive channels, to obtain the first feature to be spliced.

[0104] The second feature processing unit is used to process the first feature to be concatenated based on the convolutional network and the processing channels corresponding to the number of convolutional channels, so as to obtain the second feature to be concatenated.

[0105] The feature splicing unit is used to splice the first feature to be spliced ​​and the second feature to be spliced ​​to obtain the target feature.

[0106] Based on the above technical solutions, the autoregressive sequence generation sub-model includes a first sub-module and a second sub-module.

[0107] The first feature processing unit includes a first feature processing subunit, a second feature processing subunit, a third feature processing subunit, and a fourth feature processing subunit.

[0108] For each of the processing channels in the number of autoregressive channels: a first feature processing subunit is used to input the feature to be processed into a first sub-model to obtain a first output feature; a second feature processing subunit is used to perform residual processing on the first output feature and the feature to be processed to obtain a first residual feature; a third feature processing subunit is used to input the first residual feature into a second sub-model to obtain a second output feature; and a fourth feature processing subunit is used to perform residual processing on the second output feature and the first residual feature to obtain a first feature to be concatenated.

[0109] Based on the above technical solutions, the first submodule includes a normalization layer and a self-attention layer. The first feature processing subunit is also used to perform normalization processing on the feature to be processed based on the normalization layer to obtain a first normalized feature; and to input the first normalized feature into the self-attention layer to obtain a first output feature.

[0110] Based on the above technical solutions, the second submodule includes a normalization layer and a multilayer perceptron layer, and a third feature processing subunit, which is also used to normalize the first residual feature based on the normalization layer to obtain the second normalized feature; and to process the second normalized feature based on the multilayer perceptron layer to obtain the second output feature.

[0111] Based on the above technical solutions, for each channel in the processing channel of the number of convolutional channels: the second feature processing subunit is also used to process the first feature to be spliced ​​sequentially based on at least three convolutional layers in the convolutional network to obtain the feature to be applied; and to determine the second feature to be spliced ​​by processing the residual between the feature to be applied and the first feature to be spliced.

[0112] Based on the above technical solutions, after obtaining the target features, the device further includes a feature update module.

[0113] The feature update module is used to input the target features into the normalization layer for normalization processing.

[0114] Based on the above technical solutions, the analysis and processing includes one or more of the following: scene classification; target detection; instance segmentation; two-dimensional / three-dimensional pose estimation.

[0115] The technical solution of this disclosure first inputs the acquired image to be processed into a convolutional network to obtain the features to be processed in the image. Further, based on the target network model and a preset channel ratio, the features to be processed are processed to obtain target features. Finally, the image to be processed is analyzed and processed based on the target features. By combining the autoregressive sequence generation sub-model and the convolutional network, and dynamically adjusting the number of channels processed by the autoregressive sequence generation sub-model and the convolutional network according to the image processing requirements of the image to be processed, the advantages of the two sub-models are combined, improving the model feature extraction efficiency and simplifying the model's computational complexity.

[0116] The data processing apparatus provided in this disclosure can execute the data processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0117] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0118] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0119] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.

[0120] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0121] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0122] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0123] The electronic device provided in this embodiment and the data processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0124] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the data processing method provided in the above embodiments.

[0125] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0126] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0127] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0128] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0129] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: input the acquired image to be processed into a convolutional network to obtain the features to be processed of the image;

[0130] Based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network. The image to be processed is then analyzed and processed based on the target features.

[0131] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive a node evaluation request including at least two Internet Protocol (IP) addresses; select an IP address from the at least two IP addresses; and return the selected IP address; wherein the received IP address indicates an edge node in the content delivery network.

[0132] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0134] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0135] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0136] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0137] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0138] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0139] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A data processing method, characterized by, include: The acquired image to be processed is input into a convolutional network to obtain the features to be processed in the image; Based on the target network model and the preset channel ratio, the features to be processed are processed to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network. The image to be processed is analyzed and processed based on the target features; The step of processing the features to be processed based on the target network model and a preset channel ratio to obtain the target features includes: Based on the autoregressive sequence generation sub-model and the number of processing channels for the autoregressive channels, the features to be processed are processed to obtain the first feature to be spliced. Based on the convolutional network and the number of processing channels of the convolutional channels, the first feature to be spliced ​​is processed to obtain the second feature to be spliced. The first feature to be spliced ​​and the second feature to be spliced ​​are spliced ​​together to obtain the target feature.

2. The method according to claim 1, characterized in that, After obtaining the features to be processed in the image to be processed, the process further includes: The features to be processed are downsampled to update the features to be processed; Based on the convolutional layer and the preset channel ratio, the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network are determined.

3. The method according to claim 1, characterized in that, The autoregressive sequence generation sub-model includes a first sub-module and a second sub-module. The processing channel, based on the autoregressive sequence generation sub-model and the number of autoregressive channels, processes the features to be processed to obtain the first feature to be concatenated, including: For each of the processing channels in the number of autoregressive channels: The feature to be processed is input into the first sub-model to obtain the first output feature; The first output feature and the feature to be processed are subjected to residual processing to obtain the first residual feature; The first residual feature is input into the second sub-model to obtain the second output feature; The second output feature and the first residual feature are subjected to residual processing to obtain the first feature to be spliced.

4. The method according to claim 3, characterized in that, The first submodule includes a normalization layer and a self-attention layer. The processing of the features to be processed based on the first submodule to obtain the first output feature includes: The features to be processed are normalized based on the normalization layer to obtain the first normalized feature; The first normalized feature is input into the self-attention layer to obtain the first output feature.

5. The method according to claim 3, characterized in that, The second submodule includes a normalization layer and a multilayer perceptron layer. The step of inputting the first residual feature into the second submodule to obtain the second output feature includes: The first residual feature is normalized based on the normalization layer to obtain the second normalized feature; The second normalized feature is processed based on the multilayer perceptron layer to obtain the second output feature.

6. The method according to claim 1, characterized in that, The first feature to be concatenated is processed based on the convolutional network and the number of processing channels corresponding to the number of convolutional channels to obtain the second feature to be concatenated, including: For each channel in the processing channels of the stated number of convolution channels: The first feature to be concatenated is processed sequentially based on at least three convolutional layers in the convolutional network to obtain the feature to be applied; The second feature to be spliced ​​is determined by processing the residuals of the feature to be applied and the first feature to be spliced.

7. The method according to claim 1, characterized in that, After obtaining the target features, the process also includes: The target features are input into the normalization layer for normalization processing.

8. The method according to any one of claims 1-7, characterized in that, The analytical processing includes one or more of the following: Scene classification; object detection; instance segmentation; 2D / 3D pose estimation.

9. A data processing apparatus, characterized in that, include: An image input module is used to input the acquired image to be processed into a convolutional network to obtain the features to be processed in the image; The feature processing module is used to process the features to be processed based on the target network model and the preset channel ratio to obtain the target features. The target network model includes an autoregressive sequence generation sub-model and a convolutional network. The preset channel ratio is the ratio between the number of autoregressive channels processed corresponding to the autoregressive sequence generation sub-model and the number of convolutional channels processed corresponding to the convolutional network. The image analysis module is used to analyze and process the image to be processed based on the target features; The feature processing module includes: The first feature processing unit is used to process the feature to be processed based on the autoregressive sequence generation sub-model and the number of processing channels of the autoregressive channels to obtain the first feature to be spliced. The second feature processing unit is used to process the first feature to be spliced ​​based on the convolutional network and the number of processing channels of the convolutional channels to obtain the second feature to be spliced. The feature splicing unit is used to splice the first feature to be spliced ​​and the second feature to be spliced ​​to obtain the target feature.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method as described in any one of claims 1-8.

11. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the data processing method as described in any one of claims 1-8.