Image data processing method and device based on backbone network

By stitching multi-view images and inputting to optimized backbone networks for feature extraction, the problem of insufficient computing requirements in the existing backbone network in autonomous driving scenarios is solved, and efficient image processing and fast inference are achieved.

CN120198628APending Publication Date: 2025-06-24SHENZHEN DEEPROUTE AI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311782060.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing backbone network cannot meet the computing needs in autonomous driving scenarios, and there are problems with high redundant computing and computing power support, and performance deteriorates during quantization and deployment.

Method used

By acquiring multi-view images and performing stitching processing, a stitching tensor is obtained, and then inputting it into multiple downsampling layers and spatial channel feature extraction layers in the backbone network, downsampling and feature extraction operations of different multiples are performed, and feature tensors with downsampling and high-order semantic information are output.

Benefits of technology

The backbone network structure is optimized, the efficiency of high-resolution image processing is improved, and it is suitable for the fast reasoning requirements in autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198628A_ABST
    Figure CN120198628A_ABST
Patent Text Reader

Abstract

The invention discloses an image data processing method and device based on a backbone network, and the method comprises the steps: obtaining a multi-view image, and carrying out the splicing processing of the multi-view image, and obtaining a splicing tensor; inputting the spliced tensor into a backbone network, and performing down-sampling and feature extraction operations of different multiples through a plurality of down-sampling layers and a space channel feature extraction layer of the backbone network to obtain a corresponding down-sampling feature tensor with high-order semantic information; and outputting the obtained down-sampled feature tensor with the high-order semantic information, and completing the conversion process from the multi-view image to the feature tensor containing the high-order semantic information. According to the method, the features are extracted through the optimized backbone network structure, and the high-resolution image processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to an image data processing method and device based on a backbone network. Background Art

[0002] An autonomous driving perception model based on deep learning generally includes: a backbone network model, a multi-scale feature fusion module, and detection and segmentation heads corresponding to different tasks. Among them, the backbone network model is the uppermost part of the entire perception model and is responsible for effectively extracting features from the input data.

[0003] The designs of existing backbone networks are basically for scientific research, with a large amount of redundant calculations in the overall architecture and the design of feature extraction units, requiring extremely high computing power support; at the same time, some operators inside the backbone network have serious performance degradation during quantization and deployment, and the problem of long inference time. For scenarios such as autonomous driving that require fast inference and limited computing power of in-vehicle hardware, the current backbone network cannot be effectively deployed.

[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that, aiming at the defects of the existing technology, the present invention provides an image data processing method and device based on a backbone network to solve the problem that the existing backbone network cannot meet the computing requirements of the autonomous driving scenario.

[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0007] In a first aspect, the present invention provides an image data processing method based on a backbone network, including:

[0008] Obtain multi-view images, and perform stitching processing on the multi-view images to obtain a stitched tensor;

[0009] Input the stitched tensor into the backbone network, and perform downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial channel feature extraction layers of the backbone network to obtain corresponding downsampled feature tensors with high-order semantic information;

[0010] Output the obtained downsampled feature tensors with high-order semantic information to complete the conversion process from multi-view images to feature tensors containing high-order semantic information.

[0011] In one implementation, the obtaining multi-view images and performing stitching processing on the multi-view images to obtain a stitched tensor includes:

[0012] Obtain the multi-view images from the surround-view cameras, and splice the multi-view images into a tensor of shape N, C, H, W to obtain the spliced tensor; where N is the number of batches * the number of cameras, C is the number of channels of the input image data, and H and W respectively represent the height and width of the image.

[0013] In one implementation, input the spliced tensor into the backbone network, and perform downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network, including:

[0014] Perform downsampling on the spliced tensor through the first feature extraction unit of the backbone network to obtain a first downsampled tensor;

[0015] Perform downsampling and feature enhancement processing on the first downsampled tensor through the second feature extraction unit of the backbone network to obtain an enhanced second downsampled tensor;

[0016] Perform downsampling and feature enhancement processing on the enhanced second downsampled tensor through the third feature extraction unit of the backbone network to obtain an enhanced third downsampled tensor;

[0017] Perform downsampling and feature enhancement processing on the enhanced third downsampled tensor through the fourth feature extraction unit of the backbone network to obtain an enhanced fourth downsampled tensor;

[0018] Among them, the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor are all feature tensors with high-order semantic information.

[0019] In one implementation, the performing downsampling and feature enhancement processing on the first downsampled tensor through the second feature extraction unit of the backbone network includes:

[0020] Input the first downsampled tensor into the second feature extraction unit of the backbone network, and obtain a second downsampled tensor through the downsampling convolutional layer of the second feature extraction unit;

[0021] Pass the second downsampled tensor through the spatial feature extraction unit, add the extracted multiple features and then add them to the features of the shortcut connection to obtain the features output by the spatial feature extraction unit;

[0022] Pass the features output by the spatial feature extraction unit through the channel feature extraction unit, map the feature channels to the original S times dimension, and compress the feature channels back to C to obtain the enhanced second downsampled tensor.

[0023] In one implementation, the downsampling and feature enhancement processing of the enhanced second downsampled tensor by the third feature extraction unit of the backbone network includes:

[0024] Input the enhanced second downsampled tensor into the third feature extraction unit of the backbone network, and perform downsampling and feature enhancement processing according to the processing flow of the second feature extraction unit to obtain the enhanced third downsampled tensor.

[0025] In one implementation, the downsampling and feature enhancement processing of the enhanced third downsampled tensor by the fourth feature extraction unit of the backbone network includes:

[0026] Input the enhanced third downsampled tensor into the fourth feature extraction unit of the backbone network, and perform downsampling and feature enhancement processing according to the processing flow of the second feature extraction unit to obtain the enhanced fourth downsampled tensor.

[0027] In one implementation, the output of the downsampled feature tensor with high-order semantic information includes:

[0028] Take the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor as the inputs of the multi-scale feature fusion module and the corresponding network for the downstream task, and complete the conversion process from the multi-view image to the feature tensor containing high-order semantic information.

[0029] In a second aspect, the present invention provides an image data processing device based on a backbone network, including:

[0030] A splicing processing module for obtaining a multi-view image and performing splicing processing on the multi-view image to obtain a spliced tensor;

[0031] A backbone network module for inputting the spliced tensor into the backbone network, and performing downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial channel feature extraction layers of the backbone network to obtain the corresponding downsampled feature tensor with high-order semantic information;

[0032] An output module for outputting the obtained downsampled feature tensor with high-order semantic information to complete the conversion process from the multi-view image to the feature tensor containing high-order semantic information.

[0033] In a third aspect, the present invention provides a terminal, including: a processor and a memory, where the memory stores an image data processing program based on a backbone network, and when the image data processing program based on the backbone network is executed by the processor, it is used to implement the operations of the image data processing method based on the backbone network as described in the first aspect.

[0034] In a fourth aspect, the present invention further provides a medium, which is a computer-readable storage medium storing an image data processing program based on a backbone network. When the image data processing program based on the backbone network is executed by a processor, it is used to implement the operations of the image data processing method based on the backbone network as described in the first aspect.

[0035] The present invention adopts the above technical solutions and has the following effects:

[0036] The present invention obtains multi-view images, performs stitching processing on the multi-view images to obtain a stitching tensor, inputs the stitching tensor into a backbone network, and performs downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network to obtain corresponding downsampled feature tensors with high-order semantic information. Finally, the obtained downsampled feature tensors with high-order semantic information are output to complete the conversion process from multi-view images to feature tensors with high-order semantic information. The present invention extracts features through an optimized backbone network structure, improving the processing efficiency of high-resolution image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.

[0038] Figure 1 is a flowchart of an image data processing method based on a backbone network in an implementation manner of the present invention.

[0039] Figure 2 is a processing schematic diagram of an image data processing algorithm based on a backbone network in an implementation manner of the present invention.

[0040] Figure 3 is a functional schematic diagram of a terminal in an implementation manner of the present invention.

[0041] The implementation, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the object, technical solutions, and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] Exemplary Method

[0044] The designs of existing backbone networks are basically oriented towards scientific research, with a large amount of redundant calculations in the design of the entire architecture and feature extraction units, requiring extremely high computing power support. At the same time, some operators inside the backbone network have serious performance degradation and long inference time during quantization and deployment. For scenarios such as autonomous driving that require fast inference and limited on-vehicle hardware computing power, the current backbone networks cannot be effectively deployed.

[0045] To address the above technical problems, an image data processing method based on a backbone network is provided in an embodiment of the present invention. The method obtains multi-view images, performs stitching processing on the multi-view images to obtain a stitched tensor, and inputs the stitched tensor into the backbone network. Through multiple downsampling layers and spatial channel feature extraction layers of the backbone network, different-fold downsampling and feature extraction operations are performed to obtain corresponding downsampled feature tensors with high-order semantic information. Finally, the obtained downsampled feature tensors with high-order semantic information are output to complete the conversion process from multi-view images to feature tensors containing high-order semantic information. The embodiment of the present invention extracts features through an optimized backbone network structure, improving the processing efficiency of high-resolution images.

[0046] As Figure 1 shown, an embodiment of the present invention provides an image data processing method based on a backbone network, including the following steps:

[0047] Step S100, obtain multi-view images, and perform stitching processing on the multi-view images to obtain a stitched tensor.

[0048] In this embodiment, an autonomous driving perception model based on deep learning generally includes: a backbone network model, a multi-scale feature fusion module, and detection and segmentation heads corresponding to different tasks. Among them, the backbone network model is the uppermost part of the entire perception model, responsible for effectively extracting features from the input data; a block (internal unit of the network) is a feature extraction unit that constitutes the backbone network.

[0049] The overall architecture of the backbone network and the internal unit (block) of the backbone network in this embodiment are designed with high efficiency and light weight, with faster inference speed and less GPU (graphics processing unit) video memory occupancy, and are easy to deploy on in-vehicle chips.

[0050] At the level of the overall architecture of the backbone network: at the large-resolution feature map (the large-resolution feature map is reflected in the first feature extraction unit in the backbone network), a faster downsampling method is used to reduce the computational amount of the backbone network model; at the same time, the number of feature channels is kept constant as much as possible between stages of the backbone network and inside the internal unit of the backbone network (i.e., the light block, the spatial channel feature extraction layer), improving the parallelism of inference;

[0051] At the level of the internal units of the backbone network: A design that improves the rotational invariance of the network is adopted to effectively expand the network depth while enhancing the feature representation ability; in the design of the internal units of the backbone network in this embodiment, only conventional 1*1 convolution, 1*k / k*1 convolution, and k*k convolution are included to achieve friendly support for quantization; the spatial features and channel features are decoupled and represented, which is more conducive to the expression of features and their propagation in the network.

[0052] Specifically, in one implementation manner of this embodiment, step S100 includes the following steps:

[0053] Step S101, obtain the multi-view images from the surround-view cameras, and splice the multi-view images into a tensor with the shape of N, C, H, W to obtain the spliced tensor.

[0054] In this embodiment, multi-view images are obtained from the surround-view cameras, and then the multi-view images are input into the bird's-eye view model, and are spliced into a tensor with the shape of N, C, H, W through the bird's-eye view model. Among them, N is the number of batches * the number of cameras, C is the number of channels of the input data, the number of image channels is 3, and H and W respectively represent the height and width of the image.

[0055] It can be understood that the images to be processed in this embodiment are not only multi-view images obtained from vehicle surround-view cameras, but also images obtained from any usage scenarios, so as to complete the conversion process from multi-view images to feature tensors containing high-order semantic information, and further facilitate downstream modules or networks to implement lightweight, high-resolution, and high-efficiency image detection and segmentation tasks; for example, image processing tasks that require rapid detection and tracking of targets.

[0056] As Figure 1 shown, in one implementation manner of the embodiment of the present invention, the image data processing method based on the backbone network further includes the following steps:

[0057] Step S200, input the spliced tensor into the backbone network, and perform downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial channel feature extraction layers of the backbone network to obtain corresponding downsampled feature tensors with high-order semantic information.

[0058] In this embodiment, the backbone network includes multiple feature extraction units (stages), and each feature extraction unit (stage) includes: a downsampling layer and a spatial channel feature extraction layer; in each feature extraction unit (stage), corresponding downsampling and feature extraction operations are performed through the corresponding downsampling layer and spatial channel feature extraction layer, so as to obtain corresponding downsampled feature tensors with high-order semantic information for each feature extraction unit (stage).

[0059] In this embodiment, for the spliced tensor of N, C, H, and W, it can be sent to the first feature extraction unit (stage 1) of the backbone network, and the first feature extraction unit is used to perform fast downsampling, so as to realize the transformation of the number of channels of the spliced features.

[0060] Specifically, in one implementation manner of this embodiment, step S200 includes the following steps:

[0061] Step S201, downsample the spliced tensor through the first feature extraction unit of the backbone network to obtain a first downsampled tensor.

[0062] In this embodiment, since the starting input data (i.e., the spliced tensor of N, C, H, and W) is very large in the H and W dimensions, directly using convolution operations will be very time-consuming. Therefore, a fast downsampling operation can be adopted in the first feature extraction unit (stage 1) to avoid the problem of very time-consuming operations in turn; the fast downsampling methods in this embodiment include but are not limited to: convolution operations with large convolution kernels and large strides. After the fast downsampling operation of the first feature extraction unit (stage 1), a feature tensor of size N, C / 2, H / 4, W / 4 is output.

[0063] Step S202, downsample and perform feature enhancement processing on the first downsampled tensor through the second feature extraction unit of the backbone network to obtain an enhanced second downsampled tensor.

[0064] In this embodiment, for the feature tensor of N, C / 2, H / 4, W / 4 output by the fast downsampling of the first feature extraction unit (stage 1), it can be input into the second feature extraction unit (stage 2), and through the operations of 2-fold downsampling, spatial feature extraction, and channel feature extraction of the second feature extraction unit (stage 2), an enhanced feature tensor of N*C*H / 8*W / 8 is obtained.

[0065] Specifically, in one implementation manner of this embodiment, step S202 includes the following steps:

[0066] Step S202a, input the first downsampled tensor into the second feature extraction unit of the backbone network, and obtain a second downsampled tensor through the downsampling convolutional layer of the second feature extraction unit;

[0067] Step S202b, pass the second downsampled tensor through the spatial feature extraction unit, add the extracted multiple features and then add them to the features of the shortcut connection to obtain the features output by the spatial feature extraction unit;

[0068] In step S202c, the features output by the spatial feature extraction unit are passed through the channel feature extraction unit to map the feature channels to the original S times the dimension, and then the feature channels are compressed back to C to obtain the enhanced second downsampled tensor.

[0069] In this embodiment, a feature tensor of N, C / 2, H / 4, W / 4 is input into the second feature extraction unit (stage 2) for downsampling and enhancement processing; as Figure 2 shown, the second feature extraction unit (stage 2) includes a convolution for 2-fold downsampling and two light blocks (spatial channel feature extraction layers) for feature enhancement.

[0070] In the second feature extraction unit (stage 2), after the 2-fold downsampling convolution, the feature dimension becomes N * C * H / 8 * W / 8, and then it is sent into the light block (spatial channel feature extraction layer) for further feature extraction. The specific process is as follows:

[0071] The light block (spatial channel feature extraction layer) contains two major parts. The first part is the spatial feature extraction unit, and the second part is the channel feature extraction unit. Among them, the spatial feature extraction unit is mainly composed of 3 parallel convolutions, namely convolutions with a kernel size of k * k, k * 1, and 1 * k. During network calculation, the input feature N * C * H / 8 * W / 8 will pass through the three convolutions in parallel, and then the three obtained features are added together to obtain the added feature; finally, the added feature is added to the feature of the short-circuit connection to obtain the final output feature. This can enhance the ability of feature rotation invariance on the one hand and facilitate the stability during training on the other hand.

[0072] The second part is the channel feature extraction unit, which mainly includes two convolutions with a kernel size of 1 * 1, and there is an S-fold magnification and reduction in the channels for the two convolutions. The input feature will first pass through the first 1 * 1 convolution to map the feature channels to the original S times the dimension, and then pass through the second 1 * 1 convolution to compress the feature channels back to C. Thus, the feature extraction process inside the entire light block (spatial channel feature extraction layer) is completed.

[0073] As an implementation method, in the second feature extraction unit (stage 2), the extraction process of the light block (spatial channel feature extraction layer) can be performed twice.

[0074] Specifically, in an implementation method of this embodiment, step S200 further includes the following steps:

[0075] Step S203: Downsample and perform feature enhancement on the enhanced second downsampled tensor through the third feature extraction unit of the backbone network to obtain an enhanced third downsampled tensor.

[0076] In this embodiment, for the enhanced second downsampled tensor obtained by the downsampling and feature enhancement processing of the second feature extraction unit (stage 2), it can be input into the third feature extraction unit (stage 3) of the backbone network, and then pass through the third feature extraction unit (stage 3) once according to the same process as the second feature extraction unit (stage 2) to obtain an output feature of N*C*H / 16*W / 16.

[0077] Specifically, in one implementation manner of this embodiment, step S203 includes the following steps:

[0078] Step S203a: Input the enhanced second downsampled tensor into the third feature extraction unit of the backbone network, perform downsampling and feature enhancement processing according to the processing process of the second feature extraction unit to obtain the enhanced third downsampled tensor.

[0079] As Figure 2 shown, the structure of the third feature extraction unit (stage 3) of the backbone network is the same as that of the second feature extraction unit (stage 2), including a convolution for 2-fold downsampling and two light blocks (spatial channel feature extraction layers) for feature enhancement; therefore, in the third feature extraction unit (stage 3), through operations of 2-fold downsampling, spatial feature extraction, and channel feature extraction, an enhanced feature tensor of N*C*H / 16*W / 16 is obtained.

[0080] Step S204: Downsample and perform feature enhancement on the enhanced third downsampled tensor through the fourth feature extraction unit of the backbone network to obtain an enhanced fourth downsampled tensor.

[0081] In this embodiment, for the enhanced feature tensor of N*C*H / 16*W / 16 obtained by the downsampling and feature enhancement processing of the third feature extraction unit (stage 3), it can be input into the fourth feature extraction unit (stage 4) of the backbone network, and then pass through the fourth feature extraction unit (stage 4) once according to the same process as the second feature extraction unit (stage 2) to obtain an output feature of N*C*H / 32*W / 32.

[0082] Specifically, in one implementation manner of this embodiment, step S204 includes the following steps:

[0083] Step S204a: Input the enhanced third downsampled tensor into the fourth feature extraction unit of the backbone network, and perform downsampling and feature enhancement processing according to the processing flow of the second feature extraction unit to obtain the enhanced fourth downsampled tensor.

[0084] As Figure 2 shown, the structure of the fourth feature extraction unit (stage 4) of the backbone network is the same as that of the second feature extraction unit (stage 2), including a convolution for 2x downsampling and two light blocks (spatial channel feature extraction layers) for feature enhancement; therefore, in the fourth feature extraction unit (stage 4), through operations of 2x downsampling, spatial feature extraction, and channel feature extraction, a feature tensor of N*C*H / 32*W / 32 after enhancement is obtained.

[0085] In this embodiment, the number of channels of the second feature extraction unit (stage 2), the third feature extraction unit (stage 3), and the fourth feature extraction unit (stage 4) is all C, which can improve the computational memory access ratio during network inference and speed up the inference speed.

[0086] It is worth mentioning that the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor in this embodiment are all feature tensors with high-order semantic information. These downsampled feature tensors with high-order semantic information can be used as the input of the multi-scale feature fusion module and the following downstream task corresponding network to complete the conversion process from multi-view images to feature tensors containing high-order semantic information.

[0087] As Figure 1 shown, in an implementation manner of the embodiment of the present invention, the image data processing method based on the backbone network further includes the following steps:

[0088] Step S300: Input the obtained downsampled tensor into the multi-scale feature fusion module and the detection module, and output the detection result of the multi-view image.

[0089] In this embodiment, the outputs of the second feature extraction unit (stage 2), the third feature extraction unit (stage 3), and the fourth feature extraction unit (stage 4) of the backbone network model are collected as the input of the multi-scale feature fusion module and the following downstream task corresponding network to complete the conversion process from multi-view images to feature tensors containing high-order semantic information.

[0090] Specifically, in an implementation manner of this embodiment, step S300 includes the following steps:

[0091] Step S301: Use the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor as the inputs to the multi-scale feature fusion module and the corresponding network for the downstream task, and complete the conversion process from the multi-view image to the feature tensor containing high-order semantic information.

[0092] In this embodiment, the image backbone network is the most upstream of the entire perception model and is used to extract multi-scale features from the original image input. The second feature extraction unit (stage 2), the third feature extraction unit (stage 3), and the fourth feature extraction unit (stage 4) output representations of high-order features corresponding to different resolutions. After collecting these multi-scale features, they will also pass through a Feature Pyramid Network (FPN). Through the Feature Pyramid Network (FPN), these multi-scale features are fused into a new feature. The size of this feature is the same as the feature output by the second feature extraction unit (stage 2), but it contains multi-scale information.

[0093] It can be understood that the input image in this embodiment is a panoramic camera image in the shape of N*3*H*W. N represents the batch size, and H and W represent the height and width of the image. After the feature extraction by the first feature extraction unit (stage1), the shape of the output feature is N*C / 2*H / 4*W / 4. After the feature extraction by the second feature extraction unit (stage 2), the shape of the output feature is N*C*H / 8*W / 8; after the feature extraction by the third feature extraction unit (stage 3), the shape of the output feature is N*C*H / 16*W / 16; after the feature extraction by the fourth feature extraction unit (stage 4), the shape of the output feature is N*C*H / 32*W / 32.

[0094] Taking the output feature of the fourth feature extraction unit (stage 4) as an example, the main change in the output feature compared to the original image input (N*3*H*W) is that each point at the H*W position in the original image input is represented by three channels of red, green, and blue for a color, while now each pixel point at the H / 32*W / 32 positions is represented by C numbers. It is no longer the color representation corresponding to the three channels of red, green, and blue, but a point in a high-dimensional space with discriminative significance. This point is beneficial for discriminating where there are cars / pedestrians / obstacles in the input image, and what the shapes and categories of the cars / pedestrians / obstacles are.

[0095] Through the operations of the second feature extraction unit (stage 2), the third feature extraction unit (stage 3), and the fourth feature extraction unit (stage 4), the simple pixel representation of the input image can be gradually transformed into a feature representation that the network can understand. The features extracted by different feature extraction units (stages) contain different feature information. For example, when the network perceives the position and category of a vehicle, the discriminant information may be mainly contained in the features of the second feature extraction unit (stage 2); when the network perceives the position of a pedestrian, the corresponding discriminant information is mainly contained in the features of the second feature extraction unit (stage 2).

[0096] In this embodiment, for the final feature output of the backbone network, macroscopically speaking, given the same image input, the backbone network in this embodiment can efficiently and quickly extract features at three scales, and the discriminant information contained in these three scales of features is more abundant. The corresponding two characteristics are: fast inference speed and good performance during deployment.

[0097] Microscopically speaking, the internal unit design of the backbone network in this embodiment only includes conventional 1*1 convolutions, 1*k / k*1 convolutions, and k*k convolutions, which realizes friendly support for quantization; it adopts a decoupled representation method for spatial features and channel features, which is more conducive to feature expression and propagation in the network, so as to meet the requirements of being fast and good during deployment.

[0098] This embodiment achieves the following technical effects through the above technical solutions:

[0099] In this embodiment, by obtaining multi-view images and performing stitching processing on the multi-view images, a stitched tensor can be obtained; and the stitched tensor is input into the backbone network. Through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network, downsampling and feature extraction operations with different multiples are performed to obtain a corresponding downsampled feature tensor with high-order semantic information; finally, the obtained downsampled feature tensor with high-order semantic information is output to complete the conversion process from multi-view images to feature tensors with high-order semantic information; the embodiment of the present invention extracts features through an optimized backbone network structure, improving the processing efficiency of high-resolution images.

[0100] Exemplary device

[0101] Based on the above embodiments, the present invention further provides an image data processing device based on a backbone network, including:

[0102] A stitching processing module, configured to obtain multi-view images and perform stitching processing on the multi-view images to obtain a stitched tensor;

[0103] The backbone network module is used to input the spliced tensor into the backbone network. Through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network, different-fold downsampling and feature extraction operations are performed to obtain corresponding downsampled feature tensors with high-order semantic information.

[0104] The output module is used to output the obtained downsampled feature tensors with high-order semantic information, completing the conversion process from multi-view images to feature tensors containing high-order semantic information.

[0105] Through the above technical solutions, this embodiment achieves the following technical effects:

[0106] In this embodiment, by acquiring multi-view images and performing splicing processing on the multi-view images, a spliced tensor can be obtained; and by inputting the spliced tensor into the backbone network and performing different-fold downsampling and feature extraction operations through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network, corresponding downsampled feature tensors with high-order semantic information can be obtained; finally, the obtained downsampled feature tensors with high-order semantic information are output, completing the conversion process from multi-view images to feature tensors containing high-order semantic information. In the embodiment of the present invention, features are extracted through an optimized backbone network structure, improving the processing efficiency of high-resolution image processing.

[0107] Based on the above embodiment, the present invention also provides a terminal, and its principle block diagram can be as Figure 3 shown.

[0108] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected through a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.

[0109] When the computer program is executed by the processor, it is used to implement the operations of the image data processing method based on the backbone network.

[0110] Those skilled in the art can understand that Figure 3 the principle block diagram shown in

[0111] In one embodiment, a terminal is provided, which includes: a processor and a memory. The memory stores an image data processing program based on a backbone network. When the image data processing program based on the backbone network is executed by the processor, it is used to implement the operations of the above-mentioned image data processing method based on the backbone network.

[0112] In one embodiment, a storage medium is provided, which stores an image data processing program based on a backbone network. When the image data processing program based on the backbone network is executed by the processor, it is used to implement the operations of the above-mentioned image data processing method based on the backbone network.

[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided by the present invention can include non-volatile and volatile memories.

[0114] In summary, the present invention provides an image data processing method and device based on a backbone network. The method includes: acquiring multi-view images, and performing stitching processing on the multi-view images to obtain a stitched tensor; inputting the stitched tensor into the backbone network, and through multiple downsampling layers and spatial channel feature extraction layers of the backbone network, performing downsampling and feature extraction operations of different multiples to obtain corresponding downsampled feature tensors with high-order semantic information; outputting the obtained downsampled feature tensors with high-order semantic information to complete the conversion process from multi-view images to feature tensors containing high-order semantic information. The present invention extracts features through an optimized backbone network structure, improving the image processing efficiency of high resolution.

[0115] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. An image data processing method based on a backbone network, characterized in that Including: Obtain multi-view images, and perform stitching processing on the multi-view images to obtain a stitched tensor; Input the stitched tensor into a backbone network, and perform downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network to obtain corresponding downsampled feature tensors with high-order semantic information; Output the obtained downsampled feature tensors with high-order semantic information to complete the conversion process from multi-view images to feature tensors containing high-order semantic information.

2. The image data processing method based on a backbone network according to claim 1, wherein The obtaining of multi-view images and the performing of stitching processing on the multi-view images to obtain a stitched tensor includes: Obtain the multi-view images from a surround-view camera, and stitch the multi-view images into a tensor with a shape of N, C, H, W to obtain the stitched tensor; where N is the number of batches * the number of cameras, C is the number of channels of the input image data, and H and W respectively represent the height and width of the image.

3. The method for processing image data based on a backbone network according to claim 1, wherein The inputting of the stitched tensor into a backbone network and the performing of downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial-channel feature extraction layers of the backbone network includes: Perform downsampling on the stitched tensor through the first feature extraction unit of the backbone network to obtain a first downsampled tensor; Perform downsampling and feature enhancement processing on the first downsampled tensor through the second feature extraction unit of the backbone network to obtain an enhanced second downsampled tensor; Perform downsampling and feature enhancement processing on the enhanced second downsampled tensor through the third feature extraction unit of the backbone network to obtain an enhanced third downsampled tensor; Perform downsampling and feature enhancement processing on the enhanced third downsampled tensor through the fourth feature extraction unit of the backbone network to obtain an enhanced fourth downsampled tensor; Among them, the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor are all feature tensors with high-order semantic information.

4. The method for processing image data based on a backbone network according to claim 3, wherein The performing of downsampling and feature enhancement processing on the first downsampled tensor through the second feature extraction unit of the backbone network includes: Input the first downsampled tensor into the second feature extraction unit of the backbone network, and obtain a second downsampled tensor through the downsampling convolutional layer of the second feature extraction unit; Pass the second downsampled tensor through a spatial feature extraction unit, add the extracted multiple features, and then add the result to the features of the shortcut connection to obtain the features output by the spatial feature extraction unit; Pass the features output by the spatial feature extraction unit through a channel feature extraction unit, map the feature channels to the original S times dimension, and compress the feature channels back to C to obtain the enhanced second downsampled tensor.

5. The method for processing image data based on a backbone network according to claim 3, wherein The performing of downsampling and feature enhancement processing on the enhanced second downsampled tensor through the third feature extraction unit of the backbone network includes: Input the enhanced second downsampled tensor into the third feature extraction unit of the backbone network, and perform downsampling and feature enhancement processing according to the processing flow of the second feature extraction unit to obtain the enhanced third downsampled tensor.

6. The method for processing image data based on a backbone network according to claim 3, wherein Performing downsampling and feature enhancement processing on the enhanced third downsampled tensor through the fourth feature extraction unit of the backbone network, including: Inputting the enhanced third downsampled tensor into the fourth feature extraction unit of the backbone network, and performing downsampling and feature enhancement processing according to the processing flow of the second feature extraction unit to obtain the enhanced fourth downsampled tensor.

7. The method for processing image data based on a backbone network according to claim 3, wherein The output of the downsampled feature tensor with high-order semantic information includes: Using the enhanced second downsampled tensor, the enhanced third downsampled tensor, and the enhanced fourth downsampled tensor as inputs to the multi-scale feature fusion module and the corresponding network for downstream tasks, and completing the conversion process from multi-view images to feature tensors containing high-order semantic information.

8. An image data processing device based on a backbone network, characterized in that Including: A stitching processing module for acquiring multi-view images and performing stitching processing on the multi-view images to obtain a stitched tensor; A backbone network module for inputting the stitched tensor into the backbone network, and performing downsampling and feature extraction operations with different multiples through multiple downsampling layers and spatial channel feature extraction layers of the backbone network to obtain corresponding downsampled feature tensors with high-order semantic information; An output module for outputting the obtained downsampled feature tensor with high-order semantic information and completing the conversion process from multi-view images to feature tensors containing high-order semantic information.

9. A terminal, characterized in that, Including: A processor and a memory, where the memory stores an image data processing program based on the backbone network, and when the image data processing program based on the backbone network is executed by the processor, it is used to implement the operations of the image data processing method based on the backbone network according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image data processing program based on the backbone network, and when the image data processing program based on the backbone network is executed by a processor, it is used to implement the operations of the image data processing method based on the backbone network according to any one of claims 1-7.