Image Processing Method, Apparatus, Device, and Storage Medium
Through the combination of feature memory network and online backpropagation network, the existing multi-task model has solved the problem of low scene compatibility and low efficiency in image processing, and achieved more efficient and higher precision image processing.
Patent Information
- Application Number
- CN202411920700.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-25
AI Technical Summary
When processing pictures, the existing multitasking models have low scene compatibility and low image processing efficiency.
By fusing the key features of the original picture and historical picture into the feature memory network for feature fusion, Gaussian parameters, point cloud information and depth estimation information are obtained, and iteratively updated through the online backpropagation network to improve the accuracy of Gaussian parameters.
This improves the efficiency and accuracy of image processing and enhances the compatibility of the model for various task scenarios.
Smart Images

Figure CN119360176B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] Task Awareness refers to the ability of a system to dynamically adjust its operation strategies and parameters according to current task requirements and environmental changes during task execution, so as to adapt to different task scenarios and goals. This ability is particularly important in fields such as autonomous driving, robot control, and object detection.
[0003] In scenarios such as robots, autonomous driving, and drones, there are perception tasks such as real-time new view synthesis, monocular estimation, mapping, localization, object classification, recognition, segmentation, tracking, and feature point extraction.
[0004] Currently, multi-task models can perform limited perception tasks, have low scene compatibility, and low image processing efficiency. Summary of the Invention
[0005] The present disclosure provides an image processing method, apparatus, device, and storage medium to at least solve the problems of low scene compatibility and low image processing efficiency of existing multi-task models.
[0006] The technical solution of the present disclosure is as follows:
[0007] An embodiment of the present disclosure provides an image processing method, including:
[0008] Inputting an original image into an encoder for feature encoding to obtain key features of the original image;
[0009] Inputting the key features of the original image and the key features of historical images into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information;
[0010] Inputting the Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network for back-iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information;
[0011] Inputting the updated Gaussian parameters into predictors corresponding to each task type to obtain prediction results, where the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters.
[0012] Optionally, the feature fusion information further includes: object recognition information, object segmentation information, object tracking information, and feature point extraction information; the task types include: segmentation, detection, recognition, monocular estimation, mapping, localization, new view synthesis, and feature point extraction.
[0013] Optionally, the encoder includes: a patch embedding layer, a position encoding layer, and a plurality of encoder layers, which are serially connected between the plurality of encoder layers; the step of inputting the original picture into the encoder for feature encoding to obtain the key features of the original picture includes:
[0014] Inside the encoder, input the original picture into the patch embedding layer to obtain a patch embedding;
[0015] Input the patch embedding vector into the position encoding layer to obtain the patch embedding after position encoding;
[0016] Input the patch embedding after position encoding into the encoder layer to obtain the key features of the original picture.
[0017] Optionally, each encoder layer includes: a multi-head self-attention mechanism layer, a first residual connection and layer normalization layer, a feed-forward neural network layer, and a second residual connection and layer normalization layer; the step of inputting the patch embedding after position encoding into the encoder layer to obtain the key features of the original picture includes:
[0018] For a target encoder layer, input the patch embedding after position encoding into the multi-head self-attention mechanism layer of the target encoder layer to obtain a self-attention output feature; wherein, the target encoder layer is any one of the plurality of encoder layers;
[0019] Input the self-attention output feature and the patch embedding after position encoding into the first residual connection and layer normalization layer of the target encoder layer to obtain a first normalized feature;
[0020] Input the first normalized feature into the feed-forward neural network layer of the target encoder layer to obtain a feed-forward network output feature;
[0021] Input the feed-forward network output feature and the first normalized feature into the second residual connection and layer normalization layer of the target encoder layer to obtain a second normalized feature.
[0022] Optionally, the feature memory network includes: a plurality of serially connected basic block modules and a head module, and the step of inputting the key features of the original picture and the key features of the historical picture into the feature memory network for feature fusion to obtain feature fusion information includes:
[0023] Input the key features of the original picture and the key features of the historical picture into the plurality of basic block modules to obtain an initial fusion feature;
[0024] Input the initial fusion feature into the head module to obtain the feature fusion information.
[0025] Optionally, inputting the Gaussian parameters, the point cloud information, and the depth estimation information into an online backpropagation network for back-iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information includes:
[0026] Inputting the Gaussian parameters, the point cloud information, and the depth estimation information into an online backpropagation network to render the original picture, and obtaining a rendered picture corresponding to the original picture;
[0027] Determining a loss function according to the rendered picture and the original picture;
[0028] Performing backpropagation according to the loss function to update the Gaussian parameters, the point cloud information, and the depth estimation information, and obtaining updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
[0029] An embodiment of the present disclosure further provides an image processing apparatus, including:
[0030] An encoding module, configured to input an original picture into an encoder for feature encoding to obtain key features of the original picture;
[0031] A fusion module, configured to input the key features of the original picture and the key features of a historical picture into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information;
[0032] A rendering module, configured to input the Gaussian parameters, the point cloud information, and the depth estimation information into an online backpropagation network for iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information;
[0033] A prediction module, configured to input the updated Gaussian parameters into predictors corresponding to each task type to obtain prediction results, where the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters.
[0034] An embodiment of the present disclosure further provides an electronic device, including:
[0035] A processor;
[0036] A memory for storing executable instructions of the processor;
[0037] Wherein, the processor is configured to execute the instructions to implement the steps in the above method.
[0038] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each step in the above-mentioned method is implemented.
[0039] The embodiments of the present disclosure also provide a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, each step in the above-mentioned method is implemented.
[0040] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0041] In some embodiments of the present disclosure, the original picture is input into an encoder for feature encoding to obtain the key features of the original picture; the key features of the original picture and the key features of the historical picture are input into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information, and the present disclosure predicts Gaussian parameters through a network; inputting high-quality Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network can reduce the number of backpropagation times, improve the picture processing efficiency, and use the backpropagation method to obtain updated Gaussian parameters with higher accuracy, which can be applicable to various task models and improve the compatibility of the model for various task scenarios.
[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0044] Figure 1 It is a schematic flowchart of a picture processing method provided for an exemplary embodiment of the present disclosure;
[0045] Figure 2 It is a schematic structural diagram of a picture processing system provided for an exemplary embodiment of the present disclosure;
[0046] Figure 3 It is a schematic structural diagram of an online backpropagation network provided for an exemplary embodiment of the present disclosure;
[0047] Figure 4 It is a schematic structural diagram of a picture processing device provided for an exemplary embodiment of the present disclosure;
[0048] Figure 5 It is a schematic structural diagram of an electronic device provided for an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0051] It should be noted that the user information involved in the present disclosure includes but is not limited to: user device information and user personal information; the collection, storage, use, processing, transmission, provision, and disclosure, etc. of the user information in the present disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0052] The following describes the solutions of existing perception tasks.
[0053] First, a solution for map building and extraction of positioning feature points. The input is multiple pictures. First, the feature points of each picture are obtained, and then through feature point matching, the depth of the feature points and the pose of each picture are calculated to complete map building and positioning. This solution cannot be applied to multiple scenarios and has low universality.
[0054] Second, a monocular estimation solution. The input is a single picture. The network structure includes an encoder and a decoder, and the output is a depth map. Among them, the encoder is a feature extraction module, and the decoder is a decoding module for predicting depth according to features. This solution cannot be applied to multiple scenarios and has low universality.
[0055] Third, a new view synthesis solution. The input is multiple pictures and an initial sparse point cloud. Then, the Gaussian parameters are initialized to a preset value. For example, the opacity parameter is initialized to 0.1. Then, the parameters are updated through backpropagation. Each time backpropagation is performed, the parameters are updated once to make the parameters "a little more correct". Usually, 30,000 times of backpropagation iterations are required to obtain sufficiently correct parameters. Then, according to these parameters, through the Gaussian rendering engine, a new view picture can be rendered; the time for picture rendering is long, usually from 10 minutes to 1 hour.
[0056] In view of the above technical problems, in some embodiments of the present disclosure, the original image is input into an encoder for feature encoding to obtain the key features of the original image; the key features of the original image and the key features of the historical image are input into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information, and the present disclosure predicts the Gaussian parameters through a network; inputting the high-quality Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network can reduce the number of backpropagation times, improve the image processing efficiency, and use the backpropagation method to obtain updated Gaussian parameters with higher accuracy, which can be applicable to various task models and improve the compatibility of the model for various task scenarios.
[0057] The following will describe in detail the technical solutions provided by each embodiment of the present disclosure with reference to the accompanying drawings.
[0058] Figure 1 It is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present disclosure. As Figure 1 shown, the method includes:
[0059] S101: Input the original image into an encoder for feature encoding to obtain the key features of the original image;
[0060] S102: Input the key features of the original image and the key features of the historical image into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information;
[0061] S103: Input the Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network for backpropagation iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information;
[0062] S104: Input the updated Gaussian parameters into the predictor corresponding to each task type to obtain a prediction result, where the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters.
[0063] In this embodiment, the execution subject of the above method may be a terminal device or a server.
[0064] Among them, the terminal device includes, but is not limited to, a mobile station (MS), a mobile terminal, a mobile telephone, a handset, and portable equipment, etc. The terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or a "cellular" phone), a computer with wireless communication functions, etc. The terminal device can also be a computer with wireless transceiver functions, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. And the operating systems installed on the terminal device include, but are not limited to: IOS, Android, windows, linux, Mac OS, etc. In different networks, the terminal can be called different names. For example: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop computer, cordless phone, wireless local loop station, TV, etc. For the convenience of description, it is simply referred to as the terminal device in this embodiment.
[0065] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. Among them, the composition of the server mainly includes a processor, a hard disk, a memory, a system bus, etc., and a general computer architecture type.
[0066] In this embodiment, in some embodiments of the present disclosure, the original image is input into an encoder for feature encoding to obtain the key features of the original image; the key features of the original image and the key features of the historical image are input into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information, and the present disclosure predicts the Gaussian parameters through a network; inputting the high-quality Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network can reduce the number of backpropagation times, improve the image processing efficiency, and use the backpropagation method to obtain a higher-precision updated Gaussian parameter, which can be applicable to multiple task models and improve the compatibility of the model for various task scenarios.
[0067] In this embodiment, the original image is input into an encoder for feature encoding to obtain the key features of the original image; the key features of the original image and the key features of the historical image are input into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information, and the present disclosure predicts the Gaussian parameters through a network; inputting the high-quality Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network to render the original image to obtain the rendered image corresponding to the original image can reduce the number of backpropagation times, improve the image processing efficiency, and use the backpropagation method to update the parameters, improving the compatibility of the model for various scenarios.
[0068] It should be noted that the Gaussian parameters include: mean vector, covariance matrix, weight, amplitude, and scale parameter.
[0069] Among them, the mean vector (Mean Vector), denoted as μ, defines the central position of the Gaussian distribution. In three-dimensional space, the mean vector is usually expressed as μ = (μ_x, μ_y, μ_z), where μ_x, μ_y, and μ_z are the central coordinates of the Gaussian distribution in the x, y, and z directions respectively. The covariance matrix (Covariance Matrix), denoted as Σ, defines the shape and direction of the Gaussian distribution. In three-dimensional space, the covariance matrix is a 3x3 symmetric positive definite matrix, expressed as:
[0070] Σ = [σ_xx σ_xy σ_xz][σ_yx σ_yy σ_yz][σ_zx σ_zy σ_zz]
[0071] Among them, σ_xx, σ_yy, and σ_zz are the variances in the x, y, and z directions, while σ_xy, σ_xz, etc. are the covariances in different directions. The covariance matrix determines the expansion and rotation of the Gaussian distribution.
[0072] Among them, the weight, denoted as w, is a parameter that determines the influence or importance of the Gaussian distribution. During the rendering process, different Gaussian distributions may have different weights to represent their contribution degrees to the final image.
[0073] Among them, the amplitude, denoted as A; the amplitude parameter determines the height or intensity of the Gaussian distribution. It is a scalar value that controls the maximum value of the Gaussian distribution at its center.
[0074] Among them, the scale parameters, denoted as s_x, s_y, s_z; these parameters control the degree of expansion of the Gaussian distribution in each direction and can usually be derived from the covariance matrix.
[0075] In practical applications, these parameters jointly determine the shape, position, and influence of the Gaussian distribution, thus forming a complex distribution in three-dimensional space for representing and processing three-dimensional data points.
[0076] In this embodiment, the encoder includes: a patch embedding layer, a position encoding layer, and multiple encoder layers, which are serially connected between the multiple encoder layers.
[0077] In some embodiments of the present disclosure, the original picture is input into the encoder for feature encoding to obtain the key features of the original picture. Inside the encoder, the original picture is input into the patch embedding layer to obtain patch embeddings; the patch embedding vectors are input into the position encoding layer to obtain the patch embeddings after position encoding; the patch embeddings after position encoding are input into the encoder layer to obtain the key features of the original picture. It should be noted that the encoder in this embodiment is used to extract features from the original picture; the type of the encoder in the present disclosure is not limited and can be adjusted according to actual situations.
[0078] For example, the encoder is a vit encoder; the original picture frame t is input into the patch embedding layer, the picture is divided into patches of a fixed size, each patch is flattened into a one-dimensional vector, and the flattened patches are mapped to the embedding space through linear projection to obtain patch embeddings; among them, the shape of the original picture is H×W×C; the shape of the patch embeddings is N×D, where N is the number of patches and D is the embedding dimension. frame t represents the picture obtained at time t, and frame t-1 represents the picture obtained at time t-1, that is, the picture obtained before time t. The patch embedding vectors are input into the position encoding layer to add learnable position encoding to each patch embedding to retain position information, obtaining the patch embeddings after position encoding, where the shape of the patch embeddings after position encoding is N×D. The patch embeddings after position encoding are input into the encoder layer, and multiple encoder layers are stacked in sequence to obtain the key features of the original picture, and the shape of the key features of the original picture is N×D.
[0079] Each encoder layer includes: a multi-head self-attention mechanism layer, a first residual connection and layer normalization layer, a feed-forward neural network layer, and a second residual connection and layer normalization layer.
[0080] In some embodiments of the present disclosure, the tile embeddings after position encoding are input into the encoder layer to obtain the key features of the original picture. One implementable way is that for the target encoder layer, the tile embeddings after position encoding are input into the multi-head self-attention mechanism layer of the target encoder layer to obtain the self-attention output features; wherein, the target encoder layer is any one of the multiple encoder layers; the self-attention output features and the tile embeddings after position encoding are input into the first residual connection and layer normalization layer of the target encoder layer to obtain the first normalized features; the first normalized features are input into the feed-forward neural network layer of the target encoder layer to obtain the feed-forward network output features; the feed-forward network output features and the first normalized features are input into the second residual connection and layer normalization layer of the target encoder layer to obtain the second normalized features.
[0081] For example, the tile embeddings after position encoding are input into the multi-head self-attention mechanism layer to generate query (Q), key (K), and value (V) matrices; the attention scores are calculated and weighted and summed; multiple attention heads are calculated in parallel, and the results are concatenated and linearly transformed to obtain the self-attention output features; the shape of the self-attention output features is N×D. The self-attention output features and the tile embeddings after position encoding are input into the first residual connection and layer normalization layer of the target encoder layer, the self-attention output features and the tile embeddings after position encoding are added together, and the application layer performs a normalization operation to obtain the first normalized features; the shape of the first normalized features is N×D. The first normalized features are input into the feed-forward neural network layer of the target encoder layer, in two fully connected networks, an activation function is used in the middle to obtain the feed-forward network output features; the shape of the feed-forward network output features is N×D. The feed-forward network output features and the first normalized features are input into the second residual connection and layer normalization layer of the target encoder layer, the feed-forward network output features and the first normalized features are added together, and layer normalization processing is applied to obtain the second normalized features. Wherein, when the target encoder layer is the last encoder layer in the encoder layer, the second normalized features are the key features of the original picture.
[0082] In this embodiment, the feature memory network includes: a plurality of serially connected basic block modules and a head module. Each basic block module is composed of a traditional self-attention and cross-attention structure. The head module is composed of multiple layers of 3*3 convolutions. The input of the convolution kernel of the first layer is n and the output is 4*n. The input of the convolution kernel of the first layer is 4*n and the output is 7 + 3 + m + 1 + 20. 7 represents Gaussian parameters, 3 represents point cloud coordinates, depth estimation can be directly calculated from the point cloud coordinates, m represents the number of object segmentation categories, 1 represents the id of object tracking, and 20 represents the dimension of feature points.
[0083] In some embodiments of the present disclosure, the key features of the original image and the key features of the historical image are input into the feature memory network for feature fusion to obtain feature fusion information. One achievable way is to input the key features of the original image and the key features of the historical image into a plurality of basic block modules to obtain initial fusion features; the initial fusion features are input into the head module to obtain feature fusion information. It should be noted that frame t and feature memory t-1 features are input to the FM module together; when t is 1, frame 1 is regarded as feature memory t-1. Among them, feature memory t-1 represents the set of key features of historical images. Through the feature memory network, feature memory t can be continuously updated, and it helps feature t fuse historical features to improve the representation ability of features.
[0084] It should be noted that the feature fusion information includes but is not limited to: Gaussian parameters, point cloud information, depth estimation information, object recognition information, object segmentation information, object tracking information, and feature point extraction information.
[0085] Figure 2 It is a schematic structural diagram of an image processing system provided for an exemplary embodiment of the present disclosure. As Figure 2 shown, the original image frame t is input into the encoder to obtain the key features of the original image feature t; the key features of the original image feature t and the key features of the historical images feature t-1... feature t-n are input into the feature memory network for feature fusion to obtain Gaussian parameters, point cloud information, depth estimation information, object recognition information, object segmentation information, object tracking information, and feature point extraction information.
[0086] In some embodiments of the present disclosure, Gaussian parameters, point cloud information, and depth estimation information are input into an online backpropagation network for back-iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information. One achievable way is to input Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network to render the original image and obtain a rendered image corresponding to the original image; determine a loss function based on the rendered image and the original image; perform backpropagation according to the loss function to update the Gaussian parameters, point cloud information, and depth estimation information, and obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
[0087] For example, Figure 3 is a schematic structural diagram of an online backpropagation network provided by an exemplary embodiment of the present disclosure. As Figure 3 shown, Gaussian parameters, point cloud information, and depth estimation information are input into an online backpropagation network. According to the predicted Gaussian parameters, a rendered image corresponding to the original image is rendered. According to the rendered image and the original image, a loss function is determined. Among them, the loss function can be an L2 loss rendering loss function. Perform 10 times of backpropagation of the L2 loss for each Gaussian parameter to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
[0088] In some implementations of the present disclosure, the updated Gaussian parameters are input into predictors corresponding to each task type to obtain prediction results. Among them, the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters. The task types include: segmentation, detection, recognition, monocular estimation, mapping, localization, novel view synthesis, and feature point extraction. The present disclosure uses the backpropagation method to obtain updated Gaussian parameters with higher accuracy, which can be applied to multiple task models and improve the compatibility of the model for various task scenarios.
[0089] Figure 4 is a schematic structural diagram of an image processing device 40 provided by an exemplary embodiment of the present disclosure. As Figure 4 shown, the image processing device 40 includes: an encoding module 41, a fusion module 42, a rendering module 43, and a prediction module 44.
[0090] Among them, the encoding module 41 is used to input the original image into an encoder for feature encoding to obtain key features of the original image;
[0091] The fusion module 42 is used to input the key features of the original image and the key features of the historical image into a feature memory network for feature fusion to obtain feature fusion information. Among them, the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information;
[0092] A rendering module 43, configured to input Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network for iteration, so as to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information;
[0093] A prediction module 44, configured to input the updated Gaussian parameters into predictors corresponding to each task type to obtain prediction results, wherein the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters.
[0094] Optionally, the feature fusion information further includes: object recognition information, object segmentation information, object tracking information, and feature point extraction information; the task types include: segmentation, detection, recognition, monocular estimation, mapping, localization, novel view synthesis, and feature point extraction.
[0095] Optionally, the encoder includes: a patch embedding layer, a positional encoding layer, and a plurality of encoder layers, which are serially connected; when the encoding module 41 inputs the original picture into the encoder for feature encoding to obtain the key features of the original picture, it is configured to:
[0096] Inside the encoder, input the original picture into the patch embedding layer to obtain a patch embedding;
[0097] Input the patch embedding vector into the positional encoding layer to obtain the patch embedding after positional encoding;
[0098] Input the patch embedding after positional encoding into the encoder layer to obtain the key features of the original picture.
[0099] Optionally, each encoder layer includes: a multi-head self-attention mechanism layer, a first residual connection and layer normalization layer, a feed-forward neural network layer, and a second residual connection and layer normalization layer; when the encoding module 41 inputs the patch embedding after positional encoding into the encoder layer to obtain the key features of the original picture, it is configured to:
[0100] For a target encoder layer, input the patch embedding after positional encoding into the multi-head self-attention mechanism layer of the target encoder layer to obtain self-attention output features; wherein the target encoder layer is any one of the plurality of encoder layers;
[0101] Input the self-attention output features and the patch embedding after positional encoding into the first residual connection and layer normalization layer of the target encoder layer to obtain a first normalized feature;
[0102] Input the first normalized feature into the feed-forward neural network layer of the target encoder layer to obtain feed-forward network output features;
[0103] Input the feed-forward network output features and the first normalized feature into the second residual connection and layer normalization layer of the target encoder layer to obtain a second normalized feature.
[0104] Optionally, the feature memory network includes: a plurality of base block modules connected in series and a head module. When the fusion module 42 inputs the key features of the original picture and the key features of the historical picture into the feature memory network for feature fusion to obtain feature fusion information, it is used for:
[0105] Input the key features of the original picture and the key features of the historical picture into a plurality of base block modules to obtain initial fusion features;
[0106] Input the initial fusion features into the head module to obtain feature fusion information.
[0107] Optionally, when the rendering module 43 inputs the Gaussian parameters, point cloud information, and depth estimation information into the online backpropagation network for backpropagation iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information, it is used for:
[0108] Input the Gaussian parameters, point cloud information, and depth estimation information into the online backpropagation network to render the original picture and obtain the rendered picture corresponding to the original picture;
[0109] Determine the loss function according to the rendered picture and the original picture;
[0110] Perform backpropagation according to the loss function to perform backpropagation iteration on the Gaussian parameters, point cloud information, and depth estimation information to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
[0111] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0112] Figure 5 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. As Figure 5 shown, the electronic device includes: a memory 51 and a processor 52. In addition, the electronic device further includes a power supply component 53 and a communication component 54.
[0113] The memory 51 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method operating on the electronic device.
[0114] A memory 51, which can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0115] A communication component 54 for data transmission with other devices.
[0116] A processor 52 that can execute computer instructions stored in the memory 51 for: inputting an original picture into an encoder for feature encoding to obtain key features of the original picture; inputting the key features of the original picture and the key features of historical pictures into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information; performing backward iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information; inputting the updated Gaussian parameters into predictors corresponding to each task type to obtain prediction results, where the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters.
[0117] Optionally, the feature fusion information further includes: object recognition information, object segmentation information, object tracking information, and feature point extraction information; the task types include: segmentation, detection, recognition, monocular estimation, mapping, localization, novel view synthesis, and feature point extraction.
[0118] Optionally, the encoder includes: a patch embedding layer, a positional encoding layer, and a plurality of encoder layers, which are serially connected between the plurality of encoder layers; when the processor 52 inputs the original picture into the encoder for feature encoding to obtain the key features of the original picture, it is used for:
[0119] Inside the encoder, input the original picture into the patch embedding layer to obtain a patch embedding;
[0120] Input the patch embedding vector into the positional encoding layer to obtain the patch embedding after positional encoding;
[0121] Input the patch embedding after positional encoding into the encoder layer to obtain the key features of the original picture.
[0122] Optionally, each encoder layer includes: a multi-head self-attention mechanism layer, a first residual connection and layer normalization layer, a feed-forward neural network layer, and a second residual connection and layer normalization layer; when the processor 52 inputs the patch embedding after positional encoding into the encoder layer to obtain the key features of the original picture, it is used for:
[0123] For a target encoder layer, the position-encoded patch embeddings are input into the multi-head self-attention mechanism layer of the target encoder layer to obtain self-attention output features; wherein, the target encoder layer is any one of multiple encoder layers.
[0124] The self-attention output features and the position-encoded patch embeddings are input into the first residual connection and layer normalization layer of the target encoder layer to obtain first normalized features.
[0125] The first normalized features are input into the feed-forward neural network layer of the target encoder layer to obtain feed-forward network output features.
[0126] The feed-forward network output features and the first normalized features are input into the second residual connection and layer normalization layer of the target encoder layer to obtain second normalized features.
[0127] Optionally, the feature memory network includes: a plurality of base block modules and a head module connected in series. When the processor 52 inputs the key features of the original picture and the key features of the historical picture into the feature memory network for feature fusion to obtain feature fusion information, it is used for:
[0128] Input the key features of the original picture and the key features of the historical picture into a plurality of base block modules to obtain initial fusion features.
[0129] Input the initial fusion features into the head module to obtain feature fusion information.
[0130] Optionally, when the processor 52 inputs the Gaussian parameters, point cloud information, and depth estimation information into the online backpropagation network for backpropagation iteration to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information, it is used for:
[0131] Input the Gaussian parameters, point cloud information, and depth estimation information into the online backpropagation network to render the original picture to obtain a rendered picture corresponding to the original picture.
[0132] Determine a loss function according to the rendered picture and the original picture.
[0133] Perform backpropagation according to the loss function to perform backpropagation iteration on the Gaussian parameters, point cloud information, and depth estimation information to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
[0134] Correspondingly, an embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores the computer program and the computer program is executed by one or more processors, it causes the one or more processors to execute Figure 1 each step in the method embodiment.
[0135] Accordingly, an embodiment of the present disclosure also provides a computer program product, which includes computer programs / instructions that are executed by a processor Figure 1 for each step in the method embodiments
[0136] The above-mentioned Figure 5 The communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology and other technologies
[0137] The above-mentioned Figure 5 The power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located
[0138] The above-mentioned electronic device further includes a display screen and an audio component
[0139] The display screen includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation
[0140] The audio component is configured to output and / or input audio signals. For example, the audio component includes a Microphone (MIC), which is configured to receive external audio signals when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals
[0141] In the above method, apparatus, device, storage medium, and computer program product embodiments of the present disclosure, the original image is input into an encoder for feature encoding to obtain the key features of the original image; the key features of the original image and the key features of the historical image are input into a feature memory network for feature fusion to obtain feature fusion information, where the feature fusion information includes: Gaussian parameters, point cloud information, and depth estimation information, and the present disclosure predicts the Gaussian parameters through a network; inputting the high-quality Gaussian parameters, point cloud information, and depth estimation information into an online backpropagation network can reduce the number of backpropagation times, improve the image processing efficiency, and use the backpropagation method to obtain updated Gaussian parameters with higher accuracy, which can be applied to multiple task models and improve the compatibility of the model for various task scenarios.
[0142] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, system, or computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0143] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0144] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0146] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0147] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0148] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0149] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0150] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing an image, characterized in that: include: Input the original image into the encoder for feature encoding to obtain the key features of the original image; Inputting the key features of the original image and the key features of the historical image into a feature memory network for feature fusion to obtain feature fusion information, wherein the feature fusion information includes: Gaussian parameters, point cloud information and depth estimation information; Inputting the Gaussian parameters, the point cloud information and the depth estimation information into an online back propagation network for reverse iteration to obtain updated Gaussian parameters, updated point cloud information and updated depth estimation information; Inputting the updated Gaussian parameters into a predictor corresponding to each task type to obtain a prediction result, wherein the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters; The encoder comprises: a tile embedding layer, a position encoding layer and a plurality of encoder layers, wherein the plurality of encoder layers are serially connected; the original image is input into the encoder for feature encoding to obtain key features of the original image, including: Inside the encoder, the original image is input into the tile embedding layer, the image is divided into tiles of fixed size, each tile is flattened into a one-dimensional vector, and the flattened tiles are mapped to the embedding space by linear projection to obtain tile embedding; Inputting the tile embedding vector into the position encoding layer, adding a learnable position encoding to each tile embedding to retain the position information, and obtaining a position-encoded tile embedding; The position-encoded image block is embedded into the encoder layer, and multiple encoder layers are stacked sequentially to obtain the key features of the original image.
2. The method according to claim 1, characterized in that The feature fusion information also includes: object recognition information, object segmentation information, object tracking information and feature point extraction information; the task types include: segmentation, detection, recognition, monocular estimation, mapping, positioning, new perspective synthesis and feature point extraction.
3. The method according to claim 1, characterized in that Each of the encoder layers includes: a multi-head self-attention mechanism layer, a first residual connection and layer normalization layer, a feedforward neural network layer, and a second residual connection and layer normalization layer; the embedding of the position-encoded image blocks into the encoder layer to obtain the key features of the original image includes: For a target encoder layer, embedding the position-encoded image block into a multi-head self-attention mechanism layer input into the target encoder layer to obtain a self-attention output feature; wherein the target encoder layer is any one of the multiple encoder layers; Embedding the self-attention output feature and the position-encoded tile into the first residual connection and layer normalization layer of the target encoder layer to obtain a first normalized feature; Inputting the first normalized feature into the feedforward neural network layer of the target encoder layer to obtain a feedforward network output feature; The feedforward network output feature and the first normalized feature are input into the second residual connection and layer normalization layer of the target encoder layer to obtain a second normalized feature.
4. The method according to claim 1, characterized in that: The feature memory network includes: a plurality of serially connected basic block modules and head modules, and the key features of the original image and the key features of the historical image are input into the feature memory network for feature fusion to obtain feature fusion information, including: Inputting the original image key features and the historical image key features into a plurality of the basic block modules to obtain initial fusion features; The initial fusion features are input into the head module to obtain the feature fusion information.
5. The method according to claim 1, characterized in that The step of inputting the Gaussian parameters, the point cloud information and the depth estimation information into an online back propagation network for reverse iteration to obtain updated Gaussian parameters, updated point cloud information and updated depth estimation information comprises: Inputting the Gaussian parameters, the point cloud information, and the depth estimation information into an online back propagation network, rendering the original image, and obtaining a rendered image corresponding to the original image; Determining a loss function according to the rendered image and the original image; Back propagation is performed according to the loss function, and the Gaussian parameters, the point cloud information, and the depth estimation information are reversely iterated to obtain updated Gaussian parameters, updated point cloud information, and updated depth estimation information.
6. A picture processing device, characterized in that: include: The encoding module is used to input the original image into the encoder for feature encoding to obtain the key features of the original image; A fusion module, used for inputting the key features of the original image and the key features of the historical image into a feature memory network for feature fusion to obtain feature fusion information, wherein the feature fusion information includes: Gaussian parameters, point cloud information and depth estimation information; A rendering module, used for inputting Gaussian parameters, the point cloud information and the depth estimation information into an online back propagation network for iteration to obtain updated Gaussian parameters, updated point cloud information and updated depth estimation information; A prediction module, used for inputting the updated Gaussian parameters into a predictor corresponding to each task type to obtain a prediction result, wherein the accuracy of the updated Gaussian parameters is higher than that of the Gaussian parameters; The encoder comprises: a tile embedding layer, a position encoding layer and a plurality of encoder layers, wherein the plurality of encoder layers are serially connected; the original image is input into the encoder for feature encoding to obtain key features of the original image, including: Inside the encoder, the original image is input into the tile embedding layer, the image is divided into tiles of fixed size, each tile is flattened into a one-dimensional vector, and the flattened tiles are mapped to the embedding space by linear projection to obtain tile embedding; Inputting the tile embedding vector into the position encoding layer, adding a learnable position encoding to each tile embedding to retain the position information, and obtaining a position-encoded tile embedding; The position-encoded image block is embedded into the encoder layer, and multiple encoder layers are stacked sequentially to obtain the key features of the original image.
7. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement each step in the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method according to any one of claims 1 to 5 is implemented.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
End-to-end shipborne surveillance video dense description method and system based on dynamic feature memory
CN117746328A
3D Gaussian sputtering scene reconstruction method based on view dependence difference decoupling
CN118710792A