Method executed by electronic equipment, corresponding electronic equipment and storage medium
By extracting pixel motion embedding features and non-occlusion category label embedding features in image pairs using AI networks, the problem of insufficient scene flow estimation performance in the prior art is solved, and more accurate scene flow estimation is achieved.
Patent Information
- Application Number
- CN202311527401.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively improve performance when estimating scene flow, especially in scenarios containing complex three-dimensional structures and motion.
An artificial intelligence AI network is used to estimate the scene flow by acquiring pixel motion embedding features and non-occlusion category label embedding features in image pairs. The method includes obtaining the color image and depth image pair to be processed, and extracting motion embedding features and non-occlusion category label embedding features using an AI network to obtain a scene flow.
This method can obtain a more accurate scene flow of image pairs, improve the performance of scene flow estimation, and better meet actual needs.
Smart Images

Figure CN120014307A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing and artificial intelligence technology, and in particular to a method executed by an electronic device, a corresponding electronic device and a storage medium. Background Art
[0002] Scene flow represents the point-by-point motion between two frames, that is, how each pixel in the image changes between the previous and next frames. It involves estimating the 3D structure and 3D motion in a complex and changing scene, which can contain at least one or more of various static and dynamic objects, rigid and non-rigid objects, strong and weak textured areas, unoccluded areas and occluded areas. Scene flow has received increasing attention due to its importance in fields such as robotics, augmented reality and self-driving cars.
[0003] As a pixel-level task, scene flow estimation is beneficial to many scene understanding tasks, such as object detection, tracking, segmentation, and object pose estimation. How to improve the scene flow estimation performance is one of the key technical issues studied by technicians in this field. Summary of the invention
[0004] The purpose of the embodiment of the present disclosure is to optimize the estimation result of the scene flow. In order to achieve this purpose, the technical solution provided by the embodiment of the present disclosure is as follows:
[0005] According to one aspect of an embodiment of the present disclosure, a method performed by an electronic device is provided, the method including:
[0006] Acquire a pair of images to be processed, each image in the pair of images comprising a color image and a depth image corresponding to the color image;
[0007] Using an artificial intelligence AI network, the motion embedding features and non-occlusion category label embedding features of the pixels in the image pair are obtained, and based on the motion embedding features and non-occlusion category label embedding features (also referred to as geometric segmentation embedding features or other names), the scene flow corresponding to the image pair is obtained.
[0008] The image pair includes a first image and a second image, and the non-occluded category label embedding features corresponding to the image pair are used to characterize the category information corresponding to the pixel pair in the image pair, and the pixel pair includes any pixel in the first image and the pixel corresponding to the pixel in the second image.
[0009] Among them, the above-mentioned image pair can be a first image and a second image that are relatively continuous in the time domain, and the above-mentioned pixel pair can be interpreted as two corresponding pixels in the two images. The pixel corresponding to any pixel in the first image in the second image is also the position of any pixel point in the first image in the adjacent second image. It is very likely that relative motion has occurred between the pixels of the two adjacent images.
[0010] According to one aspect of an embodiment of the present disclosure, a method performed by an electronic device is provided, the method including:
[0011] Obtain an AI network to be trained and a training set, wherein the training samples in the training set include: a sample image pair, a scene flow truth value of the sample image pair, and a non-occluded category label mask truth value (also referred to as a geometric segmentation mask truth value, wherein each sample image in the sample image pair includes a color image and a corresponding depth image, and the mask truth value represents the true category corresponding to a pixel pair in the sample image pair, wherein the pixel pair includes any pixel in a first sample image in the sample image pair and a pixel corresponding to the pixel in a second sample image in the sample image pair;
[0012] Repeating the training operation on the AI network to be trained based on the training set until the training end condition is met to obtain a trained AI network, wherein the training operation includes:
[0013] For each of the sample image pairs, using an AI network, obtain motion embedding features and non-occlusion category label embedding features corresponding to the sample image pair, obtain the predicted scene flow corresponding to the sample image pair based on the motion embedding features and non-occlusion category label embedding features, and obtain the non-occlusion category label mask prediction value of the sample image pair based on the non-occlusion category label embedding features, wherein the non-occlusion category label embedding features characterize the category information corresponding to the pixel pair in the sample image pair;
[0014] Based on the true value of the scene flow and the predicted scene flow of each sample image pair, a first training loss is determined; based on the true value of the non-occluded category label mask and the predicted value of the non-occluded category label mask corresponding to each sample image pair, a second training loss is determined; based on the first training loss and the second training loss, a total training loss is determined; and based on the total training loss, the model parameters of the AI network are adjusted.
[0015] An embodiment of the present disclosure further provides an electronic device, which includes at least one processor, and the at least one processor is configured to execute the method provided in any embodiment of the present disclosure.
[0016] An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in any embodiment of the present disclosure is executed.
[0017] Compared with the existing related technologies, the solution provided by the embodiment of the present disclosure can obtain a more accurate scene flow of the image pair and better meet the actual needs. The optional implementation methods of the present application and the corresponding effects will be described in detail below in conjunction with the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a method performed by an electronic device provided in an embodiment of the present disclosure;
[0019] Figure 2 A schematic diagram of the structure of an AI network provided in an embodiment of the present disclosure;
[0020] Figure 3 A schematic diagram of a network structure of a feature extraction part provided in an embodiment of the present disclosure;
[0021] Figure 4a A schematic diagram of a partial network structure of an AI network provided in an embodiment of the present disclosure;
[0022] Figure 4b A schematic diagram of the principle of supervised training of an AI network provided in an embodiment of the present disclosure;
[0023] Figure 5 A flowchart of a method performed by an electronic device provided in an embodiment of the present disclosure;
[0024] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered as exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures may be omitted.
[0026] The terms and expressions used in the following specification and claims are not limited to their dictionary meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purposes only and not for the purpose of limiting the present disclosure as defined in the appended claims and their equivalents.
[0027] It should be understood that the singular forms "a", "an", and "the" may also include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. When we refer to an element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or the one element and the other element may establish a connection relationship through an intermediate element. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling.
[0028] The term "include" or "may include" refers to the presence of the corresponding disclosed functions, operations or components that can be used in various embodiments of the present disclosure, rather than limiting the presence of one or more additional functions, operations or features. In addition, the term "include" or "have" may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components or combinations thereof.
[0029] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" may be implemented as parameter A including A1 or A2 or A3, or may be implemented as parameter A including at least two of the three items A1, A2, and A3.
[0030] Unless defined differently, all terms (including technical terms or scientific terms) used in the present disclosure have the same meanings as understood by those skilled in the art described in the present disclosure. Common terms as defined in dictionaries are interpreted as having meanings consistent with the context in the relevant technical field, and should not be interpreted ideally or overly formally unless clearly defined in the present disclosure.
[0031] At least some functions of the device or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model, such as at least one module among multiple modules of the device or electronic device can be implemented by an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory and processor.
[0032] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).
[0033] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.
[0034] Here, providing by learning means obtaining a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device or electronic device itself in which the AI according to the embodiment is executed, and / or can be implemented by a separate server / system.
[0035] The AI model may include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.
[0036] A learning algorithm is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0037] The method provided in the present disclosure may involve one or more technical fields such as speech, language, image, video or data intelligence.
[0038] Optionally, when it comes to the field of speech or language, in a method performed by an electronic device, the method may receive a speech signal as an analog signal via an audio acquisition device (e.g., a microphone), and convert the speech portion into computer-readable text using an automatic speech recognition (ASR) model. The user's speech intention can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence-specific processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model can be obtained through training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform the desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.
[0039] Optionally, when it comes to the field of images or videos, in a method performed in an electronic device, the method can obtain output data for identifying relevant information in an image or image pair by using image data as input data of an artificial intelligence model. The artificial intelligence model can be obtained by training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform the desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. The method of the present disclosure may relate to the field of visual understanding of artificial intelligence technology, which is a technology for identifying and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / positioning, or image enhancement.
[0040] Optionally, when it comes to the field of intelligent data processing, in a method performed in an electronic device, the method may use an artificial intelligence model to recommend information / perform data processing. The processor of the electronic device may perform preprocessing operations on the data to convert it into a form suitable for use as an input to the artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operating rule or artificial intelligence model configured to perform the desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. Inference prediction is a technique for logical reasoning and prediction by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning or recommendation.
[0041] Optionally, the method relates to the field of image processing. The method provided by the embodiment of the present disclosure can use the image pair as input data of the artificial intelligence network / model to obtain the scene flow corresponding to the image pair.
[0042] The solution provided by the embodiments of the present disclosure is theoretically applicable to any application scenario that requires scene flow estimation, and these application scenarios may include but are not limited to object detection, segmentation, posture trajectory, etc. For example, the solution can be applied to an autonomous driving scenario. By using the solution provided by the embodiments of the present disclosure, a more accurate scene flow estimation result can be obtained based on the image pairs collected during autonomous driving, and the current autonomous driving environment can be obtained based on the estimated scene flow. For example, based on the estimated scene flow, an autonomous driving car can infer the arbitrary movement of multiple independent objects, thereby enabling safer and more reliable autonomous driving.
[0043] The solution provided by the embodiments of the present disclosure may be executed by any electronic device, including but not limited to a user terminal, a server, etc., wherein the server may be a physical server, a cloud server, an independent server, or a server cluster.
[0044] The following describes several optional embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be pointed out that in the absence of conflicts or contradictions in the implementation of different implementations, the following implementations can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different implementations will not be described repeatedly.
[0045] Figure 1 The flowchart of a method performed by an electronic device according to an embodiment of the present disclosure is shown. The electronic device may be any electronic device, such as a terminal device or a server. Figure 1 As shown, the method may include:
[0046] Step S110: obtaining a pair of images to be processed, each image in the pair of images comprising a color image and a depth image corresponding to the color image;
[0047] Step S120: Using the AI network, obtain the motion embedding features and non-occlusion category label embedding features of the pixels in the image pair, and obtain the scene flow corresponding to the image pair based on the motion embedding features and non-occlusion category label embedding features.
[0048] In the disclosed embodiments, there is no limitation on the specific acquisition method of the image pair to be processed. Optionally, the image pair to be processed may be two frames of images that are continuous in time, each frame of image includes a color image and a depth image corresponding to the color image, that is, each frame of image has both the color information of the image and the depth information of the image, wherein the temporal continuity may be relative continuity or absolute continuity, for example, it may be two adjacent frames of image in a video stream acquired by an image acquisition device, or it may be two adjacent frames of image in an image sequence obtained after downsampling the video stream at a certain sampling rate. Optionally, the color image may be a color image in RGB (Red, Green, Blue) color mode, then each image in the image pair to be processed is a color and depth (RGB-D, RGB-Depth) image, optionally, it may be acquired by an RGB-D camera.
[0049] For the convenience of description, in the embodiment of the present disclosure, the two images in the image pair to be processed are respectively referred to as the first image and the second image, wherein the first image may be an image located before the second image in time. The first image includes a first color image and a first depth image corresponding to the first color image, and the second image includes a second color image and a second depth image corresponding to the second color image.
[0050] For an image pair containing color information and depth information, the method provided by the embodiment of the present disclosure can use a trained AI network to obtain the scene flow corresponding to the image pair, that is, the change of each point between the previous and next frames, for example, 3D motion information, and can also obtain the 2D operation information corresponding to the image pair, namely, optical flow.
[0051] The scene flow estimates not only the background area, but also the rigid dynamic object area, and even the non-rigid dynamic object, occlusion and out-of-boundary area. In the disclosed embodiment, in order to be able to more accurately distinguish the various areas with consistent geometry in the image pair, for example, geometric consistency and instance consistency (for example, the same pose information), so that the scene flow estimation is accurate in the entire image area, in the process of obtaining the scene flow using the AI network, the idea of non-occlusion category labels is introduced, and the scene flow corresponding to the image pair is estimated by simultaneously extracting the motion embedding features and non-occlusion category label embedding features of the pixels in the image (also referred to as geometric segmentation embedding features, which are not limited in this disclosure). Among them, the motion embedding feature can also be called the rigid-motion embedding feature, which can be used to soft-group the objects in the image pair into different rigid objects / objects (for example, different rigid objects in the image are divided into different groups). The non-occluded category label embedding feature is a feature extracted by an AI network based on the idea of non-occluded category labels, which can characterize the category information corresponding to the pixel pair in the image pair, where the pixel pair includes any pixel in the first image and the pixel corresponding to the pixel in the second image. Objects belonging to the same category (such as pixels of a moving object in the image, or pixels in the non-occluded background area of the image, etc.) should have geometric consistency and the posture should be the same or basically the same. The AI network can be trained based on training samples so that the AI network can learn how to distinguish objects of different categories from the samples.
[0052] Among them, the category information corresponding to a pixel pair can also be called the category corresponding to any pixel of the pixel pair in the image pair, that is, the non-occluded category label embedding feature corresponding to the image pair can characterize the category corresponding to the pixel in any image in the image pair, for example, the category corresponding to any pixel in the first image, for example, a pixel is a pixel of a moving object or a pixel of the background area, etc.
[0053] According to the solution provided by the embodiments of the present disclosure, the AI network can soft-group the pixels in the image pair into non-occluded and non-out-of-boundary background areas with consistent geometry, multiple non-occluded and non-out-of-boundary dynamic object areas with consistent geometry, and other areas with consistent geometry based on the motion embedding features and non-occluded category label embedding features of the pixels in the image pair. The areas with geometric consistency should have the same or substantially the same motion field / pose. Therefore, a more accurate scene flow estimation result can be obtained through more accurate soft grouping.
[0054] In practical applications, any pixel in the first image of the image pair (such as the earlier image in the image pair) may or may not have a matching pixel in the second image, for example, the corresponding pixel does not exist in the second image or is occluded due to operation. The solution provided by the embodiment of the present disclosure can distinguish pixels of various categories that are not occluded in the image pair based on motion embedding features and non-occluded category label embedding features, for example, pixels belonging to various moving objects in the image pair, and pixels in the unoccluded background area (or static area / non-moving area, etc.).
[0055] The embodiments of the present application do not make a sole limitation on the network structure of the AI network used in the method provided in the embodiments of the present disclosure. Optionally, the network structure of the AI network can be a neural network model constructed based on a RAFT-3D (Recurrent All-Pairs Field Transforms-three dimensional, a deep neural network for optical flow estimation) network, which can be obtained by optimizing and improving the RAFT-3D network. The optional embodiments of the AI network provided in the embodiments of the present disclosure will be described in detail later.
[0056] It should be noted that the names of the various features described in the embodiments of the present disclosure may also be named by other names. For example, the above-mentioned non-occluded category label embedding feature may also be called a geometric segmentation feature, a segmentation feature, or a first feature or other names, such as a category mask feature or a geometric segmentation mask feature. This feature characterizes the category information corresponding to the pixel pair in the image pair. Therefore, combining this feature can more accurately classify the pixels in the image. The two features of the non-occluded category label embedding feature and the rigid motion embedding feature extracted by the trained AI network can be combined from the two dimensions of rigid objects and object category information to more accurately realize the soft grouping of objects in the image pair. All pixels in the image pair can be softly classified into several groups, and all pixels in each group have the same three-dimensional motion information, i.e., posture.
[0057] Optionally, obtaining the scene flow corresponding to the image pair based on the motion embedding feature and the non-occlusion category label embedding feature includes:
[0058] Based on the motion field corresponding to the image pair, use the AI network to obtain the motion embedding features and non-occluded category label embedding features corresponding to the image pair;
[0059] Fuse the motion embedding features and the non-occluded category label embedding features to obtain hybrid embedding features;
[0060] Based on the hybrid embedded features, the motion field is updated to obtain the target motion field with the goal of minimizing the reprojection error between pixel pairs in the image pair;
[0061] Based on the target motion field, the scene flow corresponding to the image pair is obtained.
[0062] Optionally, in actual implementation, a first motion field (initial motion field / initial pose) corresponding to the image pair may be initialized first; and a target motion field corresponding to the image pair is obtained by iteratively updating the initial motion field. The process may include:
[0063] Initialize the first motion field (three-dimensional motion field) corresponding to the image pair;
[0064] Based on the first motion field, a target motion field is obtained by performing at least one motion field update operation, wherein the current motion field based on which the first update operation is performed is the first motion field, and the target motion field is the motion field obtained by the last motion field update operation; the above-mentioned motion field update operation may include the following steps:
[0065] Based on the current motion field, the motion embedding features and non-occluded category label embedding features corresponding to the image pair are obtained; the motion embedding features and non-occluded category label embedding features are fused to obtain a hybrid embedding feature; based on the hybrid embedding feature, the current motion field is updated with the goal of minimizing the reprojection error between pixel pairs in the image pair, so as to obtain the current motion field based on which the next update operation is performed.
[0066] That is, each update operation will obtain an updated motion field, based on which the motion embedding features and non-occluded class label embedding features corresponding to the image pair can be obtained again to update the motion field again. The target motion field is obtained through a preset number of iterative update operations.
[0067] The method of embedding the non-occluded category label for initializing the first motion field of the image pair is not limited in the embodiments of the present disclosure. In theory, it can be a three-dimensional motion field obtained by any initialization method. Optionally, the method of initializing the first motion field can be the method of initializing the motion field when the RAFT-3D network is currently used for scene flow estimation (initializing SE3 (Special Euclidean Group, Special Euclidean Group) motion, also known as dense pose-level motion or pose). The AI network can continuously iteratively optimize the initial motion field of the image pair based on the motion embedding features and non-occluded category label embedding features obtained from the image pair to obtain the target motion field of the image pair. Among them, the target motion field can be decomposed into rotation and translation components, and the scene flow can be obtained by projecting the target motion field onto the image of the image pair.
[0068] In the disclosed embodiment, there is no unique limitation on the specific method of fusing motion embedding features and non-occlusion category label embedding features, and may include but is not limited to splicing motion embedding features and non-occlusion category label embedding features to obtain hybrid embedding features. The hybrid embedding features can be used to soft-group pixels into various types of areas with consistent geometry. The hybrid embedding features of pixel point pairs (pixel pairs) with geometrically consistent points should be as similar as possible. Therefore, the motion field can be updated and optimized based on the similarity / matching degree between the hybrid embedding features of the pixel point pairs, thereby obtaining the target motion field.
[0069] Optionally, the updating of the motion field based on the hybrid embedding feature includes:
[0070] For each pixel in the first image, determine a neighborhood point set corresponding to the pixel in the second image, and determine a matching degree between the pixel and each pixel in the neighborhood point set based on the hybrid embedding feature of the pixel and the hybrid embedding feature of each pixel in the neighborhood point set;
[0071] According to the matching degree between each pixel in the first image and each pixel in its corresponding neighborhood point set, the motion field is updated to obtain the target motion field with the goal of minimizing the reprojection error between the pixel in the first image and each pixel in its neighborhood point set in the image pair.
[0072] In the implementation of the present disclosure, a pixel refers to a pixel point. It can be understood that in actual implementation, a feature map with a smaller resolution can be obtained by performing feature extraction on the image. A pixel can also refer to a feature point in the feature map. Of course, the feature point in the feature map can also be upsampled to obtain an image of the same size as the image. A feature point in the feature map can correspond to at least one pixel point in the image.
[0073] For each pixel in the first image of the image pair to be processed (it can also be each pixel in the feature map of the first image), the pixel corresponding to the pixel can be found in the second image (through the projection and / or back-projection process of the spatial point in the camera coordinate system, the pixel corresponding to the first pixel in the first image in the second image can be calculated), wherein the above-mentioned neighborhood point set can be the pixels found and their surrounding neighborhood (generally, the neighborhood range can be determined according to a set neighborhood radius).
[0074] For each pixel, the pose of the pixel and the pixels in its corresponding neighborhood should be the same or have a small difference. If the difference is too large, it means that the current motion field is not accurate enough and needs to be updated and optimized. Based on this, in the embodiment of the present disclosure, for each pixel, the motion field of the image pair can be updated by calculating the matching degree between the hybrid embedded features of the pixel and the pixels in the corresponding neighborhood.
[0075] Optionally, the degree of matching between pixel pairs can be obtained by calculating the distance (such as L2 distance) between the mixed embedded features of each pixel and its neighborhood point set, and the update of the motion field can be constrained based on the matching degree so that pixels of the same object can have the same three-dimensional motion information / pose, so that the reprojection error between corresponding pixel pairs is theoretically minimized.
[0076] Optionally, for each pixel, a cost function can be constructed based on the reprojection error between the pixel and each pixel in its neighborhood point set, and the degree of matching between the pixel and each pixel in its neighborhood point set, with the goal of minimizing the algebraic function to achieve the update of the pose information corresponding to the pixel. The value of the cost function can be determined by the weighted reprojection error between the pixel and each pixel in its neighborhood point set, and the weight of the reprojection error corresponding to each pixel in the neighborhood point set is determined by the degree of matching corresponding to the pixel. In each iterative update process, the initial value of the reprojection error corresponding to the cost function is based on the reprojection error between pixel pairs calculated based on the current motion field on which the update operation is based.
[0077] In an optional embodiment of the present disclosure, the method may further include:
[0078] Based on the similarity between pixels in at least one color image in the image pair to be processed, obtaining a weight for motion field adjustment corresponding to the image pair;
[0079] The step of obtaining the scene flow corresponding to the image pair based on the target motion field may include:
[0080] The target motion field is weightedly adjusted based on the weights, and the scene flow corresponding to the image pair is obtained based on the adjusted motion field.
[0081] In practical applications, there may be pixels that are blocked and / or outside the boundary in the two frames of the image pair to be processed (for example, a pixel exists in the first image but not in the second image). It is relatively difficult to estimate the motion of these pixels, and the prediction result error of the scene flow may be large. In view of this problem, in the above optional embodiment of the present disclosure, after obtaining the target motion field of the image pair by continuous iterative updating, the above scheme for further optimizing the target motion field is also provided. It has been found through research that the scene flow and the image itself have a high degree of structural similarity. Based on this, the scheme provided by the embodiment of the present disclosure recommends propagating the high-quality scene flow prediction in the unblocked and non-cross-border pixels to the blocked and out-of-bounds pixels, and further optimization of the scene flow can be achieved by measuring the self-similarity of the features. Specifically, the weight for adjusting the target motion field, i.e., the above weight, can be obtained through the autocorrelation between the pixel information of the image itself, and the target motion field is weighted updated using the weight, so as to obtain a more accurate scene flow based on the updated motion field.
[0082] Optionally, the AI network includes an attention encoder (or attention module / network, etc.), which can be used to obtain the correlation between pixels in at least one color image in the image pair, and determine the weight for motion field adjustment based on the correlation.
[0083] Optionally, the weight for adjusting the motion field based on the correlation includes at least one of the following:
[0084] determining weights for motion field adjustment based on correlations between pixels in either color image of the image pair;
[0085] Based on the correlation between pixels in the first color image, a first weight is determined, and based on the correlation between pixels in the second color image, a second weight is determined, and by fusing the first weight and the second weight, a weight for motion field adjustment is obtained, wherein the first color image and the second color image are two color images in an image pair.
[0086] The embodiment of the present application does not make a sole limitation on the method of fusing the first weight and the second weight, and may include but is not limited to calculating the weight of the first weight and the second weight to obtain the weight for adjusting the motion field.
[0087] Among them, the above-mentioned attention encoder can be an encoder based on the self-attention mechanism, and the input of the encoder can be any color image in the image pair to be processed (for example, the first color image or the second color image). The encoder can learn the image features of the color image, and based on the image features, learn the correlation between the pixels in the color image, so as to obtain the weights for motion field adjustment, that is, the self-attention weights learned by the attention encoder. Of course, the attention encoder can also be used to process the two color images in the image pair separately to obtain the weights corresponding to each color image, and then after fusing the two weights (such as averaging), the fused weights are used to adjust the target motion field.
[0088] It can be understood that in actual implementation, the motion field and the weights used for adjusting the motion field can both be in the form of graphs, that is, the above-mentioned weights can be weight graphs, the graph size of the motion field is the same as the size of the weight graph, and the weighted adjustment of the motion field can be performed by using the weight value of each position point in the weight graph to weight the posture information of the corresponding position in the motion field.
[0089] As an optional solution of the present disclosure, based on the motion field corresponding to the image pair, using an AI network to obtain the motion embedding features corresponding to the image pair and the non-occluded category label embedding features may include:
[0090] Using a feature encoder in AI, extracting first image features of the first image and second image features of the second image respectively;
[0091] Based on the correlation between the first image feature and the second image feature, a correlation volume corresponding to the image pair is constructed;
[0092] Using the context encoder in the AI network, extracting context features and an initial hidden state corresponding to the first image;
[0093] Based on the context features, the initial hidden state, the motion field corresponding to the image pair and the correlation volume, the update network in the AI network is used to obtain the motion embedding features and non-occluded category label embedding features corresponding to the image pair.
[0094] As can be seen from the foregoing description, based on the initialized first motion field, the target motion field corresponding to the image pair can be obtained by iteratively updating the motion field multiple times (for example, a preset number of times) or updating to meet a preset condition (such as the motion field meeting a convergence condition). The motion field based on which the first iterative update is based is the initial motion field, and the motion field based on which each iterative update except the first one is the updated motion field obtained by the previous iterative update.
[0095] In the present formula embodiment, the motion embedding features and non-occluded category label embedding features obtained by each iterative update are obtained based on the motion field based on the current update. Optionally, each iterative update can be implemented based on the correlation between pixel pairs of image pairs (for example, correlation volume), contextual features of the first image, and the initial hidden state obtained by feature extraction of the first image and the initialized first motion field. Optionally, the iterative update process can adopt a motion field update method similar to that of the RAFT-3D network, but different from RAFT-3D, in the present disclosed embodiment, the embedding features output by the update network of the AI network include non-occluded category label embedding features in addition to motion embedding features, and the update of the motion field is not based solely on motion embedding features, but on a mixed embedding feature of motion embedding features and non-occluded category label embedding features.
[0096] In addition, in the disclosed embodiment, the image features are obtained by extracting the color image and the depth image. That is, the image features of each image in the image pair can be extracted based on the color image and the depth image in each image, rather than just extracting features from the color image. By fusing colors and color images, features with better feature expression capabilities can be obtained, so that more image information is integrated into the extracted image features. After obtaining the image features of each image, the correlation between the image features of the two images can be calculated to obtain a 4D correlation volume (4D correlation volume) of the two images. Optionally, the method of calculating the correlation volume corresponding to the two images in the current RAFT-3D network can be used to obtain the correlation volume by calculating the dot product between all feature pairs of the first image feature and the second image feature. The difference is that the image features of the two images used to calculate the correlation volume are image features that fuse the image information of the color image and the depth image, rather than just features extracted from the color image.
[0097] The context features and the initial hidden state may be features of the first color image and the first depth image, and may be obtained by extracting semantic and context information of the image from the first color image and the first depth image using a context encoder.
[0098] After initializing the first motion field, the initial flow field, twist field, and depth residual corresponding to the image pair can be calculated based on the first motion field, and the corresponding correlation features (which may include the correlation features between each pixel and each pixel in the corresponding neighborhood point set) can be found from the correlation volume based on the 2D pixel coordinates obtained according to the initial flow field transformation. The context features, initial hidden features, initial flow field, twist field, depth residual, and correlation features can be used as input information for updating the network (such as a gated recurrent unit) in the AI network, and a new hidden state is generated by updating the network. The correction items corresponding to the scene flow (for example, including the correction items of the flow field), motion embedding features, non-occluded category label embedding features, and confidence features can be obtained according to the new hidden state. Based on the obtained modification items, motion embedding features, non-occluded category label embedding features, and confidence features, the first motion field can be updated to obtain a new motion field after an iterative update.
[0099] Among them, the confidence feature can also be understood as a confidence feature map (Confidence map) used to update the motion field. The eigenvalues in the confidence feature represent the confidence of the matching degree between the corresponding pixel pairs. The lower the confidence, the less matching the pixel pairs are, and the lower the possibility of belonging to the same group. The confidence can be used to constrain the reprojection error between pixel pairs and correct the reprojection error.
[0100] The correction item corresponding to the scene flow includes a correction item of the flow field, which can be used to update the flow field based on the current iteration. Optionally, the correction item can be used to update the current motion field, specifically, to correct the reprojection error corresponding to the pixel pair calculated based on the current motion field, and a new flow field can be obtained based on the updated motion field.
[0101] After obtaining the new motion field based on the correction terms, confidence features and hybrid embedding features corresponding to the scene flow, similarly, a new flow field, distortion field, depth residual can be calculated based on the new motion field, and the correlation features can be found from the correlation volume, so as to obtain the input information for iterative update again, including the above-mentioned context features, the new hidden state generated by the last iterative update, the new flow field, distortion field, depth residual, and at least one or more combinations of the re-found correlation features, and these input information are input into the update network again to obtain the updated motion field again.
[0102] By continuously repeating the above-mentioned iterative update process until the set number of times is reached, the target motion field is obtained. Among them, in the embodiment of the present disclosure, in each iterative update process, optionally, the motion embedding features and the non-occluded category label embedding features can be spliced to obtain hybrid embedding features, and the motion field is updated based on the hybrid embedding features and the confidence features. Different from the current RAFT-3D network, in the embodiment of the present disclosure, the embedding features used to update the motion field not only include motion embedding features, but also incorporate non-occluded category label embedding features, which can more accurately soft-group the pixels in the image, thereby improving the accuracy of the final estimated scene flow.
[0103] The disclosed embodiment provides an end-to-end AI network for scene flow estimation, which can combine the characteristics of geometry and segmentation, and proposes a solution to optimize scene flow estimation based on non-occluded category label embedding features (also referred to as geometric segmentation features). In an optional embodiment of the present disclosure, an attention mechanism is also provided to propagate scene flow estimation from unoccluded and non-out-of-boundary areas to occluded and out-of-boundary areas (for example, weighted updates to motion fields based on image autocorrelation) to optimize the overall scene flow between two frames.
[0104] The specific network architecture of the AI network provided in the embodiment of the present disclosure can be selected according to the needs. In order to better understand and illustrate the principle and effect of the solution provided in the embodiment of the present disclosure, the optional method of the present disclosure is described in detail below in combination with an optional specific AI network structure. The AI network in this optional method can be an improved RAFT-3D network, which can be called a RAFT-3D++ network / system. Figure 2 A schematic diagram of the structure of an AI network provided by an embodiment of the present disclosure is shown. Figure 2 As shown, the network architecture of the AI network may include feature extraction and related modules (e.g., the encoding part in the figure), a multi-task convolutional gate recurrent unit, which may also be called a multi-head convolutional gate recurrent unit / module (multi-head ConvGRUmodule), a differentiable dense pose-level network layer, and an attention mechanism ( Figure 2 The three-dimensional motion propagation module shown in FIG. Figure 3 An optional network structure diagram of feature extraction and related modules is shown. Figure 4a The figure shows the working principle of the multi-task convolutional gate recurrent unit and the differentiable dense pose-level network layer. This AI network is an end-to-end scene flow estimation network that can combine the characteristics of geometry and segmentation to optimize scene flow estimation. Figures 2 to 4a , the AI network provided in the embodiment of the present disclosure is described in detail.
[0105] 1. Feature extraction and related modules
[0106] The disclosed embodiment can use three independent networks to extract features, such as Figure 3 The relevant encoder (eg, feature encoder), context encoder, and attention encoder shown in . This part can use three encoding branches to extract three types of features, including relevant features, context features, and attention features.
[0107] Optionally, the correlation encoder extracts dense image features according to a set resolution, such as extracting a 128-dimensional feature vector at a resolution of 1 / 8, for constructing a four-dimensional correlation volume (4D correlation volume). Specifically, feature extraction can be performed on the two images in the image pair respectively to extract the image features to construct a four-dimensional correlation body (e.g., correlation volume). The specific network structure of the correlation encoder is not limited in this embodiment of the application. Optionally, the correlation encoder can be composed of 6 residual blocks, of which 2 are 1 / 2 resolution, 2 are 1 / 4 resolution, and 2 are 1 / 8 resolution.
[0108] The context encoder can extract semantic and context information from the first image. Similarly, the specific network structure of the relevant encoder is not limited in the present embodiment of the application. Optionally, the context encoder can use a pre-trained ResNet50 with a jump connection (for example, a 50-layer residual neural network) to extract context features at a resolution of 1 / 8. The resolution of the feature map extracted by the context encoder and the relevant encoder is the same.
[0109] Optionally, the correlation encoder and context encoder of the AI network provided in the embodiment of the present disclosure may adopt the feature encoder and context encoder in the current RAFT-3D.
[0110] Unlike RAFT-3D, in addition to the correlation encoder and the context encoder, the feature extraction network of the AI network provided by the embodiment of the present disclosure may also include an attention encoder, such as Figure 3 As shown, optionally, the input of the encoder can be any color image in the image pair, such as the first color image of the first image (for example, the previous image in the image pair), and the attention encoder is used to extract the attention feature, which can be used for subsequent three-dimensional motion propagation, that is, for weighted update of the motion field. It can be understood that the attention encoder is also a feature extraction network for extracting image features of color images (also referred to as attention features, etc.), and the network structure of the encoder is not limited in the embodiments of the present disclosure. The resolution of the features extracted by the encoder is the same as that of the features extracted by the context encoder and the related encoder. Optionally, the attention encoder can adopt the same network structure as the related encoder, but does not share network weights.
[0111] Figure 2 The four-dimensional correlation volume shown in is the correlation volume described above. The four-dimensional correlation volume is constructed by calculating the dot product between the feature vector pairs generated by the correlation encoder.
[0112] Given a motion field, i.e., a 3D motion, each pixel x = (u, v) in the first image can be found by using the projection function and the back-projection function to find the corresponding pixel x′ = π (Tπ -1 (x)), where π(·) and π -1 (·) respectively represent the projection and back-projection functions, x = (u, v) represents the 2D pixel coordinates, T represents SE3 motion (i.e., motion field, the motion field / pose corresponding to the image pair), T specifically represents the motion information / pose corresponding to each pixel x in the first image to the second image, and x′ is the estimated coordinate corresponding to the pixel x calculated by coordinate transformation based on SE3 motion. For each x, based on its current estimated coordinate x′, a set of relevant features can be indexed from the relevant quantity (correlation volume) through the search module.
[0113] Figure 2 The 3D motion initialization module shown in can be used to initialize hidden states, context features, attention features, and SE3 motion T∈SE(3) H×W (First motion field / initial motion field) Specifically, the initial hidden state and context features are extracted through the context encoder, and the attention feature is extracted through the attention encoder (the feature is a three-dimensional motion propagation module). When SE3 motion T is initialized, it can be initialized to a matrix with unit rotation and zero translation.
[0114] 2. Multi-head convolution gate recurrent unit
[0115] The multi-head convolutional gated recurrent unit is the update network described above. The multi-head convolutional gated recurrent unit (Multi-head ConvGRU) is a recurrent convolutional gated recurrent unit for multiple tasks. Similar to the way RAFT-3D estimates scene flow, the flow field, distortion field, depth residual, correlation features, context features, and hidden states are used as inputs to the multi-task convolutional gated recurrent unit. The updated hidden state is then used to generate the corresponding corrected feature r of the scene flow (i.e., the correction term in the previous article), the corresponding confidence feature w, and the motion embedding feature v M Unlike RAFT-3D, in the disclosed embodiment, the AI network also predicts the non-occluded category label embedding feature v G , and the non-occluded category label embedding feature v G There is corresponding supervision, which will be introduced in detail in the supervision part (AI network training part later). In the embodiment of the present disclosure, the motion embedding feature vM and non-occluded category label embedding feature v G The combination is named hybrid embedding feature v H (It may also be other names, such as soft grouping feature or others).
[0116] Unlike RAFT-3D, the solution provided by the embodiment of the present disclosure can improve the matching quality of pixel pairs by learning hybrid embedding features to satisfy the three-dimensional geometric consistency of the pixels under different viewpoints and the instance consistency of the pixels and the neighborhood. In order to learn the hybrid embedding features, the embodiment of the present disclosure can calculate the true value of the non-occluded classification label mask, which can divide the image area into a static non-occluded area and multiple dynamic non-occluded areas, that is, a non-occluded background area and a non-occluded moving instance object area.
[0117] like Figure 4a As shown, the output features of the convolutional gate recurrent unit of the AI network provided by the embodiment of the present disclosure may include four parts, namely, the motion embedding feature v M , non-occluded category label embedding feature v G , confidence feature w and correction feature r, the motion can be embedded into feature v M and non-occluded category label embedding feature v G The mixed embedding feature v is obtained by feature concatenation H , used as one of the inputs to the motion field update layer (pose-level network layer).
[0118] Taking the first iteration as an example, the input of the multi-task convolution gate recurrent unit includes the flow field, distortion field, depth residual and related features found from the four-dimensional correlation volume obtained based on the initial SE3 motion T, as well as the context features and hidden states extracted by the context encoder. Based on these inputs, the multi-task convolution gate recurrent unit can generate a new hidden state (one of the inputs of the multi-task convolution gate recurrent unit in the next iteration), and generate the above-mentioned correction feature r (for correction), confidence feature w, and motion embedding feature v based on the new hidden state. M and non-occluded category label embedding feature v M The generated features are used as the input of the pose-level network layer, and the updated SE3 motion T is obtained through the pose-level network layer. The updated SE3 motion T is used to obtain the flow field, distortion field, depth residual and related features as the input of the multi-head convolutional gate recurrent unit in the second iteration. These features are combined with the above-mentioned context features generated by the context encoder and the new hidden state generated in the first iteration, and are input into the multi-task convolutional gate recurrent unit to generate a new hidden state again.
[0119] 3. Differentiable dense pose-level network layer (Dense-SE3 Layer)
[0120] The role of the pose-level network layer is to update the SE3 motion T. The network layer can be updated and optimized using the Gauss-Newton iteration method. The goal of the update optimization is to divide the pixels in the image into different object groups, and the pixels of each group share the same SE3 motion T. In the disclosed embodiment, a differentiable dense pose-level network layer can be used to update the SE3 motion T.
[0121] Specifically, a cost function E can be defined based on the reprojection error δ (δ i ):
[0122]
[0123]
[0124]
[0125] Among them, σ is the sigmoid function, i represents each pixel, and j represents the neighborhood point set N corresponding to i. i Each pixel in the set, i and each pixel in the set constitute a pixel pair (i, j), α ij It indicates the affinity between point pairs, that is, the degree of matching. is the Mahalanobis distance corresponding to the point pair, which can be called the energy function. i represents the SE3 motion information / pose corresponding to pixel i, T j represents the SE3 motion information corresponding to pixel j, X j Represents the three-dimensional coordinate information of pixel j, r j represents the corrected feature corresponding to pixel j, w j Indicates the confidence feature corresponding to pixel j, for example, a weight feature. Represents the SE3 motion increment corresponding to pixel i. Based on the Gauss-Newton iteration method, the updated motion information corresponding to each pixel, i.e., the updated motion field, is obtained by minimizing the above cost function.
[0126] Specifically, in the disclosed embodiment, an enhanced projection function π() is used to map a 3D point X = (X, Y, Z) to its projected pixel coordinates and depth x = (u, v, d). The above cost function shows that for each pixel i, an SE3 motion T is required to describe the neighborhood j∈N i However, it is not the case that j∈N iEach j in belongs to the same moving object, which is the purpose of the embedding vector. Only pixel pairs (i, j) with similar embeddings contribute significantly to the cost function. In order to obtain a more accurate embedding vector, the present disclosure embodiment designs two embedding feature vectors: motion embedding feature v M and non-occluded category label embedding feature v G , and concatenate these two embeddings to get the mixed embedding vector v H .
[0127] The hybrid embedding feature vector is used to soft-group pixels into non-occluded and non-out-of-bounds background regions with consistent geometry, multiple non-occluded and non-out-of-bounds dynamic object regions with consistent geometry, and other regions with consistent motion. For any pixel pair (i, j), given the corresponding two hybrid embedding vectors and The affinity α of the corresponding pixel pair can be calculated based on the L2 distance (ij) ∈[0,1] (that is, the degree of matching of the pixel pairs described above).
[0128] Based on the above expression, to minimize the cost function E δ (δ i ) is used to update the SE3 motion T, obtain the updated SE3 motion T, and perform the next iteration based on the new SE3 motion T. Figure 2 In the schematic diagram shown, the target motion field can be obtained through N iterative updates, that is, the number of iterations reaches a preset number or the update of the motion field meets the preset convergence condition. When the preset number of update iterations is used as a condition, the target motion field is the SE3 motion T obtained by the Nth iterative update, where N is a positive integer and the value can be set according to experimental values and / or empirical values.
[0129] 4. Three-dimensional motion propagation
[0130] Differentiable dense pose-level network layers (e.g., Dense-SE3 Layer) implicitly assume that the corresponding pixels are visible in both input frames and that the scene flow prediction is accurate. Only with this assumption can the Dense-SE3 Layer accurately and robustly estimate the SE3 motion T of the entire image. However, this assumption is invalid for occluded and out-of-bounds pixels in the two input frames. To solve this problem, by observing that the scene flow and the image itself have a high degree of structural similarity, the disclosed embodiment proposes to propagate the high-quality scene flow prediction in unoccluded and non-out-of-bounds pixels to occluded and out-of-bounds pixels by measuring the self-similarity of the features, which can be effectively implemented by a simple self-attention layer.
[0131]
[0132] Among them, F represents the attention feature, that is, the image feature of the color image extracted by the attention encoder, and the correction weight of the target motion field T (which can be understood as the SE3 motion obtained in the last iteration) is obtained by calculating the self-attention weight of the image feature The target motion field T is adjusted using the weight to obtain the adjusted motion field in is a normalization factor to avoid large values after the dot product operation.
[0133] Getting the final playground Afterwards, you can Decomposed into rotation and translation components, they can be projected onto the images of the image pair to obtain scene flow and optical flow.
[0134] Based on the RAFT-3D++ network provided in the embodiment of the present disclosure, the two frames of color images I input to the network can be 1 ,I 2 and the depth image d 1 ,d 2 The features extracted in are used to construct a 4D correlation volume (a 4D correlation volume is constructed by computing the visual similarity between all pairs of pixels) and initialize the correlation state. The SE3 motion T is initialized to find the identity of each pixel when the model starts to execute. The context features and hidden states are initialized for iteration of the Multi-task Conv GRU module. The attention features of the attention module are initialized to propagate the SE3 motion. During each iteration, the lookup module uses the current SE3 motion estimate to index from the correlation volume, and the correlation features and hidden states are used to generate estimates of the corresponding corrected features, corresponding confidence features, and hybrid embedded features of the scene flow. These estimates are inserted into the differentiable dense pose-level network layer (Dense-SE3 Layer), which is a least squares optimization layer that uses geometric constraints to generate updates to the SE3 motion estimate.
[0135] During each iteration, the current estimate of SE3 motion is used to build an index from the 4D correlation volume, and the Multi-head ConvGRU is iterated to obtain the relevant features and generate the corresponding correction features of the scene flow, the corresponding confidence features and the hybrid embedding features, and then the differentiable dense pose-level network layer (Dense-SE3 Layer) generates the update of SE3 motion. The attention module propagates the SE3 motion in unoccluded and non-out-of-bounds pixels to occluded and out-of-bounds pixels by measuring feature self-similarity.
[0136] Figure 4aFigure 1 shows the process of each iteration of RAFT-3D++: During each iteration, the lookup module uses the current SE3 motion estimate to index from the 4D correlation volume (correlation volume), uses the correlation features, context features and hidden states to generate an estimate r of the corresponding corrected features of the scene flow, the corresponding confidence features w and the hybrid embedded features v H The hybrid embedding feature includes the non-occluded category label embedding feature v G and motion embedding feature v M These estimates are inserted into the differentiable dense pose-level network layer (Dense-SE3 Layer), which uses geometric constraints to generate updates for the SE3 motion T. After successive iterations, these estimates are inserted into the 3D motion propagation module (also known as the attention module), which propagates the SE3 motion in unoccluded and non-out-of-bounds pixels to occluded and out-of-bounds pixels by measuring feature self-similarity, that is, propagating the 3D motion information of areas that can be predicted more accurately and relatively easily to the 3D motion information of areas that are more difficult to predict accurately.
[0137] Through multiple iterative updates, a dense, robust and accurate SE3 motion is finally recovered, which can be decomposed into rotation and translation components. The SE3 motion can be projected onto the image to recover the scene optical flow and scene flow.
[0138] The AI network provided in the embodiments of the present disclosure is a neural network obtained by training with a training set. Figure 5 A method performed by an electronic device according to an embodiment of the present disclosure is shown. Figure 5 As shown, the method may include:
[0139] Step S510: obtaining an AI network to be trained and a training set, wherein the training samples in the training set include: sample image pairs, scene flow truth values and geometric segmentation mask truth values (non-occluded category label mask truth values) of the sample image pairs, wherein each sample image in the sample image pair includes a color image and a corresponding depth image, and the mask truth value represents the true category of the pixel points in the sample image pair, that is, the true category corresponding to the pixel pair in the sample image pair, wherein the pixel pair includes any pixel in the first sample image in the sample image pair and the pixel corresponding to the pixel in the second sample image in the sample image pair;
[0140] Step S520: Repeat the training operation on the AI network to be trained based on the training set until the training end condition is met to obtain a trained AI network, wherein the training operation includes:
[0141] For each sample image pair, use the AI network to obtain the motion embedding features and non-occlusion category label embedding features corresponding to the sample image pair, obtain the predicted scene flow corresponding to the sample image pair based on the motion embedding features and non-occlusion category label embedding features, and obtain the non-occlusion category label mask prediction value of the sample image pair based on the non-occlusion category label embedding features; wherein the non-occlusion category label embedding features represent the category information corresponding to the pixel pairs in the sample image pair;
[0142] Based on the true value of the scene flow and the predicted scene flow of each sample image pair, the first training loss is determined; based on the true value of the non-occluded category label mask and the predicted value of the non-occluded category label mask corresponding to each sample image pair, the second training loss is determined; based on the first training loss and the second training loss, the total training loss is determined; based on the total training loss, the model parameters of the AI network are adjusted.
[0143] It can be understood that the method is a training method for an AI network, and the electronic device executing the method is the same as the Figure 1 The electronic devices used in the methods shown in the examples may be different or the same.
[0144] The specific method of obtaining training samples in the training set is not limited in the embodiments of the present disclosure. The training set includes a large number of training samples, each of which includes a sample image pair with a supervised training label, the label including the scene flow true value and the geometric segmentation mask true value corresponding to the sample image pair, wherein the geometric segmentation mask true value represents the true category of the object / region in the image pair, and the objects / regions of the same category have the same posture. For example, a sample image pair contains three types of objects (such as two moving / dynamic non-occluded objects and a non-occluded background area). The true value of the geometric segmentation mask of the sample image pair may include three mask images, one mask image corresponding to one type of object. The mask image includes two values, such as 1 and 0. The size of the mask image is the same as the size of the sample image. A pixel with a value of 1 (for example, corresponding to a pixel pair) indicates that the pixel belongs to the object, and a pixel with a value of 0 indicates that the pixel does not belong to the object. Therefore, the true pixel grouping result of the sample image pair can be known based on the true value of the geometric segmentation mask of the sample image pair. Therefore, the AI network can be supervised trained based on the true value, so that the trained AI network can accurately extract the non-occluded category label embedding features corresponding to the image pair to be processed, so that a more accurate scene flow estimation result can be obtained based on the feature.
[0145] It can be understood that the specific process of processing the sample image pair through the AI network is the same as the process of performing image processing on the image pair to be processed through the AI network described in the previous article, and will not be repeated here.
[0146] During the training stage of the AI network, for any sample image pair, after obtaining the non-occluded category label embedding feature corresponding to the sample image pair through the AI network, the non-occluded category label mask prediction value of the sample image pair can be obtained based on the embedded feature, so that the mask loss, i.e., the second training loss, can be calculated based on the difference between the non-occluded category label mask prediction value and the non-occluded category label mask true value.
[0147] Optionally, the feature value of the non-occluded category label embedding feature corresponding to each pixel can be used as the non-occluded category label mask prediction value corresponding to the pixel. As an optional method, for any sample image pair, based on the non-occluded category label embedding feature, the geometric segmentation mask prediction value of the sample image pair is obtained, including:
[0148] Determine pixels of each category in the sample image pair based on the true value of the non-occluded category label mask of the sample image pair;
[0149] For each category, based on the non-occluded category label embedding features of each pixel belonging to the category, determine the average non-occluded category label feature corresponding to the category;
[0150] For each pixel in the first sample image, a non-occluded category label mask prediction value corresponding to the category of the pixel is determined based on a difference between a non-occluded category label embedding feature corresponding to the pixel and an average non-occluded category label feature of the category to which the pixel belongs.
[0151] The larger the difference corresponding to a pixel, the lower the possibility that the pixel value belongs to the category, and the smaller the predicted value of the non-occluded category label mask corresponding to the category of the pixel. Optionally, the true value of the mask corresponding to a pixel can be 0 or 1, and the predicted value can also be 0 or 1, or a value normalized to [0,1]. The larger the corresponding difference, the smaller the predicted value.
[0152] In the disclosed embodiment, the total training loss of the AI network is obtained by combining the mask loss corresponding to each sample image pair with the scene flow prediction loss of each sample image pair, i.e., the first training loss, so that the parameters of the model can be adjusted based on the total loss. The adjusted AI network is trained continuously until the training end condition is met (such as the total training loss converges or the number of training times reaches the set number of times, etc.), and a trained AI network is obtained.
[0153] Similarly, in the training phase, for any sample image pair, based on the motion embedding features and the non-occlusion category label embedding features, obtaining the predicted scene flow corresponding to the sample image pair may include:
[0154] Based on the motion field corresponding to the image pair, the AI network is used to obtain the motion embedding features and non-occluded category label embedding features corresponding to the sample image pair; the motion embedding features and the non-occluded category label embedding features are fused (such as splicing) to obtain hybrid embedding features; based on the hybrid embedding features, the motion field corresponding to the sample image pair is updated with the goal of minimizing the reprojection error between pixel pairs in the sample image pair to obtain a target motion field; based on the target motion field, the predicted scene flow corresponding to the sample image pair is obtained.
[0155] Similarly, the initial motion field corresponding to the sample image pair can be obtained first; based on the initial motion field, the motion field is iteratively updated using an AI network, and the motion field obtained from each update is used as the motion field based on which the next iterative update is based. After a preset number of iterative updates, the scene flow corresponding to the sample image pair is obtained, that is, the scene flow of the sample image pair predicted by the AI network. It can be understood that the specific process of obtaining the scene flow corresponding to the sample image pair in the training phase is the same as the process of obtaining the scene flow corresponding to the image pair to be processed through the trained AI network summarized in the previous article, except that one is for the image pair to be processed and the other is for the sample image pair.
[0156] Optionally, for any sample image pair, the method may further include:
[0157] Based on the similarity between pixels in any sample color image in the sample image pair, a weight for motion field adjustment (also referred to as motion field adjustment weight or scene flow correction weight, etc.) of the sample image pair is obtained;
[0158] The target motion field corresponding to the sample image pair is weightedly adjusted based on the above weights, and the predicted scene flow of the sample image pair is obtained based on the adjusted motion field.
[0159] Among them, the scene flow correction weight can be implemented by using the self-attention mechanism. The image features obtained by feature extraction of the color image are used as the input features of the attention mechanism, and the self-attention weight corresponding to the color image is calculated as the correction weight.
[0160] In the embodiment of the present disclosure, there is no unique limitation on the method for obtaining the true value of the geometric segmentation mask of the sample image pair, and it can be obtained by manual marking. In order to improve the efficiency of sample acquisition and reduce labor costs, in the optional embodiment of the present disclosure, for any sample image pair, the true value of the non-occluded category label mask of the sample image pair is determined by the following method:
[0161] Obtaining instance object segmentation results of the sample image pair;
[0162] Determining a first photometric error between matching pixel pairs in the sample image pair according to the scene flow true value of the sample image pair;
[0163] Based on the first photometric error and the instance object segmentation result, determine the true value of the non-occluded category label mask corresponding to each instance object in the sample image (hereinafter referred to as m obj );
[0164] Based on the true value of the non-occluded category label mask corresponding to each instance object, the true value of the non-occluded category label mask of the sample image pair is obtained.
[0165] In this optional solution, the instance object segmentation result of the sample image pair can indicate which instance objects are in the sample image pair. The instance objects may include dynamic objects, static objects, and the background in the image can also be regarded as an instance object. Based on the instance object segmentation result, it can be known which pixels in the image pair belong to the same object.
[0166] Since the scene flow truth value of the sample image pair is known, the first photometric error between the pixel pairs in the sample image pair (for example, the breadth error between a pixel and the pixel corresponding to its projection position) can be calculated based on the scene flow truth value. It is also known which pixels in the sample image pair belong to the same instance object. Therefore, based on the first photometric error, the geometric segmentation mask truth value corresponding to each instance object can be determined, that is, which pixels belong to which instance object. Specifically, for any pixel, if the pixel belongs to an instance object and the first photometric error corresponding to the pixel is less than the first threshold, it can be determined that the pixel does belong to the instance object, and the mask truth value corresponding to the pixel on the mask truth map corresponding to the instance object can be determined to be 1, otherwise it is 0. Using this scheme, the geometric segmentation mask truth value corresponding to each instance object can be obtained.
[0167] As an optional method, the geometric segmentation mask true value of each instance object obtained above can be used as the geometric segmentation mask true value of the sample image pair for training the AI network. However, in practical applications, although there may be multiple instance objects in an image pair, some instance objects may always be static in the two images, such as the background in the image, static houses, etc. Considering this problem, in another optional solution of the present disclosure, for any sample image pair, the geometric segmentation mask true value of the sample image pair is determined by the following method:
[0168] Obtain the true value of the motion field corresponding to the sample image pair;
[0169] Determine a second photometric error and a depth error between matching pixel pairs in the sample image pair based on a true value of the motion field corresponding to the sample image pair;
[0170] Based on the second depth error and the depth error, determine the true value of the non-occluded category label mask corresponding to the non-occluded background area in the sample image pair (hereinafter referred to as mstatic );
[0171] By fusing the true value of the non-occluded category label mask corresponding to the non-occluded background area and the true value of the non-occluded category label mask corresponding to each instance object, the true value of the geometric segmentation mask of the sample image pair is obtained.
[0172] Since the motion field true value and depth image of the sample image pair are known, the second optical error and depth error that incorporate the true position and depth information of the pixel pair can be more accurately calculated based on the motion field true value and depth image of the image pair. Based on the optical error and depth error, it is possible to accurately distinguish which areas are the non-occluded and non-out-of-boundary background areas in the image pair, and obtain the mask true value corresponding to the background area.
[0173] Optionally, for any pixel pair, if the second optical error corresponding to the pixel is less than the second threshold and the depth error corresponding to the pixel is less than the third threshold, the pixel pair is considered to belong to the non-occluded and non-out-of-boundary background area, and the mask true value corresponding to the pixel pair can be determined to be 1, otherwise it is determined to be 0. Based on this method, a mask true value map of the non-occluded and non-out-of-boundary background area can be obtained.
[0174] After obtaining the mask truth map corresponding to each instance object and the mask truth map of the non-occluded and non-out-of-boundary background area, these truth maps can be fused to obtain the final geometric segmentation mask truth map of the sample image for training the AI network. Specifically, if the pixels that are all 1 in the mask truth map of one or some instance objects are also all 1 in the mask truth map of the above background area, then the mask truth maps of these instance objects can be removed, and these instance objects belong to the above background area class. Through the above fusion, the final mask truth map can be obtained.
[0175] As an example, suppose there are 5 instance objects in a sample image pair. After obtaining the mask truth maps of the 5 instance objects and the mask truth map of the background area, suppose that the mask truth maps of two instance objects are covered by the mask truth map of the background area (for example, the area with pixel 1 is covered), then the final mask truth map is the mask truth map of 4 categories, namely, the mask truth map of the background area and the mask truth maps of the other 3 instance objects. The other 3 objects are each regarded as a category, and the background area is regarded as 1 category.
[0176] The training method of the AI network provided by the embodiment of the present disclosure can jointly optimize the multi-task loss (first training loss, second training loss) so that the scene flow estimation performance of the trained AI network is better and can better meet the needs of practical applications. Among them, the training method can combine ground truth optical flow and depth changes to supervise the learning of the network, and also define geometric segmentation masks and use them to supervise the network. Figure 4b A schematic diagram of the principle of supervised training provided by an embodiment of the present disclosure is shown. Figure 4b The relevant contents of the multi-task loss involved in the embodiments of the present disclosure are further described in detail.
[0177] 1. Scene flow loss is the first training loss, corresponding to Figure 4b Scene flow supervision in
[0178] The AI network provided in the embodiment of the present disclosure can output a series of T 1 , T 2 , ..., T K . Where K represents the number of instance objects in the image pair, T K Represents the SE3 motion of the Kth instance object, i.e., the pose, that is, the transformation information. For each transformation T K , the optical flow and depth changes caused by the corresponding instance object can be calculated Where x represents a pixel in the image pair, which also represents the dense pixel correspondence, and the scene flow loss L flow The L1 distance calculation can be used, which is calculated as follows:
[0179] L flow =||f gt -f est || 1
[0180] Among them, f gt is the scene flow truth value of the sample, f est is the scene flow predicted by the network, that is Figure 4b In the above, the updated three-dimensional motion predicted by the network (which can be understood as the target motion field corresponding to the sample image pair) is projected onto the sample image to obtain the scene flow.
[0181] 2. Mask loss is the second training loss, corresponding to Figure 4b The mask supervision in , this loss can also be called the non-occluded category label mask loss
[0182] The disclosed embodiment may use per-pixel non-occluded category label mask loss to help the AI network learn the non-occluded category label embedding feature v G ∈RD×H×W , where D, H, and W represent the number of channels, height, and width of the image in the image pair, respectively. Given a non-occluded class label mask truth value First, we calculate the average embedding t corresponding to each category k∈{1, 2, ..., K} k ∈R D That is, the average non-occluded category label mask feature:
[0183]
[0184] Assume that the average non-occluded class label embedding t k Represents the mean feature of K truth masks (also understood as K categories), where k represents a category (non-occluded static background area with consistent geometry or non-occluded dynamic object area with consistent geometry or occluded / out-of-bounds area). Represents the sum of all values in the mask truth map of category k, that is, the number of pixels whose true value is 1, Indicates the non-occluded category label embedding features corresponding to the sample image predicted by the AI network, represents the true value (0 or 1) of the mask corresponding to any pixel belonging to category k, t k It is the mean of the non-occluded category label embedding features corresponding to all pixels belonging to category k.
[0185] Then, features can be embedded based on the non-occluded category labels And t corresponding to each category k Perform the discrimination task and calculate the mask L mask , calculated as follows:
[0186]
[0187] L mask =||-(m gt log(m est )+(1-m gt )log(1-m est ))|| 1
[0188] Specifically, for each category k, the feature t can be embedded according to the average non-occluded category label of the category k And the non-occluded category label embedding feature of each point belonging to this category obtained through the network The difference between them gives the mask prediction value of each pixel corresponding to the category. In this option, and t k The difference can be calculated based on the L2 distance between the two.
[0189] Then, we can calculate the true value m of the non-occluded category label mask corresponding to each sample image pair gt and the predicted value m est The cross entropy loss between them gives the mask loss.
[0190] In the embodiment of the present disclosure, the true value of the non-occluded label mask corresponding to the sample image pair can be based on the color image I of the given sample image pair. 1 , I 2 , depth image D 1 , D 2 , real scene flow gt , and the actual two camera poses T 1 , T 2 , and the instance object label truth value O gt Calculated. Among them, each intensity / luminosity consistency can be calculated through a pair of color images and the real scene flow, and the occluded area and the non-occluded area can be distinguished based on the consistency of the intensity. gt , the occluded and non-occluded areas of each instance object can be distinguished.
[0191] Optionally, the photometric error of each instance object k can be calculated, and the mask truth value corresponding to each instance object can be determined based on the photometric error. If the error is lower than a threshold, the pixel of the instance object is considered to be non-occluded. Based on this, the non-occluded mask truth value m corresponding to the sample image pair is obj It can be calculated in this way, without considering whether the instance object is static or dynamic:
[0192] Based on the two color images I in the sample image pair 1 , I 2 , scene flow gt (scene flow truth value) and instance object o gt The true value of (characterizing the instance object segmentation result of the sample image pair), the mask truth value m of the non-occluded and non-out-of-bounds object region (non-occluded instance object region) with consistent geometry obj , the photometric error (first photometric error) can be calculated by the following formula, where k represents the instance object index.
[0193] E p (x)=||I i (x)-I j (x+f gt )|| 2
[0194]
[0195] Among them, I 1(x) represents the pixel value of pixel x in the first sample color image, I 2 (x+f gt ) is the pixel value of the pixel corresponding to the pixel x in the second sample image calculated based on the scene flow true value, ||.|| 2 represents the L2 distance, Th 1 is the first threshold, o gt ∈}1,...,k,...K} represents the instance object corresponding to index k. The above formula means: if a pixel corresponds to E p is less than the first threshold, and the pixel is the pixel of the instance object corresponding to index k, and the mask truth map m corresponding to the instance object obj The true value of the mask at the position corresponding to the pixel is k, otherwise it is 0.
[0196] Through the above calculation method, the mask truth map corresponding to each instance object can be obtained.
[0197] In addition to the above data, considering that the static instances (static areas) in the image have global motion, which is consistent with the change of the camera pose, in order to distinguish dynamic and static examples, in the embodiment of the present disclosure, given the ground truth camera pose, the photometric error and depth error of each pose can also be calculated. If the motion of an instance object coexists with the camera motion, then the instance object may be static, dynamic or other (such as being occluded or out of bounds). Based on this, the depth D can also be obtained in the embodiment of the present disclosure. 1 , D 2 and the poses T of the two cameras 1 , T 2 True value, based on these data, the photometric error (second photometric error) and depth error can be calculated. The specific calculation formula is as follows:
[0198]
[0199]
[0200] in, represents the depth of pixel x' corresponding to pixel x calculated based on camera pose and depth, represents the depth value calculated in the second sample depth image of the second sample image according to the estimated pixel coordinate of x′. For other parameters in the calculation formula, please refer to the corresponding explanations in the previous text.
[0201] After calculating the above photometric error and depth error, the mask truth map of the non-occluded background area can be calculated. Specifically, the truth value m of the mask of the non-occluded background / static area with consistent geometric shape static Calculated by photometric error and depth error:
[0202]
[0203] Among them, Th 2 is the second threshold, Th 3 is the third threshold. The second threshold and the third threshold can be determined according to experimental values and / or empirical values. If the depth error corresponding to a pixel pair is small, it can be understood that the pixel pair is stationary in the image pair, and if the photometric error corresponding to the pixel pair is small, it can be understood that the pixel pair is not occluded or out of bounds. Therefore, if the depth error and photometric error corresponding to the pixel pair are small, the pixel pair can be considered to be a non-occluded and non-out-of-bounds stationary pixel pair, and belongs to the background.
[0204] Through two masks m obj and m static , we can get the real non-occluded category label mask m gt :
[0205]
[0206] It can be seen that the scene can be roughly divided into three categories. One category is the independently moving non-occluded moving instance objects (it can be understood that in practice, this category may be multiple categories, and a non-occluded moving instance (such as a moving person or object) is a category). The other is the non-occluded static area, which consists of a non-occluded background and a non-occluded static instance object. All pixels in this class have consistent global motion. The last category is the occluded area (that is, m in the above formula gt = 0), where the pixel cannot be seen from two different views at the same time. Based on the mask m gt We can learn non-occluded dynamic instance embeddings and use the average embedding to assist the non-occluded category labeling mask m gt Supervise it.
[0207] During the training process, after calculating the scene flow loss and mask loss based on the scene flow true value and mask true value corresponding to each training sample in the training set, and the scene flow prediction value and mask prediction value corresponding to each sample image predicted by the AI network, the total training loss can be obtained based on these two losses. Optionally, the total loss can be the weighted sum of the scene flow loss and the mask loss. The weights of the two losses can be configured according to actual needs, or set according to empirical values or experimental values. Optionally, the hyperparameter weights w1 and w2 are set to 1.0 and 0.1 respectively. The total training loss L can be expressed as follows:
[0208] L=w1*L flow +w2*L mask
[0209] L flow and Lmask They represent the scene flow loss and mask loss corresponding to all training samples used in each training operation.
[0210] The scene flow estimation scheme provided by the embodiment of the present disclosure can ensure the matching accuracy as much as possible to perform nonlinear optimization, so that the pose estimation will be more accurate and more robust. The scene flow estimated by this method can not only estimate the background area, but also estimate the rigid dynamic object area, and even non-rigid dynamic objects, occlusions and out-of-boundary areas. In order to make the scene flow estimation accurate in the entire image area, the embodiment of the present disclosure designs different estimation methods for different areas. First, a geometric segmentation mask is designed, which can accurately distinguish non-occluded and non-out-of-boundary background areas, as well as each non-occluded and non-out-of-boundary rigid object area, and use this geometric segmentation mask to supervise the non-occluded category label embedding features, and estimate the SE3 motion of these areas by inputting the mixed embedding features into the Dense-SE3 Layer. Then, the attention module is used to propagate the SE3 motion of the unoccluded and non-out-of-boundary areas to the occluded and out-of-boundary areas, thereby optimizing the SE3 motion of the occluded and out-of-boundary areas.
[0211] The AI network of the disclosed embodiment can be called RAFT-3D++, which can estimate the 3D motion information of an image pair (such as a pair of RGB-D video frames) at the pixel level including color information and depth information. The dense 3D motion learned by the existing scene flow estimation method does not take into account the correct correspondence between dense pixels, such as 3D geometric consistency, instance (also referred to as instance object, etc.) consistency (because the rigid motion embedding feature is not supervised during the training stage, it may be easier to break the three-dimensional geometric consistency and instance consistency, resulting in the pixel matching quality of the occluded area), which is very important for improving the matching quality of pixel pairs in the end-to-end AI network framework. Based on this, the disclosed embodiment provides a new AI network, which can be implemented based on RAFT-3D. The A network provided in the disclosed embodiment can, on the one hand, group pixels with similar embedded features by learning a hybrid embedded feature representation. This ability of the AI network can be learned based on the mask true value of the training sample, and the mask value of the pixel can classify the pixels on the mask according to the three-dimensional geometric consistency under different viewpoints and its instance consistency with the neighborhood. With the help of the mask, the pixel can find a high-quality correct match to minimize the reprojection error that occurs regardless of the dynamic rigid objects or static background in the scene. On the other hand, an attention mechanism (e.g., an attention encoder and a three-dimensional motion propagation module) is also introduced in the embodiment of the present disclosure. Based on this attention mechanism, reliable and accurate 3D motion estimated from pixels on the mask of non-occluded and non-out-of-bounds categories can be propagated to more challenging occluded or textureless areas. Experiments show that compared with RAFT-3D, the scheme proposed in the present disclosure can effectively improve the performance of scene flow estimation.
[0212] An embodiment of the present disclosure also provides an electronic device, which includes a processor and, optionally, may also include a transceiver and / or a memory coupled to the processor, and the processor is configured to execute the steps of the method provided in any optional embodiment of the present disclosure.
[0213] Figure 6 A schematic diagram of the structure of an electronic device applicable to an embodiment of the present invention is shown in FIG. Figure 6 As shown, Figure 6The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. The electronic device can be a user terminal device or a server.
[0214] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0215] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in 6, but it does not mean that there is only one bus or one type of bus.
[0216] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0217] The memory 4003 is used to store the computer program for executing the embodiment of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0218] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0219] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.
[0220] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in the text.
[0221] It should be understood that, although the flowchart of the embodiment of the present disclosure indicates each operation step by arrows, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios with different execution times, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present disclosure does not limit this.
[0222] The above text and drawings are provided only as examples to help readers understand the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the contents disclosed herein, it is obvious to those skilled in the art that the embodiments and examples shown can be changed without departing from the scope of the present disclosure, and other similar implementation means based on the technical ideas of the present disclosure are adopted, which also fall within the protection scope of the embodiments of the present disclosure.
Claims
1. A method performed by an electronic device, characterized in that: The method comprises: Acquire a pair of images to be processed, each image in the pair of images comprising a color image and a depth image corresponding to the color image; Using an artificial intelligence (AI) network, obtaining motion embedding features and non-occlusion category label embedding features corresponding to the image pair, and obtaining a scene flow corresponding to the image pair based on the motion embedding features and non-occlusion category label embedding features; The image pair includes a first image and a second image, the non-occluded category label embedding feature is used to characterize the category information corresponding to the pixel pair in the image pair, and the pixel pair includes any pixel in the first image and the pixel corresponding to the pixel in the second image.
2. The method according to claim 1, characterized in that Obtaining a scene flow corresponding to the image pair, comprising: Based on the motion field corresponding to the image pair, using the AI network, obtaining motion embedding features and non-occluded category label embedding features corresponding to the image pair; Fusion of the motion embedding feature and the non-occluded category label embedding feature to obtain a hybrid embedding feature; Based on the hybrid embedded features, the motion field is updated with the goal of minimizing the reprojection error between pixel pairs in the image pair to obtain a target motion field; Based on the target motion field, a scene flow corresponding to the image pair is obtained.
3. The method according to claim 2, characterized in that Updates to the playing field include: For each pixel in the first image, determine a neighborhood point set corresponding to the pixel in the second image, and determine a matching degree between the pixel and each pixel in the neighborhood point set based on the hybrid embedding feature of the pixel and the hybrid embedding feature of each pixel in the neighborhood point set; The motion field is updated according to the matching degree between each pixel in the first image and each pixel in its corresponding neighborhood point set, with the goal of minimizing the reprojection error between the pixel in the first image and each pixel in its neighborhood point set in the image pair.
4. The method according to claim 2 or 3, characterized in that: The method further comprises: Based on the similarity between pixels in at least one color image in the image pair, obtaining a weight for motion field adjustment corresponding to the image pair; The obtaining the scene flow corresponding to the image pair based on the target motion field includes: The target motion field is weightedly adjusted based on the weight, and the scene flow corresponding to the image pair is obtained based on the adjusted target motion field.
5. The method according to claim 4, characterized in that Based on the similarity between pixels in at least one color image in the image pair, obtaining a weight for motion field adjustment corresponding to the image pair, comprising: Using the attention encoder in the AI network, the correlation between pixels in at least one color image in the image pair is obtained, and based on the correlation, the weight for motion field adjustment is determined.
6. The method according to claim 5, characterized in that The determining the weight for adjusting the motion field based on the correlation comprises at least one of the following: Determining the weights for motion field adjustment based on correlations between pixels in any color image in the image pair; Based on the correlation between pixels in the first color image, a first weight is determined, and based on the correlation between pixels in the second color image, a second weight is determined, and the weight for motion field adjustment is obtained by fusing the first weight and the second weight, wherein the first color image and the second color image are two color images in the image pair.
7. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining the motion embedding features and the non-occluded category label embedding features corresponding to the image pair based on the motion field corresponding to the image pair using the AI network includes: Using a feature encoder in the AI, respectively extracting a first image feature of the first image and a second image feature of the second image; Based on the correlation between the first image feature and the second image feature, construct a correlation volume corresponding to the image pair; Using a context encoder in the AI network, extracting context features and an initial hidden state corresponding to the first image; Based on the context features, the initial hidden state, the motion field corresponding to the image pair and the correlation volume, an update network in the AI network is used to obtain the motion embedding features and the non-occluded category label embedding features corresponding to the image pair.
8. A method performed by an electronic device, characterized in that: The method comprises: Obtain an AI network to be trained and a training set, wherein the training samples in the training set include: a sample image pair, a scene flow truth value of the sample image pair, and a non-occluded category label mask truth value, wherein each sample image in the sample image pair includes a color image and a corresponding depth image, and the mask truth value represents the true category corresponding to a pixel pair in the sample image pair, wherein the pixel pair includes any pixel in a first sample image in the sample image pair and a pixel corresponding to the pixel in a second sample image in the sample image pair; Repeating the training operation on the AI network to be trained based on the training set until the training end condition is met to obtain a trained AI network, wherein the training operation includes: For each of the sample image pairs, using an AI network, obtain motion embedding features and non-occlusion category label embedding features corresponding to the sample image pair, obtain the predicted scene flow corresponding to the sample image pair based on the motion embedding features and non-occlusion category label embedding features, and obtain the non-occlusion category label mask prediction value of the sample image pair based on the non-occlusion category label embedding features, wherein the non-occlusion category label embedding features characterize the category information corresponding to the pixel pair in the sample image pair; Based on the true value of the scene flow and the predicted scene flow of each sample image pair, a first training loss is determined; based on the true value of the non-occluded category label mask and the predicted value of the non-occluded category label mask corresponding to each sample image pair, a second training loss is determined; based on the first training loss and the second training loss, a total training loss is determined; and based on the total training loss, the model parameters of the AI network are adjusted.
9. The method according to claim 8, characterized in that For any sample image pair, based on the non-occluded category label embedding feature, obtaining a non-occluded category label mask prediction value of the sample image pair includes: Based on the true value of the non-occluded category label mask and the non-occluded category label embedding feature of the sample image pair, determine the average non-occluded category label feature corresponding to each category in the sample image pair; For each pixel in the first sample image, based on the difference between the non-occluded category label embedding feature corresponding to the pixel and the non-occluded category label feature of the category to which the pixel belongs, determine the non-occluded category label mask prediction value corresponding to the category of the pixel.
10. The method according to claim 8 or 9, characterized in that: For any sample image pair, the true value of the non-occluded category label mask of the sample image pair is determined as follows: Obtaining instance object segmentation results of the sample image pair; Determining a first photometric error between matching pixel pairs in the sample image pair according to the scene flow true value of the sample image pair; Based on the first photometric error and the instance object segmentation result, determining the true value of the non-occluded category label mask corresponding to each instance object in the sample image; Based on the true value of the non-occluded category label mask corresponding to each instance object, the true value of the non-occluded category label mask of the sample image pair is obtained.
11. The method according to claim 10, characterized in that For any sample image pair, the true value of the non-occluded category label mask of the sample image pair is determined as follows: Obtain the true value of the motion field corresponding to the sample image pair; Determine a second photometric error and a depth error between matching pixel pairs in the sample image pair based on a true value of the motion field corresponding to the sample image pair; Determine, based on the second depth error and the depth error, a true value of a non-occluded category label mask corresponding to a non-occluded background region in the sample image pair; By fusing the true value of the non-occluded category label mask corresponding to the non-occluded background area and the true value of the non-occluded category label mask corresponding to each instance object, the true value of the non-occluded category label mask of the sample image pair is obtained.
12. An electronic device, characterized in that: The method comprises at least one processor configured to execute the method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is executed.