Image Processing Method, Apparatus, Computer Device, and Storage Medium
By performing image stitching and multi-scale feature fusion learning on adjacent image frames in the video, the problem of insufficient accuracy of optical flow estimation is solved, and the accuracy of optical flow estimation and the accuracy of video processing tasks are improved.
Patent Information
- Application Number
- CN202110894043.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-08-04
AI Technical Summary
In the prior art, the accuracy of optical flow estimation is insufficient, which affects the accuracy of video processing tasks.
By performing image stitching processing on the target image frame and reference image frame in the video to be processed, combined with multi-scale feature learning, feature fusion learning in the time domain and airspace, the target fusion features are obtained, and optical flow estimation is performed based on this.
Improves the accuracy of optical flow estimation and improves the accuracy of video processing tasks.
Smart Images

Figure CN114299105B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, specifically to the field of computer vision technologies in artificial intelligence, and particularly to an image processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the rapid progress of Internet technologies, artificial intelligence technology, an important branch of Internet technologies, has developed vigorously. And computer vision technology in artificial intelligence technology is the basis for image processing tasks and video processing tasks. Optical flow estimation is a classic problem studied in computer vision technology and is the basis for solving many problems in video processing tasks. It is usually used to study the motion problem between two consecutive image frames adjacent in playback time in a video. It is not difficult to see that accurate optical flow estimation can greatly improve the accuracy of video processing tasks. Therefore, how to improve the accuracy of optical flow estimation has become a current research hotspot. Summary of the Invention
[0003] Embodiments of this application provide an image processing method, apparatus, computer device, and storage medium, which can improve the accuracy of optical flow estimation.
[0004] On the one hand, embodiments of this application provide an image processing method, which includes:
[0005] Obtain a target image frame and a reference image frame of the target image frame from a video to be processed, where the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed;
[0006] Perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image;
[0007] According to the requirements of multi-scale feature learning, perform feature fusion learning in the time domain and spatial domain on the stitched image to obtain a target fusion feature;
[0008] Perform optical flow estimation on the target image frame based on the target fusion feature to obtain target optical flow information of the target image frame.
[0009] On the other hand, embodiments of this application provide an image processing apparatus, which includes:
[0010] An obtaining unit, configured to obtain a target image frame and a reference image frame of the target image frame from a video to be processed, where the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed;
[0011] A processing unit, configured to perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image;
[0012] The processing unit is further configured to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to the requirements of multi-scale feature learning to obtain target fusion features;
[0013] The processing unit is further configured to perform optical flow estimation on the target image frame based on the target fusion features to obtain the target optical flow information of the target image frame.
[0014] In one implementation, when the processing unit is configured to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to the requirements of multi-scale feature learning to obtain target fusion features, the following steps are specifically executed:
[0015] Obtain a feature fusion network, where the feature fusion network includes N feature learning branches, one feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1;
[0016] Call each feature learning branch in the feature fusion network to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale;
[0017] Perform feature fusion processing on the fusion features learned by each feature learning branch to obtain target fusion features.
[0018] In one implementation, the N feature learning branches include a first feature learning branch. When the first feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0019] Perform fusion convolution processing on the stitched image in the time domain and the spatial domain according to the first receptive field to obtain a first convolution feature, where the first receptive field is used to describe the feature learning scale corresponding to the first feature learning branch;
[0020] Perform downsampling processing on the first convolution feature to obtain the fusion feature learned by the first feature learning branch.
[0021] In one implementation, the N feature learning branches include a second feature learning branch. When the second feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0022] Call the first shallow residual learning module to perform fusion residual learning on the stitched image in the time domain and the spatial domain to obtain a first residual feature;
[0023] Perform fusion convolution processing on the first residual feature in the time domain and the spatial domain according to the second receptive field to obtain a second convolution feature; the second receptive field and the receptive field involved in the first shallow residual learning module jointly describe the feature learning scale corresponding to the second feature learning branch;
[0024] Perform downsampling processing based on the second convolutional feature to obtain the fused feature learned by the second feature learning branch.
[0025] In one implementation, the N feature learning branches include a third feature learning branch. When the third feature learning branch performs feature fusion learning in the time domain and spatial domain on the spliced image according to the corresponding feature learning scale, the following steps are specifically executed:
[0026] Call the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain the first residual feature;
[0027] Call the second shallow residual learning module to perform fusion residual learning on the first residual feature in the time domain and spatial domain to obtain the second residual feature; The receptive fields involved in the first shallow residual learning module and the receptive fields involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch;
[0028] Perform downsampling processing based on the second residual feature to obtain the fused feature learned by the third feature learning branch.
[0029] In one implementation, when performing downsampling processing, the downsampling methods adopted by each feature learning branch among the N feature learning branches are different from each other.
[0030] In one implementation, the target fused feature includes feature maps of multiple channels, and the target optical flow information is represented by a vector; The processing unit, when used to perform optical flow estimation on the target image frame based on the target fused feature to obtain the target optical flow information of the target image frame, is specifically used to execute the following steps:
[0031] Upsample the target fused feature according to the image size of the target image frame to obtain the upsampled fused feature;
[0032] Perform dimensionality reduction processing on the number of channels of the upsampled fused feature to obtain the fused feature after dimensionality reduction processing, and the number of channels of the fused feature after dimensionality reduction processing matches the vector dimension of the target optical flow information;
[0033] Perform activation processing on the fused feature after dimensionality reduction processing to obtain the target optical flow information of the target image frame.
[0034] In one implementation, when the processing unit is used to perform dimensionality reduction processing on the number of channels of the upsampled fused feature to obtain the fused feature after dimensionality reduction processing, the following steps are specifically executed:
[0035] Perform feature calibration processing on the upsampled fused feature to obtain the fused feature after feature calibration;
[0036] Perform dimensionality reduction on the fused features after feature calibration to obtain the fused features after dimensionality reduction.
[0037] In one implementation, the processing unit is further configured to perform the following steps:
[0038] Generate an optical flow visualization image of the target image frame based on the target optical flow information and the target image frame;
[0039] Perform image super-resolution processing on the target image frame according to the optical flow visualization image to obtain a super-resolution image of the target image frame, and the resolution of the super-resolution image is higher than that of the target image frame.
[0040] In one implementation, the target fused feature is obtained through a feature fusion network, and the video to be processed is a sample video for training the feature fusion network; when the acquisition unit is used to acquire the target image frame and the reference image frame of the target image frame from the video to be processed, it is specifically configured to perform the following steps:
[0041] Perform scene detection on each image frame in the video to be processed to determine the scene to which each image frame belongs;
[0042] Among multiple image frames in the same scene, acquire the target image frame and the reference image frame of the target image frame; wherein, the target image frame is any image frame other than the first image frame among the multiple image frames.
[0043] In one implementation, the processing unit is further configured to perform the following steps:
[0044] Obtain the marked optical flow information corresponding to the target image frame;
[0045] Determine the loss value of the feature fusion network based on the difference between the target optical flow information and the marked optical flow information of the target image frame;
[0046] Optimize the network parameters of the feature fusion network in the direction of reducing the loss value.
[0047] On the other hand, an embodiment of the present application provides a computer device, which includes:
[0048] A processor, adapted to implement a computer program; and, a computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the above-mentioned image processing method.
[0049] On the other hand, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is read and executed by the processor of the computer device, the computer device is enabled to perform the above-mentioned image processing method.
[0050] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image processing method.
[0051] In an embodiment of the present application, according to the requirements of multi-scale feature learning, feature fusion learning in the time domain and the spatial domain can be performed on a spliced image of two consecutive image frames adjacent in playing time in a video. Then, based on the target fusion features obtained from the feature fusion learning, optical flow estimation can be performed on the image frame with a later playing time among the two consecutive image frames to obtain the optical flow information of the image frame. As can be seen from the above, among the target fusion features learned by the feature fusion learning, on the one hand, the features of the spliced image at multiple scales are fused, and on the other hand, the features of the spliced image in the time domain and the spatial domain are fused. Using the target fusion features that perform multi-dimensional (i.e., multi-scale, time domain dimension, and spatial domain dimension) feature fusion for optical flow estimation can greatly improve the accuracy of optical flow estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a schematic flowchart of an image processing solution provided by an embodiment of the present application;
[0054] Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0055] Figure 3a is a schematic diagram of an upsampling process provided by an embodiment of the present application;
[0056] Figure 3b is a schematic diagram of the architecture of an optical flow estimation model provided by an embodiment of the present application;
[0057] Figure 4 is a schematic flowchart of another image processing method provided by an embodiment of the present application;
[0058] Figure 5a is a schematic diagram of the structure of a feature fusion network provided by an embodiment of the present application;
[0059] Figure 5bIt is a schematic structural diagram of a shallow residual learning module provided by an embodiment of the present application;
[0060] Figure 5c It is a schematic diagram of an optical flow visualization image provided by an embodiment of the present application;
[0061] Figure 5d It is a schematic diagram of an image super-resolution scenario provided by an embodiment of the present application;
[0062] Figure 6 It is a schematic flowchart of another image processing method provided by an embodiment of the present application;
[0063] Figure 7 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application;
[0064] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0066] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. Artificial intelligence software technology mainly includes several major directions such as computer vision (CV) technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0067] Among them, computer vision technology is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses cameras and computers to replace the human eye to identify and measure targets, and further performs graphic processing to make the computer process images that are more suitable for human eye observation or transmission to instrument detection. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (three-dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation and other technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0068] Based on the image processing technology in the above-mentioned computer vision technology, an image processing solution is proposed in the embodiments of the present application to achieve optical flow estimation for two consecutive image frames with adjacent playback times in a video and improve the accuracy of optical flow estimation. In a specific implementation, the image processing solution can be executed by a computer device, which can be a terminal or a server; the terminal mentioned here can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV, etc., but is not limited thereto; the server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0069] To facilitate understanding of the image processing solution proposed in the embodiments of the present application, the following first explains terms such as optical flow estimation and optical flow information involved in the image processing solution:
[0070] Optical flow estimation refers to the motion information technology used to study the associated pixel points between two consecutive image frames adjacent in playback time in a video. Optical flow estimation obtains optical flow information, which can be used to reflect the motion information between the associated pixel points of two consecutive image frames adjacent in playback time in the video. Among them, associated pixel points refer to two pixel points with matching pixel values (for example, the pixel values can be the same) in two consecutive image frames; for example, two consecutive image frames can include a target image frame and a reference image frame, and the reference image frame is the previous image frame of the target image frame. If the pixel value of the reference pixel point in the reference image frame matches the pixel value of the target pixel point in the target image frame, then the reference pixel point and the target pixel point are mutually associated pixel points. More specifically, the optical flow information can include the motion information of each target pixel point in the target image frame relative to the reference pixel point that is mutually associated with each target pixel point in the reference image frame, and the motion information can include the displacement direction and the displacement magnitude.
[0071] The representation methods of optical flow information can include vector representation and color image representation. In the vector representation, the optical flow information can be represented by two-dimensional vectors. The first dimension of the vector represents the displacement magnitude of each target pixel point in the target image frame relative to the reference pixel point that is mutually associated with each target pixel point in the reference image frame in the horizontal direction (i.e., the X-axis direction); the second dimension of the vector represents the displacement magnitude of each target pixel point in the target image frame relative to the reference pixel point that is mutually associated with each target pixel point in the reference image frame in the vertical direction (i.e., the Y-axis direction). In the color image representation, the optical flow information can be represented by a color optical flow image. Different colors in the color optical flow image represent different displacement directions, and the depth of the color represents different displacement magnitudes; for example, the first target pixel point in the target image frame is displayed as dark red in the color optical flow image, and the second target pixel point in the target image frame is displayed as light red in the color optical flow image. This can indicate that the displacement direction of the first target pixel point in the target image frame relative to the associated first reference pixel point in the reference image frame is the same as the displacement direction of the second target pixel point in the target image frame relative to the associated second reference pixel point in the reference image frame, and the displacement magnitude of the first target pixel point in the target image frame relative to the associated first reference pixel point in the reference image frame is different from the displacement magnitude of the second target pixel point in the target image frame relative to the associated second reference pixel point in the reference image frame.
[0072] Based on the above description, the general principle of the image processing solution proposed in the embodiments of the present application will be elaborated below in combination with Figure 1 :
[0073] For any two consecutive image frames adjacent at any playback time in the video (taking the above-mentioned target image frame and reference image frame as an example), the two consecutive image frames can be processed by image stitching to obtain a stitched image; then, multiple (for example, two or more) feature learning branches of the feature fusion network can be called (for example Figure 1 the first feature learning branch, the second feature learning branch, the Nth feature learning branch (N is a positive integer greater than 1), etc. shown), and according to the respective feature learning scales corresponding to each feature learning branch, perform feature fusion learning on the stitched image in the time domain and spatial domain to obtain the fusion features learned by each feature learning branch (for example Figure 1 the first fusion feature, the second fusion feature, the Nth fusion feature, etc. shown), then the fusion features learned by each feature learning branch can be processed by feature fusion to obtain the target fusion feature, so that the optical flow estimation can be performed based on the target fusion feature to obtain the target optical flow information.
[0074] It can be seen that in the feature fusion learning process of this image processing scheme, not only the features of the stitched image in the time domain and spatial domain are fused at each feature learning scale, but also the features of the stitched image at each feature learning scale are fused. The multi-dimensional feature fusion network enables the learned target fusion feature to accurately reflect the features of the stitched image, thus greatly improving the accuracy of the optical flow estimation process based on the target fusion feature.
[0075] Based on the above description, the following will be combined with Figure 2 , Figure 4 and Figure 6 to introduce the image processing scheme provided by the embodiments of the present application in more detail.
[0076] The embodiments of the present application propose an image processing method, and this image processing method can be executed by the aforementioned computer device. In the embodiments of the present application, this image processing method mainly introduces the image stitching process and the optical flow estimation process based on the target fusion feature. As Figure 2 shown, this image processing method may include the following steps S201-S204:
[0077] S201, obtain a target image frame and a reference image frame of the target image frame from the video to be processed.
[0078] Among them, the video to be processed can be any type of video, such as a movie video, a variety show video, a self-media video, a game video, and so on. The so-called movie video refers to a video that is recorded of the performance process of people and / or animals and the surrounding environment in a specified shooting scene according to a pre-produced script, and is later added with audio, special effects, etc.; the variety show video refers to a video that combines multiple art forms and has entertainment; the self-media video refers to a video that ordinary people shoot a certain scene with a camera device and publish it through channels such as the Internet, such as a vlog (video blog); the game video refers to a video that is screen-recorded of the game screen displayed on the terminal screen of any player user during the process of one or more player users playing the target game, or the game screen displayed on the terminal screen of the viewing user who watches the game process of any player user.
[0079] Specifically, the video to be processed may include multiple consecutive image frames with adjacent playing times. When it is necessary to perform optical flow estimation on the image frames in the video to be processed, any two consecutive image frames with adjacent playing times can be obtained from the video to be processed. Among them, any two consecutive image frames with adjacent playing times may include a target image frame and a reference image frame of the target image frame. The target image frame is any image frame in the multiple image frames included in the video to be processed except the first image frame, and the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed.
[0080] S202, perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image.
[0081] After obtaining the target image frame and the reference image frame of the target image frame from the video to be processed, image stitching processing can be performed on the target image frame and the reference image frame to obtain a stitched image; the image stitching processing can include any one of direct image stitching processing or indirect image stitching processing. The following introduces these two image stitching processing methods respectively:
[0082] (1) For the image stitching processing method of direct image stitching, the target image frame can include images of multiple channels, and the reference image frame can include images of multiple channels. The number of image channels included in the target image frame is the same as the number of image channels included in the reference image frame. The images of multiple channels included in the target image frame can be directly stitched with the images of multiple channels included in the reference image frame according to the image channel dimension to obtain a stitched image. The number of image channels included in the stitched image is equal to the sum of the number of image channels included in the target image frame and the number of image channels included in the reference image frame. Specifically, an image stitching network (abbreviated as Concat) can be obtained, and the image stitching network can be called to stitch the images of multiple channels included in the target image frame with the images of multiple channels included in the reference image frame according to the image channel dimension to obtain a stitched image; among them, the image stitching network can use the concatenate function to implement the image stitching processing.
[0083] For example, when the color mode of the target image frame and the reference image frame is the RGB (Red, Green, Blue) mode, the target image frame includes images of 3 channels, namely the image of the R channel (i.e., the image of the red channel), the image of the G channel (i.e., the image of the green channel), and the image of the B channel (i.e., the image of the blue channel); the reference image frame includes images of 3 channels, namely the image of the R channel, the image of the G channel, and the image of the B channel. The image of the R channel in the target image frame can be stitched with the image of the R channel in the reference image frame to obtain two images of the R channel included in the stitched image; similarly, the image of the G channel in the target image frame can be stitched with the image of the G channel in the reference image frame to obtain two images of the G channel included in the stitched image, and the image of the B channel in the target image frame can be stitched with the image of the B channel in the reference image frame to obtain two images of the B channel included in the stitched image; that is to say, the stitched image can include images of 6 channels, namely two images of the R channel, two images of the G channel, and two images of the B channel.
[0084] Compared with the image stitching processing method of direct image stitching processing, the difference between the image stitching processing method of indirect image stitching processing and the image stitching processing method of direct image stitching processing lies in that: before stitching the images of multiple channels included in the target image frame and the images of multiple channels included in the reference image frame according to the image channel dimension to obtain a stitched image, it is necessary to perform image normalization processing on the images of each channel included in the target image frame to obtain the normalized images of each channel in the target image frame, and it is necessary to perform image normalization processing on the images of each channel included in the reference image frame to obtain the normalized images of each channel in the reference image frame. Then, an image stitching network can be obtained, and the image stitching network can be called to stitch the normalized images of multiple channels in the target image frame and the normalized images of multiple channels in the reference image frame according to the image channel dimension to obtain a stitched image. Among them, performing image normalization processing on the image of any channel means: performing normalization processing on the pixel values of each pixel point in the image of this channel, that is, mapping the pixel values of each pixel point to a preset interval (such as the preset interval [0, 1], the preset interval [-1, 1], etc.), so that the pixel values of each pixel point after normalization jointly form the normalized image of this channel. For example, the value range of the pixel values of each pixel point in the image of this channel is [0, 255], and the pixel values of each pixel point are normalized to the preset interval [0, 1] through normalization processing.
[0085] S203, According to the requirements of multi-scale feature learning, perform feature fusion learning on the stitched image in the time domain and spatial domain to obtain the target fusion feature.
[0086] After performing image stitching processing on the target image frame and the reference image frame to obtain a stitched image, according to the requirements of multi-scale feature learning, feature fusion learning can be performed on the stitched image in the time domain and spatial domain to obtain the target fusion feature. Specifically, a feature fusion network (abbreviated as FFB) can be obtained. The feature fusion network can include N feature learning branches. One feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1. Then, each feature learning branch in the feature fusion network can be called to perform feature fusion learning on the stitched image in the time domain and spatial domain according to the corresponding feature learning scale. After that, the feature fusion network can be called to perform feature fusion processing on the fusion features learned by each feature learning branch to obtain the target fusion feature.
[0087] It should be noted that before performing feature fusion learning in the time domain and spatial domain on the spliced image according to the requirements of multi-scale feature learning to obtain the target fusion feature, the spliced image can be first subjected to feature extraction processing to obtain the preliminary features of the spliced image, and then, according to the requirements of multi-scale feature learning, the preliminary features of the spliced image are subjected to feature fusion learning in the time domain and spatial domain to obtain the target fusion feature. Specifically, a feature extraction network (abbreviated as InputBlock) can be obtained, and the feature extraction network is called to perform feature learning on the spliced image to obtain the preliminary features of the spliced image; then, a feature fusion network can be obtained, and each feature learning branch in the feature fusion network is called to perform feature fusion learning in the time domain and spatial domain on the preliminary features of the spliced image according to the corresponding feature learning scale, and feature fusion processing can be performed on the fusion features learned by each feature learning branch to obtain the target fusion feature.
[0088] Among them, the feature extraction network can be formed by stacking one or more convolutional layers and activation layers in a cycle; in other words, the feature extraction network can include one or more groups of convolutional sub-networks; one or more groups of convolutional sub-networks are connected in series. Among one or more groups of convolutional sub-networks: the output end of the first group of convolutional sub-networks is connected to the input end of the second group of convolutional sub-networks, the output end of the second group of convolutional sub-networks is connected to the input end of the third group of convolutional sub-networks, and so on. The output end of the penultimate group of convolutional sub-networks is connected to the input end of the last group of convolutional sub-networks. Each group of convolutional sub-networks can include an activation layer and one or more convolutional layers, and the activation layer and one or more convolutional layers are connected in series; in any group of convolutional sub-networks: the output end of the first convolutional layer is connected to the input end of the second convolutional layer, the output end of the second convolutional layer is connected to the input end of the third convolutional layer, and so on. The output end of the penultimate convolutional layer is connected to the input end of the last convolutional layer, and the output end of the last convolutional layer is connected to the input end of the activation layer. Among them, the convolutional layer (Convolutional Layer) is composed of several convolutional units and can be used to extract different features of the input. The activation layer can be used to enhance the nonlinear characteristics of the decision function and the entire network. It uses an activation function to normalize the feature values of each unit in the feature map to a specified interval (such as the specified interval (0, 1), the specified interval (-1, 1), etc.). The activation function can include the ReLU (Rectified Linear Unit) function, the LReLU (Leaky ReLU) function, the Tanh (hyperbolic tangent) function, the Sigmoid function, etc. In this embodiment of the application, the activation function adopted by the activation layer of the feature extraction network is taken as an example of the ReLU function for illustration, but this does not limit this embodiment of the application. In an actual feature extraction scenario, the activation function adopted by the activation layer in the feature extraction network can also be the Tanh function, the Sigmoid function, etc.
[0089] Based on the above description of the structure of the feature extraction network, taking the feature extraction network including a group of convolutional sub-networks, where the convolutional sub-network includes a convolutional layer and an activation layer as an example, the process of using the feature extraction network to perform feature extraction processing on the spliced image can include: using the convolutional layer of the feature extraction network to perform fusion convolution processing on the spliced image in the time domain and the spatial domain to obtain the fusion convolution features of the spliced image, and using the activation layer of the feature extraction network to perform activation processing on the fusion convolution features to obtain the preliminary features of the spliced image.
[0090] S204, perform optical flow estimation on the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame.
[0091] After performing feature fusion learning in the time domain and spatial domain on the spliced image according to the requirements of multi-scale feature learning to obtain the target fusion feature, the optical flow of the target image frame can be estimated based on the target fusion feature to obtain the target optical flow information of the target image frame. Specifically, the process of estimating the optical flow of the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame can include the following steps:
[0092] (1) Upsample the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature. Specifically, an upsampling network can be obtained, and the upsampling network can be called to upsample the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature.
[0093] In one implementation, the upsampling network may include an upsampling layer. The upsampling layer can be used to upsample the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature. The upsampled fusion feature may include feature maps of multiple channels, and the feature map sizes of the feature maps of each channel included in the upsampled fusion feature match (e.g., are the same as) the image size of the target image frame. Specifically, the process of using the upsampling layer to upsample the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature can include: ① The target fusion feature may include feature maps of multiple channels. Obtain the number of channels of the feature maps in the target fusion feature, and obtain the feature map size of the feature maps in the target fusion feature (including the width parameter before upsampling and the height parameter before upsampling). ② According to the number of channels of the feature maps in the target fusion feature, the feature map size of the feature maps in the target fusion feature, and the image size of the target image frame (including the width parameter of the target image frame and the height parameter of the target image frame), determine the number of channels of the feature maps in the upsampled fusion feature. ③ According to the number of channels of the feature maps in the target fusion feature, the feature map size of the feature maps in the target fusion feature, the image size of the target image frame, and the number of channels of the feature maps in the upsampled fusion feature, transform the features in the channel dimension of the target fusion feature into the spatial dimension of the upsampled fusion feature. The spatial dimension can be determined by the width parameter of the target image frame and the height parameter of the target image frame. For example, the spatial dimension may be equal to the product of the width parameter of the target image frame and the height parameter of the target image frame.
[0094] Among them, the process of determining the number of channels of the feature map in the fused feature after upsampling according to the number of channels of the feature map in the target fused feature, the size of the feature map in the target fused feature, and the size of the target image frame may include: determining the upsampling ratio according to the size of the feature map in the target fused feature and the size of the target image frame, and determining the number of channels of the feature map in the fused feature after upsampling based on the determined upsampling ratio and the number of channels of the feature map in the target fused feature. Specifically, the upsampling ratio is equal to the ratio between the size of the target image frame and the size of the feature map in the target fused feature. The size of the feature map in the target fused feature is equal to the product of the width parameter and the height parameter of the feature map in the target fused feature. The size of the target image frame is equal to the product of the width parameter and the height parameter of the target image frame. The number of channels of the feature map in the fused feature after upsampling is equal to the ratio between the number of channels of the feature map in the target fused feature and the upsampling ratio.
[0095] For example, the upsampling layer may implement upsampling using the DepthToSpace algorithm. The process of implementing upsampling using the DepthToSpace algorithm is as Figure 3a shown. The target fused feature includes a feature map with 4 channels, and the size of the feature map of each channel is 2×2 (i.e., the width parameter before upsampling is 2, and the height parameter before upsampling is 2). Here, it is necessary to upsample the target fused feature to obtain a fused feature after upsampling with a feature map size of 4×4. It can be calculated that the upsampling ratio is (4×4) / (2×2) = 4 times, and the number of channels of the feature map in the fused feature after upsampling is 4 / 4 = 1, that is, the fused feature after upsampling includes a feature map with 1 channel, and the size of the feature map of this 1 channel is 4×4. Thus, the features in the channel dimension of the target fused feature can be transformed into the spatial dimension of the fused feature after upsampling. The feature map of the target fused feature before transformation can be seen in Figure 3a the left schematic diagram, and the transformation result can be seen in Figure 3a the right schematic diagram.
[0096] In another implementation, the upsampling network may include a convolutional layer and an upsampling layer, and the target fusion feature may include feature maps of multiple channels. The process of performing upsampling processing on the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature may include: the convolutional layer of the upsampling network may be used to perform dimensionality reduction processing on the number of channels of the target fusion feature to obtain a reference fusion feature, and the number of channels of the feature maps included in the reference fusion feature is less than the number of channels of the feature maps included in the target fusion feature; then, the upsampling layer of the upsampling network may be used to perform upsampling processing on the reference fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature, and the feature map sizes of the feature maps of each channel included in the upsampled fusion feature match (e.g., are the same as) the image size of the target image frame. This process is similar to the process of using the upsampling layer of the upsampling network to perform upsampling processing on the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature. For details, refer to the above upsampling process of the target fusion feature and will not be elaborated here.
[0097] (2) Perform dimensionality reduction processing on the number of channels of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing. Specifically, a channel dimensionality reduction network may be obtained, and the channel dimensionality reduction network may be called to perform dimensionality reduction processing on the number of channels of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing; wherein, the channel dimensionality reduction network includes a convolutional layer, that is, the channel dimensionality reduction network realizes dimensionality reduction processing of the number of channels through the convolutional layer; the target optical flow information may be represented by a vector, and the number of channels of the fusion feature after dimensionality reduction processing matches (e.g., is the same as) the vector dimension of the target optical flow vector. For example, the target optical flow information is represented by a two-dimensional vector, and the fusion feature after dimensionality reduction processing includes feature maps of 2 channels, that is, the number of channels of the fusion feature after dimensionality reduction processing is 2.
[0098] It should be noted that before performing dimensionality reduction processing on the number of channels of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing, feature calibration processing may also be performed on the upsampled fusion feature to obtain the fusion feature after feature calibration processing. Specifically, a feature calibration network may be obtained, and the feature calibration network may be called to perform feature calibration processing on the upsampled fusion feature to obtain the fusion feature after feature calibration processing, that is, the feature calibration network may be used to perform feature calibration processing on the upsampled fusion feature first to obtain the fusion feature after feature calibration processing, and then the channel dimensionality reduction network may be used to perform dimensionality reduction processing on the number of channels of the fusion feature after feature calibration processing to obtain the fusion feature after dimensionality reduction processing.
[0099] Specifically, the feature calibration network may include a convolutional layer and an activation layer. The process of using the feature calibration network to perform feature calibration on the fused features after upsampling to obtain the fused features after feature calibration may include: using the convolutional layer of the feature calibration network to perform fused convolution on the fused features after upsampling to obtain intermediate fused features; using the activation layer of the feature calibration network to perform activation on the intermediate fused features to obtain the fused features after feature calibration. In the embodiments of the present application, the activation function used by the activation layer of the feature calibration network is taken as an example of the LReLU function. Other activation functions may also be used in the activation layer involved in the feature calibration network, such as the ReLU function, the Tanh function, and so on. Since the above upsampling process is a transformation of the features in the channel dimension to the spatial dimension in the target fused features, this will cause the fused features after upsampling to not accurately describe the features of the stitched image. Through the feature calibration network, the fused features after upsampling can be calibrated, so that the fused features after feature calibration can more accurately describe the features of the stitched image, improving the accuracy of the fused features after feature calibration, and further improving the accuracy of optical flow estimation.
[0100] (3) Perform activation processing on the fused features after dimensionality reduction to obtain the target optical flow information of the target image frame.
[0101] After performing dimensionality reduction on the fused features after upsampling to obtain the fused features after dimensionality reduction, activation processing can be performed on the fused features after dimensionality reduction to obtain the target optical flow information of the target image frame. Specifically, a feature activation network can be obtained. The feature activation network includes an activation layer, and the activation layer of the feature activation network is called to perform activation processing on the fused features after dimensionality reduction to obtain the target optical flow information of the target image frame. In the embodiments of the present application, the activation function used by the activation layer of the feature activation network is taken as an example of the Tanh function. Other activation functions may also be used in the activation layer involved in the feature activation network, such as the ReLU function, the LReLU function, and so on.
[0102] It should be noted that the image stitching network, feature extraction network, feature fusion network, upsampling network, feature calibration network, channel dimensionality reduction network, and feature activation network mentioned in the above steps S201 to S204 may be integrated into different models respectively. For example, the image stitching network is integrated into the image stitching model, and the feature fusion network is integrated into the feature fusion model, and so on. Or, the image stitching network, feature extraction network, feature fusion network, upsampling network, feature calibration network, channel dimensionality reduction network, and feature activation network mentioned in the above steps S201 to S204 may be integrated into the same model, such as being integrated into the optical flow estimation model. Figure 3b Taking the above networks as being integrated into the optical flow estimation network as an example for introduction, asFigure 3b As shown in Figure 3b , the optical flow estimation model 30 includes an image stitching network 301, a feature extraction network 302, a feature fusion network 303, an upsampling network 304, a feature calibration network 305, a channel dimensionality reduction network 306, and a feature activation network 307. The output end of the image stitching network 301 is connected to the input end of the feature extraction network 302, the output end of the feature extraction network 302 is connected to the input end of the feature fusion network 303, the output end of the feature fusion network 303 is connected to the input end of the upsampling network 304, the output end of the upsampling network 304 is connected to the input end of the feature calibration network 305, the output end of the feature calibration network 305 is connected to the input end of the channel dimensionality reduction network 306, and the output end of the channel dimensionality reduction network 306 is connected to the input end of the feature activation network 307. The input of the optical flow estimation model 30 is two consecutive image frames adjacent in playback time in the video to be processed (i.e., Figure 3b the target image frame and the reference image frame shown in Figure 3b ), and the output of the optical flow estimation model 30 is the target optical flow information of the target image frame.
[0103] It should be noted that the steps S201 - S204 mentioned in the embodiments of the present application can be executed during the network training and optimization process of the feature fusion network, or can be executed during the actual application process of the feature fusion network. For the actual application process of the feature fusion network, the execution process of step S201 can be referred to the foregoing content; for the training and optimization process of the feature fusion network, the execution process of step S201 can be referred to the foregoing content, and the computer device can also obtain the target image frame and the reference image frame from the video to be processed in combination with scene detection.
[0104] In the embodiments of the present application, during the process of stitching the target image frame and the reference image frame, image normalization processing can be performed on the images of each channel in the target image frame and the reference image frame, and then the normalized images of each channel are stitched. The image normalization processing maps larger pixel values in the image (for example, pixel values belonging to the interval [0, 255]) to smaller values within a preset interval (for example, the preset interval can be [0, 1]). This can greatly reduce the data volume involved in the image stitching process, as well as the subsequent feature fusion process and optical flow estimation process, thereby improving the efficiency of optical flow estimation. In addition, after upsampling the target fusion features output by the feature fusion network of the optical flow estimation model, feature calibration processing can be performed on the upsampled fusion features, so that the fusion features after feature calibration can more accurately describe the features of the stitched image, thereby improving the accuracy of optical flow estimation.
[0105] The embodiments of the present application also propose an image processing method, which can be executed by the aforementioned computer device. In the embodiments of the present application, this image processing method mainly introduces the structure of the feature fusion network and the feature fusion process. As Figure 4 shown, this image processing method may include the following steps S401 - S406:
[0106] S401, Obtain a target image frame and a reference image frame of the target image frame from the video to be processed.
[0107] In the embodiments of the present application, a feature fusion network is involved; the steps S401 - S406 mentioned in the embodiments of the present application can be executed during the network training and optimization process of the feature fusion network, or can be executed during the actual application process of the feature fusion network. When steps S401 - S406 are executed during the network optimization process of the feature fusion network, the video to be processed can be understood as a sample video for training the feature fusion network; when steps S401 - S406 are executed during the actual application process of the feature fusion network, the video to be processed can be understood as a service video that needs to perform service processing through optical flow estimation, and this is not limited.
[0108] In a specific implementation, regardless of whether the video to be processed is a sample video or a service video, when the computer device executes step S401, it can select any image frame from the remaining image frames of the video to be processed except the first image frame as the target image frame, and use the previous image frame of the selected image frame as the reference image frame. In another specific implementation, when the video to be processed is a sample video, the computer device can also combine scene detection to obtain the target image frame and the reference image frame from the video to be processed, so that the obtained target image frame and reference image frame belong to the same scene, to avoid large differences between the target image frame and the reference image frame caused by scene switching, which affects the accuracy of the subsequent obtained target optical flow information, and further avoids affecting the training and optimization effect of the network.
[0109] The execution process of step S401 in the embodiments of the present application is the same as the execution process of step S201 in the Figure 2 embodiment shown above. For the specific execution process, reference can be made to the specific description of step S201 in the Figure 2 embodiment shown above, and details are not described here again.
[0110] S402, Perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image.
[0111] The execution process of step S402 in the embodiments of the present application is the same as the execution process of step S202 in the Figure 2 embodiment shown above. For the specific execution process, reference can be made to the specific description of step S202 in the Figure 2The specific description of step S202 in the illustrated embodiment will not be elaborated here.
[0112] S403. Obtain a feature fusion network.
[0113] The feature fusion network may include N feature learning branches and a feature fusion layer. One feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1. Each feature learning branch may be used to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to its corresponding feature learning scale. The feature fusion layer may be used to perform feature fusion processing on the fusion features learned by each feature learning branch to obtain target fusion features. Figure 5a FIG. is a schematic structural diagram of a feature fusion network provided by an embodiment of the present application. The feature fusion network 50 includes 3 feature learning branches and a feature fusion layer 504. The 3 feature learning branches are respectively a first feature learning branch 501, a second feature learning branch 502, and a third feature learning branch 503.
[0114] S404. Invoke each feature learning branch in the feature fusion network to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale.
[0115] Among the N feature learning branches of the feature fusion network, there may be a first feature learning branch. The first feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, which may include: performing fusion convolution processing on the stitched image in the time domain and the spatial domain according to the first receptive field to obtain a first convolution feature, where the first receptive field is used to describe the feature learning scale corresponding to the first feature learning branch; performing downsampling processing based on the first convolution feature to obtain the fusion feature learned by the first feature learning branch.
[0116] Specifically, referring to Figure 5a the shown first feature learning branch 501, the first feature learning branch 501 may include a convolutional layer, an activation layer, and a pooling layer; the output end of the convolutional layer is connected to the input end of the activation layer, and the output end of the activation layer is connected to the input end of the pooling layer. Among them:
[0117] ① The process of performing fusion convolution processing on the spliced image in the time domain and spatial domain according to the first receptive field to obtain the first convolution feature can be implemented by the convolution layer of the first feature learning branch 501, that is, the convolution layer of the first feature learning branch 501 can be called to perform fusion convolution processing on the spliced image in the time domain and spatial domain according to the first receptive field to obtain the first convolution feature; the first receptive field refers to the size of the eigenvalue region of the feature values of each unit in the feature map used to calculate the first convolution feature in the feature map of the spliced image, and the first receptive field is determined according to the convolution kernel size of the convolution layer of the first feature learning branch 501; for example, if the convolution kernel size of the convolution layer of the first feature learning branch 501 is 3×3, then the first receptive field is 3×3, and the size of the eigenvalue region of the feature values of each unit in the feature map used to calculate the first convolution feature in the feature map of the spliced image is 3×3.
[0118] ② The process of performing downsampling processing on the basis of the first convolution feature to obtain the fusion feature learned by the first feature learning branch can be implemented by the activation layer and pooling layer of the first feature learning branch 501. The activation layer of the first feature learning branch 501 can be called to perform activation processing on the first convolution feature to obtain the first activation feature, and the pooling layer of the first feature learning branch 501 can be called to perform downsampling processing on the first activation feature to obtain the fusion feature learned by the first feature learning branch.
[0119] Among the N feature learning branches of the feature fusion network, there may be a second feature learning branch. The second feature learning branch performs feature fusion learning on the spliced image in the time domain and spatial domain according to the corresponding feature learning scale, which may include: calling the first shallow residual learning (abbreviated as SRB1) module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain the first residual feature; performing fusion convolution processing on the first residual feature in the time domain and spatial domain according to the second receptive field to obtain the second convolution feature; the second receptive field and the receptive field involved in the first shallow residual learning module jointly describe the feature learning scale corresponding to the second feature learning branch; performing downsampling processing on the basis of the second convolution feature to obtain the fusion feature learned by the second feature learning branch.
[0120] Specifically, referring to Figure 5a the second feature learning branch 502 shown, the second feature learning branch 502 may include a first shallow residual learning module, a convolution layer, an activation layer, and a pooling layer; the output end of the first shallow residual learning module is connected to the input end of the convolution layer, the output end of the convolution layer is connected to the input end of the activation layer, and the output end of the activation layer is connected to the input end of the pooling layer. Among them:
[0121] ① The structure of the first shallow residual learning module can be referred to Figure 5b Figure 5bIt is a schematic structural diagram of a shallow residual learning module provided by an embodiment of the present application. The first shallow residual learning module may include a convolutional layer and an activation layer. It should be noted that the so-called shallow residual learning module can be understood as a residual learning module that contains a convolutional layer with a number less than or equal to a threshold number. Correspondingly, the so-called deep residual learning module can be understood as a residual learning module that contains a convolutional layer with a number greater than the threshold number. For Figure 5b For the first shallow residual learning module shown, the process of calling the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and the spatial domain to obtain the first residual feature may include: calling the convolutional layer of the first shallow residual learning module to perform fusion convolutional processing on the spliced image in the time domain and the spatial domain to obtain the residual convolutional feature of the first shallow residual module; performing fusion processing on the spliced image and the residual convolutional feature of the first shallow residual module to obtain the residual fusion feature of the first shallow residual learning module; calling the activation layer of the first shallow residual module to perform activation processing on the residual fusion feature of the first shallow residual module to obtain the first residual feature.
[0122] ② The process of performing fusion convolutional processing on the first residual feature in the time domain and the spatial domain according to the second receptive field to obtain the second convolutional feature can be implemented by the convolutional layer of the second feature learning branch 502, that is, the convolutional layer of the second feature learning branch 502 can be called to perform fusion convolutional processing on the first residual feature in the time domain and the spatial domain according to the second receptive field to obtain the second convolutional feature. The second receptive field refers to the size of the eigenvalue region of the feature values of each unit in the feature map for calculating the second convolutional feature in the feature map of the first residual feature. The second receptive field is determined according to the convolutional kernel size of the convolutional layer of the second feature learning branch 501. It should be noted that the second receptive field and the receptive field involved in the first shallow residual learning module together describe the feature learning scale corresponding to the second feature learning branch, where the receptive field involved in the first shallow residual learning module is determined according to the convolutional kernel size of the convolutional layer in the first shallow residual learning module.
[0123] ③ The process of performing downsampling processing on the second convolutional feature to obtain the fusion feature learned by the second feature learning branch can be implemented by the activation layer and the pooling layer of the second feature learning branch 502. The activation layer of the second feature learning branch 502 can be called to perform activation processing on the second convolutional feature to obtain the second activation feature, and the pooling layer of the second feature learning branch 502 can be called to perform downsampling processing on the second activation feature to obtain the fusion feature learned by the second feature learning branch.
[0124] Among the N feature learning branches of the feature fusion network, a third feature learning branch may be included. The third feature learning branch performs feature fusion learning in the time domain and spatial domain on the spliced image according to the corresponding feature learning scale, and may include: calling a first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain a first residual feature; calling a second shallow residual learning (abbreviated as SRB2) module to perform fusion residual learning on the first residual feature in the time domain and spatial domain to obtain a second residual feature; the receptive fields involved in the first shallow residual learning module and the receptive fields involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch; performing downsampling processing on the second residual feature to obtain the fusion feature learned by the third feature learning branch.
[0125] Specifically, referring to Figure 5a the third feature learning branch 503 shown in, the third feature learning branch 503 may include a first shallow residual learning module, a second shallow residual learning module, a convolutional layer, and an activation layer; the output end of the first shallow residual learning module is connected to the input end of the second shallow residual learning module, the output end of the second shallow residual learning module is connected to the input end of the convolutional layer, and the output end of the convolutional layer is connected to the input end of the activation layer. Among them:
[0126] ① The structure of the first shallow residual learning module can be referred to Figure 5b , and the process of calling the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain a first residual feature may include: calling the convolutional layer of the first shallow residual learning module to perform fusion convolutional processing on the spliced image in the time domain and spatial domain to obtain the residual convolutional feature of the first shallow residual module; performing fusion processing on the spliced image and the residual convolutional feature of the first shallow residual module to obtain the residual fusion feature of the first shallow residual learning module; calling the activation layer of the first shallow residual module to perform activation processing on the residual fusion feature of the first shallow residual module to obtain a first residual feature.
[0127] ② The structure of the second shallow residual learning module is similar to that of the first shallow residual learning module, and can be referred to Figure 5bThe structure of the first shallow residual learning module; the process of calling the second shallow residual learning module to perform fused residual learning on the first residual feature in the time domain and the spatial domain to obtain the second residual feature may include: calling the convolutional layer of the second shallow residual learning module to perform fused convolutional processing on the first residual feature in the time domain and the spatial domain to obtain the residual convolutional feature of the second shallow residual module; performing fusion processing on the first residual feature and the residual convolutional feature of the second shallow residual module to obtain the residual fusion feature of the second shallow residual learning module; calling the activation layer of the second shallow residual module to perform activation processing on the residual fusion feature of the second shallow residual module to obtain the second residual feature. It should be noted that the receptive fields involved in the first shallow residual learning module and the receptive fields involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch, wherein, the receptive field involved in the second shallow residual learning module is determined according to the convolutional kernel size of the convolutional layer in the second shallow residual learning module.
[0128] ③ The process of performing downsampling processing on the second residual feature to obtain the fused feature learned by the third feature learning branch can be implemented by the activation layer and the convolutional layer of the third feature learning branch 503. The convolutional layer of the third feature learning branch 503 can be called to perform downsampling processing on the second residual feature to obtain the residual feature after downsampling processing, and the activation layer of the third feature learning branch 503 can be called to perform activation processing on the residual feature after downsampling processing to obtain the fused feature learned by the third feature learning branch.
[0129] S405, perform feature fusion processing on the fused features learned by each feature learning branch to obtain the target fused feature.
[0130] The process of performing feature fusion processing on the fused features learned by each feature learning branch to obtain the target fused feature can be implemented by the feature fusion layer in the feature fusion network (for example Figure 5aIt is implemented by the feature fusion layer 504) in the shown feature fusion network 50. That is to say, the feature fusion layer in the feature fusion network can be called to perform feature fusion processing on the fusion features learned by each feature learning branch to obtain the target fusion feature. Specifically, the fusion features learned by each feature learning branch all include feature maps of multiple channels, and the number of feature map channels included in the fusion features learned by each feature learning branch is the same. Calling the feature fusion layer in the feature fusion network to perform feature fusion processing on the fusion features learned by each feature learning branch means: splicing the feature maps in the fusion features learned by each feature learning branch according to the channel dimension to obtain the target fusion feature; where the number of feature map channels included in the target fusion feature is equal to the sum of the number of feature map channels included in the fusion features learned by each feature learning branch. For example, the fusion feature learned by the first feature learning branch includes feature maps of 50 channels, the fusion feature learned by the second feature learning branch includes feature maps of 50 channels, and the fusion feature learned by the third feature learning branch includes feature maps of 50 channels. Then, the target fusion feature further fused from the fusion features learned by the three feature learning branches includes feature maps of 150 channels.
[0131] It should be noted that Figure 5a the structure of the shown feature fusion network is only for illustration. In actual application scenarios, the structure of the feature fusion network can also be in other forms. For example, the feature fusion network can include 4 feature learning branches, 7 feature learning branches, and so on. In addition, Figure 5a the sharing of the first shallow residual learning module by the shown second feature learning branch 502 and third feature learning branch 503 is only for illustration. In actual application scenarios, the second feature learning branch 502 and the third feature learning branch 503 can each use a shallow residual learning module. The first shallow residual learning module and the second shallow residual learning module can be the same module or different modules. That the first shallow residual learning module and the second shallow residual learning module are the same module means: the number of convolutional layers used by the first shallow residual learning module and the second shallow residual learning module is the same, the convolutional kernels of the convolutional layers are the same, and the activation functions of the activation layers used are the same. That the first shallow residual learning module and the second shallow residual learning module are different modules means: the number of convolutional layers used by the first shallow residual learning module and the second shallow residual learning module is different, or the convolutional kernels of the convolutional layers are different, or the activation functions of the activation layers used are different, and so on. In the embodiments of the present application, the activation function used by the activation layer involved in the feature fusion network is taken as an example of the LReLU function for illustration. The activation layer involved in the feature fusion network can also use other activation functions, such as the ReLU function, the Tanh function, and so on.
[0132] It should also be noted that when each feature learning branch among the N feature learning branches performs downsampling processing, the downsampling methods adopted are different from each other. For example Figure 5a in the feature fusion network shown, the downsampling method adopted by the first feature learning branch 501 is to perform downsampling processing using a max pooling layer, the downsampling method adopted by the second feature learning branch is to perform downsampling processing using an average pooling layer, and the downsampling method adopted by the third feature learning branch is to perform downsampling processing using a convolutional layer. Through downsampling processing, the position offset caused by pixel movement between the target image frame and the reference image frame can be eliminated, facilitating the alignment of the fusion features learned by each feature learning branch when the fusion features are stitched together. Additionally, by setting different downsampling methods, both large displacements and small displacements of pixel points between the target image frame and the reference image frame can be taken into account, further improving the accuracy of the feature fusion process, thereby improving the accuracy of optical flow estimation.
[0133] S406, perform optical flow estimation on the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame.
[0134] In the embodiment of the present application, the execution process of step S406 is the same as the execution process of step S204 in the above Figure 2 shown embodiment. The specific execution process can refer to the description of step S204 in the above Figure 2 shown embodiment and will not be elaborated here.
[0135] It should be noted that if the aforementioned video to be processed is a service video, after performing optical flow estimation on the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame, the computer device can also generate an optical flow visualization image of the target image frame based on the target optical flow information and the target image frame. Among them, the optical flow visualization image is similar to a color optical flow image. In the optical flow visualization image, different colors in the optical flow information represent different displacement directions, and the depth of the color represents different displacement magnitudes. Figure 5c is a schematic diagram of an optical flow visualization image provided by an embodiment of the present application. As Figure 5c shown, the target image frame 505 and the reference image frame 506 undergo optical flow estimation to obtain the target optical flow information. Based on the target optical flow information and the target image frame 505, an optical flow visualization image 507 can be generated. From the optical flow visualization image 507, it can be seen that the person in the target image frame 505 has moved relative to the person in the reference image frame 506, and the building in the target image frame 505 is stationary relative to the building in the reference image frame 506.
[0136] Among them, the target optical flow information may include the motion information of each target pixel point in the target image frame relative to the associated reference pixel point in the reference image frame. The motion information may include the displacement direction and the displacement magnitude. The process of generating the optical flow visualization image of the target image frame based on the target optical flow information and the target image frame may include: performing normalization processing on the pixel values of each pixel point in the target image frame to obtain a normalized image, performing offset transformation processing on each pixel point in the normalized image according to the displacement direction and the displacement magnitude indicated by the target optical flow information to obtain a transformed image of the target image frame, and performing color rendering on the transformed image of the target image frame to obtain the optical flow visualization image. It should be noted that the above process of generating the optical flow visualization image of the target image frame based on the target optical flow information and the target image frame can be specifically implemented by using the dense_image_warp algorithm in TensorFlow. TensorFlow is a symbolic mathematics system based on data flow transformation and is widely used in the programming implementation of various machine learning algorithms. Dense_image_warp is an image affine transformation algorithm.
[0137] The optical flow visualization image can be applied to scenarios such as image super-resolution, motion segmentation, and motion estimation. In the image super-resolution scenario, the target image frame can be subjected to image super-resolution processing based on the optical flow visualization image to obtain a super-resolution image of the target image frame. The resolution of the super-resolution image is higher than that of the target image frame. As Figure 5d shown, Figure 5d FIG. is a schematic diagram of an image super-resolution scenario provided by an embodiment of the present application. Figure 5d In the figure, the resolution of the super-resolution image 509 is higher than the resolution 508 of the target image frame. The image resolution is related to the clarity of the image. The higher the resolution of the image, the clearer the image, and the lower the resolution of the image, the more blurred the image. In the motion segmentation scenario, the moving object (for example, the person in the optical flow visualization image shown in Figure 5c FIG., which may be other animals, cars, etc. that generate motion) existing in the target image frame can be determined based on the optical flow visualization image, so that an image block containing the moving object can be segmented from the target image frame. In the motion estimation scenario, the moving object existing in the target image frame and the reference image frame can be determined based on the optical flow visualization image. Thus, a first image block containing the moving object can be segmented from the target image frame, and a second image block containing the moving object can be segmented from the reference image frame. By performing matching analysis on the first image block and the second image block, the relative displacement of the moving object in the first image block and the moving object in the second image block can be determined.
[0138] In the embodiments of the present application, a feature fusion network can be used to perform multi-scale feature fusion learning on the spliced image. For any feature learning branch in the feature fusion network, the fused feature learned by the feature learning branch fuses the features of the spliced image in the time domain and the spatial domain. For each feature learning branch in the feature fusion network, the target fused feature output by the feature fusion network fuses the fused features learned by each feature learning branch at their respective corresponding feature learning scales. Using a feature fusion network with multiple dimensions (i.e., the above-mentioned time domain and spatial domain dimensions, as well as the feature learning scale dimension) to learn the features of the spliced image can improve the accuracy of the learned target fused feature, thereby improving the accuracy of optical flow estimation. In addition, different downsampling methods can be used for downsampling processing in each feature learning branch of the feature fusion network, which can take into account both large displacements and small displacements between pixel points in the target image frame and the reference image frame, and can effectively eliminate the position offset caused by pixel point movement between the target image frame and the reference image frame, enabling the fused features learned by each feature learning branch to be aligned, which helps to further fuse the fused features learned by each feature learning branch and further improve the accuracy of optical flow estimation.
[0139] Based on the above description, embodiments of the present application also propose an image processing method, which can be executed by the computer device mentioned above. In the embodiments of the present application, the image processing method is mainly described by taking the video to be processed as a sample video for training the feature fusion network as an example.
[0140] As Figure 6 shown, the image processing method may include the following steps S601-S607:
[0141] S601, obtain a target image frame and a reference image frame of the target image frame from the video to be processed.
[0142] Among them, the target image frame and the reference image frame are two consecutive image frames in the video to be processed with adjacent playing times and matching scenes. The process of obtaining the target image frame and the reference image frame of the target image frame from the video to be processed may include: performing scene detection on each image frame in the video to be processed to determine the scene to which each image frame belongs; obtaining the target image frame and the reference image frame of the target image frame from multiple image frames in the same scene; where the target image frame is any image frame except the first image frame among the multiple image frames, and the reference image frame is the previous image frame adjacent to the target image frame among the multiple image frames.
[0143] For example, the first image frame, the second image frame, and the third image frame belong to the same scene. The first image frame is the first image frame among the last three image frames, the second image frame is the subsequent image frame adjacent to the first image frame, and the third image frame is the subsequent image frame adjacent to the second image frame. The second image frame can be selected as the target image frame from the above three image frames, and the first image frame can be used as the reference image frame; or the third image frame can be selected as the target image frame from the above three image frames, and the second image frame can be used as the reference image frame.
[0144] It should be noted that if the trained feature fusion network is applied to the image super-resolution scenario, image enhancement processing can also be performed on the target image frame and the reference image frame to obtain the enhanced target image frame and the enhanced reference image frame. Then, the enhanced target image frame and the enhanced reference image frame can be used for image stitching processing to obtain a stitched image. Among them, the image enhancement processing can include at least one of the following: adding Gaussian noise to the target image frame and the reference image frame, performing Gaussian blur processing on the target image frame and the reference image frame, adding decompression noise to the target image frame and the reference image frame, etc. By performing image enhancement processing on the target image frame and the reference image frame, the generalization ability of the feature fusion network can be improved.
[0145] S602. Perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image.
[0146] S603. According to the requirements of multi-scale feature learning, perform feature fusion learning in the time domain and spatial domain on the stitched image to obtain the target fusion feature.
[0147] S604. Based on the target fusion feature, perform optical flow estimation on the target image frame to obtain the target optical flow information of the target image frame.
[0148] During the training process of the feature fusion network, the execution processes of steps S602 to S604 are similar to the relevant steps in the application process of the optical flow estimation model in the above Figure 2 or Figure 4 shown embodiments. For details, reference can be made to the descriptions of the relevant steps in the above Figure 2 or Figure 4 embodiments. For example, the execution process of step S602 can be referred to the description of step S202 in the above Figure 2 shown embodiment, the execution process of step S603 can be referred to the descriptions of steps S403 to S405 in the above Figure 4 shown embodiment, and the execution process of step S604 can be referred to the description of step S204 in the above Figure 2 shown embodiment, which will not be elaborated here.
[0149] S605, Obtain the labeled optical flow information corresponding to the target image frame.
[0150] During the training process of the optical flow estimation model, the labeled optical flow information corresponding to the target image frame can be obtained, and the labeled optical flow information is the annotation data of the target image frame.
[0151] S606, Determine the loss value of the feature fusion network based on the difference between the target optical flow information and the labeled optical flow information of the target image frame.
[0152] S607, Optimize the network parameters of the feature fusion network in the direction of reducing the loss value.
[0153] In steps S606 to S607, after obtaining the labeled optical flow information corresponding to the target image frame, the loss value of the feature fusion network can be determined based on the difference between the target optical flow information and the labeled optical flow information of the target image frame, and then the network parameters of the feature fusion network can be optimized in the direction of reducing the loss value. Specifically, after obtaining the labeled optical flow information corresponding to the target image frame, the loss function of the feature fusion network can be obtained, and then the loss value of the feature fusion network under the loss function can be determined based on the difference between the target optical flow information and the labeled optical flow information of the target image frame, so that the network parameters of the feature fusion network can be optimized in the direction of reducing the loss value.
[0154] Among them, the "direction of reducing the loss value" mentioned in the embodiments of the present application refers to: the network optimization direction with the goal of minimizing the loss value; through this direction of network optimization, the loss value generated again by the feature fusion network after each optimization needs to be less than the loss value generated by the feature fusion network before optimization. For example, if the loss value of the feature fusion network calculated this time is 0.85, then after optimizing the feature fusion network in the direction of reducing the loss value, the loss value generated by optimizing the feature fusion network should be less than 0.85. In addition, the loss function involved in the embodiments of the present application can be the L1 norm loss function or the L2 norm loss function, but this does not constitute a limitation on the embodiments of the present application. The loss function adopted in the embodiments of the present application can also be other loss functions, such as the cross-entropy loss function, etc.
[0155] It should be noted that, as can be seen from the foregoing, the feature fusion network can be integrated into a single model, or the feature fusion network can be integrated with an image stitching network, a feature extraction network, an upsampling network, a feature calibration network, a channel dimension reduction network, a feature activation network, etc. into the same optical flow estimation model. When the feature fusion network is integrated into a single model, for example, the feature fusion network is integrated into a feature fusion model, the loss value of the feature fusion network can be used to optimize the model parameters of the feature fusion model (i.e., the network parameters of the feature fusion network) in the direction of reducing the loss value, so as to realize the training of the feature fusion network. When the feature fusion network is integrated with an image stitching network, a feature extraction network, an upsampling network, a feature calibration network, a channel dimension reduction network, a feature activation network, etc. into an optical flow estimation model, the loss value of the feature fusion network can be used to optimize the model parameters of the optical flow estimation model in the direction of reducing the loss value, so as to realize the training of the optical flow estimation model.
[0156] In the embodiments of the present application, scene detection needs to be performed when obtaining training data. The training data obtained for training the feature fusion network are two consecutive image frames that are adjacent in playback time and match in scene in the video. Through scene detection, the two image frames in the training data belong to the same scene, which avoids the large difference between the target image frame and the reference image frame due to the target image frame and the reference image frame belonging to different scenes, affecting the accuracy of the subsequent obtained target optical flow information, and further avoiding affecting the training optimization effect of the network. In addition, if the trained feature fusion network is applied to the image super-resolution scene, image enhancement processing can also be performed on the target image frame and the reference image frame during the training process of the feature fusion network. Through image enhancement processing, the generalization performance of the feature fusion network can be improved, so that the feature fusion network can achieve a good feature learning and fusion effect for different types of images. For example, for some images containing noise, the feature fusion network can not only perform feature fusion processing, but also eliminate the noise in the image to be processed, which can strengthen the image super-resolution ability of the trained feature fusion network.
[0157] Based on the description of the related embodiments of the above image processing method, the embodiments of the present application also propose an image processing device, which can be a computer program (including program code) running in a computer device. Specifically, the image processing device can execute Figure 2 、 Figure 4 or Figure 6 the method steps in the image processing method shown; please refer to Figure 7 , the image processing device can run the following units:
[0158] An acquisition unit 701, configured to acquire a target image frame and a reference image frame of the target image frame from a video to be processed, where the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed;
[0159] A processing unit 702, configured to perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image;
[0160] The processing unit 702 is further configured to perform feature fusion learning in the time domain and the spatial domain on the stitched image according to the requirements of multi-scale feature learning to obtain a target fusion feature;
[0161] The processing unit 702 is further configured to perform optical flow estimation on the target image frame based on the target fusion feature to obtain target optical flow information of the target image frame.
[0162] In one implementation, when the processing unit 702 is configured to perform feature fusion learning in the time domain and the spatial domain on the stitched image according to the requirements of multi-scale feature learning to obtain a target fusion feature, the following steps are specifically performed:
[0163] Obtain a feature fusion network, where the feature fusion network includes N feature learning branches, one feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1;
[0164] Call each feature learning branch in the feature fusion network to perform feature fusion learning in the time domain and the spatial domain on the stitched image according to the corresponding feature learning scale;
[0165] Perform feature fusion processing on the fusion features learned by each feature learning branch to obtain a target fusion feature.
[0166] In one implementation, the N feature learning branches include a first feature learning branch. When the first feature learning branch performs feature fusion learning in the time domain and the spatial domain on the stitched image according to the corresponding feature learning scale, the following steps are specifically performed:
[0167] Perform fusion convolution processing on the stitched image in the time domain and the spatial domain according to a first receptive field to obtain a first convolution feature, where the first receptive field is used to describe the feature learning scale corresponding to the first feature learning branch;
[0168] Perform downsampling processing on the first convolution feature to obtain the fusion feature learned by the first feature learning branch.
[0169] In one implementation, the N feature learning branches include a second feature learning branch. When the second feature learning branch performs feature fusion learning in the time domain and the spatial domain on the stitched image according to the corresponding feature learning scale, the following steps are specifically performed:
[0170] Call the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and the spatial domain, and obtain the first residual feature;
[0171] Perform fusion convolution processing on the first residual feature in the time domain and the spatial domain according to the second receptive field, and obtain the second convolution feature; The second receptive field and the receptive field involved in the first shallow residual learning module jointly describe the feature learning scale corresponding to the second feature learning branch;
[0172] Perform downsampling processing based on the second convolution feature to obtain the fusion feature learned by the second feature learning branch.
[0173] In one implementation, the N feature learning branches include a third feature learning branch. When the third feature learning branch performs feature fusion learning on the spliced image in the time domain and the spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0174] Call the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and the spatial domain, and obtain the first residual feature;
[0175] Call the second shallow residual learning module to perform fusion residual learning on the first residual feature in the time domain and the spatial domain, and obtain the second residual feature; The receptive field involved in the first shallow residual learning module and the receptive field involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch;
[0176] Perform downsampling processing based on the second residual feature to obtain the fusion feature learned by the third feature learning branch.
[0177] In one implementation, the downsampling methods adopted by each feature learning branch among the N feature learning branches are different from each other.
[0178] In one implementation, the target fusion feature includes feature maps of multiple channels, and the target optical flow information is represented by a vector; The processing unit 702, when used to perform optical flow estimation on the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame, is specifically used to execute the following steps:
[0179] Perform upsampling processing on the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature;
[0180] Perform dimensionality reduction processing on the channel number of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing, and the channel number of the fusion feature after dimensionality reduction processing matches the vector dimension of the target optical flow information;
[0181] Perform activation processing on the fusion feature after dimensionality reduction processing to obtain the target optical flow information of the target image frame.
[0182] In one implementation, the processing unit 702 is configured to perform dimensionality reduction on the number of channels of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction. Specifically, it is configured to perform the following steps:
[0183] Perform feature calibration on the upsampled fusion feature to obtain the fusion feature after feature calibration;
[0184] Perform dimensionality reduction on the number of channels of the fusion feature after feature calibration to obtain the fusion feature after dimensionality reduction.
[0185] In one implementation, the processing unit 702 is further configured to perform the following steps:
[0186] Generate an optical flow visualization image of the target image frame based on the target optical flow information and the target image frame;
[0187] Perform image super-resolution processing on the target image frame according to the optical flow visualization image to obtain a super-resolution image of the target image frame, and the resolution of the super-resolution image is higher than that of the target image frame.
[0188] In one implementation, the target fusion feature is obtained through a feature fusion network, and the video to be processed is a sample video for training the feature fusion network; the acquisition unit 701 is configured to obtain the target image frame and the reference image frame of the target image frame from the video to be processed. Specifically, it is configured to perform the following steps:
[0189] Perform scene detection on each image frame in the video to be processed to determine the scene to which each image frame belongs;
[0190] Among multiple image frames in the same scene, obtain the target image frame and the reference image frame of the target image frame; wherein, the target image frame is any image frame other than the first image frame among the multiple image frames.
[0191] In one implementation, the processing unit 702 is further configured to perform the following steps:
[0192] Obtain the marked optical flow information corresponding to the target image frame;
[0193] Determine the loss value of the feature fusion network based on the difference between the target optical flow information and the marked optical flow information of the target image frame;
[0194] Optimize the network parameters of the feature fusion network in the direction of reducing the loss value.
[0195] According to an embodiment of the present application, Figure 2 、 Figure 4 or Figure 6 The method steps involved in the method shown may be performed byFigure 7 executed by respective units in the image processing apparatus shown. For example, Figure 2 the step S201 shown in Figure 7 can be executed by the acquisition unit 701 shown in Figure 2 and the steps S202 - S204 shown in Figure 7 can be executed by the processing unit 702 shown in. Another example, Figure 4 the step S401 shown in Figure 7 can be executed by the acquisition unit 701 shown in Figure 4 and the steps S402 - S406 shown in Figure 7 can be executed by the processing unit 702 shown in. Yet another example, Figure 6 the step S601 and the step S605 shown in Figure 7 can be executed by the acquisition unit 701 shown in Figure 6 and the steps S602 - S604, as well as the steps S606 - S607 shown in Figure 7 can be executed by the processing unit 702 shown in.
[0196] According to another embodiment of the present application, Figure 7 the respective units in the image processing apparatus shown can be respectively or all combined into one or several other units to form, or a certain one (or some) of the units can also be further split into multiple smaller units in terms of function to form, which can achieve the same operations without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, based on the image processing apparatus, other units may also be included. In practical applications, these functions can also be assisted by other units to be realized, and can be realized by the cooperation of multiple units.
[0197] According to another embodiment of the present application, it is possible to run, on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read - only storage medium (ROM), a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 2 , Figure 4 or Figure 6 to construct the image processing apparatus shown in Figure 7 and to implement the image processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer - readable storage medium, and be loaded into the above - mentioned computing device through the computer - readable storage medium and run therein.
[0198] In the embodiments of the present application, according to the requirements of multi-scale feature learning, feature fusion learning in the time domain and spatial domain can be performed on the spliced image of two consecutive image frames with adjacent playing times in the video. Then, based on the target fusion features obtained from the feature fusion learning, optical flow estimation can be performed on the image frame with a later playing time among the two consecutive image frames to obtain the optical flow information of this image frame. As can be seen from the above, among the target fusion features learned by the feature fusion learning, on the one hand, the features of the spliced image at multiple scales are fused, and on the other hand, the features of the spliced image in the time domain and spatial domain are fused. Using the target fusion features that perform multi-dimensional (i.e., multi-scale, time domain dimension, and spatial domain dimension) feature fusion for optical flow estimation can greatly improve the accuracy of optical flow estimation.
[0199] Based on the descriptions of the above method embodiments and apparatus embodiments, the embodiments of the present application further provide a computer device. Please refer to Figure 8 , this computer device at least includes a processor 801, an input interface 802, an output interface 803, and a computer-readable storage medium 804. Among them, the processor 801, the input interface 802, the output interface 803, and the computer-readable storage medium 804 can be connected through a bus or other means.
[0200] The computer-readable storage medium 804 can be stored in the memory of the computer device. The computer-readable storage medium 804 is used to store a computer program, and the computer program includes computer instructions. The processor 801 is used to execute the program instructions stored in the computer-readable storage medium 804. The processor 801 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and it is suitable for implementing one or more computer instructions, specifically suitable for loading and executing one or more computer instructions to implement the corresponding method flow or corresponding function.
[0201] The embodiments of the present application also provide a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the computer device. And, in this storage space, one or more computer instructions suitable for being loaded and executed by a processor are also stored. These computer instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0202] In one implementation, one or more computer instructions stored in the computer-readable storage medium 804 can be loaded and executed by the processor 801 to implement the corresponding steps of the above-mentioned Figure 2 、 Figure 4 or Figure 6 shown image processing methods. Specifically, one or more computer instructions in the computer-readable storage medium 804 are loaded and executed by the processor 801 to perform the following steps:
[0203] Obtain a target image frame and a reference image frame of the target image frame from the video to be processed. The reference image frame is the previous image frame adjacent to the target image frame in the video to be processed;
[0204] Perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image;
[0205] According to the requirements of multi-scale feature learning, perform feature fusion learning in the time domain and spatial domain on the stitched image to obtain a target fusion feature;
[0206] Based on the target fusion feature, perform optical flow estimation on the target image frame to obtain the target optical flow information of the target image frame.
[0207] In one implementation, when one or more computer instructions in the computer-readable storage medium 804 are loaded and executed by the processor 801 to perform feature fusion learning in the time domain and spatial domain on the stitched image according to the requirements of multi-scale feature learning to obtain a target fusion feature, the following steps are specifically performed:
[0208] Obtain a feature fusion network. The feature fusion network includes N feature learning branches. One feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1;
[0209] Call each feature learning branch in the feature fusion network, and perform feature fusion learning on the spliced image in the time domain and spatial domain according to the corresponding feature learning scale;
[0210] Perform feature fusion processing on the fusion features learned by each feature learning branch to obtain the target fusion feature.
[0211] In one implementation, the N feature learning branches include a first feature learning branch. When the first feature learning branch performs feature fusion learning on the spliced image in the time domain and spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0212] Perform fusion convolution processing on the spliced image in the time domain and spatial domain according to the first receptive field to obtain a first convolution feature, and the first receptive field is used to describe the feature learning scale corresponding to the first feature learning branch;
[0213] Perform downsampling processing based on the first convolution feature to obtain the fusion feature learned by the first feature learning branch.
[0214] In one implementation, the N feature learning branches include a second feature learning branch. When the second feature learning branch performs feature fusion learning on the spliced image in the time domain and spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0215] Call the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain a first residual feature;
[0216] Perform fusion convolution processing on the first residual feature in the time domain and spatial domain according to the second receptive field to obtain a second convolution feature; the second receptive field and the receptive field involved in the first shallow residual learning module jointly describe the feature learning scale corresponding to the second feature learning branch;
[0217] Perform downsampling processing based on the second convolution feature to obtain the fusion feature learned by the second feature learning branch.
[0218] In one implementation, the N feature learning branches include a third feature learning branch. When the third feature learning branch performs feature fusion learning on the spliced image in the time domain and spatial domain according to the corresponding feature learning scale, the following steps are specifically executed:
[0219] Call the first shallow residual learning module to perform fusion residual learning on the spliced image in the time domain and spatial domain to obtain a first residual feature;
[0220] Invoke the second shallow residual learning module to perform fusion residual learning on the first residual feature in the time domain and spatial domain to obtain the second residual feature; the receptive fields involved in the first shallow residual learning module and the receptive fields involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch;
[0221] Perform downsampling processing based on the second residual feature to obtain the fusion feature learned by the third feature learning branch.
[0222] In one implementation, when each of the N feature learning branches performs downsampling processing, the downsampling methods used are different from each other.
[0223] In one implementation, the target fusion feature includes feature maps of multiple channels, and the target optical flow information is represented by a vector; when one or more computer instructions in the computer-readable storage medium 804 are loaded and executed by the processor 801 to perform optical flow estimation on the target image frame based on the target fusion feature to obtain the target optical flow information of the target image frame, it is specifically used to execute the following steps:
[0224] Upsample the target fusion feature according to the image size of the target image frame to obtain the upsampled fusion feature;
[0225] Perform dimensionality reduction processing on the channel number of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing, and the channel number of the fusion feature after dimensionality reduction processing matches the vector dimension of the target optical flow information;
[0226] Perform activation processing on the fusion feature after dimensionality reduction processing to obtain the target optical flow information of the target image frame.
[0227] In one implementation, when one or more computer instructions in the computer-readable storage medium 804 are loaded and executed by the processor 801 to perform dimensionality reduction processing on the channel number of the upsampled fusion feature to obtain the fusion feature after dimensionality reduction processing, it is specifically used to execute the following steps:
[0228] Perform feature calibration processing on the upsampled fusion feature to obtain the fusion feature after feature calibration;
[0229] Perform dimensionality reduction processing on the channel number of the fusion feature after feature calibration to obtain the fusion feature after dimensionality reduction processing.
[0230] In one implementation, one or more computer instructions in the computer-readable storage medium 804 are loaded by the processor 801 and are further used to execute the following steps:
[0231] Generate an optical flow visualization image of the target image frame based on the target optical flow information and the target image frame;
[0232] Performing image super-resolution processing on the target image frame according to the optical flow visualization image to obtain a super-resolution image of the target image frame, where the resolution of the super-resolution image is higher than that of the target image frame.
[0233] In one implementation, the target fusion feature is obtained through a feature fusion network, and the video to be processed is a sample video for training the feature fusion network; when one or more computer instructions in the computer-readable storage medium 804 are loaded and executed by the processor 801 to obtain the target image frame and the reference image frame of the target image frame from the video to be processed, it is specifically used to perform the following steps:
[0234] Performing scene detection on each image frame in the video to be processed to determine the scene to which each image frame belongs;
[0235] Among multiple image frames in the same scene, obtaining the target image frame and the reference image frame of the target image frame; wherein, the target image frame is any image frame other than the first image frame among the multiple image frames.
[0236] In one implementation, one or more computer instructions in the computer-readable storage medium 804 are loaded by the processor 801 and are further used to perform the following steps:
[0237] Obtaining the marked optical flow information corresponding to the target image frame;
[0238] Based on the difference between the target optical flow information and the marked optical flow information of the target image frame, determining the loss value of the feature fusion network;
[0239] Optimizing the network parameters of the feature fusion network in the direction of reducing the loss value.
[0240] In the embodiments of the present application, according to the requirements of multi-scale feature learning, feature fusion learning in the time domain and spatial domain can be performed on the spliced image of two consecutive image frames adjacent in playback time in the video, and then based on the target fusion feature obtained by the feature fusion learning, optical flow estimation can be performed on the image frame with a later playback time among the two consecutive image frames to obtain the optical flow information of the image frame. As can be seen from the above, in the target fusion feature learned by the feature fusion learning, on the one hand, the features of the spliced image at multiple scales are fused, and on the other hand, the features of the spliced image in the time domain and spatial domain are fused. Using the target fusion feature that performs multi-dimensional (i.e., multi-scale, and time domain dimension and spatial domain dimension) feature fusion for optical flow estimation can greatly improve the accuracy of optical flow estimation.
[0241] It should be noted that, according to one aspect of the present application, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative manners in the aspects of the above-mentioned Figure 2 , Figure 4 or Figure 6 illustrated method embodiments of the image processing method.
[0242] As mentioned above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining a target image frame and a reference image frame of the target image frame from a video to be processed, where the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed; Performing image stitching processing on the target image frame and the reference image frame to obtain a stitched image; Obtaining a feature fusion network, where the feature fusion network includes N feature learning branches, one feature learning branch corresponding to one feature learning scale, and N is an integer greater than 1; Invoking each feature learning branch in the feature fusion network to perform feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale; Performing feature fusion processing on the fusion features learned by each feature learning branch to obtain a target fusion feature; Performing optical flow estimation on the target image frame based on the target fusion feature to obtain target optical flow information of the target image frame.
2. The method according to claim 1, wherein The N feature learning branches include a first feature learning branch, and the first feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, including: Performing fusion convolution processing on the stitched image in the time domain and the spatial domain according to a first receptive field to obtain a first convolution feature, where the first receptive field is used to describe the feature learning scale corresponding to the first feature learning branch; Performing downsampling processing based on the first convolution feature to obtain the fusion feature learned by the first feature learning branch.
3. The method according to claim 1, wherein The N feature learning branches include a second feature learning branch, and the second feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, including: Invoking a first shallow residual learning module to perform fusion residual learning on the stitched image in the time domain and the spatial domain to obtain a first residual feature; Performing fusion convolution processing on the first residual feature in the time domain and the spatial domain according to a second receptive field to obtain a second convolution feature; the second receptive field and the receptive field involved in the first shallow residual learning module jointly describe the feature learning scale corresponding to the second feature learning branch; Performing downsampling processing based on the second convolution feature to obtain the fusion feature learned by the second feature learning branch.
4. The method according to claim 1, wherein The N feature learning branches include a third feature learning branch, and the third feature learning branch performs feature fusion learning on the stitched image in the time domain and the spatial domain according to the corresponding feature learning scale, including: Invoking a first shallow residual learning module to perform fusion residual learning on the stitched image in the time domain and the spatial domain to obtain a first residual feature; Invoking a second shallow residual learning module to perform fusion residual learning on the first residual feature in the time domain and the spatial domain to obtain a second residual feature; the receptive field involved in the first shallow residual learning module and the receptive field involved in the second shallow residual learning module jointly describe the feature learning scale corresponding to the third feature learning branch; Performing downsampling processing based on the second residual feature to obtain the fusion feature learned by the third feature learning branch.
5. The method according to any one of claims 1-4, characterized in that, When each of the N feature learning branches performs downsampling processing, the downsampling methods adopted are different from each other.
6. The method according to claim 1, wherein The target fusion feature includes feature maps of multiple channels, and the target optical flow information is represented by a vector; the method for estimating the target optical flow information of the target image frame based on the target fusion feature includes: Performing upsampling processing on the target fusion feature according to the image size of the target image frame to obtain an upsampled fusion feature; Performing dimensionality reduction processing on the number of channels of the upsampled fusion feature to obtain a fusion feature after dimensionality reduction, and the number of channels of the fusion feature after dimensionality reduction matches the vector dimension of the target optical flow information; Performing activation processing on the fusion feature after dimensionality reduction to obtain the target optical flow information of the target image frame.
7. The method according to claim 6, wherein The performing dimensionality reduction processing on the number of channels of the upsampled fusion feature to obtain a fusion feature after dimensionality reduction includes: Performing feature calibration processing on the upsampled fusion feature to obtain a fusion feature after feature calibration; Performing dimensionality reduction processing on the number of channels of the fusion feature after feature calibration to obtain a fusion feature after dimensionality reduction.
8. The method according to claim 1, characterized in that, The method further includes: Generating an optical flow visualization image of the target image frame based on the target optical flow information and the target image frame; Performing image super-resolution processing on the target image frame according to the optical flow visualization image to obtain a super-resolution image of the target image frame, and the resolution of the super-resolution image is higher than that of the target image frame.
9. The method according to claim 1, wherein The target fusion feature is obtained through a feature fusion network, and the video to be processed is a sample video for training the feature fusion network; The obtaining the target image frame and the reference image frame of the target image frame from the video to be processed includes: Performing scene detection on each image frame in the video to be processed to determine the scene to which each image frame belongs; Among multiple image frames in the same scene, obtaining the target image frame and the reference image frame of the target image frame; wherein, the target image frame is any image frame other than the first image frame among the multiple image frames.
10. The method according to claim 9, characterized in that The method further includes: Obtaining the marked optical flow information corresponding to the target image frame; Determining the loss value of the feature fusion network based on the difference between the target optical flow information and the marked optical flow information of the target image frame; Optimizing the network parameters of the feature fusion network in the direction of reducing the loss value.
11. An image processing apparatus, characterized in that, The image processing device includes: An obtaining unit, configured to obtain a target image frame and a reference image frame of the target image frame from a video to be processed, where the reference image frame is the previous image frame adjacent to the target image frame in the video to be processed; A processing unit, configured to perform image stitching processing on the target image frame and the reference image frame to obtain a stitched image; The processing unit is further configured to obtain a feature fusion network, where the feature fusion network includes N feature learning branches, one feature learning branch corresponds to one feature learning scale, and N is an integer greater than 1; call each feature learning branch in the feature fusion network to perform feature fusion learning on the spliced image in the time domain and the spatial domain according to the corresponding feature learning scale; perform feature fusion processing on the fusion features learned by each feature learning branch to obtain target fusion features; The processing unit is further configured to perform optical flow estimation on the target image frame based on the target fusion features to obtain target optical flow information of the target image frame.
12. A computer device, characterized in that, The computer device includes: a processor adapted to implement a computer program; and, a computer-readable storage medium storing a computer program, where the computer program is adapted to be loaded and executed by the processor to perform the image processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, where the computer program is adapted to be loaded and executed by a processor to perform the image processing method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the image processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
End-to-end optical flow estimation method based on multi-stage loss
CN110111366A