Method, device and equipment for determining optical flow motion vector map and readable storage medium
By dividing the image exposure area and performing optical flow calculation in the terminal device, the problem of insufficient accuracy of optical flow estimation in traditional HDR video synthesis is solved, and accurate processing and stable display of high dynamic range images are achieved.
Patent Information
- Application Number
- CN202411792052.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Traditional HDR video compositing techniques suffer from low accuracy in optical flow estimation, leading to overexposure in bright areas or loss of detail in dark areas. Existing terminal devices struggle to accurately process high dynamic range images.
By acquiring first and second images with different brightness, the images are divided into low-exposure, normal-exposure, and high-exposure regions according to the transmission channel of each pixel. The optical flow motion vector is determined based on the region type, and optical flow calculation and information fusion are performed using traditional algorithms and AI solutions.
It improves the accuracy of optical flow motion vector graphics, ensuring detail retention and color rendering under different exposure conditions, and enhances the stability of HDR videos and user experience.
Smart Images

Figure CN119603566B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video processing technology, specifically relating to a method, apparatus, device, and readable storage medium for determining optical flow motion vector graphics. Background Technology
[0002] With the development of terminal devices, more users are taking photos and viewing photos and previewing videos on mobile devices. In real-world scenes, objects have a very high dynamic range. Most terminal devices support 24-bit color representation, but a single pixel value is only 8 bits of data. Therefore, the dynamic range that a single photo taken by a typical terminal device can cover is extremely limited. Using too short an exposure time will make low-brightness areas in the scene appear dark and introduce noise, while using too long an exposure time will cause high-brightness areas to be overexposed and lose detail. To achieve rich color detail and brightness gradation in images, High Dynamic Range (HDR) technology, which combines multiple frames with alternating exposures, has become the mainstream solution.
[0003] Among them, the optical flow estimation accuracy involved in traditional HDR video synthesis technology is relatively low. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and readable storage medium for determining optical flow motion vector diagrams, which can achieve more accurate optical flow estimation.
[0005] In a first aspect, embodiments of this application provide a method for determining an optical flow motion vector map, executed by a terminal, including:
[0006] Obtain a first image and a second image from the input image; wherein the first image and the second image have different brightness;
[0007] Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are divided into multiple regions, and the region types include low exposure region, normal exposure region and high exposure region.
[0008] Based on the region type to which each region belongs, determine the optical flow motion vector diagrams of the first image and the second image.
[0009] Secondly, embodiments of this application provide a device for determining an optical flow motion vector diagram, comprising:
[0010] A first processing module is used to acquire a first image and a second image from an input image; wherein the first image and the second image have different brightness;
[0011] The second processing module is used to divide the first image and the second image into multiple regions according to the transmission channel of each pixel in the first image and the second image, respectively. The region types include low exposure region, normal exposure region and high exposure region.
[0012] The third processing module is used to determine the optical flow motion vector diagrams of the first image and the second image according to the region type to which each region belongs.
[0013] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0014] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0016] In a sixth aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to perform the steps of the method as described in the first aspect.
[0017] In this embodiment, after acquiring the first and second images from the input image, the terminal can use the transmission channels of each pixel in the first and second images to perform region division of the first and second images, distinguishing between low-exposure areas, normal-exposure areas, and high-exposure areas. Then, based on the region type to which each region belongs, the optical flow motion vector diagrams of the first and second images are determined. Because different exposure levels of the images are considered when determining the optical flow motion vector diagrams, the accuracy of the optical flow motion vector diagrams is improved. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the process of determining the optical flow motion vector diagram according to an embodiment of this application;
[0019] Figure 2 This is one of the schematic diagrams of the input image;
[0020] Figure 3 This is the second illustration of the input image;
[0021] Figure 4 This is a schematic diagram of image channel processing;
[0022] Figure 5 This is a schematic diagram of the channel diagram segmentation;
[0023] Figure 6 This is a schematic diagram showing the division of the overexposed area in Image 1;
[0024] Figure 7 This is a schematic diagram of the optical flow estimation process;
[0025] Figure 8 This is a schematic diagram of the generation of an optical flow motion vector diagram;
[0026] Figure 9 This is a schematic diagram of the terminal hardware structure;
[0027] Figure 10 This is a schematic diagram of the HDR-MV-LITE module structure;
[0028] Figure 11 This is a schematic diagram of the terminal interface;
[0029] Figure 12 This is a schematic diagram of the HDR video processing workflow;
[0030] Figure 13 This is a schematic diagram of the offline HDR video processing workflow;
[0031] Figure 14 This is a schematic diagram of the module structure of the device for determining the optical flow motion vector diagram according to an embodiment of this application;
[0032] Figure 15 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0033] Figure 16 This is a schematic diagram of the structure of an electronic device according to another embodiment of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0035] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0036] The method for determining the optical flow motion vector diagram provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0037] like Figure 1 As shown, a method for determining an optical flow motion vector diagram according to an embodiment of this application is executed by a terminal and includes:
[0038] Step 101: Obtain a first image and a second image from the input image; wherein the first image and the second image have different brightness.
[0039] Here, the input image is used to estimate the optical flow of the current shooting scene. For this input image, two images with brightness differences can be selected as the first image and the second image, respectively, to obtain its optical flow motion vector map.
[0040] Step 102: Based on the transmission channel of each pixel in the first image and the second image, divide the first image and the second image into multiple regions respectively. The region types include low exposure region, normal exposure region and high exposure region.
[0041] In this step, the transmission channels of each pixel in the first image and the second image are used to divide the regions of the first image and the second image, so as to distinguish the low exposure area, normal exposure area and high exposure area, laying the foundation for the next step.
[0042] Step 103: Determine the optical flow motion vector diagrams of the first image and the second image according to the region type to which each region belongs.
[0043] By dividing the regions of the first image and the second image in step 102, this step can determine the optical flow motion vector diagrams of the first image and the second image for low exposure regions, normal exposure regions, and high exposure regions.
[0044] Thus, following steps 101-103 above, after acquiring the first and second images from the input image, the terminal can use the transmission channels of each pixel in the first and second images to perform region division of the first and second images, distinguishing between low-exposure areas, normal-exposure areas, and high-exposure areas. Furthermore, based on the region type to which each region belongs, the optical flow motion vector diagrams of the first and second images are determined. By considering different exposure levels of the images when determining the optical flow motion vector diagrams, the accuracy of the optical flow motion vector diagrams is improved.
[0045] It should be understood that the method in the application embodiments is applicable to HDR synthesis applications of video frames with alternating long and short exposures, providing high-precision optical flow estimation for terminal image / video applications. For example, the input image may include multiple frames, and the exposure ratio can support alternating long and short exposure sequences, such as... Figure 2 The three-frame image shown (short-long-short) can also support long, medium, and short image sequences, such as... Figure 3 The image shown is a set of five frames (increasing in length from shortest to longest). From this input image, the terminal can select each set (two images with different brightness levels) as the first image and the second image.
[0046] In this embodiment, the optical flow motion vector diagram is also called the optical flow vector diagram.
[0047] Optionally, the step of dividing the first image and the second image into multiple regions according to the transmission channel of each pixel in the first image and the second image includes:
[0048] Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are processed separately by channel to obtain the channel map of each channel;
[0049] The channel map is segmented step by step based on the segmentation strategy until the variance of the pixel values of the resulting segmented parts is less than a first threshold.
[0050] The region type to which each segment belongs is determined based on a pixel threshold.
[0051] Based on the region type and the location information of the segmented portion, multiple regions of the first image and the second image are determined.
[0052] In this way, more efficient region (also known as block) division can be achieved through channel pixels.
[0053] Channel segmentation involves extracting and stitching pixels from the same transmission channel based on different transmission channel types. Taking the first image as an example, if the image is arranged in an RGGB channel pattern, containing one R channel, one B channel, and two G channels, after channel segmentation processing, a channel map of the R channel (averaging the values of the two channels), G channel, and B channel can be obtained, as shown below. Figure 4 As shown in the figure (different green fills are used to distinguish the two G channels).
[0054] The segmentation strategy can include a hierarchical segmentation pattern, such as N*M segmentation at the first level, N*N segmentation at the second level, etc. Of course, the segmentation pattern can be the same at each level. For example... Figure 5 The channel map of the G channel shown is divided into two parts in the first segmentation (first-level segmentation) by 2*4. Since the variance of the pixel value of a certain segment is not less than the first threshold, it needs to be segmented again (second-level segmentation). The second segmentation is divided into two parts by 1*2, resulting in two segments (the red box part and the blue box part). At this time, the variance of the pixel value of all segments is less than the first threshold, and no further segmentation is needed.
[0055] After one segmentation, to determine whether to perform another segmentation, it is necessary to determine whether the pixels in the segmented part are stable, that is, whether the variance is less than the first threshold. If the variance is less than the first threshold, it means that the pixels in this segmented part are stable. If the variance is not less than the first threshold, it means that the pixels in this segmented part are unstable.
[0056] The segmented areas can be categorized into low-exposure, normal-exposure, and high-exposure regions.
[0057] After determining the region type of each segmented part, for the target image of the region division, such as the first image, the segments belonging to the same region type in each channel image of the first image are merged according to their positions in the first image to obtain multiple regions of the first image.
[0058] Optionally, the pixel threshold includes a second threshold and a third threshold, wherein the second threshold is greater than the third threshold;
[0059] The step of determining the region type of each segmented part based on a pixel threshold includes:
[0060] If the average pixel value within the current segment is greater than or equal to the second threshold, then the region type of the current segment is a high-exposure region.
[0061] If the average pixel value within the current segment is less than the second threshold and greater than or equal to the third threshold, then the region type of the current segment is a normal exposure region.
[0062] If the average pixel value within the current segment is less than the third threshold, then the region type of the current segment is a low-exposure region.
[0063] Here, to distinguish between low-exposure areas, normal-exposure areas, and high-exposure areas, pixel thresholds, including a second threshold and a third threshold, are set. After the channel map is segmented, the region type is determined based on the average pixel value within each segment and the magnitude of the second and third thresholds.
[0064] As an optional implementation, when the target image for region partitioning is image 1 in the input RAW domain image, such as Figure 6 As shown, image 1 is first processed by channel segmentation to obtain an R-channel image (the channel map of the R channel), a G-channel image (the channel map of the G channel), and a B-channel image (the channel map of the B channel). Then, the single-channel image is segmented step-by-step. After each segmentation, the variance of the pixel values in the segmented portion is checked against a first threshold. If the variance is less than the first threshold, the pixel values within that segment are considered stable and belong to the same region type. If the variance is not less than (greater than or equal to) the first threshold, the pixel values within that segment are considered to have significant differences, and there may be at least one type of sub-region within that segment. In this case, the segment size is reduced (e.g., by half), and the variance of the pixel values in the next segment is calculated. This process continues until the variance is less than the first threshold, at which point the channel map segmentation is complete. When calculating the variance of the pixel values in the segmented portion, the mean pixel value of the segment can also be calculated. Thus, after the channel map segmentation is completed, the mean pixel value of the segment is compared to the pixel threshold to determine low-exposure areas, normal-exposure areas, and high-exposure areas. After the above process is completed, the low exposure area, normal exposure area and high exposure area can be merged and divided separately. Taking the high exposure area as an example, the segmented parts of the R, G and B channel images that belong to the high exposure area are selected and used as the division result of the high exposure area of image 1.
[0065] Furthermore, considering that for a given region type, the movement vectors of associated channels can be used to characterize the movement features of other channels, for example, in a RAW domain image, the pixel arrangement is RGGB, where R and B each represent one channel, and G represents data collected from both channels. Therefore, in a RAW domain image pixel, the value on the G channel is higher than the values on the R and B channels. Additionally, in a RAW domain image, RGGB pixels are arranged sequentially, and pixel arrangements across different channels exhibit similar image edge features.
[0066] Optionally, in this embodiment, determining the optical flow motion vector diagrams of the first image and the second image based on the region type to which each region belongs includes:
[0067] Optical flow information is obtained by using the associated channel pixels based on each region type.
[0068] The optical flow information is fused to determine the optical flow motion vector diagrams of the first image and the second image.
[0069] Therefore, for the first and second images, optical flow calculation (also known as off-flow estimation) needs to be performed using channel pixels associated with the same region type to obtain the optical flow information for this region type; after obtaining the optical flow information corresponding to each region type, they are then fused to determine the optical flow motion vector map.
[0070] Optionally, the step of calculating optical flow information using associated channel pixels based on each of the region types includes:
[0071] Based on the first channel of each pixel in the first region of the first image, a first channel image is obtained;
[0072] A second channel map is obtained based on the second channel of each pixel in the second region of the second image;
[0073] Optical flow estimation is performed based on the first channel map and the second channel map to determine the optical flow information of the first region type;
[0074] Wherein, the region type to which the first region and the second region belong is the first region type, and the first region type is any one of the multiple region types, and the transmission channel type to which the first channel and the second channel belong is associated with the first region type.
[0075] If the first region is a high-exposure region, since the pixel values in the R and B channels are low, the transmission channel types associated with the high-exposure region are R and B channels. Therefore, the first channel is R and B channels, the second channel is R and B channels, and the pixels in the R and B channels are used for optical flow estimation. If the first region is a low-exposure region, since the pixel values in the G channel are low, the transmission channel type associated with the low-exposure region is G channel. Therefore, the first channel is G channel, the second channel is G channel, and the pixels in the G channel are used for optical flow estimation. If the first region is a normal-exposure region, and the transmission channel types associated with the normal-exposure region are R, B, and G channels, then the first channel is all channels, the second channel is all channels, and the pixels in all four channels are used for optical flow estimation.
[0076] In this embodiment, the algorithms used for optical flow calculation include, but are not limited to, traditional schemes (Lucas-Kanade, Horn-Schunck, etc.) and AI schemes (such as FlowNet, Spynet, RAFT, etc.).
[0077] Of course, it should also be noted that optical flow calculations are performed on a single channel. Therefore, assuming the first image is a long-exposure image L1 and the second image is a short-exposure image S1, as follows... Figure 7 As shown, after obtaining four channels for each pixel of the first image and the second image respectively, and obtaining four channel maps, optical flow estimation of the channels is performed based on the same transmission channel type. Then, based on each region type and the transmission channel type associated with the region type, the optical flow information corresponding to each region type is determined.
[0078] Thus, the process of fusion of optical flow information is as follows: for regions of the same type, the channel maps are stitched together by combining optical flow information to generate an optical flow motion vector map, such as... Figure 8 As shown in the figure (different green fills are used to distinguish the two G channels). The fusion of optical flow information can also be understood as stitching together the optical flow density maps of the three channels into a single optical flow image, i.e., an optical flow motion vector image.
[0079] After obtaining the optical flow motion vector image, optical flow registration (warp registration) can be performed. For example, the optical flow motion vector image can be used to register a long exposure image, and then the registered image can be fused with a short exposure image to finally generate a combined long and short exposure HDR image.
[0080] It should be noted that, in this embodiment, the hardware components of the terminal shown are as follows: Figure 9 As shown, it includes (a) audio / video systems, such as cameras (image sensors), audio devices, and display devices; (b) chip processing systems, such as central processing units (CPUs), neural network processing units (NPUs), and graphics processing units (GPUs); (c) memory, such as system memory and local cache; and (d) communication buses. Of course, it may also include other hardware components such as communication interfaces. Software components include the kernel layer, system layer, application framework layer, and application layer. Software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, and other mature storage media in the field. This storage medium resides in memory, and the processor executes the instructions in the memory, combined with its hardware implementation.
[0081] The camera converts optical images into electronic signals (raw image data). The processor (CPU, NPU, GPU) coordinates, schedules, and manages the system, handling image processing and AI neural network processing, providing systematic computational support for the algorithm model. System memory temporarily stores data processed by the CPU and exchanges cached data with the CPU. Local cache acts as a buffer for CPU data exchange, addressing the speed difference between CPU and memory. The processor, memory, and communication structure are connected via a communication bus; memory stores execution instructions, and the processor executes calculations and retrieves the necessary preceding and following frame image data. The display device presents the digital image signals on the screen, ensuring the LCD monitor functions correctly. The chip processing system, in the HDR video (High Dynamic Range Video Module-Lite, HDR-VM-Lite) module, processes the data captured by the camera and outputs HDR-fused RGB images and HDR video. Details of the HDR-VM-Lite module are as follows... Figure 10 As shown, it mainly includes four sub-modules: (a) HDR video frame acquisition sub-module; (b) HDR single-frame synthesis sub-module; (c) storage sub-module; and (d) HDR video single-frame processing sub-module. The HDR-MV-LITE module takes continuous long and short exposure RAW domain images as input and outputs RAW domain HDR composite images. The HDR video acquisition strategy (the strategy used by the image sensor to acquire images) is located in the HDR video frame acquisition module.
[0082] The HDR-VM-Lite module is used for functions such as taking photos, recording videos, making video calls, and live streaming on the terminal.
[0083] The HDR single-frame synthesis submodule includes (a) optical flow estimation unit 1, (b) long and short exposure registration unit, and (c) long and short exposure synthesis unit.
[0084] The optical flow estimation unit 1 is used to estimate the optical flow motion vector between frames before and after long and short exposures, such as the method for determining the optical flow motion vector diagram in the embodiments of this application.
[0085] The long and short exposure registration unit is used to align pixels in long and short exposure images based on optical flow motion information. It can be calculated using the following formula: I(x',y')=I(x+Δx,y+Δy). Here, I represents the pixel value, (x',y') represents the pixel position in the target image, (x,y) represents the pixel position in the original image, and (Δx,Δy) represents the optical flow motion vector calculated by optical flow estimation unit 1.
[0086] The long and short exposure synthesis unit processes long and short exposure images respectively and outputs a fusion weight map. The fusion weight can be calculated based on factors such as exposure ratio, pixel brightness, regional similarity, and temporal similarity, or it can be output through an AI neural network.
[0087] The HDR video single-frame processing submodule includes an optical flow estimation unit 2, a preprocessing unit, an HDR multi-frame feature extraction unit, an HDR multi-frame feature association unit, an HDR single-frame feature generation unit, and an artifact removal unit. The optical flow estimation unit 2 can also apply the optical flow motion vector diagram determination method of this application embodiment. The optical flow estimation result is used for HDR multi-frame feature fusion; after extracting the HDR video frame features, alignment operations are performed on the feature layer. The final output is a jitter-free, color-stable HDR video.
[0088] Among them, the HDR video frame acquisition submodule automatically adjusts the fusion strategy, acquisition interval, and sensitivity according to the light intensity and object movement of the external environment to ensure optimal front-end long and short frame input in different scenarios.
[0089] In one embodiment, the HDR-VM-Lite module takes a RAW domain image as input. The RAW domain image is directly acquired by the front-end camera sensor, and its pixel values are proportional to exposure time, brightness, etc., constituting linear data. Traditional optical flow estimation is mostly applied to non-linear feature images in the RGG domain after Demosaic processing, and is not suitable for linear feature images in the RAW domain. Furthermore, compared to the RGB domain, the pixel arrangement in the Raw domain is RGGB, therefore Raw domain pixels have higher values in the G (green) channel and relatively lower values in the B (blue) and R (red) channels. The optical flow motion vector diagram determination method of this application, while ensuring real-time algorithm performance, improves the accuracy of inter-frame optical flow estimation and reduces HDR video fusion anomalies, thereby enhancing the user's terminal shooting and viewing experience.
[0090] Specifically, when a user enables HDR video mode on their device, such as by clicking... Figure 11 The HDR mode function button shown indicates that the terminal camera will capture video frames at a fixed frame rate with alternating long and short exposures. The subsequent HDR video processing flow is as follows: Figure 12 As shown:
[0091] The terminal can determine at least one fusion strategy and input the target image (a set of images) into the HDR single-frame synthesis submodule. The HDR single-frame synthesis submodule selects alternating long and short exposure frames from DDR(a) memory and performs corresponding HDR single-frame fusion. At the same time, the optical flow estimation unit 1 in the HDR single-frame synthesis submodule applies the optical flow motion vector determination method of the embodiment of this application, calculates the optical flow information between the input long and short exposure frames based on the first and second images in the target image, determines the optical flow motion vector, and stores the calculation result in DDR(c) memory. The HDR single-frame synthesis submodule saves one copy of the long and short exposure fusion result in DDR(b) memory for use by the subsequent HDR video single-frame processing submodule, and directly inputs another copy into the HDR video single-frame processing submodule. The HDR video single-frame processing submodule reads the original long and short exposure data, optical flow data (such as optical flow motion vector), and HDR fusion image, and finally outputs an HDR image with the same texture, color, motion, and other features as the historical frames, as a frame in the HDR output video.
[0092] In the HDR video processing workflow, optical flow calculations are performed on the SoC GPU, while other functions of the HDR single-frame synthesis submodule are implemented on the SoC CPU; other functions of the HDR video single-frame processing submodule are implemented on the SoC NPU. This fully utilizes the performance of each chip on the SoC to output HDR video data in real time, while ensuring the stability of the HDR video.
[0093] In this way, real-time HDR video output can be achieved on the terminal side. By combining the optical flow information of HDR video frames and relying on an AI neural network model to improve the stability of HDR video frames, real-time online HDR video output can be achieved. This solution can be used to reduce HDR video jitter, and HDR video can accurately reflect the color range of real-world scenes.
[0094] Alternatively, specifically, after taking HDR photos using the terminal, the user can set up the HDR offline processing function. Once set up, the terminal's background processes will initiate the HDR video compositing effect when the current process is not busy. At this time, the original long and short exposure data captured by the front-end camera has been cleared from the DDR (Memory Transfer Record), so frame interpolation is required on the already composited HDR video. Simultaneously, the optical flow model is invoked to calculate the optical flow vector information of the preceding and following frames of the HDR video. The offline processing flow is as follows: Figure 13 As shown:
[0095] The HDR video is first split into frames, and the results are stored in DDR(a) memory. The split images are then fed into optical flow estimation unit 1, which outputs optical flow motion vector information (such as an optical flow vector image) between consecutive frames. The HDR video single-frame processing submodule is then invoked. This submodule self-trains the HDR video. After training, the trained model is used to regenerate each frame of the HDR video from the split HDR data, removing HDR video jitter. In this way, offline HDR video output can be achieved on the terminal side, and self-training can be performed specifically for HDR videos to better match the model with the HDR video content, resulting in more stable HDR videos.
[0096] It should be noted that the terminal uses a CPU-based control strategy to allocate corresponding GPU resources for each optical flow calculation, accelerating the optical flow calculation results and predicting the optical flow motion vector field between long and short exposures. Taking Lucas-Kanade optical flow estimation as an example, before calculating the optical flow motion information (such as the optical flow motion vector map) for long and short exposures, both images need to be converted to a uniform grayscale image. Then, the optical flow information of the two images is calculated according to the Lucas-Kanade method. After obtaining the optical flow motion vector information, pixel A in the previous frame can be moved to a new position according to the motion vector, and pixel B in the current frame can be moved to the position in the next frame according to the motion vector.
[0097] The method for determining an optical flow motion vector map provided in this application can be executed by a device for determining an optical flow motion vector map. This application uses an example of a device for determining an optical flow motion vector map executing the method to illustrate the device for determining an optical flow motion vector map provided in this application.
[0098] like Figure 14 As shown, an embodiment of this application provides a device 1400 for determining an optical flow motion vector diagram, comprising:
[0099] The first processing module 1410 is used to acquire a first image and a second image from the input image; wherein the first image and the second image have different brightness.
[0100] The second processing module 1420 is used to divide the first image and the second image into multiple regions according to the transmission channel of each pixel in the first image and the second image, respectively. The region types include low exposure region, normal exposure region and high exposure region.
[0101] The third processing module 1430 is used to determine the optical flow motion vector diagrams of the first image and the second image according to the region type to which each region belongs.
[0102] Optionally, the second processing module is further configured to:
[0103] Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are processed separately by channel to obtain the channel map of each channel;
[0104] The channel map is segmented step by step based on the segmentation strategy until the variance of the pixel values of the resulting segmented parts is less than a first threshold.
[0105] The region type to which each segment belongs is determined based on a pixel threshold.
[0106] Based on the region type and the location information of the segmented portion, multiple regions of the first image and the second image are determined.
[0107] Optionally, the third processing module is further configured to:
[0108] Obtain optical flow information for each of the aforementioned region types;
[0109] The optical flow information is fused to determine the optical flow motion vector diagrams of the first image and the second image.
[0110] After acquiring the first and second images from the input image, the device can use the transmission channels of each pixel in the first and second images to divide the first and second images into regions, distinguishing between low-exposure areas, normal-exposure areas, and high-exposure areas. Then, based on the region type of each region, it determines the optical flow motion vector diagrams of the first and second images. By considering different exposure levels of the images when determining the optical flow motion vector diagrams, the accuracy of the optical flow motion vector diagrams is improved.
[0111] The device for determining the optical flow motion vector diagram in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0112] The device for determining the optical flow motion vector diagram in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0113] The device for determining optical flow motion vector diagrams provided in this application embodiment can achieve... Figures 1-13 The various processes implemented in the embodiment of the method for determining the optical flow motion vector diagram shown are not described in detail here to avoid repetition.
[0114] Optionally, such as Figure 15 As shown, this application embodiment also provides an electronic device 1500, including a processor 1501 and a memory 1502. The memory 1502 stores a program or instructions that can run on the processor 1501. When the program or instructions are executed by the processor 1501, they implement the various steps of the above-described method embodiment for determining optical flow motion vector diagrams and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0115] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0116] Figure 16 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0117] The electronic device 1600 includes, but is not limited to, components such as: radio frequency unit 1601, network module 1602, audio output unit 1603, input unit 1604, sensor 1605, display unit 1606, user input unit 1607, interface unit 1608, memory 1609, and processor 1610.
[0118] Those skilled in the art will understand that the electronic device 1600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 16 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0119] The processor 1610 is used to acquire a first image and a second image from the input image; wherein the first image and the second image have different brightness.
[0120] Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are divided into multiple regions, and the region types include low exposure region, normal exposure region and high exposure region.
[0121] Based on the region type to which each region belongs, determine the optical flow motion vector diagrams of the first image and the second image.
[0122] After acquiring a first image and a second image from the input image, this electronic device can use the transmission channels of each pixel in the first and second images to divide the first and second images into regions, distinguishing between low-exposure areas, normal-exposure areas, and high-exposure areas. Furthermore, based on the region type of each region, it determines the optical flow motion vector diagrams for the first and second images. By considering different exposure levels of the images when determining the optical flow motion vector diagrams, the accuracy of the optical flow motion vector diagrams is improved.
[0123] Optionally, the processor 1610 is also used for:
[0124] Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are processed separately by channel to obtain the channel map of each channel;
[0125] The channel map is segmented step by step based on the segmentation strategy until the variance of the pixel values of the resulting segmented parts is less than a first threshold.
[0126] The region type to which each segment belongs is determined based on a pixel threshold.
[0127] Based on the region type and the location information of the segmented portion, multiple regions of the first image and the second image are determined.
[0128] Optionally, the pixel threshold includes a second threshold and a third threshold, wherein the second threshold is greater than the third threshold;
[0129] The 1610 processor is also used for:
[0130] If the average pixel value within the current segment is greater than or equal to the second threshold, then the region type of the current segment is a high-exposure region.
[0131] If the average pixel value within the current segment is less than the second threshold and greater than or equal to the third threshold, then the region type of the current segment is a normal exposure region.
[0132] If the average pixel value within the current segment is less than the third threshold, then the region type of the current segment is a low-exposure region.
[0133] Optionally, the processor 1610 is also used for:
[0134] Optical flow information is obtained by using the associated channel pixels based on each region type.
[0135] The optical flow information is fused to determine the optical flow motion vector diagrams of the first image and the second image.
[0136] Optionally, the processor 1610 is also used for:
[0137] Based on the first channel of each pixel in the first region of the first image, a first channel image is obtained;
[0138] A second channel map is obtained based on the second channel of each pixel in the second region of the second image;
[0139] Optical flow estimation is performed based on the first channel map and the second channel map to determine the optical flow information of the first region type;
[0140] Wherein, the region type to which the first region and the second region belong is the first region type, and the first region type is any one of the multiple region types, and the transmission channel type to which the first channel and the second channel belong is associated with the first region type.
[0141] It should be understood that, in this embodiment, the input unit 1604 may include a graphics processing unit (GPU) 16041 and a microphone 16042. The GPU 16041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1606 may include a display panel 16061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1607 includes at least one of a touch panel 16071 and other input devices 16072. The touch panel 16071 is also called a touch screen. The touch panel 16071 may include a touch detection device and a touch controller. Other input devices 16072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0142] The memory 1609 can be used to store software programs and various data. The memory 1609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1609 may include volatile memory or non-volatile memory, or it may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0143] Processor 1610 may include one or more processing units; optionally, processor 1610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1610.
[0144] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described method for determining optical flow motion vector diagrams and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0145] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0146] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described method embodiment for determining optical flow motion vector diagrams, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0147] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0148] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described method embodiment for determining optical flow motion vector diagrams, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0151] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for determining an optical flow motion vector diagram, characterized in that, Executed by the terminal, including: Obtain a first image and a second image from the input image; wherein the first image and the second image have different brightness; Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are divided into multiple regions, and the region types include low exposure region, normal exposure region and high exposure region. Based on the region type to which each region belongs, determine the optical flow motion vector diagrams of the first image and the second image; The step of determining the optical flow motion vector diagrams of the first image and the second image according to the region type to which each region belongs includes: Optical flow information is obtained by using the associated channel pixels based on each region type. The optical flow information is fused to determine the optical flow motion vector diagrams of the first image and the second image.
2. The method according to claim 1, characterized in that, The step of dividing the first image and the second image into multiple regions based on the transmission channel of each pixel in the first image and the second image includes: Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are processed separately by channel to obtain the channel map of each channel; The channel map is segmented step by step based on the segmentation strategy until the variance of the pixel values of the resulting segmented parts is less than a first threshold. The region type to which each segment belongs is determined based on a pixel threshold. Based on the region type and the location information of the segmented portion, multiple regions of the first image and the second image are determined.
3. The method according to claim 2, characterized in that, The pixel threshold includes a second threshold and a third threshold, wherein the second threshold is greater than the third threshold; The step of determining the region type of each segmented part based on a pixel threshold includes: If the average pixel value within the current segment is greater than or equal to the second threshold, then the region type of the current segment is a high-exposure region. If the average pixel value within the current segment is less than the second threshold and greater than or equal to the third threshold, then the region type of the current segment is a normal exposure region. If the average pixel value within the current segment is less than the third threshold, then the region type of the current segment is a low-exposure region.
4. The method according to claim 1, characterized in that, The step of calculating optical flow information based on each region type using associated channel pixels includes: Based on the first channel of each pixel in the first region of the first image, a first channel image is obtained; A second channel map is obtained based on the second channel of each pixel in the second region of the second image; Optical flow estimation is performed based on the first channel map and the second channel map to determine the optical flow information of the first region type; Wherein, the region type to which the first region and the second region belong is the first region type, and the first region type is any one of the multiple region types, and the transmission channel type to which the first channel and the second channel belong is associated with the first region type.
5. A device for determining an optical flow motion vector diagram, characterized in that, include: A first processing module is used to acquire a first image and a second image from an input image; wherein the first image and the second image have different brightness; The second processing module is used to divide the first image and the second image into multiple regions according to the transmission channel of each pixel in the first image and the second image, respectively. The region types include low exposure region, normal exposure region and high exposure region. The third processing module is used to determine the optical flow motion vector diagrams of the first image and the second image according to the region type to which each region belongs; The third processing module is also used for Optical flow information is obtained by using the associated channel pixels based on each region type. The optical flow information is fused to determine the optical flow motion vector diagrams of the first image and the second image.
6. The apparatus according to claim 5, characterized in that, The second processing module is also used for: Based on the transmission channel of each pixel in the first image and the second image, the first image and the second image are processed separately by channel to obtain the channel map of each channel; The channel map is segmented step by step based on the segmentation strategy until the variance of the pixel values of the resulting segmented parts is less than a first threshold. The region type to which each segment belongs is determined based on a pixel threshold. Based on the region type and the location information of the segmented portion, multiple regions of the first image and the second image are determined.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method for determining an optical flow motion vector diagram as described in any one of claims 1-4.
8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method for determining an optical flow motion vector diagram as described in any one of claims 1-4.
Citation Information
Patent Citations
Image processing method and related electronic equipment
CN113824873A
Image processing method and device, electronic equipment and storage medium
CN114418908A