Image processing method and electronic equipment
By combining the ToF sensor and the RGB sensor and utilizing multi-frame image processing technology and target models, the low resolution problem of the ToF sensor is solved, high-resolution image processing is achieved, and the image quality of the ToF sensor is improved.
Patent Information
- Application Number
- CN202510572607.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-19
AI Technical Summary
Existing ToF sensors have low resolution, making it difficult to achieve high-resolution complex scene recognition applications.
By combining the ToF sensor and the RGB sensor and using multi-frame image processing technology, the mapping relationship of the object change is determined, and the depth data is mapped to the RGB image coordinate system for fusion, and the target model is used to improve the image resolution.
The image resolution obtained by the ToF sensor is improved, high-resolution image processing is achieved, computing resources are low, and the implementation method is simple.
Smart Images

Figure CN120672573A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of electronic devices, and in particular to an image processing method and electronic device. Background Art
[0002] Time of Flight (ToF) sensors used in electronic devices often have very sparse sampling points. This limited functionality limits ToF sensors to a few applications, such as detecting user presence or recognizing simple gestures. Complex scene recognition applications requiring high resolution struggle. Consequently, images generated by current ToF sensors suffer from low resolution. Summary of the Invention
[0003] The present disclosure provides an image processing method and an electronic device to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided an image processing method, comprising:
[0005] Acquire a first image based on a first acquisition device, and acquire a second image based on a second acquisition device, wherein the first image includes information representing depth data;
[0006] Determining a first mapping relationship of object changes in the second image based on the plurality of frames of the second image, wherein the first mapping relationship is used to characterize the motion changes of the object in the second image of consecutive frames;
[0007] Processing the plurality of frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; wherein the second mapping relationship is used to represent a mapping relationship between depth data in the coordinate system of the first image and depth data in the coordinate system of the second image;
[0008] The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
[0009] In one embodiment, determining a first mapping relationship of object changes in the second image based on the second image of consecutive frames includes:
[0010] Acquire a current frame image and a previous frame image in a second image of continuous frames;
[0011] Based on objects present in the current frame image and the previous frame image, performing image segmentation on the current frame image and the previous frame image respectively to obtain object segmentation results of the current frame image and the previous frame image;
[0012] Based on the object segmentation results of the current frame image and the previous frame image, determining objects that exist in both the current frame image and the previous frame image;
[0013] Determining, based on an object present in both the current frame image and the previous frame image, a motion change of the object in the current frame image compared with the previous frame image;
[0014] Based on the motion change of the object in the current frame image compared with the previous frame image, a first mapping relationship of the object change in the second image is determined.
[0015] In one possible implementation manner, before fusing the multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image, the method further includes:
[0016] Acquiring a second mapping relationship between the first image and the second image; the acquiring the second mapping relationship between the first image and the second image includes:
[0017] Acquire a first initial image of the object captured by a first capturing device;
[0018] Acquire a second initial image of the object captured by a second capturing device;
[0019] determining a positional relationship between the first acquisition device and the second acquisition device based on a first position parameter of the object in the first initial image and a second position parameter of the object in the second initial image;
[0020] Based on the positional relationship between the first acquisition device and the second acquisition device, a second mapping relationship between the first image acquired by the first acquisition device and the second image acquired by the second acquisition device is determined.
[0021] In one possible implementation manner, the fusing of multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image to obtain the third image includes:
[0022] Based on the second mapping relationship, projecting the depth data of the multiple frames of the first image into the coordinate system of the second image;
[0023] Based on the first mapping relationship, projecting the depth data of the multiple frames of the first image in the coordinate system of the second image into the first image of the target frame;
[0024] The depth data in the first image projected onto the target frame is superimposed to obtain a third image; the resolution of the third image is greater than the resolution of the first image.
[0025] In one possible implementation manner, after obtaining the third image, the method further includes:
[0026] Mapping pixels in the third image to a pixel grid in the second image;
[0027] The pixel values of the positions where the pixels of the third image are missing in the pixel grid of the second image are assigned to zero, thereby obtaining a third image having the resolution of the second image.
[0028] In one embodiment, inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image includes:
[0029] determining an image edge of the object in the second image based on the object present in the second image;
[0030] Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image;
[0031] Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair;
[0032] Inputting the segmentation map pair into a target model to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image; the target model is used to set the resolution of the segmentation map of the corresponding third image in the segmentation map pair according to the resolution of the segmentation map of the second image;
[0033] The segmented images having the resolution of the second image are stitched together to obtain a fourth image having a resolution greater than that of the first image; wherein the resolution of the fourth image is greater than that of the third image.
[0034] In one embodiment, inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image includes:
[0035] determining an image edge of the object in the second image based on the object present in the second image;
[0036] Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image;
[0037] Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair;
[0038] Inputting the segmentation map pair into a target model, wherein the target model sets the resolution of the region with zero pixel value in the corresponding segmentation map in the third image based on the resolution of the segmentation map of the first image, so as to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image;
[0039] The segmented images having the resolution of the second image are stitched together to obtain a fourth image.
[0040] In one embodiment, the acquiring the first image based on the first acquisition device includes:
[0041] Setting the first acquisition device on the mobile device;
[0042] During the operation of the mobile device, the first acquisition device acquires a first image at multiple acquisition points; wherein the first acquisition device acquires multiple frames of the first image at each acquisition point; the first image includes depth data, and at least two frames of the first image acquired at the multiple acquisition points include different depth data.
[0043] According to a second aspect of the present disclosure, there is provided an electronic device, including:
[0044] at least one processor; and
[0045] a memory communicatively coupled to the at least one processor;
[0046] The first acquisition device,
[0047] The second acquisition device,
[0048] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform
[0049] Acquire a first image based on a first acquisition device, and acquire a second image based on a second acquisition device, wherein the first image includes information representing depth data;
[0050] Determining a first mapping relationship of object changes in the second image based on the plurality of frames of the second image, wherein the first mapping relationship is used to characterize the motion changes of the object in the second image of consecutive frames;
[0051] Processing the plurality of frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; wherein the second mapping relationship is used to represent a mapping relationship between depth data in the coordinate system of the first image and depth data in the coordinate system of the second image;
[0052] The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
[0053] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute
[0054] Acquire a first image based on a first acquisition device, and acquire a second image based on a second acquisition device, wherein the first image includes information representing depth data;
[0055] Determining a first mapping relationship of object changes in the second image based on the plurality of frames of the second image, wherein the first mapping relationship is used to characterize the motion changes of the object in the second image of consecutive frames;
[0056] Processing the plurality of frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; wherein the second mapping relationship is used to represent a mapping relationship between depth data in the coordinate system of the first image and depth data in the coordinate system of the second image;
[0057] The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
[0058] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0060] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0061] Figure 1 A schematic diagram of the implementation flow of the image processing method according to an embodiment of the present disclosure is shown;
[0062] Figure 2 A schematic diagram of capturing depth images using multiple ToF image acquisition points according to an embodiment of the present disclosure is shown;
[0063] Figure 3A schematic diagram of multiple acquisition points of a ToF image according to an embodiment of the present disclosure is shown;
[0064] Figure 4 A schematic diagram of multiple frames of RGB images according to an embodiment of the present disclosure is shown;
[0065] Figure 5 A schematic diagram of a process for constructing a first mapping relationship according to an embodiment of the present disclosure is shown;
[0066] Figure 6 A schematic diagram showing the positional relationship between the ToF sensor and the RGB sensor according to an embodiment of the present disclosure is shown;
[0067] Figure 7 A schematic diagram of a process for constructing a second mapping relationship according to an embodiment of the present disclosure is shown;
[0068] Figure 8 A schematic diagram of fusing multiple frames of first images according to an embodiment of the present disclosure is shown;
[0069] Figure 9 A schematic diagram of segmentation of an RGB image and a third image according to an embodiment of the present disclosure is shown;
[0070] Figure 10 A schematic diagram of splicing segmentation images of the third image in an embodiment of the present disclosure is shown;
[0071] Figure 11 A schematic structural diagram of an electronic device according to an embodiment of the present disclosure is shown;
[0072] Figure 12 A schematic diagram of the structure of another electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0073] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.
[0074] In related technologies, Time of Flight (ToF) depth sensors are mostly low-resolution sensors. The resolution of a ToF sensor is generally about 1 / 10 of that of a red, green, blue (RGB) image.
[0075] The image processing method and electronic device provided in this application can improve the resolution of images obtained by extremely sparse ToF sensors, with low computing resource usage and simple implementation.
[0076] An image processing method and electronic device provided by the present application are described below with reference to the accompanying drawings.
[0077] like Figure 1 As shown, the present application provides an image processing method, comprising:
[0078] S101, capturing a first image using a first capturing device and capturing a second image using a second capturing device, wherein the first image includes information representing depth data;
[0079] The first acquisition device may be a depth sensor, for example, a ToF sensor. The ToF sensor emits near-infrared light or laser pulses, collects the time difference between the light being emitted and the light being reflected back by the object, calculates the distance based on the speed of light, and generates depth information. The first image acquired by the ToF sensor in this application includes depth data. The resolution of a ToF sensor is typically low, generally not exceeding 640x480 pixels.
[0080] The second acquisition device can be an RGB sensor. An RGB sensor can detect the color of an object and convert it into an electrical signal output. It is based on the light intensity detection of three basic colors: red, green, and blue. By separately detecting the intensity of red, green, and blue light reflected or transmitted by the object, the object's color is determined based on the ratio of these three colors. An RGB sensor can accurately distinguish subtle differences in various colors and respond quickly to color changes. Its high resolution, typically 1920x1080 pixels, allows it to provide high-definition image quality.
[0081] For example, in this application, in the same scenario, a ToF sensor can be used to capture a first image of an object. The first image contains depth data, but the resolution of the first image is low. An RGB sensor can be used to capture a second image of the object. The resolution of the second image is higher than that of the first image.
[0082] S102, determining a first mapping relationship for a change of an object in the second image based on multiple frames of the second image, wherein the first mapping relationship is used to represent a motion change of the object in the second image of consecutive frames;
[0083] It is understandable that in the present application, the RGB sensor is fixed at a preset position, and the object may be a moving object. Since the object moves, it is necessary to align multiple frames of the second image first. Subsequently, the first image is aligned based on the first mapping relationship of the aligned second image. Aligning the first image is used to ensure the accuracy of the subsequent fusion of the first image data. Within a preset time period, the RGB sensor captures the moving object to obtain multiple frames of the second image. Through the changes of the object in the multiple frames of the second image, a first mapping relationship of the change of the object in the second image can be obtained. The first mapping relationship can be a motion mapping, which characterizes the motion change of the object in the second image. It is understandable that the object can also be a stationary object.
[0084] For example, the object may be a rectangular object. The RGB sensor captures multiple RGB images of the rectangular object in motion, and compares changes in a current RGB image with changes in a previous RGB image of the same object in the multiple RGB images to obtain a motion map of the object changes in the RGB images.
[0085] S103: Process the multiple frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; the second mapping relationship is used to represent a mapping relationship between the depth data in the coordinate system of the first image and the coordinate system of the second image;
[0086] It will be appreciated that in this application, the first and second acquisition devices simultaneously capture the subject, and therefore, the changes in the subject in the second image are similar to those in the first image. In other words, the first mapping relationship for the changes in the subject in the second image can be applied to the first image, thereby eliminating the effects of subject movement in the fusion of multiple frames of the first image.
[0087] The second mapping relationship between the first image and the second image can be a depth-RGB mapping obtained by the positional relationship between the first acquisition device and the second acquisition device. Based on the depth-RGB mapping, the depth data in the first image can be converted to the coordinate system of the RGB image. And based on the first mapping relationship in the coordinate system of the RGB image, the change of the object can be known, and then multiple frames of the first image are fused to obtain a third image. It can be understood that because the position of the object in the third image is different, the depth data in each third image is different. Therefore, by fusing multiple first images, more depth data is obtained, so that the resolution of the third image is higher than that of the first image.
[0088] For example, based on the depth-RGB mapping between the ToF sensor and the RGB sensor, the depth data of multiple depth images captured by the ToF sensor is converted to the coordinate system of the RGB image. Then, based on the motion mapping of the object changes in the RGB image, the multiple depth images are further fused with the first depth image of the multiple depth images as a reference, thereby obtaining a third image with more depth data. Because the third image has more depth data than the first image, the resolution of the third image is higher than that of the first image.
[0089] S104: Input the third image and the second image into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
[0090] The target model in this application is trained in advance, and the target model can set the resolution of the third image according to the resolution of the second image.
[0091] For example, the third image and the RGB image are input into the target model, and the target model sets the resolution of the third image according to the resolution of the RGB image, thereby obtaining a fourth image having the same resolution as the RGB image. The resolution of the fourth image is further improved relative to the resolution of the third image.
[0092] The image processing method provided by the present application obtains a first image with depth data through a first acquisition device and obtains a second image with RGB data through a second acquisition device. A second mapping relationship is obtained by changes in objects in multiple second images. A second mapping relationship is obtained by the positional relationship between the first acquisition device and the second acquisition device and the second image. Based on the second mapping relationship, the depth data in the first image is converted to the coordinate system of the second image. Based on the first mapping relationship, multiple frames of first images in the second image coordinate system are fused, that is, the depth data is fused to obtain a third image after fusion of the depth data. At this time, the resolution of the third image is initially improved relative to the resolution of the first image. The third image and the second image are then input into a target model, and the target model sets the resolution of the third image according to the resolution of the second image, thereby obtaining a fourth image with the same resolution as the second image. The resolution of the fourth image is further improved relative to the resolution of the third image. The image processing method provided by the present application improves the resolution of the image obtained by the ToF sensor, and the method for improving the resolution consumes low computing resources and is simple to implement.
[0093] In some embodiments, the capturing of the first image based on the first capturing device includes:
[0094] Setting the first acquisition device on the mobile device;
[0095] During the operation of the mobile device, the first acquisition device acquires a first image at multiple acquisition points; wherein the first acquisition device acquires multiple frames of the first image at each acquisition point; the first image includes depth data, and at least two frames of the first image acquired at the multiple acquisition points include different depth data.
[0096] In this application, the first acquisition device is mounted on a motion device, and the motion device can make the first acquisition device move at the sub-pixel level. Sub-pixels refer to virtual pixels obtained by calculation between two physical pixels. Since there is a certain distance between physical pixels (for example, 5.2 microns), these pixels are connected together at the macro level, but at the micro level, there are smaller "sub-pixels" between them. This enables the first acquisition device to collect depth data of multiple frames of the first image in a short period of time. Figure 2 As shown, the mobile device can be a motion device, for example, an xy linear motor can be selected (the xy linear motor is a vibration motor commonly used in smart terminals), and its motion track is as follows Figure 2 As shown. Figure 3 As shown, the motor can move in four directions: up, down, left, and right. The frequency is fixed, and the amplitude of the movement is determined by the control current / voltage. By adjusting the control current / voltage, the first acquisition device can be positioned at different locations between the two acquisition points to collect depth data at different locations.
[0097] For example, a ToF sensor is mounted on a vibration motor to collect ToF images within a preset time period. Figure 3 As shown, multiple acquisition points can be set in this application. When the vibration motor moves to the acquisition point, the ToF image captures the ToF image of the object, and each acquisition point can capture multiple frames of ToF images. Because the object also moves, the depth data of the multiple frames of ToF images collected by each acquisition point is different. This can obtain more depth data during fusion, improving the resolution of the fused image.
[0098] In some embodiments, determining a first mapping relationship of object changes in the second image based on the second image of consecutive frames includes:
[0099] Acquire a current frame image and a previous frame image in a second image of continuous frames;
[0100] Based on objects present in the current frame image and the previous frame image, performing image segmentation on the current frame image and the previous frame image respectively to obtain object segmentation results of the current frame image and the previous frame image;
[0101] Based on the object segmentation results of the current frame image and the previous frame image, determining objects that exist in both the current frame image and the previous frame image;
[0102] Determining, based on an object present in both the current frame image and the previous frame image, a motion change of the object in the current frame image compared with the previous frame image;
[0103] Based on the motion change of the object in the current frame image compared with the previous frame image, a first mapping relationship of the object change in the second image is determined.
[0104] In the present application, a second image of a continuous frame is first captured by a second acquisition device. The motion of the object in the second image can be determined by comparing the changes in the object in the previous frame with the next frame. In this way, the present application segments the objects in both the current frame and the previous frame to obtain segmentation results for the two images. By comparing the objects that exist in both the segmentation results and checking the motion changes of the objects that exist in both, a first mapping relationship representing the changes in the objects in the second image can be obtained using 3D rotation and displacement projection methods.
[0105] The second image is an RGB image. When performing object segmentation on an RGB image, object segmentation can be performed by performing edge detection on the objects in the RGB image. Edge detection uses existing edge detection methods, such as the OpenCV image edge detection method. Specifically, object edges are extracted by detecting the extreme points of the first-order derivative or the zero-crossing points of the second-order derivative of the RGB image. Edges are then screened by detecting features such as brightness, color, and texture of pixels at the edge junctions of the objects to form closed-loop edges, completing object segmentation.
[0106] For example, Figure 4 and Figure 5 As shown, first, the edge of each object in each frame of the RGB image is extracted, and the RGB image is divided into multiple object images according to the edge of each object in the RGB image. Because the time interval between the acquisition of two consecutive frames of RGB images is short (for example, 50ms), new objects that appear / disappear during the acquisition process can be ignored. According to the edge changes of the same object in the previous and next two frames (such as the images in the nth frame and the n+1th frame), the movement of the object is calculated, and then the 3D rotation and displacement projection method is used to construct a motion mapping of each frame of RGB image to the previous frame of RGB image. Among them, the 3D rotation and displacement projection method can be implemented using the existing method, which will not be repeated in this application.
[0107] In some embodiments, before fusing the multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image, the method further includes:
[0108] Acquiring a second mapping relationship between the first image and the second image; the acquiring the second mapping relationship between the first image and the second image includes:
[0109] Acquire a first initial image of the object captured by a first capturing device;
[0110] Acquire a second initial image of the object captured by a second capturing device;
[0111] determining a positional relationship between the first acquisition device and the second acquisition device based on a first position parameter of the object in the first initial image and a second position parameter of the object in the second initial image;
[0112] Based on the positional relationship between the first acquisition device and the second acquisition device, a second mapping relationship between the first image acquired by the first acquisition device and the second image acquired by the second acquisition device is determined.
[0113] In the present application, the positions of the first acquisition device and the second acquisition device are different, and it is necessary to first align the acquisition data of the first acquisition device and the second acquisition device to obtain a second mapping relationship between the first image and the second image. The way to obtain the second mapping relationship between the first image and the second image is to first acquire a first initial image of the object through the first acquisition device and acquire a second initial image of the calibration object through the second acquisition device. Therefore, based on the position of the calibration object, the conversion relationship, that is, the position relationship, between the position of the first acquisition device and the position of the second acquisition device is determined. According to the position relationship between the first acquisition device and the second acquisition device, the second mapping relationship between the first image acquired by the first acquisition device and the second image acquired by the second acquisition device can be determined.
[0114] In the present application, the first acquisition device and the second acquisition device are aligned to collect data, that is, the measurement points in the first acquisition device and the pixel points in the second acquisition device are aligned. Specifically, the same plane rectangle is used as the calibration object. The calibration object is collected simultaneously by the first acquisition device and the second acquisition device, based on the positional relationship between the calibration object in the initial image collected by the first acquisition device and the initial image collected by the second acquisition device. The positional relationship includes the rotation and displacement of the calibration object. That is, the position of the calibration object in the first image can be obtained by how the rotation and position of the calibration object are obtained in the second image. In this way, the first mapping relationship between the sampling points of the first acquisition device and the pixel points of the second acquisition device is determined.
[0115] For example, Figure 6 As shown in the figure, when capturing images of an object, the RGB sensor and the ToF sensor have different installation positions, angles, and capture ranges. Figure 7As shown, in this application, a ToF sensor is placed at each acquisition position, where the acquisition position can be pre-set. A smooth calibration object is placed in front of the ToF sensor and RGB sensor. This application aligns the edges of the calibration object with the outermost sampling points of the ToF sensor and records the acquisition results of the ToF and RGB sensors. Aligning the edges of the calibration object with the edges of the ToF sensor sampling points facilitates the subsequent determination of the positional relationship between the ToF and RGB sensors. By placing the calibration object at the edge of the ToF sensor sampling points, the position of the calibration object is clearly visible in the resulting ToF image; simply aligning the calibration object in the RGB image with the object in the ToF image is sufficient. If the edges of the calibration object are uncertain, the position of the calibration object in the ToF image must first be determined, and then the calibration object in the RGB image must be aligned with the object in the ToF image. The acquisition range of the RGB sensor is larger than that of the ToF sensor, and the RGB image captured by the RGB sensor can include the entire calibration object. The positional relationship between the two sensors is inferred based on the pixel position of the calibration object in the RGB image. Rotation and displacement are then used to construct the positional relationship between each ToF sensor sampling point and each pixel of the RGB sensor, i.e., the depth-RGB mapping. Thus, a depth-RGB mapping of the ToF image captured by the ToF sensor and the RGB image captured by the RGB sensor is obtained.
[0116] In some embodiments, fusing multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image to obtain a third image includes:
[0117] Based on the second mapping relationship, projecting the depth data of the multiple frames of the first image into the coordinate system of the second image;
[0118] Based on the first mapping relationship, projecting the depth data of the multiple frames of the first image in the coordinate system of the second image into the first image of the target frame;
[0119] The depth data in the first image projected onto the target frame is superimposed to obtain a third image; the resolution of the third image is greater than the resolution of the first image.
[0120] In the present application, based on the second mapping relationship, the depth data of multiple frames of first images can be mapped to the coordinate system of the second image, and then the depth data of the first images in the coordinate system of the second image can be fused based on the first mapping relationship with the target frame as the reference, thereby superimposing the depth data in the multiple first images to obtain a third image with the depth data superimposed. Because the depth data of the multiple first images are fused, the resolution of the third image is improved, making the resolution of the third image higher than that of the first image.
[0121] For example, Figure 8 As shown, the depth data of multiple frames of ToF images are input into the depth-RGB mapping, thereby projecting the depth data of each frame of the ToF image into the RGB coordinate system of the RGB image. Then, based on the motion mapping of the RGB image previously obtained, the depth data in the RGB coordinate system is mapped to the first frame of the ToF image. The depth data of all frames of the ToF image are thus superimposed, resulting in a third image with a preliminary increased resolution. The third image is essentially a sparse depth map after the preliminary increased resolution.
[0122] In some embodiments, after obtaining the third image, the method further includes:
[0123] Mapping pixels in the third image to a pixel grid in the second image;
[0124] The pixel values of the positions where the pixels of the third image are missing in the pixel grid of the second image are assigned to zero, thereby obtaining a third image having the resolution of the second image.
[0125] In the present application, after obtaining the third image, although the resolution of the third image is improved compared to the resolution of the first image, it is still lower than the resolution of the second image. Therefore, in the present application, the pixels in the third image are mapped to the pixel grid of the second image, and the pixel values of the positions in the pixel grid of the second image where the pixels of the third image are missing are assigned to zero, thereby obtaining a third image with the resolution of the second image.
[0126] For example, the pixels of the third image may be arranged in a 10×10 matrix, while the pixel grid of the RGB image may be arranged in a 1000×1000 matrix. The 10×10 pixels of the third image are mapped to the 1000×1000 matrix. Except for the mapped 10×10 pixels, the pixel values of the unmapped pixels in the 1000×1000 matrix are assigned to zero. Thus, the third image has the resolution of the RGB image.
[0127] In some embodiments, inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image includes:
[0128] determining an image edge of the object in the second image based on the object present in the second image;
[0129] Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image;
[0130] Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair;
[0131] Inputting the segmentation map pair into a target model to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image; the target model is used to set the resolution of the segmentation map of the corresponding third image in the segmentation map pair according to the resolution of the segmentation map of the second image;
[0132] The segmented images having the resolution of the second image are stitched together to obtain a fourth image having a resolution greater than that of the first image; wherein the resolution of the fourth image is greater than that of the third image.
[0133] In the present application, an edge detection method is used to determine the edges of objects in the second image, and the objects in the second image are segmented according to the edges of the objects, thereby obtaining a segmentation map of each object in the second image. Because the positions of the objects in the second image and the third image are completely corresponding, that is, they can overlap. Therefore, the objects in the third image are segmented based on the edges of the objects in the second image, and a segmentation map of the objects in the third image is obtained. Thus, the segmentation map of the objects in the second image and the segmentation map of the objects in the third image constitute a segmentation map pair. The segmentation map pair is input into the target model, and the resolution of the segmentation map of the third image is set according to the resolution of the segmentation map of the second image based on the segmentation map pair target model, thereby obtaining a segmentation map with the resolution of the second image. Finally, the segmentation maps of the third image with the resolution of the second image are spliced to obtain a fourth image. The fourth image has the resolution of the second image.
[0134] The target model in this application can adopt a neural network model. For example, the neural network model can be composed of a plurality of depth prediction submodules connected together. The main body of the depth prediction submodule can be composed of two downsampling convolution layers and two upsampling deconvolution layers. Other networks or submodules can also be used, and this application does not limit this. The target model provided in this application can predict a high-resolution segmented image from the second image and the third image. Before output, a convolution layer is used to merge channels to splice the segmentation map and output the final fourth image. It can be understood that the target model used in this application is a pre-trained model, which has the characteristics of high accuracy and low complexity. The target model in this application can also be composed of other structures, which this application does not limit here.
[0135] For example, Figure 9As shown, the objects in the RGB image are segmented according to the edges to obtain the RGB segmentation map of each object. Based on the overlap between the RGB image and the third image, the objects in the third image are segmented by edges to obtain the third image segmentation map. It can be understood that the RGB segmentation map and the segmentation map of the third image that can overlap form a segmentation map pair. The segmentation map pair is input into the target model, and the target model sets the resolution of the third image segmentation map in the segmentation map pair according to the resolution of the RGB segmentation map. Figure 10 As shown, finally, the third image segmentation maps with improved resolution are spliced and the junctions are smoothed to form a complete high-resolution depth map, i.e., the fourth image.
[0136] In some embodiments, inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image includes:
[0137] determining an image edge of the object in the second image based on the object present in the second image;
[0138] Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image;
[0139] Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair;
[0140] Inputting the segmentation map pair into a target model, wherein the target model sets the resolution of the region with zero pixel value in the corresponding segmentation map in the third image based on the resolution of the segmentation map of the first image, so as to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image;
[0141] The segmented images having the resolution of the second image are stitched together to obtain a fourth image.
[0142] In the present application, the segmentation map of the object in the second image and the segmentation map of the object in the third image constitute a segmentation map pair. In the above embodiment, the positions with pixels in the third image have been mapped to the pixel grid in the second image. The pixel values of the positions where the pixels of the third image are missing in the pixel grid are assigned to zero, and a third image with the resolution of the second image is obtained. The segmentation map pair is input into the target model, and the target model sets the pixel values of the third image in the segmentation map pair to zero according to the pixel values of the RGB segmentation map. Finally, the segmentation maps of the third image with increased resolution are spliced, and the junctions are smoothed to form a complete high-resolution depth map, i.e., the fourth image.
[0143] For example, the RGB segmentation map and the overlapping third image segmentation map form a segmentation map pair. The segmentation map pair is input into the target model, which sets the pixel values of the third image segmentation map in the segmentation map pair to zero based on the pixel values of the RGB segmentation map. Finally, the segmentation maps of the third image with increased resolution are spliced together, and the boundaries are smoothed to form a fourth image.
[0144] The data processing method provided in this application can enhance the resolution of extremely sparse ToF sensors to obtain higher-resolution images.
[0145] like Figure 11 As shown, the present application provides an electronic device, including: at least one processor 1101; and
[0146] a memory 1102 communicatively connected to the at least one processor 1101;
[0147] The first acquisition device 1103,
[0148] The second acquisition device 1104,
[0149] in,
[0150] The memory 1102 stores instructions that can be executed by the at least one processor 1101. The instructions are executed by the at least one processor 1101 to enable the at least one processor 1101 to execute
[0151] Acquire a first image using the first acquisition device 1103 and acquire a second image using the second acquisition device 1104, wherein the first image includes information representing depth data;
[0152] Determining a first mapping relationship of object changes in the second image based on the plurality of frames of the second image; wherein the first mapping relationship is used to characterize the motion changes of the object in the second image of consecutive frames;
[0153] Processing the plurality of frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; wherein the second mapping relationship is used to represent a mapping relationship between depth data in the coordinate system of the first image and depth data in the coordinate system of the second image;
[0154] The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
[0155] The electronic device provided in the present application captures a first image through a first acquisition device 1103 and a second acquisition device 1104, and determines a first mapping relationship of object changes in the second image based on multiple frames of the second image; processes the multiple frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; the second mapping relationship is used to represent the mapping relationship of depth data in the coordinate system of the first image to the coordinate system of the second image; the third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent the resolution of the third image set according to the resolution of the second image.
[0156] Exemplarily, the first acquisition device 1103 is a ToF sensor, the second acquisition device 1104 is an RGB sensor, and the processor 1101 controls the first acquisition device 1103 and the second acquisition device 1104 .
[0157] The electronic device provided by this application also includes:
[0158] The mobile device 1105 is used to move the first collection device 1103 to the collection point according to the location of the collection point.
[0159] The mobile device 1105 may be a sports device.
[0160] Exemplarily, a ToF sensor is mounted on a moving device and captures ToF images of a moving object at pre-set capture points. An RGB sensor captures RGB images of the moving object at a fixed position. Based on motion mapping of multiple RGB image frames and depth-RGB mapping between the ToF and RGB images, the multiple ToF image frames are fused to generate a third image. The RGB and third images are input into a target model, which sets the resolution of the corresponding area in the third image based on the resolution of the RGB image to generate a fourth image. The fourth image has the resolution of the RGB image, thereby improving the resolution of the ToF image.
[0161] It should be noted that the image processing device of the embodiment of the present application solves the problem in a similar principle to the aforementioned image processing method. Therefore, the implementation process, implementation principle, and beneficial effects of the image processing device can all be referred to the description of the implementation process, implementation principle, and beneficial effects of the aforementioned method, and the repeated parts will not be repeated.
[0162] An embodiment of the present application provides an electronic device, including:
[0163] at least one processor; and
[0164] a memory communicatively connected to the at least one processor; wherein,
[0165] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in any one of the above embodiments.
[0166] An embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in any of the above embodiments.
[0167] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0168] Figure 12 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0169] like Figure 12 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0170] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0171] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the image processing method. For example, in some embodiments, the image processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the image processing method by any other appropriate means (e.g., by means of firmware).
[0172] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0173] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0174] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0176] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0177] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0178] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0179] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0180] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. An image processing method, comprising: Acquire a first image based on a first acquisition device, and acquire a second image based on a second acquisition device, wherein the first image includes information representing depth data; Determining a first mapping relationship of object changes in the second image based on the multiple frames of the second image; The first mapping relationship is used to represent the motion change of the object in the second image of the continuous frames; Processing the plurality of frames of the first image based on the first mapping relationship and the second mapping relationship between the first image and the second image to obtain a third image; The second mapping relationship is used to represent a mapping relationship between depth data in the first image coordinate system and depth data in the second image coordinate system; The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
2. The method according to claim 1, wherein determining a first mapping relationship of object changes in the second image based on the second image of consecutive frames comprises: Acquire a current frame image and a previous frame image in a second image of continuous frames; Based on objects present in the current frame image and the previous frame image, performing image segmentation on the current frame image and the previous frame image respectively to obtain object segmentation results of the current frame image and the previous frame image; Determine objects that exist in both the current frame image and the previous frame image based on object segmentation results of the current frame image and the previous frame image; Determining, based on an object present in both the current frame image and the previous frame image, a motion change of the object in the current frame image compared with the previous frame image; Based on the motion change of the object in the current frame image compared with the previous frame image, a first mapping relationship of the object change in the second image is determined.
3. The method according to claim 1, before fusing the multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image, further comprising: Acquire a second mapping relationship between the first image and the second image; The acquiring a second mapping relationship between the first image and the second image includes: Acquire a first initial image of the object captured by a first capturing device; Acquire a second initial image of the object captured by a second capturing device; determining a positional relationship between the first acquisition device and the second acquisition device based on a first position parameter of the object in the first initial image and a second position parameter of the object in the second initial image; Based on the positional relationship between the first acquisition device and the second acquisition device, a second mapping relationship between the first image acquired by the first acquisition device and the second image acquired by the second acquisition device is determined.
4. The method according to claim 1, wherein fusing multiple frames of first images based on the first mapping relationship and the second mapping relationship between the first image and the second image to obtain a third image comprises: Based on the second mapping relationship, projecting the depth data of the multiple frames of the first image into the coordinate system of the second image; Based on the first mapping relationship, projecting the depth data of the multiple frames of the first image in the coordinate system of the second image into the first image of the target frame; The depth data in the first image projected onto the target frame is superimposed to obtain a third image; the resolution of the third image is greater than the resolution of the first image.
5. The method according to claim 1, further comprising: after obtaining the third image; Mapping pixels in the third image to a pixel grid in the second image; The pixel values of the positions where the pixels of the third image are missing in the pixel grid of the second image are assigned to zero, thereby obtaining a third image having the resolution of the second image.
6. The method according to claim 1, wherein inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image comprises: determining an image edge of the object in the second image based on the object present in the second image; Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image; Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair; Inputting the segmentation map pair into a target model to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image; the target model is used to set the resolution of the segmentation map of the corresponding third image in the segmentation map pair according to the resolution of the segmentation map of the second image; The segmented images having the resolution of the second image are stitched together to obtain a fourth image having a resolution greater than that of the first image; wherein the resolution of the fourth image is greater than that of the third image.
7. The method according to claim 5, wherein inputting the third image and the second image into a target model to obtain a fourth image having a resolution greater than a resolution of the first image comprises: determining an image edge of the object in the second image based on the object present in the second image; Segmenting the second image based on image edges of objects in the second image to obtain a segmentation map of the first image; Segmenting the third image based on image edges of the second image to obtain a segmentation map of the third image; the segmentation map of the third image and the segmentation map of the same area in the second image constitute a segmentation map pair; Inputting the segmentation map pair into a target model, wherein the target model sets the resolution of the region with zero pixel value in the corresponding segmentation map in the third image based on the resolution of the segmentation map of the first image, so as to obtain a segmentation map having the resolution of the second image corresponding to each segmentation map of the third image; The segmented images having the resolution of the second image are stitched together to obtain a fourth image.
8. The method according to claim 1, wherein acquiring the first image based on the first acquisition device comprises: Setting the first acquisition device on the mobile device; During the operation of the mobile device, the first acquisition device acquires a first image at multiple acquisition points; wherein the first acquisition device acquires multiple frames of the first image at each acquisition point; the first image includes depth data, and at least two frames of the first image acquired at the multiple acquisition points include different depth data.
9. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The first acquisition device, The second acquisition device, in, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform Acquire a first image based on a first acquisition device, and acquire a second image based on a second acquisition device, wherein the first image includes information representing depth data; Determining a first mapping relationship of object changes in the second image based on the plurality of frames of the second image; wherein the first mapping relationship is used to characterize the motion changes of the object in the second image of consecutive frames; Processing the plurality of frames of the first image based on the first mapping relationship and a second mapping relationship between the first image and the second image to obtain a third image; wherein the second mapping relationship is used to represent a mapping relationship between depth data in the coordinate system of the first image and depth data in the coordinate system of the second image; The third image and the second image are input into a target model to obtain a fourth image having a resolution greater than that of the first image; the target model is used to represent that the resolution of the third image is set according to the resolution of the second image.
10. The electronic device according to claim 9, further comprising: The mobile device is used to move the first collection device to the collection point according to the location of the collection point.