Image processing method and related apparatus

By capturing images simultaneously using camera devices with different pixel sizes and combining this with image processing methods, the problem of low camera frame rate was solved, achieving high dynamic range and high resolution images, thus improving the user experience.

CN119052659BActive Publication Date: 2025-11-07HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310621936.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2025-11-07
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

The current technology uses cameras with low frame rates, resulting in low video frame rates and affecting user experience.

Method used

By using a first camera device and a second camera device with different unit pixel sizes to capture images of the same exposure time at the same moment, a target image is obtained based on the first image and the second image, and the frame rate of the image is improved by utilizing images with different brightness dynamic ranges.

Benefits of technology

The increased frame rate ensures high dynamic range and resolution of images, enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052659B_ABST
    Figure CN119052659B_ABST
Patent Text Reader

Abstract

The application provides an image processing method and related device, which can be used in the technical field of image processing. In the technical scheme provided by the application, different unit pixel size camera devices are used to capture images with the same exposure time at the same time to obtain images with different brightness dynamic ranges, and the images with different brightness dynamic ranges are processed to obtain images with higher brightness dynamic ranges. In the application, the same exposure time is used for capturing images by different unit pixel size camera devices to obtain images with different brightness dynamic ranges. Compared with obtaining images with different brightness dynamic ranges by different exposure times, more images can be captured in the same time, so that more images with high brightness dynamic ranges can be obtained, that is, the frame rate can be improved, and the user experience can be finally improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an image processing method and related apparatus. BACKGROUND

[0002] Video perspective technology refers to a technology of collecting surrounding environment images by a camera and displaying the images on a display screen. The video perspective technology has high requirements on the resolution, frame rate and image quality of the camera, and the dynamic range in the image quality is a main factor affecting the sense of reality. The image dynamic range refers to a luminance range of a scene that can be captured by an image.

[0003] Since the dynamic range of the camera is far less than the dynamic range that can be perceived by the human eye, a current method for improving the image dynamic range is high dynamic range (HDR) imaging technology, that is, a high dynamic range image is obtained by fusing pictures of different exposure ranges, so that the captured image or video is closer to the human perception.

[0004] However, when the above method is used to improve the image dynamic range, it is found that the frame rate of the camera is low, thereby resulting in a low video frame rate. SUMMARY

[0005] The present application provides an image processing method and related apparatus, which can improve the video frame rate and further improve the user experience.

[0006] In a first aspect, the present application provides an image processing method, in which the first camera device and the second camera device are controlled to capture images of the same exposure time at the same time to obtain a first image and a second image, the unit pixel size of the first camera device is different from the unit pixel size of the second camera device, the first image is an image of the same exposure time captured by the first camera device at the same time, and the second image is an image of the same exposure time captured by the second camera device at the same time; a target image is obtained based on the first image and the second image, and the luminance dynamic range of the target image is higher than the luminance dynamic range of the first image and the luminance dynamic range of the second image.

[0007] In this method, the images of different luminance dynamic ranges are obtained by the camera devices with different unit pixel sizes for the same exposure time, and the image of high dynamic luminance range is obtained based on the images of different luminance dynamic ranges. Compared with obtaining the images of different luminance dynamic ranges by different exposure times, the exposure time of the image can be saved, and thus the frame rate of the image can be improved.

[0008] In the present application, the target image is acquired based on the first image and the second image, which can be understood as that the image of high luminance dynamic range is acquired based on the images of low luminance dynamic range. The low luminance dynamic range and the high luminance dynamic range are relative concepts, which mainly refer to the relative height and low of the luminance dynamic range between the first image, the second image and the target image.

[0009] In the present application, the implementation manner of acquiring the image of high luminance dynamic range based on the images of low luminance dynamic range is not limited.

[0010] In some possible implementation manners, the target image can be acquired based on the first image and the second image, which can include: matching the first image and the second image based on the similarity between the first image and the second image to obtain a first matching result; acquiring a disparity map of the first image and a disparity map of the second image based on the first matching result; acquiring a depth map of the first image based on the disparity map of the first image; acquiring a depth map of the second image based on the disparity map of the second image; performing matching processing on the first image and the second image based on the depth map of the first image and the depth map of the second image to obtain a second matching result; and performing luminance fusion processing on the first image and the second image according to the second matching result to obtain the target image.

[0011] In the implementation manner, the matching accuracy can be improved, and thus the image quality of the target image can be improved.

[0012] In some possible implementation manners, the matching of the first image and the second image based on the similarity between the first image and the second image can include: performing target processing on the first image and the second image respectively to obtain a third image and a fourth image, and the target processing includes binocular stereo rectification processing; and matching the first image and the second image based on the similarity between the third image and the fourth image.

[0013] In some possible implementation manners, the matching processing of the first image and the second image based on the depth map of the first image and the depth map of the second image can include: mapping the first image into the coordinate system of the first camera device based on the conversion relationship between the depth map of the first image and the coordinate system of the first image, to obtain a first three-dimensional image; converting the first three-dimensional image into the coordinate system of the second camera device based on the conversion relationship between the coordinate system of the first camera device and the coordinate system of the second camera device, to obtain a second three-dimensional image; mapping the second image into the coordinate system of the second camera device based on the conversion relationship between the depth map of the second image and the coordinate system of the second image, to obtain a third three-dimensional image; and determining the second matching result of the first image and the second image based on the matching result of the second three-dimensional image and the third three-dimensional image.

[0014] In a second aspect, the present application provides an image processing apparatus, which can include various functional modules for implementing the method in the first aspect. For example, the apparatus includes an interaction module and a processing module.

[0015] In some implementations, the modules can be implemented by software and / or hardware. For example, the interaction module can be implemented by a communication interface, and the processing module can be implemented by a processor executing program codes stored in a memory. Optionally, the memory can also be included.

[0016] It can be understood that the processing apparatus provided in the second aspect can also be a chip system.

[0017] In a third aspect, the present application provides a computer readable storage medium, which stores program codes for apparatus execution, and the program codes include instructions for implementing the method in the first aspect.

[0018] In a fourth aspect, the present application provides a computer program product containing instructions, which, when the computer program product is run on an apparatus, causes the apparatus to implement the method in the first aspect.

[0019] In a fifth aspect, the present application provides an image processing apparatus, which includes a first camera device, a second camera device and a processing unit, the unit pixel sizes of the first camera device and the second camera device are different, and the processing unit is configured to implement the method in the first aspect or any possible implementation manner thereof.

[0020] In a sixth aspect, the present application provides a terminal device, which includes the image processing apparatus in the fifth aspect. For example, the terminal device can be a mediated reality (MR) glasses.

[0021] It can be understood that the effects obtained by the second aspect to the sixth aspect can refer to the description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 An exemplary architecture diagram of an image processing system according to an embodiment of the present application;

[0023] Figure 2 A schematic diagram of an MR glasses according to an embodiment of the present application;

[0024] Figure 3 An exemplary flowchart of an image processing method according to an embodiment of the present application;

[0025] Figure 4 An exemplary diagram of disparity and depth according to an embodiment of the present application;

[0026] Figure 5 An exemplary structural diagram of an image processing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0028] In order to clearly describe the technical solutions in the embodiments of the present application, in the embodiments of the present application, the same items or similar items with basically the same functions and effects are distinguished by using "first", "second", and the like. Those skilled in the art can understand that "first", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different.

[0029] It should be noted that in the embodiments of the present application, "exemplary" or "for example" is used to mean an example, an illustration, or a description. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Rather, "exemplary" or "for example" is used as a specific manner to present the relevant concept.

[0030] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0031] Figure 1 An exemplary architecture diagram of an image processing system according to an embodiment of the present application. As shown in the figure, the image processing system 100 can include a camera 101, a camera 102, a processor 103, a memory 104, a display 105, and a display 106. Figure 1

[0032] The camera 101, the camera 102, the processor 103, the memory 104, the display 105, and the display 106 communicate with each other through internal connection paths.

[0033] ​The camera 101 and the camera 102 have basic functions such as video shooting and / or still image capturing, and the like. After an image is captured by a lens, the image is processed and converted into a digital signal recognizable by a processor by a photosensitive component circuit and a control component in the camera.

[0034] The camera 101 and the camera 102 each include an image sensor having photosensitive elements thereon. The photosensitive elements can divide a light image on a light receiving surface thereof into a plurality of small units and convert the light image into usable electrical signals.

[0035] For photosensitive elements having the same number of effective pixels, the larger the size of the photosensitive elements, the larger the unit area of each pixel, and the better the photosensitive performance, and thus more image details can be recorded. The unit area of each pixel can also be referred to as a unit pixel area, a single pixel size, or a unit pixel size. One way to calculate the unit pixel area is: unit pixel area = total number of pixels / physical photosensitive screen area of the image sensor.

[0036] The unit pixel size of the camera 101 is different from the unit pixel size of the camera 102, and thus the photosensitivity of the camera 101 is different from the photosensitivity of the camera 102, and further, the dynamic range of an image captured by the camera 101 is different from the dynamic range of an image captured by the camera 102.

[0037] In some implementations, the unit pixel size of the camera 101 is greater than the unit pixel size of the camera 102, the photosensitivity of the camera 101 is higher than the photosensitivity of the camera 102, and the dynamic range of an image captured by the camera 101 is greater than the dynamic range of an image captured by the camera 102.

[0038] In other implementations, the unit pixel size of the camera 101 is smaller than the unit pixel size of the camera 102, the photosensitivity of the camera 101 is lower than the photosensitivity of the camera 102, and the dynamic range of an image captured by the camera 101 is lower than the dynamic range of an image captured by the camera 102.

[0039] The memory 104 is configured to store instructions. Optionally, the memory 104 can include read-only memory and random access memory, and provide instructions and data to the processor 103. A portion of the memory can also include non-volatile random access memory.

[0040] For example, the memory 104 can also store device type information.

[0041] The processor 103 can be configured to execute instructions stored in the memory 104, and when the processor 103 executes the instructions stored in the memory 103, the processor 103 can be configured to perform the steps and / or processes of any one of the method embodiments.

[0042] It should be understood that, in the embodiments of the present application, the processor 103 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0043] The display 105 and the display 106 can receive signals sent by the processor, form images and display the images. For example, the display 105 displays images obtained by the processor 103 processing images captured by the camera 101, and the display 106 displays images obtained by the processor 103 processing images captured by the camera 102.

[0044] It can be understood that, Figure 1 The structure of the image processing system shown is only an example, and the image processing system proposed in the present application can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. For example, the image processing system proposed in the present application can also include a communication interface, a power supply interface, or more cameras, etc.

[0045] In some embodiments of the present application, a terminal device is also proposed, and the terminal device proposed in the embodiments of the present application includes Figure 1 The image processing system shown. As an example, the terminal device can be a smartphone, XR glasses, MR glasses, a camera, etc.

[0046] Figure 2 A schematic diagram of the MR glasses of one embodiment of the present application. Figure 2 The MR glasses 200 shown can include Figure 1 The image processing system shown, Figure 2 Only the camera 101 and the camera 102 in the image processing system 100 are shown in the image processing system.

[0047] The camera 101 and the camera 102 can be referred to as binocular cameras of the MR glasses. The placement positions of the camera 101 and the camera 102 in the MR glasses 200 are as shown in Figure 2 The camera 101 can be referred to as a left-eye camera, and the camera 102 can be referred to as a right-eye camera.

[0048] As an example, the MR glasses 200 are equipped with a camera with a size of 2.0 micrometers (μm) on the left side and a camera with a size of 1.0 μm on the right side.

[0049] As an example, the MR glasses include a left mirror and a right mirror. The left mirror includes a display 105 for displaying an image processed after being captured by a left eye camera, and the right mirror includes a display 106 for displaying an image processed after being captured by a right eye camera.

[0050] Figure 3 An example flowchart of an image processing method according to an embodiment of the present application. The image processing method can be implemented by the processor 103 in the image processing system described above. As shown in the flowchart, the image processing method can include S310 and S320. Figure 3

[0051] S310, controlling a first camera and a second camera to capture images with a same exposure time at a same time, to obtain a first image and a second image, a unit pixel size of the first camera being different from a unit pixel size of the second camera.

[0052] As an example, the first camera can be the camera 101, and the second camera can be the camera 102.

[0053] In some implementations, the unit pixel size of the first camera is larger than the unit pixel size of the second camera; in other implementations, the unit pixel size of the first camera is smaller than the unit pixel size of the second camera.

[0054] As an example, the processor can send an instruction to the first camera and the second camera, instructing the first camera and the second camera to capture images.

[0055] In some possible implementations, the processor can instruct the first camera and the second camera with a same starting time and a same exposure time. In this way, the first camera and the second camera can start capturing images at the same time and capture images with the same exposure time.

[0056] In this embodiment, for convenience of description, the image captured by the first camera is referred to as a first image, and the image captured by the second camera is referred to as a second image.

[0057] Because the first camera and the second camera capture the first image and the second image at the same time, the first image and the second image are images of the same scene.

[0058] ​Because the unit pixel size of the first camera is different from the unit pixel size of the second camera, the sensitivities of the first camera and the second camera are different; because the exposure time of the first image and the second image is the same, the dynamic range of the first image is different from the dynamic range of the second image.

[0059] For example, if the unit pixel size of the first camera is greater than the unit pixel size of the second camera, the sensitivity of the first camera is greater than the sensitivity of the second camera, the dynamic range of the first image is greater than the dynamic range of the second image, the first image contains more details of the high-light area in the scene, and the second image contains more details of the low-light area in the scene.

[0060] In other implementations, the unit pixel size of the first camera is less than the unit pixel size of the second camera, the sensitivity of the first camera is less than the sensitivity of the second camera, the dynamic range of the first image is less than the dynamic range of the second image, the second image contains more details of the high-light area in the scene, and the first image contains more details of the low-light area in the scene.

[0061] S320, obtaining a target image based on the first image and the second image, the brightness dynamic range of the target image being higher than the brightness dynamic range of the first image and the brightness dynamic range of the second image.

[0062] This step can be understood as obtaining an image with a higher dynamic range based on the first image and the second image, and the image with the higher dynamic range is referred to as a target image.

[0063] For example, after the first camera and the second camera capture the first image and the second image under the control of the processor, the processor can perform image fusion processing based on the first image and the second image to obtain the target image.

[0064] In this embodiment, the target image can include one or more. For example, based on the first image, the first image is fused in brightness using the second image to obtain a first target image; based on the second image, the second image is fused in brightness using the first image to obtain a second target image.

[0065] In this embodiment, because the unit pixel size of the first camera and the unit pixel size of the second camera are different, the first camera and the second camera can capture images with different dynamic ranges in the same exposure time, and the purpose of obtaining an image with a higher dynamic range based on images with different dynamic ranges can be achieved.

[0066] Compared with the technical solution of acquiring images of different dynamic ranges by long and short exposure time, the technical solution can avoid more time consumed by long exposure time, so that images of different dynamic ranges can be acquired in shorter time, more images of different dynamic ranges can be acquired in a certain time, i.e. more high dynamic range images can be acquired in a certain time, and finally the frame rate can be improved.

[0067] In addition, when the target image contains multiple target images, for example, the first target image and the second target image, because the first image and the second image can be acquired by the same exposure time, the shooting time of the first image and the second image can be ensured to be the same, so that the time sequence of the first target image and the second target image can be ensured to be aligned.

[0068] In addition, because the first image and the second image are respectively acquired by two independent cameras, the resolution of the first image and the second image can be ensured, so that the resolution of the high dynamic range target image finally acquired can be ensured.

[0069] The embodiment does not limit the method or way of acquiring a higher dynamic range image based on the first image and the second image. Some implementation manners of acquiring a target image of a higher dynamic range based on the first image and the second image are introduced below.

[0070] In some possible implementation manners, the following operation can be included: performing matching processing on the first image and the second image to obtain a matching result; and performing luminance fusion processing on the first image and the second image according to the matching result to obtain a target image.

[0071] The matching result is used to indicate which pixel point in the first image matches which pixel point in the second image, and which pixel point in the first image matches which pixel point in the second image can be understood as which pixel point in the first image and which pixel point in the second image are the same point in the scene.

[0072] As an example, the matching result can be a conversion relationship between the coordinates of the pixel points in the first image and the coordinates of the pixel points in the second image, and the pixel points in the first image and the pixel points in the second image that satisfy the coordinate conversion relationship can be considered to match.

[0073] The luminance fusion processing on the first image and the second image according to the matching result can be understood as the luminance of the matched pixel points in the first image and the second image is fused, and the obtained pixel points constitute the target image.

[0074] For example, the brightness of the pixel points in the first image is fused using the brightness of the matching pixel points in the second image to obtain a first target image, and the brightness of the pixel points in the second image is fused using the brightness of the matching pixel points in the first image to obtain a second target image.

[0075] Taking the first camera as a left-eye camera on the MR glasses and the second camera as a right-eye camera on the MR glasses as an example, the first target image can be an image displayed by a left lens on the MR glasses, and the second target image can be an image displayed by a right lens on the MR glasses.

[0076] In some possible implementation manners, the operation of performing matching processing on the first image and the second image can include: matching the first image and the second image based on the similarity of the first image and the second image to obtain a first matching result; obtaining a disparity map of the first image relative to the second image and a disparity map of the second image relative to the first image based on the first matching result; obtaining a depth map of the first image based on the disparity map of the first image relative to the second image; obtaining a depth map of the second image based on the disparity map of the second image relative to the first image; and matching the first image and the second image based on the depth map of the first image and the depth map of the second image to obtain a second matching result, which can be used for brightness fusion.

[0077] The matching of the first image and the second image based on the similarity of the first image and the second image, the obtaining of the spatial depth of the pixel points in the first image based on the first matching result, and the matching of the first image and the second image based on the spatial depth of the pixel points in the first image can refer to related technologies in the prior art.

[0078] In this implementation manner, after the matching of the first image and the second image based on the similarity of the first image and the second image, the matching is further performed based on the depth, which can improve the matching accuracy, thereby improving the fusion quality and further improving the quality of the target image. The matching based on the depth can improve the efficiency of the matching based on the depth, thereby improving the efficiency of obtaining the target image.

[0079] In this implementation manner, the disparity map of the first image relative to the second image can be understood as a set of differences between the coordinates of the pixel points in the first image on the x axis and the coordinates of the matching pixel points in the second image on the x axis, and the disparity map of the second image relative to the first image can be understood as a set of differences between the coordinates of the pixel points in the second image on the x axis and the coordinates of the matching pixel points in the first image on the x axis. The matching pixel points are the pixel points indicated by the first matching result.

[0080] In this implementation, the depth map of the first image can be understood as a set of distances between scene points corresponding to pixel points in the first image and the first camera device, and the depth map of the second image can be understood as a set of distances between scene points corresponding to pixel points in the second image and the second camera device.

[0081] Figure 4 An example diagram of disparity and depth for an embodiment of the present application. In the diagram, f is the focal length of the camera device; xl represents the offset of a pixel point Pl in the first image relative to the center of the first image, or the coordinate of the pixel point on the x-axis; xr represents the offset of a pixel point Pr in the second image relative to the center of the second image, or the coordinate of the pixel point on the x-axis; b represents the baseline distance between the two camera devices; P represents a point on the scene, Pl represents the pixel point corresponding to P in the first image; Pr represents the pixel point corresponding to P in the second image; Z represents the Z-axis, X represents the X-axis, and Y is not shown; and z represents the depth of P.

[0082] From the diagram, it can be seen that the disparity d of the first image is xl-xr, and z=b*f / d. Figure 4

[0083] Before matching the first image and the second image based on the similarity of the first image and the second image, the first image can be subjected to target processing, which can include one or more of image equalization processing, distortion correction processing, and binocular stereo correction processing; and the second image can be subjected to target processing. For the sake of convenience, the image obtained by subjecting the first image to target processing can be referred to as a third image, and the image obtained by subjecting the first image to target processing can be referred to as a fourth image.

[0084] It can be understood that when the target processing includes multiple processing of image equalization processing, distortion correction processing, and binocular stereo correction processing, the present embodiment does not limit the execution order of the multiple processing. As an example, when the target processing includes image equalization processing, distortion correction processing, and binocular stereo correction processing, the execution order is image equalization processing, distortion correction processing, and binocular stereo correction processing in turn.

[0085] In this way, the brightness of the first image and the second image can be adjusted to be closer, thereby improving the matching accuracy of the first image and the second image, and further improving the quality of the target image.

[0086] In this way, the brightness of the first image and the second image can be adjusted to be closer, thereby improving the matching accuracy of the first image and the second image, and further improving the quality of the target image.

[0087] ​The binocular stereo correction processing is performed on the first image and the second image respectively, and the matching efficiency of the first image and the second image can be improved, and then the efficiency of obtaining the target image can be improved.

[0088] In the embodiment, the binocular stereo correction processing is performed on the first image and the second image respectively, so that the imaging origin coordinates of the corrected first image and the second image are consistent, the optical axes are parallel, and the planes are coplanar. In this way, any pixel point on the corrected first image and the matching point on the corrected second image are in the same row of pixel points, so that the matching pixel point in the other image can be obtained by performing one-dimensional search on the row of pixel points in the corrected first image or the corrected second image, and the matching efficiency can be improved.

[0089] In some possible implementation manners of the embodiment, after the disparity map of the first image or the second image is obtained, the time domain smoothing and the hole filling operation can be performed on the disparity map, and then the depth map of the first image or the second image is obtained based on the disparity map obtained by the operation. The depth map obtained in this way is more accurate, the matching result of the first image and the second image is more accurate, and the quality of the target image obtained by fusion is higher.

[0090] The time domain smoothing operation can include: aligning a previous frame disparity map and a current frame disparity map through a front-rear frame pose relationship, comparing depth value changes of corresponding poses, and setting a threshold to filter out noise points with large depth value changes. For example, in the embodiment, a disparity map of a previous image captured by the first camera (or the second camera) at a time after the first image (or the second image) can be obtained, the disparity map of the previous image is aligned with the disparity map of the first image (or the second image) based on the pose relationship between the first image (or the second image) and the previous image, the depth value transformation is compared, and the pixel points with large depth value changes in the disparity map of the first image (or the second image) are filtered out based on the set change threshold. The pixel point can be referred to as a noise point.

[0091] The hole filling operation can include: establishing a global optimization equation according to a neighborhood pixel constraint relationship of the disparity map, solving the missing disparity value, and simultaneously performing a spatial smoothing operation on the existing disparity value to finally obtain a smooth dense disparity map. For example, in the embodiment, a global optimization equation can be established based on the neighborhood pixel constraint relationship of the disparity map of the first image (or the second image), the missing disparity value is solved, and a spatial smoothing operation is performed on the existing disparity value to finally obtain a smooth dense disparity map of the first image (or the second image). The dense disparity map can be used as the final disparity map of the first image (or the second image), and the depth map of the first image (or the second image) is determined based on the final disparity map and image matching is performed.

[0092] In some possible implementation manners of the embodiment, matching the first image and the second image based on the depth map of the first image and the depth map of the second image can include: mapping the first image into the coordinate system of the first camera device based on the depth map of the first image and a conversion relationship between the coordinate system of the first image and the coordinate system of the first camera device, to obtain a first three-dimensional image; converting the first three-dimensional image into the coordinate system of the second camera device based on a conversion relationship between the coordinate system of the first camera device and the coordinate system of the second camera device, to obtain a second three-dimensional image; mapping the second image into the coordinate system of the second camera device based on the depth map of the second image and a conversion relationship between the coordinate system of the second image and the coordinate system of the second camera device, to obtain a third three-dimensional image; and determining the second matching result of the first image and the second image based on a matching result of the second three-dimensional image and the third three-dimensional image.

[0093] In some implementation manners, determining the second matching result of the first image and the second image based on the matching result of the second three-dimensional image and the third three-dimensional image can include: finding two three-dimensional space points in the second three-dimensional image that match the two three-dimensional space points in the third three-dimensional image based on the second three-dimensional image and the third three-dimensional image, wherein a point in the two three-dimensional space points that belongs to the second three-dimensional image is referred to as a first three-dimensional space point, and a point in the two three-dimensional space points that belongs to the third three-dimensional image is referred to as a second three-dimensional space point; finding a three-dimensional space point in the first three-dimensional image that maps to the first three-dimensional space point; continuing to find a pixel point in the first image that maps to the three-dimensional space point in the first three-dimensional image; finding a pixel point in the second image that maps to the three-dimensional space point in the third three-dimensional image; and determining that the pixel point in the first image and the pixel point in the second image are matching pixel points.

[0094] In some possible implementation manners of the embodiment, matching the first image and the second image based on the depth map of the first image and the depth map of the second image can include: mapping the first image into the coordinate system of the first camera device based on the depth map of the first image and a conversion relationship between the coordinate system of the first image and the coordinate system of the first camera device, to obtain a first three-dimensional image; converting the first three-dimensional image into the coordinate system of the second camera device based on a conversion relationship between the coordinate system of the first camera device and the coordinate system of the second camera device, to obtain a second three-dimensional image; mapping the second image into the coordinate system of the second camera device based on the depth map of the second image and a conversion relationship between the coordinate system of the second image and the coordinate system of the second camera device, to obtain a third three-dimensional image; and determining the second matching result of the first image and the second image based on a matching result of the second three-dimensional image and the third three-dimensional image.

[0095] In some implementations, determining the second matching result of the first image and the second image based on the matching result of the fifth three-dimensional image and the sixth three-dimensional image can include: based on the fifth three-dimensional image and the sixth three-dimensional image, finding two three-dimensional space points in the fifth three-dimensional image that match two three-dimensional space points in the sixth three-dimensional image, the point in the two three-dimensional space points that belongs to the fifth three-dimensional image is referred to as a third three-dimensional space point, and the point in the two three-dimensional space points that belongs to the sixth three-dimensional image is referred to as a fourth three-dimensional space point; finding a three-dimensional space point in the fourth three-dimensional image that maps to the third three-dimensional space point, and then finding a pixel point in the second image that maps to the three-dimensional space point in the fourth three-dimensional image; finding a pixel point in the first image that maps to the three-dimensional space point in the sixth three-dimensional image; and determining that the pixel point in the first image and the pixel point in the second image are matching pixel points.

[0096] In some implementations, determining the second matching result of the first image and the second image based on the matching result of the fifth three-dimensional image and the sixth three-dimensional image can include: based on the fifth three-dimensional image and the sixth three-dimensional image, finding two three-dimensional space points in the fifth three-dimensional image that match two three-dimensional space points in the sixth three-dimensional image, the point in the two three-dimensional space points that belongs to the fifth three-dimensional image is referred to as a third three-dimensional space point, and the point in the two three-dimensional space points that belongs to the sixth three-dimensional image is referred to as a fourth three-dimensional space point; finding a three-dimensional space point in the fourth three-dimensional image that maps to the third three-dimensional space point, and then finding a pixel point in the second image that maps to the three-dimensional space point in the fourth three-dimensional image; finding a pixel point in the first image that maps to the three-dimensional space point in the sixth three-dimensional image; and determining that the pixel point in the first image and the pixel point in the second image are matching pixel points.

[0097] In some implementations, determining the second matching result of the first image and the second image based on the matching result of the fifth three-dimensional image and the sixth three-dimensional image can include: based on the fifth three-dimensional image and the sixth three-dimensional image, finding two three-dimensional space points in the fifth three-dimensional image that match two three-dimensional space points in the sixth three-dimensional image, the point in the two three-dimensional space points that belongs to the fifth three-dimensional image is referred to as a third three-dimensional space point, and the point in the two three-dimensional space points that belongs to the sixth three-dimensional image is referred to as a fourth three-dimensional space point; finding a three-dimensional space point in the fourth three-dimensional image that maps to the third three-dimensional space point, and then finding a pixel point in the second image that maps to the three-dimensional space point in the fourth three-dimensional image; finding a pixel point in the first image that maps to the three-dimensional space point in the sixth three-dimensional image; and determining that the pixel point in the first image and the pixel point in the second image are matching pixel points.

[0098] In some possible implementation manners of the embodiment, the brightness fusion processing is performed on the first image and the second image based on the matching result, including: for the common view area, a weight is calculated based on the brightness of the pixel point in the first image and the brightness of the matched pixel point in the second image, and then fusion is performed according to the weight; for the non-common view area, brightness adjustment is performed on the non-common view area according to the overall brightness change of the common view area, and an excessive area is added. The common view area can be understood as an area formed by the matched pixel points.

[0099] One implementation manner of the brightness fusion in the embodiment includes: combining three weight information of contrast C i,j,k , saturation S i,j,k , and exposure E i,j,k , and the overall higher weight tends to be given to the pixel with high contrast, high saturation, and good exposure. At this time, one expression for calculating the fusion weight is as follows:

[0100] W i,j,k =C i,j,k wc *S i,j,k *E i,j,k

[0101] Where i, j, and k represent the row coordinate (x-axis coordinate), the column coordinate (y-axis coordinate), respectively.

[0102] One implementation manner of the brightness adjustment in the embodiment includes: respectively counting the average adjustment proportion Ratio gray(I) of different gray values of the fusion area 0-255 in the first image and the second image, and the average brightness adjustment proportion Ratio avg of the overall fusion area, then respectively traversing the non-fusion area in the first image and the second image, and the brightness value in the non-fusion area is adjusted according to the current gray value. One calculation manner of the brightness value I adjust of the pixel point in the non-fusion area is as follows:

[0103] I adjust =d*(w*Ratio gray(I) +(1-w)*Ratio avg )

[0104] Where w is the weight of the average local and global, and the value range is 0 to 1, and the default is 0.5, which can be adjusted according to the effect.

[0105] In one implementation manner of the brightness transition of the embodiment, the calculation manner of the brightness value I transition of the transition area is as follows:

[0106] I transition =d*I merge +(1-d)*I adjust

[0107] wherein d is the weight of two regions, the value of the transition region is 0-1, the farther the distance from the fusion region, the smaller the value; the transition region size is generally 1% of the full image resolution, which can be adjusted according to the actual effect (the boundary is obvious, and the transition range is increased). merge The brightness of the matched pixel point in the co-view region is obtained.

[0108] In some embodiments of the present application, after obtaining the target image, the target image can be displayed on the display screen, for example, the first target image is displayed on the left mirror of the MR glasses, and the second target image is displayed on the right mirror of the MR glasses.

[0109] In some embodiments of the present application, an image processing apparatus is also provided, which is used to implement the image processing method in any of the foregoing method embodiments.

[0110] Figure 5 An exemplary structural diagram of the image processing apparatus of an embodiment of the present application is shown in FIG. 5. As shown in FIG. 5, the image processing apparatus 500 can include a control module 501 and a processing module 502. Figure 5

[0111] As an example, the image processing apparatus 500 can be used to implement the processing method of the embodiment shown in FIG. 3. For example, the control module 501 can be used to perform S310, and the processing module 502 can be used to perform S320. Figure 3

[0112] The present application also provides an image processing apparatus, which can include a processor 103. Optionally, the image processing apparatus can also include a memory 104.

[0113] Optionally, the image processing apparatus can also include a camera 101 and a camera 102.

[0114] The present application also provides a computer storage medium, which stores an image processing program. When the image processing program is executed by a processor, the steps of the image processing program method according to any of the above embodiments are implemented.

[0115] The specific embodiments of the computer storage medium of the present application are basically the same as the above-mentioned embodiments of the image processing program method of the present application, and are not repeated here.

[0116] The present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the image processing method of the present application according to any of the above embodiments are implemented, which are not repeated here.

[0117] ​​The above application embodiment serial numbers are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0118] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and a necessary general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) as described above, and includes a plurality of instructions for causing a terminal device (which can be a TWS earphone or the like) to execute the methods described in the various embodiments of the present application.

Claims

1. An image processing method, characterized by, The method comprises the following steps: controlling the first camera and the second camera to capture images with the same exposure time at the same time, to obtain a first image and a second image, the unit pixel size of the first camera being different from the unit pixel size of the second camera, the first image being an image with the same exposure time captured by the first camera at the same time, and the second image being an image with the same exposure time captured by the second camera at the same time; acquiring a target image based on the first image and the second image, the brightness dynamic range of the target image being higher than the brightness dynamic range of the first image and the brightness dynamic range of the second image.

2. The method of claim 1, wherein, The step of acquiring the target image based on the first image and the second image comprises the following steps: matching the first image and the second image based on the similarity between the first image and the second image, to obtain a first matching result; acquiring a disparity map of the first image and a disparity map of the second image based on the first matching result; acquiring a depth map of the first image based on the disparity map of the first image; acquiring a depth map of the second image based on the disparity map of the second image; performing matching processing on the first image and the second image based on the depth map of the first image and the depth map of the second image, to obtain a second matching result; performing brightness fusion processing on the first image and the second image according to the second matching result, to obtain the target image.

3. The method of claim 2, wherein, The step of matching the first image and the second image based on the similarity between the first image and the second image comprises the following steps: performing target processing on the first image and the second image respectively, to obtain a third image and a fourth image, the target processing comprising binocular stereo rectification processing; matching the first image and the second image based on the similarity between the third image and the fourth image.

4. The method according to claim 2 or 3, characterized in that, The step of performing matching processing on the first image and the second image based on the depth map of the first image and the depth map of the second image comprises the following steps: mapping the first image into the coordinate system of the first camera based on the conversion relationship between the depth map of the first image and the coordinate system of the first image and the coordinate system of the first camera, to obtain a first three-dimensional image; converting the first three-dimensional image into the coordinate system of the second camera based on the conversion relationship between the coordinate system of the first camera and the coordinate system of the second camera, to obtain a second three-dimensional image; mapping the second image into the coordinate system of the second camera based on the conversion relationship between the depth map of the second image and the coordinate system of the second image and the coordinate system of the second camera, to obtain a third three-dimensional image; determining the second matching result of the first image and the second image based on the matching result of the second three-dimensional image and the third three-dimensional image.

5. An image processing apparatus characterized by comprising: The method comprises modules for performing the method as claimed in any one of claims 1 to 4.

6. An image processing apparatus characterized by comprising: An apparatus comprising a processor coupled with a memory storing instructions that, when executed by the processor, cause the apparatus to perform the method of any one of claims 1 to 4.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed, causes the method of any one of claims 1 to 4 to be performed.

8. A computer program product, characterised in that, A computer program that, when run, causes the method of any one of claims 1 to 4 to be performed.

9. An image processing apparatus characterized by comprising: The first camera device, the second camera device and the processing unit, a unit pixel size of the first camera device is different from a unit pixel size of the second camera device: The processing unit is configured to: control the first camera device and the second camera device to capture images of a same exposure duration at a same time to obtain a first image and a second image, and obtain a target image based on the first image and the second image, the unit pixel size of the first camera device being different from the unit pixel size of the second camera device, the first image being the image of the same exposure duration captured by the first camera device at the same time, the second image being the image of the same exposure duration captured by the second camera device at the same time, and a luminance dynamic range of the target image being higher than luminance dynamic ranges of the first image and the second image.

10. The apparatus of claim 9, wherein, The processing unit is configured to, when obtaining the target image based on the first image and the second image: match the first image and the second image based on similarity between the first image and the second image to obtain a first matching result, and obtain a disparity map of the first image and a disparity map of the second image based on the first matching result; obtain a depth map of the first image based on the disparity map of the first image; obtain a depth map of the second image based on the disparity map of the second image; perform matching processing on the first image and the second image based on the depth map of the first image and the depth map of the second image to obtain a second matching result; perform luminance fusion processing on the first image and the second image based on the second matching result to obtain the target image.

11. The apparatus of claim 10, wherein, The processing unit is configured to match the first image and the second image based on similarity between the first image and the second image, and specifically configured to: perform target processing on the first image and the second image respectively to obtain a third image and a fourth image, the target processing including binocular rectification processing, and match the first image and the second image based on similarity between the third image and the fourth image.

12. The apparatus of claim 10 or 11, wherein, The processing unit is configured to perform matching processing on the first image and the second image based on the depth map of the first image and the depth map of the second image, and specifically configured to: map the first image into a coordinate system of the first camera device based on a depth map of the first image and a conversion relationship between the coordinate system of the first image and the coordinate system of the first camera device, to obtain a first three-dimensional image; convert the first three-dimensional image into a coordinate system of the second camera device based on a conversion relationship between the coordinate system of the first camera device and the coordinate system of the second camera device, to obtain a second three-dimensional image; map the second image into the coordinate system of the second camera device based on a depth map of the second image and a conversion relationship between the coordinate system of the second image and the coordinate system of the second camera device, to obtain a third three-dimensional image; determine the second matching result of the first image and the second image based on a matching result of the second three-dimensional image and the third three-dimensional image.

13. A terminal device, comprising: An image processing device as claimed in any one of claims 9 to 12.

14. The terminal device according to claim 13, characterized by The terminal device is a mediated reality (MR) glasses.

Citation Information

Patent Citations

  • Image processing method, device, electronic device, and storage medium

    CN109218627A

  • Image processing method and device, storage medium and electronic equipment

    CN110381263A