Image processing method, electronic device, and computer program product
By pre-processing the images captured by the dual cameras, the problem of dizziness during dual-camera 3D photography is solved, generating clear three-dimensional images and avoiding dizziness caused by excessive parallax.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2024-12-02
- Publication Date
- 2026-06-02
AI Technical Summary
When using dual cameras for 3D photography, if the target object is close enough to the camera or when taking macro photos, the resulting image may have excessive binocular parallax, causing the user to feel dizzy.
By acquiring at least two images from different perspectives, it is determined whether the subject meets preset conditions, such as being too close to the camera or having too large a parallax. Images that meet the conditions are then processed using preset methods, such as blurring, occlusion, or replacement, to generate a three-dimensional image.
It effectively reduces the dizziness experienced by the human eye when viewing 3D images, while preserving the complete content of the image, allowing users to clearly view the image within the entire field of view.
Smart Images

Figure CN122137945A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an image processing method, electronic device, and computer program product. Background Technology
[0002] The carriers of naked-eye 3D (3D) technology typically include smart terminals such as tablets and mobile phones. To enrich the functionality of these terminals and broaden their business scope, 3D photography and video recording have become required features for these terminals. With the development of smart terminals, in order to meet user needs, terminals are usually equipped with dual cameras, whether front-facing or rear-facing, which can be used to achieve 3D photography to generate 3D images for display on naked-eye 3D screens.
[0003] However, when using dual cameras to take 3D photos and generate 3D images, the following problems usually exist: when the target object is close enough to the camera or when taking macro photos, the binocular parallax of the image generated by the object close to the camera is too large, which exceeds the human brain's ability to form three-dimensional images, causing users to feel dizzy. Summary of the Invention
[0004] This application provides an image processing method, an electronic device, and a computer program product.
[0005] This application provides an image processing method, which includes:
[0006] Acquire at least two images from different perspectives of the same shooting scene, wherein the at least two images are respectively captured by at least two camera devices, and each camera device captures one image;
[0007] If the subject in at least one of the at least two images meets the preset conditions, the subject in any one of the at least two images that meets the preset conditions shall be subject to preset processing.
[0008] Generate a three-dimensional image based on the pre-processed image and other images.
[0009] This application provides an electronic device, including: one or more processors; and a memory storing one or more computer programs thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement any of the image processing methods in this application.
[0010] This application provides a computer program product, which includes a computer program that, when executed by a processor, implements any of the image processing methods described in this application.
[0011] According to the image processing method, electronic device and computer program product provided in the embodiments of this application, by performing preset processing on any one of at least two images in which the subject meets the preset conditions, the dizziness caused by the human eye when viewing three-dimensional stereoscopic images can be effectively reduced, while the human eye can still view the image content within the entire field of view.
[0012] Further details regarding the above embodiments and other aspects of this application, as well as their implementations, are provided in the accompanying drawings, detailed description, and claims. Attached Figure Description
[0013] In the accompanying drawings of the embodiments of this application:
[0014] Figure 1 This diagram illustrates a flowchart of an image processing method provided in an embodiment of this application.
[0015] Figure 2 This diagram illustrates a flowchart of another image processing method provided in an embodiment of this application.
[0016] Figure 3 This diagram illustrates a flowchart of yet another image processing method provided in an embodiment of this application.
[0017] Figure 4 A schematic diagram showing the parallax between the same subject in images corresponding to any two camera devices;
[0018] Figure 5 This diagram illustrates a preset processing method for objects in an image.
[0019] Figure 6 This diagram illustrates a three-dimensional image generated from two images.
[0020] Figure 7 A schematic diagram of a three-dimensional stereoscopic image switching display is shown;
[0021] Figure 8 This is a block diagram illustrating the composition of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the technical solutions of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0023] The present application will be described more fully below with reference to the accompanying drawings; however, the embodiments shown may be embodied in different forms, and the present application should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that this application will be thorough and complete, and will enable those skilled in the art to fully understand the scope of the application.
[0024] The accompanying drawings of the embodiments of this application are used to provide a further understanding of the embodiments of this application and constitute a part of the specification. They are used together with the detailed embodiments to explain this application and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the description of the detailed embodiments with reference to the accompanying drawings.
[0025] This application can be described with reference to plan views and / or cross-sectional views using the ideal schematic diagram of this application. Therefore, the example illustrations can be modified according to manufacturing techniques and / or tolerances.
[0026] Where there is no conflict, the various embodiments of this application and the features thereof may be combined with each other.
[0027] The terminology used in this application is for describing specific embodiments only and is not intended to limit the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated enumerated entries. The singular forms "a" and "the" as used herein are also intended to include the plural forms unless the context clearly indicates otherwise. The terms "comprising," "made of," etc., as used herein, specify the presence of the stated feature, integral, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof.
[0028] Unless otherwise specified, all terms used in this application (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined in this application.
[0029] This application is not limited to the embodiments shown in the accompanying drawings, but includes modifications to the configuration based on the manufacturing process. Therefore, the areas illustrated in the drawings are schematic, and the shapes of the areas shown in the drawings illustrate the specific shapes of the areas of the element, but are not intended to be limiting.
[0030] Please see Figure 1 , Figure 1 This illustration shows a flowchart of an image processing method provided in an embodiment of this application. The embodiment of this application provides an image processing method, such as... Figure 1As shown, the image processing method in this application embodiment includes, but is not limited to, the following steps.
[0031] Step S11: Obtain at least two images from different perspectives of the same shooting scene. The at least two images are captured by at least two camera devices, with each camera device capturing one image.
[0032] In this embodiment, there are no special restrictions on the placement of the at least two camera devices, as long as they can capture at least two images from different perspectives in the same shooting scene. For example, in the same shooting scene, at least two camera devices can be placed adjacent to each other and the lenses of at least two camera devices can be on the same horizontal line, thereby using at least two camera devices to capture at least two images from different perspectives.
[0033] Step S12: If the subject in at least one of the at least two images meets the preset conditions, perform preset processing on the subject in any one of the at least one images that meets the preset conditions.
[0034] In this embodiment, after acquiring at least two images, it is detected whether there is a subject in each image that meets preset conditions. If at least one of the at least two images contains a subject that meets the preset conditions, then the subject in that image that meets the preset conditions is subjected to preset processing, so that the subject in that image after preset processing has a visual difference from the corresponding subject in other images. The subject can be one or more.
[0035] The preset processing may include, but is not limited to: blurring, occlusion or replacement, for example, using black blocks, gray blocks, white blocks or blocks of the same color as the background in the image to occlude a subject in an image that meets the preset conditions.
[0036] In some embodiments, if no subject exists in any of the at least two images that meet the preset conditions, no further processing is performed, and a three-dimensional image is directly generated based on the at least two images.
[0037] Step S13: Generate a three-dimensional image based on the pre-processed image and other images.
[0038] According to the image processing method of this application embodiment, by performing preset processing on any one of the at least two images captured by at least two camera devices where the subject meets preset conditions, the subject that meets the preset conditions in the image will have a visual difference from the corresponding subject in other images. This will allow the generated three-dimensional stereoscopic image to be displayed on a naked-eye 3D display screen. For subjects that do not meet the preset conditions, the user's left and right eyes will see a clear stereoscopic image in the brain when viewing the display screen. For subjects that meet the preset conditions, one eye will see the image with preset processing while the other eye will see a clear image. Subjects that meet the preset conditions will form a clear planar image in the brain. This can effectively reduce the dizziness caused by viewing three-dimensional stereoscopic images. Since the preset processing is only performed on the image where the subject meets the preset conditions, while the other images retain the complete content of the image, the human eye can still view the image content within the entire field of view.
[0039] In some embodiments, the distance parameter between the object in the image and the corresponding camera device is obtained to determine whether the object in the image meets the preset conditions. Figure 2 This document illustrates a flowchart of another image processing method provided in an embodiment of this application, as shown below. Figure 2 As shown, the image processing methods include:
[0040] Step S21: Obtain at least two images from different perspectives of the same shooting scene. The at least two images are captured by at least two camera devices, with each camera device capturing one image.
[0041] Step S22: Obtain the distance between the object in each image and the corresponding camera device.
[0042] Step S23: If at least one image contains a subject that meets the preset conditions, perform preset processing on any one of the at least one images containing a subject that meets the preset conditions. The preset conditions include: the distance between the subject in the image and the camera device corresponding to the image is less than a distance setting value.
[0043] Step S24: Generate a three-dimensional image based on the pre-processed image and other images.
[0044] In some embodiments, preset conditions can be applied to each subject in the image, or preset conditions can be applied to some subjects in the image. For example, preset conditions can be applied to subjects in the image that are relatively close to the camera device.
[0045] In some embodiments, in step S22, the distance between the object in the image and the corresponding camera device can be obtained by a structured light depth ranging method or a binocular vision ranging method.
[0046] Structured light depth ranging is a non-contact active optical 3D measurement technology. Its basic principle is to project a beam of coded light onto the surface of the object being measured. When the surface topography changes, the distribution of the coded light is modulated by the object's height. A camera then captures images of the object's surface within its field of view, and demodulates these images to reconstruct the 3D topography containing the object's height information. This yields the distance distribution data from the object's 3D surface to the camera lens surface. This structured light depth ranging method can quickly and accurately acquire depth information, enabling 3D reconstruction.
[0047] Binocular vision ranging methods typically use binocular camera equipment, that is, two camera devices simultaneously capture the same scene, obtaining two images. The 3D image is reconstructed by calculating the differences between the two images. The distance between the binocular camera devices is approximately the distance between human eyes, which can simulate the stereoscopic vision effect when the human eye observes an object. The principle of binocular vision ranging methods mainly relies on two camera devices capturing the same scene from different positions. By comparing the positional differences of corresponding points in the two images, the binocular disparity for each object in the image is calculated, and thus the 3D position of the object is estimated. Binocular disparity refers to the difference in image position of the same object from two different viewpoints (such as the two eyes of a person or two camera devices). Binocular disparity is related to the distance from the object to the observer (camera device); the closer the distance, the greater the disparity. The calculation process of the binocular disparity algorithm includes: when two camera devices simultaneously capture the same scene, they capture images of the object from their respective viewpoints. Through image processing techniques, the positions of corresponding points in the two images can be found. By calculating the difference between them on the X-axis, the corresponding disparity is obtained. Based on the correspondence between parallax and distance, the distance between an object and the observer can be estimated by calculating the parallax magnitude. Binocular vision ranging methods are commonly used in computer vision and robotics, such as 3D vision inspection and obstacle detection in autonomous vehicles. In these applications, very high spatial positioning accuracy can be achieved by accurately calculating parallax.
[0048] In some embodiments, the step of obtaining the distance between the subject in each image and the corresponding camera device, i.e., step S22 above, may further include: obtaining the distance between the subject in the field of view of each camera device and the lens surface of the corresponding camera device by means of a structured light depth ranging method, thereby obtaining the distance between the subject in each image and the corresponding camera device.
[0049] In some embodiments, the step of obtaining the distance between the subject in each image and the corresponding camera device, i.e., step S22 above, may further include: obtaining the disparity between the same subject in the images corresponding to any two camera devices respectively by using a binocular visual ranging method; and determining the distance between the subject in each image and the corresponding camera device based on the disparity corresponding to the subject in the image captured by the camera device and a pre-set correspondence between disparity and distance.
[0050] In some embodiments, the parallax parameter between the same subject in the images corresponding to any two camera devices is obtained to determine whether the subject in the image meets the preset conditions. Figure 3 This illustration shows a flowchart of another image processing method provided in an embodiment of this application, such as... Figure 3 As shown, the image processing methods include:
[0051] Step S31: Obtain at least two images from different perspectives of the same shooting scene. The at least two images are captured by at least two camera devices, with each camera device capturing one image.
[0052] Step S32: Obtain the parallax size between the same subject in the images corresponding to any two camera devices.
[0053] Step S33: If the subject in at least one image meets the preset conditions, perform preset processing on the subject in any one of the at least one images that meets the preset conditions. The preset conditions include: the parallax of the subject is greater than the parallax setting value.
[0054] It is understandable that if the parallax between the same subject in two images is greater than the parallax setting value, it means that the subject in each of the two images meets the preset conditions.
[0055] Step S34: Generate a three-dimensional image based on the pre-processed image and other images.
[0056] In some embodiments, the step of obtaining the parallax size between the same subject in the images corresponding to any two camera devices, i.e., step S32 above, may further include: obtaining the parallax size between the same subject in the images corresponding to any two camera devices through a binocular parallax algorithm.
[0057] In some embodiments, the step of obtaining the disparity between the same subject in the images corresponding to any two camera devices, i.e., step S32 above, may further include: obtaining the distance between the subject in each image and the corresponding camera device; and determining the disparity between the same subject in the images corresponding to any two camera devices based on the distance between the subject and the corresponding camera device and a pre-defined correspondence between disparity and distance. The distance between the subject and the corresponding camera device can be obtained using the structured light depth ranging method described above.
[0058] In some embodiments, a pre-defined relationship is established between the parallax of the same object in images captured by two camera devices and the distance from the object to the camera device. The closer the object is to the camera device, the greater the parallax; the farther the object is from the camera device, the smaller the parallax. Figure 4 This diagram illustrates the parallax between the same subject in images from any two camera devices. Figure 4 The image shows a scene with people and mountains, with the people in front and the mountains behind. At the bottom are the images from cameras 1 and 2. In the image from camera 1 on the right, the people are on the left side of the mountain, while in the image from camera 2 on the left, the people are on the right side. This is because the two cameras (1 and 2) have different viewing angles. By comparing the images from cameras 1 and 2, it can be determined that the mountains are farther away from cameras 1 and 2, resulting in a smaller parallax, while the people are closer, resulting in a larger parallax. In other words, the parallax is smaller when the human eye observes the mountains and larger when observing the people. It can be understood that the point of minimum parallax, denoted as 0, is at infinity in front of cameras 1 and 2; the point of maximum parallax, denoted as Xm, is at the lens of the camera in front of cameras 1 and 2.
[0059] In some embodiments, a parallax setting value X less than Xm is calibrated. When the objects in the images captured by the two camera devices have a parallax between X and Xm, that is, the parallax of the objects is greater than the parallax setting value X, it means that when a person views an object with a parallax between X and Xm on a naked-eye 3D display, they will experience dizziness. When the objects in the images captured by the two camera devices have a parallax between 0 and X, that is, the parallax of the objects is less than the parallax setting value X, it means that when a person views an object with a parallax between 0 and X on a naked-eye 3D display, a clear stereoscopic image can be formed in the human brain, and dizziness will not occur.
[0060] In some embodiments, when the parallax of an object in the field of view of the calibrated camera is at a parallax setting value X, the distance from the object to the lens surface of the camera is L. Using L as the distance setting value, when the distance between an object in the field of view and the dual camera lenses is less than this distance setting value L, it indicates that when the images of this object taken by the two camera devices are displayed on the naked-eye 3D display screen, the parallax formed by the object in the images entering the left and right eyes of the human is too large, exceeding the ability of the human brain to form a stereoscopic image, which will cause dizziness in the human eye; when the distance between an object in the field of view and the lens surface of the camera device is greater than this distance setting value L, it indicates that when the images of this object taken by the two camera devices are displayed on the naked-eye 3D display screen, the object in the images entering the left and right eyes of the human can form a clear stereoscopic image.
[0061] In some embodiments, when the distance between the subject and the corresponding camera device in at least one of at least two images is less than a distance set value L, preset processing is performed on the subject in one of the at least two images whose distance is less than the distance set value L, while no processing is performed on the other images, so as to preserve the complete image content in the other images. When performing preset processing on the subject in the image, the surface with a distance set value L from the lens surface of the camera device is used as the contour surface. Preset processing is performed on the portion of the subject in the image that is less than L. For example, the subject in the image can be processed into black, gray, or white. It is even possible to extract the overall tone of the background in which the subject is located and use the background tone as the preset processing color, so that the content after preset processing can be ignored by the human eye.
[0062] Figure 5 This diagram illustrates a preset processing of a subject in an image, for example, such as... Figure 5 As shown, in Image 1 captured by camera device 1 and Image 2 captured by camera device 2, the distance between the subject (person) and the corresponding camera device is less than the distance setting value L, and the distance between the subject (mountain) and the corresponding camera device is greater than the distance setting value L. That is, in both Images 1 and 2, there are subjects that meet the preset conditions (the distance between the subject and the corresponding camera device is less than the distance setting value L). The subject (person) in Image 1 is processed by a preset process, which makes the subject (person) black, while Image 2 retains the complete image content.
[0063] In practical applications, combined with Figure 5When a human eye views a glasses-free 3D display screen, the images observed by the left and right eyes are images 1 and 2 captured by camera device 1 and camera device 2, respectively. Before 3D display, an image is generated based on the distance between the object being photographed and the lens surface of the corresponding camera device, which is less than a set distance value L. For example, the object (person) in image 1 captured by camera device 1 is pre-processed to be less than the set distance value L. Then, a three-dimensional stereoscopic image is generated based on the pre-processed image 1 and image 2 and displayed on the glasses-free 3D display screen. Since the part of the object in an image that is less than the set distance value L from the camera device is pre-processed, the large parallax of the part of the object that is less than the set distance value L from the camera device in the image displayed on the glasses-free 3D display screen can be effectively avoided when it enters the left and right eyes, thereby improving the dizziness caused by the human eye.
[0064] Figure 6 This diagram illustrates a three-dimensional image generated from two images, combined with... Figure 5 and Figure 6 As shown, the right eye observes the complete image 2, while the left eye observes the near-field object (person) in image 1, which is pre-processed and eliminated in the brain's imaging. The distant object (mountain) in both the right and left eyes can be seen simultaneously, thus forming a clear three-dimensional image in the brain. However, the near-field object (person) can only be seen by the right eye; the near-field object seen by the left eye is ignored in the brain due to pre-processing, resulting in a clear two-dimensional image in the brain.
[0065] This effectively ensures that when a user views stereoscopic images captured by camera devices 1 and 2 on a glasses-free 3D display, the image seen in one eye is complete and clear, while the image in the other eye is partially processed. This means that when viewing stereoscopic images captured by camera devices 1 and 2 on a glasses-free 3D display, for objects at a distance greater than a set distance L from the corresponding camera device, both eyes can see a clear stereoscopic image; for objects at a distance less than the set distance L, one eye sees a complete and clear image, while the other eye sees a partially processed image, resulting in a flat image in the brain. Therefore, when viewing stereoscopic images captured by camera devices 1 and 2 on a glasses-free 3D display, the image seen by the human eye does not contain objects with large parallax, and the user will not experience dizziness.
[0066] In some embodiments, when the parallax of the subject in at least one of at least two images is greater than a parallax setting value X, preset processing is applied to the subject in one of the at least two images whose parallax is greater than the parallax setting value X, while no processing is applied to the other images, thus preserving the complete image content in the other images. When performing preset processing on the subject in the image, the surface with the corresponding parallax value X is used as the contour surface. Preset processing is applied to the portion of the subject in the image where the parallax is greater than the parallax setting value X. For example, the subject in the image can be processed into black, gray, or white. Alternatively, the overall hue of the background where the subject is located can be extracted and used as the preset processing color, so that the human eye can ignore the pre-processed content.
[0067] In practical applications, combined with Figure 5 When a human eye views a glasses-free 3D display screen, the images observed by the left and right eyes are images 1 and 2 captured by camera device 1 and camera device 2, respectively. Before 3D display, an image with a parallax value greater than the parallax setting value X corresponding to the photographed object is pre-processed. For example, the object (person) with a parallax value greater than the parallax setting value X corresponding to image 1 captured by camera device 1 is pre-processed. Then, a three-dimensional stereoscopic image is generated based on the pre-processed image 1 and image 2 and displayed on the glasses-free 3D display screen. When displayed on the glasses-free 3D display screen, the part of the image with a large parallax value seen by the left eye is pre-processed, while the right eye can see a complete and clear image.
[0068] For a subject (mountain) with a parallax smaller than the parallax setting value X, when the stereoscopic image formed by images 1 and 2 is displayed on a glasses-free 3D screen, both eyes can see the subject (mountain), and the subject (mountain) viewed by the user forms a clear stereoscopic image in the brain. For a subject (person) with a parallax larger than the parallax setting value X, when the stereoscopic image formed by images 1 and 2 is displayed on a glasses-free 3D screen, one eye sees a partially pre-processed image 1, while the other eye sees a complete and clear image 2, and the subject (person) viewed by the user forms a clear planar image in the brain. Therefore, the human eye can view the entire field of view captured by cameras 1 and 2, and the dizziness caused by viewing images exceeding binocular parallax can be effectively reduced.
[0069] In some embodiments, the step of performing preset processing on a subject that meets preset conditions in any one of at least one image may further include: performing preset processing on the subject that meets preset conditions in an image if there is a subject in an image that meets preset conditions.
[0070] In some embodiments, the step of performing preset processing on a subject that meets preset conditions in any one of at least one images may further include: when there are subjects that meet preset conditions in at least two images, performing texture recognition on the surface of the subjects that meet preset conditions in at least two images to obtain the texture density of the surface of the subjects that meet preset conditions in at least two images; comparing the texture density of the surface of the subjects that meet preset conditions in at least two images, and performing preset processing on the subject that meets preset conditions in the image with the smallest corresponding texture density among the at least two images.
[0071] In some embodiments, when at least two images contain objects that meet preset conditions, texture recognition is performed on the surfaces of the objects that meet the preset conditions in each of the at least two images. Based on the texture density of the surfaces of the objects that meet the preset conditions in the at least two images, a target image requiring preset processing is determined, and the objects that meet the preset conditions in that image are then subjected to preset processing. When there are multiple objects that meet the preset conditions in an image, the texture density of the surfaces of the objects that meet the preset conditions in the image is the sum of the texture densities of the surfaces of the multiple objects that meet the preset conditions. The higher the texture density of a surface, the richer the details in the image of the subject, and the greater its observational value for the user. Conversely, the lower the texture density of the surface of the subject, the less detail it contains. Therefore, when determining the image that needs to be pre-processed, the texture density of the surface of the subject that meets the pre-conditions in at least two images is compared. The image with the lowest corresponding texture density among the at least two images is taken as the target image, and the subject in that image that meets the pre-conditions is pre-processed. Meanwhile, the image with the corresponding higher texture density is retained, so that the user can view an image with richer object details.
[0072] In some embodiments, the step of performing preset processing on a subject that meets preset conditions in any one of at least one images may further include: if a subject in at least two images meets preset conditions, obtaining the pixel area corresponding to the subject that meets preset conditions in each of the at least two images; comparing the pixel area of the subject that meets preset conditions in the at least two images, and performing preset processing on the subject that meets preset conditions in the image with the smallest corresponding pixel area in the at least two images.
[0073] In some embodiments, when at least two images contain objects that meet preset conditions, the pixel area corresponding to the object meeting the preset conditions in each of the at least two images is obtained, and a target image requiring preset processing is determined based on the pixel area of the object meeting the preset conditions in the at least two images, so that the object meeting the preset conditions in that image can be subjected to preset processing; when there are multiple objects meeting the preset conditions in an image, the pixel area corresponding to the object meeting the preset conditions in the image is the sum of the pixel areas corresponding to the multiple objects; the pixel area corresponding to the object The larger the area, the richer the details in the image of the subject, and the greater its observational value for the user. Conversely, the smaller the area of the pixel region corresponding to the subject, the less detail the image contains. Therefore, when determining the image that needs to be pre-processed, the pixel region areas corresponding to the subject that meet the pre-conditions in at least two images are compared. The image with the smallest corresponding pixel region area among the at least two images is taken as the target image, and the subject in that image that meets the pre-conditions is pre-processed. Meanwhile, the image with the larger corresponding pixel region area is retained, so that the user can view an image with richer object details.
[0074] In some embodiments, the step of performing preset processing on a subject that meets preset conditions in any one of at least one images may further include: in response to a preset action command input by a user, performing preset processing on a subject that meets preset conditions in an image captured by a camera device corresponding to the preset action command in at least one image.
[0075] It is understandable that if there is no subject in the image captured by the camera device corresponding to the preset action command input by the user, then the image does not need to be processed.
[0076] In practical applications, when the human eye views stereoscopic images captured by two camera devices through a naked-eye 3D display, the image captured by one camera device undergoes pre-processing while retaining the content of the image captured by the other camera device. The image viewed by one eye is the complete image, while the image viewed by the other eye is the pre-processed image. The image viewed by the other eye may miss some details of objects. In order to make it easier for users to see more details of objects in the image, some action commands can be preset, corresponding to the images captured by different camera devices. When there are at least two images of the subject that meet the preset conditions, the user can specify and switch the image to be pre-processed and the displayed stereoscopic image by inputting the preset action commands.
[0077] Figure 7A schematic diagram of a three-dimensional image switching display is shown, such as... Figure 7 As shown, users can switch between images requiring preset processing and display the corresponding generated 3D stereoscopic images by inputting different preset action commands. For example, image 1 captured by camera device 1 corresponds to preset action command 1, and image 2 captured by camera device 2 corresponds to preset action command 2. When both images 1 and 2 contain objects that meet preset conditions, when the user inputs preset action command 1, the objects in image 1 that meet the preset conditions are processed, the complete image content of image 2 is preserved, and a 3D stereoscopic image is generated based on the pre-processed images 1 and 2 for 3D display. When the user inputs preset action command 2, the objects in image 2 that meet the preset conditions are processed, the complete image content of image 1 is preserved, and a 3D stereoscopic image is generated based on the pre-processed images 2 and 1 for 3D display.
[0078] The preset action commands can be specific actions or gestures input by the user through the camera device. For example, a specific action can be the closing or opening of the eyes. Taking the closing of the eyes as an example, the closing of the left eye can be preset action command 1, and the closing of the right eye can be preset action command 2. Figure 7 and Figure 5 Preset action command 1 corresponds to image 1 captured by camera device 1, and preset action command 2 corresponds to image 2 captured by camera device 2. When the closing action of the left eye is detected, i.e., preset action command 1, the complete image 2 captured by camera device 2 is retained, and the subject in image 1 captured by camera device 1 that meets the preset conditions is subjected to preset processing; when the closing action of the right eye is detected, i.e., preset action command 2, the complete image 1 captured by camera device 1 is retained, and the subject in image 2 captured by camera device 2 that meets the preset conditions is subjected to preset processing.
[0079] It should be noted that this application exemplifies the description of the closing action of the human eye and the correspondence between the closing action of the human eye and the preset action command. However, in the embodiments of this application, the specific action described above may not be limited to the closing or opening action of the human eye, and the correspondence between the specific action and the preset action command may not be limited to the correspondence between the closing action of the left and right eyes of the human eye and the preset action command. In the embodiments of this application, the specific action may also be configured as other actions according to actual needs, and the correspondence between the specific action and the preset action command may be set.
[0080] For example, a gesture could be a hand swiping from left to right or from right to left. The hand swiping from left to right could be preset gesture command 2, and the hand swiping from right to left could be preset gesture command 1. Figure 7 and Figure 5 Preset action command 1 corresponds to image 1 captured by camera device 1, and preset action command 2 corresponds to image 2 captured by camera device 2. When a hand is detected moving from left to right (i.e., preset action command 2), the complete image 1 captured by camera device 1 is retained, while the subject in image 2 captured by camera device 2 that meets the preset conditions is processed using preset methods. When a hand is detected moving from right to left (i.e., preset action command 1), the complete image 2 captured by camera device 2 is retained, while the subject in image 1 captured by camera device 1 that meets the preset conditions is processed using preset methods.
[0081] It should be noted that this application exemplifies the use of hand gestures and the correspondence between hand gestures and preset action commands to describe preset action commands. However, in the embodiments of this application, the above-mentioned hand gestures are not limited to the above-mentioned hand gestures, and the correspondence between hand gestures and preset action commands is not limited to the above-mentioned correspondence between hand gestures and preset action commands. In the embodiments of this application, the hand gestures can also be configured as other hand gestures and the correspondence between hand gestures and preset action commands can be set according to actual needs.
[0082] In some embodiments, the preset action instruction may be information instructions or button instructions entered by the user on the operation interface or display screen. This application does not impose any special restrictions on the specific implementation form of the preset action instruction.
[0083] In some embodiments, at least two camera devices include a first camera device and a second camera device, the first camera device and the second camera device constitute a binocular camera device, and at least two images include an image captured by the first camera device and an image captured by the second camera device.
[0084] In some application scenarios, the above image processing method can be applied to a terminal, which includes a central processing unit (CPU) and a binocular camera device, which includes a first camera and a second camera. When the terminal takes a 3D photo, the CPU activates the binocular camera device to acquire images captured by the first and second cameras. Simultaneously, it activates a structured light depth detection device to acquire the distance between each object in each image and its corresponding camera. The CPU executes the image processing method described in the above embodiment to obtain a three-dimensional image. The CPU then sends the images captured by the first and second cameras from the three-dimensional image to the pixels on the display screen responsible for left-eye viewing and right-eye viewing, respectively, for 3D display. The left and right eyes see different images, generating a stereoscopic image in the human brain.
[0085] In some embodiments, the terminal for performing the image processing method of the above embodiments and the display screen for performing naked-eye 3D display can be set up independently, or the terminal includes the display screen.
[0086] In some embodiments, the terminal described above is a glasses-free 3D display terminal, including but not limited to mobile phones, tablets, and other terminals.
[0087] Among them, naked-eye 3D display is a new type of display technology, mainly based on the parallax characteristics of the human eye. It displays two slightly different images to the left and right eyes, causing the brain to synthesize a three-dimensional image with a stereoscopic effect. This technology simulates the stereoscopic vision mechanism of the human eye, allowing viewers to perceive the depth and stereoscopic effect of the image. Naked-eye 3D technology utilizes this principle, employing different techniques to project images from different perspectives, allowing the left and right eyes to receive different images, thus creating a stereoscopic effect. Specific implementation methods include: Slit grating technology: Adding slit gratings in front of the screen, the vertical fine stripes formed by these stripes separate the visible images for the left and right eyes, allowing the viewer to see a 3D image. Lens technology: Utilizing the refraction principle of lenses, projecting corresponding pixels for the left and right eyes separately, achieving image separation. The advantage of this technology is that it does not block light, improving brightness. Directional light source technology: Precisely controlling two sets of screens to project images to the left and right eyes respectively, achieving stereoscopic display.
[0088] It should be clarified that this application is not limited to the specific configurations and processes described in the above embodiments and shown in the figures. For the sake of convenience and brevity, detailed descriptions of known methods are omitted here and will not be repeated.
[0089] Figure 8 This is a block diagram illustrating the composition of an electronic device provided in an embodiment of this application.
[0090] like Figure 8 As shown, the electronic device includes: one or more processors 801 and a memory 802; the memory 802 stores one or more computer programs, which, when executed by one or more processors 801, enable the one or more processors 801 to implement any of the image processing methods described in the above embodiments.
[0091] In some embodiments, the electronic device further includes an I / O interface (read / write interface) 803, which is connected between the processor 801 and the memory 802 and enables information interaction between the memory 802 and the processor 801. The I / O interface 803 includes, but is not limited to, a data bus.
[0092] Among them, the processor is a device with data processing capabilities, including but not limited to the central processing unit (CPU); the memory is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH).
[0093] This application also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the image processing methods described in the above embodiments.
[0094] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements any of the image processing methods described in the above embodiments.
[0095] Those skilled in the art will understand that all or some of the steps, systems, and devices disclosed above, as functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0096] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components working together.
[0097] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; read-only optical disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cartridges, magnetic tapes, disk storage or other magnetic storage; and any other media that can be used to store desired information and can be accessed by a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0098] This application has disclosed exemplary embodiments, and although specific terminology has been used, it is used and should be interpreted only in a general illustrative sense and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this application as set forth by the appended claims.
Claims
1. An image processing method, comprising: Acquire at least two images from different perspectives of the same shooting scene, wherein the at least two images are respectively captured by at least two camera devices, and each camera device captures one image; If the subject in at least one of the at least two images meets the preset conditions, the subject in any one of the at least two images that meets the preset conditions shall be subject to preset processing. Generate a three-dimensional image based on the pre-processed image and other images.
2. The image processing method according to claim 1, wherein, After acquiring at least two images from different perspectives of the same shooting scene, the image processing method further includes: acquiring the distance between the subject in each image and the corresponding camera device; The preset conditions include: the distance between the object being photographed and the corresponding camera device is less than a set distance value.
3. The image processing method according to claim 1, wherein, After acquiring at least two images from different perspectives of the same shooting scene, the image processing method further includes: acquiring the parallax size between the same subject in the images corresponding to any two camera devices; The preset conditions include: the parallax of the photographed object is greater than the parallax setting value.
4. The image processing method according to claim 2, wherein, The step of obtaining the distance between the subject in each image and the corresponding camera device includes: The distance between the object in the field of view of each camera device and the lens surface of the corresponding camera device is obtained by using structured light depth ranging method, thus obtaining the distance between the object in each image and the corresponding camera device.
5. The image processing method according to claim 3, wherein, The step of obtaining the parallax between the same subject in the images corresponding to any two camera devices includes: The parallax difference between the same subject in the images corresponding to any two camera devices is obtained by using a binocular parallax algorithm.
6. The image processing method according to claim 1, wherein, The preset processing of the photographed object that meets the preset conditions in any one of the at least one images includes: If an object in an image meets the preset conditions, then the object in that image that meets the preset conditions is subjected to preset processing. If the subject in at least two images satisfies the preset conditions, perform texture recognition on the surface of the subject in at least two images that satisfies the preset conditions, and obtain the texture density of the surface of the subject in at least two images that satisfies the preset conditions. Compare the texture density of the surface of the photographed object that meets the preset conditions in at least two images, and perform preset processing on the photographed object that meets the preset conditions in the image with the smallest corresponding texture density among at least two images.
7. The image processing method according to claim 1, wherein, The preset processing of the photographed object that meets the preset conditions in any one of the at least one images includes: In response to a preset action command input by the user, the objects in the images captured by the camera device corresponding to the preset action command in the at least one image that meet the preset conditions are subjected to preset processing.
8. The image processing method according to claim 1, wherein, The at least two camera devices include a first camera device and a second camera device, the first camera device and the second camera device forming a binocular camera device, and the at least two images include an image captured by the first camera device and an image captured by the second camera device.
9. An electronic device, wherein, include: One or more processors; A memory having stored thereon one or more computer programs that, when executed by one or more processors, cause the one or more processors to implement the image processing method as described in any one of claims 1 to 8.
10. A computer program product comprising a computer program that, when executed by a processor, implements the image processing method as described in any one of claims 1 to 8.