Model training method, focusing processing method and electronic equipment
By simulating moiré patterns to train a model and combining it with a preset algorithm and laser ranging, the problem of autofocus failure caused by moiré patterns was solved, achieving accurate focusing and clear image capture in moiré pattern scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-10-25
- Publication Date
- 2026-05-05
AI Technical Summary
When shooting images involving electronic screens, moiré patterns interfere with the distribution of light intensity signals, causing autofocus failure, blurry images, and a poor user experience.
The model is trained by simulating moiré images, generating a large number of enhanced training images. The model parameters are adjusted to learn the phase difference in moiré scenes, and combined with a preset algorithm and laser ranging, accurate focusing is achieved.
Even with moiré patterns present, accurate focusing can still be achieved, improving image clarity and enhancing the user experience.
Smart Images

Figure CN121985226A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a model training method, a focusing processing method, and an electronic device. Background Technology
[0002] In daily life and work, people often use their mobile phones to take pictures. In some cases, they will use their mobile phones to take pictures of electronic screens to record important information or preserve wonderful moments.
[0003] With the development of technology, most smartphones now come equipped with autofocus. One type of autofocus is based on detecting the light intensity signal distribution corresponding to two types of pixels in an image, calculating the phase difference of the distribution to determine whether the image is in focus, and adjusting the motor according to the phase difference to achieve focus.
[0004] However, when an electronic screen is present in the image captured by the user, moiré patterns are likely to appear. Moiré patterns can affect the distribution of light intensity signals, leading to abnormal phase differences in calculations and making autofocus difficult. Summary of the Invention
[0005] This application provides a model training method and a focusing processing method, applied in the field of terminal technology. The model training method yields a model capable of focusing on moiré patterns. Applying this model to the focusing processing method ensures accurate focus even when moiré patterns appear on an electronic screen, preventing image blur and improving the user experience.
[0006] Firstly, embodiments of this application propose a model training method, including:
[0007] From the multiple original training images contained in the first training set, select multiple images to be processed;
[0008] The multiple images to be processed are subjected to moiré overlay processing to obtain multiple enhanced training images;
[0009] For any one of the multiple augmented training images, the augmented training image is input into the first model so that the first model outputs the predicted target phase difference corresponding to the augmented training image;
[0010] The model parameters of the first model are adjusted based on the predicted target phase difference corresponding to the enhanced training image and the actual target phase difference corresponding to the enhanced training image.
[0011] In this implementation, a subset of images is randomly selected from multiple original training images for moiré enhancement. The first model is then trained using image data containing these enhanced images. This allows the first model to learn the corresponding pattern even when moiré patterns are present in the image, and after inference, output a predicted target phase difference that is consistent with or close to the actual target phase difference. This avoids the problem of the first model outputting a predicted target phase difference that differs significantly from the actual target phase difference simply because the input image contains moiré patterns.
[0012] In one possible implementation, the process of superimposing moiré patterns on the multiple images to be processed to obtain multiple enhanced training images includes:
[0013] For any one of the multiple images to be processed, generate a moiré pattern image corresponding to the image to be processed;
[0014] The moiré pattern image and the image to be processed are superimposed to obtain the enhanced training image corresponding to the image to be processed.
[0015] In one possible implementation, generating the moiré image corresponding to the image to be processed includes:
[0016] For the image to be processed, a first sine wave and a second sine wave are generated;
[0017] The first sine wave and the second sine wave are superimposed to generate a moiré pattern image corresponding to the image to be processed.
[0018] This implementation describes how to simulate a moiré pattern and superimpose the simulated moiré pattern onto the image to be processed. In this implementation, two sine waves are used to simulate the moiré pattern. It should be noted that, to ensure the randomness of the moiré pattern, the position of the sine waves and their periods or frequencies in each direction can be randomly determined each time the pattern is generated. Furthermore, this implementation does not impose specific restrictions on whether to use sine waves or other shapes, or whether to use one or multiple patterns to superimpose and simulate the moiré pattern; the choice can be made based on the actual situation.
[0019] Compared to manually acquiring real images with moiré patterns, this implementation can easily obtain a large number of moiré-enhanced training images in a short time. Furthermore, it can add moiré patterns to both clear and blurry images. Manually acquiring real images with moiré patterns, however, may not be convenient due to the difficulty in obtaining clear images with moiré patterns. In other words, this implementation can obtain a large and comprehensive number of training images, which is crucial for deep learning models.
[0020] Secondly, embodiments of this application propose a focusing processing method, including:
[0021] Determine the first focus frame in the first preview image;
[0022] The first preview image and the first focus frame are input into the first model so that the first model outputs the first target phase difference corresponding to the first focus frame, wherein the first model is trained according to the first aspect or any implementation of the first aspect;
[0023] The position of the first motor in the terminal device is adjusted according to the first target phase difference.
[0024] In this implementation, the first model is trained on moiré-enhanced image data, learning to accurately output phase difference even in scenes with moiré patterns. Therefore, when applied to real-world scenarios, the first model can correctly process images regardless of whether moiré patterns appear in the first preview image, and adjust the position of the first motor in the terminal device based on the obtained first target phase difference, ensuring the camera is in focus relative to the first focus frame. This results in a clear first preview image for the user, and subsequent photos or videos are also clear, significantly improving the user experience.
[0025] In one possible implementation, before inputting the first preview image and the position of the focus frame into the first model, the method further includes:
[0026] Multiple second preview images are processed according to a preset algorithm to obtain the second target phase difference corresponding to the second focus frame in each of the multiple second preview images. The acquisition time of the multiple second preview images is before the acquisition time of the first preview image.
[0027] The first preview image is processed according to the preset algorithm to obtain the phase difference of the third target corresponding to the first focus frame in the first preview image;
[0028] Based on the second target phase difference corresponding to each of the multiple second preview images and the third target phase difference corresponding to the first preview image, detect whether the third target phase difference is abnormal; and / or, based on the sub-target phase difference corresponding to each of the multiple sub-regions included in the first focus frame, detect whether the third target phase difference is abnormal.
[0029] In this implementation, before using the first model to perform inference and prediction on the first preview image, a preset algorithm, namely a traditional autofocus algorithm, is used to calculate the phase difference. Understandably, the terminal device continuously performs focusing during shooting, and the preset algorithm can be used to calculate the corresponding phase differences for the first preview image and multiple previous second preview images. Simultaneously, since a focus frame on a preview image contains multiple sub-regions, the phase difference for each sub-region can also be calculated. Based on the phase differences of multiple preview images, or multiple phase differences within a single preview image, anomaly detection can be performed. Based on the detection results, more strategies can be implemented. For example, when the phase difference calculated by the preset algorithm does not show any anomalies, the phase difference calculated by the preset algorithm can be used directly for focusing, thus reducing the overall focusing time required by the terminal device.
[0030] In one possible implementation, detecting whether the third target phase difference is abnormal based on the second target phase difference corresponding to each of the plurality of second preview images and the third target phase difference corresponding to the first preview image includes:
[0031] The phase differences of the second target corresponding to each of the multiple second preview images and the phase differences of the third target corresponding to the first preview image are sorted according to the image acquisition time to obtain a first sequence;
[0032] In the first sequence, a first number of target phase differences located at wave crests and a second number of target phase differences located at wave troughs are determined;
[0033] The sum of the first quantity and the second quantity is compared with a first preset threshold to detect whether there is an anomaly in the third target phase difference.
[0034] This implementation describes a time-domain anomaly detection process for a preset algorithm. When moiré patterns appear in an image, the phase difference calculated by the preset algorithm will significantly deviate from zero, or it can be understood as having a larger absolute value compared to the phase difference under normal conditions. Then, by comparing the target phase differences of multiple second preview images, if there are excessively large or small phase differences, these can be considered peaks or troughs. The sum of the number of peaks and troughs is compared with a first preset threshold. If the sum of the number of peaks and troughs is greater than the first preset threshold, then it can be determined that the current preset algorithm has a time-domain anomaly. There can be various methods for identifying peaks or troughs; this implementation does not impose specific limitations.
[0035] In one possible implementation, detecting whether the phase difference of the third target is abnormal based on the phase differences of the sub-targets corresponding to the multiple sub-regions included in the first focus frame includes:
[0036] For the phase difference of each sub-target corresponding to the multiple sub-regions included in the first focusing frame, calculate the standard deviation;
[0037] The standard deviation is compared with a second preset threshold to detect whether the third target phase difference is abnormal.
[0038] This implementation describes a process for detecting anomalies in a preset algorithm in the spatial domain. For multiple sub-regions within the first focus frame of a preview image, the phase difference of each sub-region can be calculated. Understandably, since these sub-regions are spatially very close, their corresponding phase differences should also be relatively close. Therefore, anomalies can be determined by comparing the phase differences of multiple sub-regions within a focus frame. To determine the closeness or dispersion of multiple values, the standard deviation of these values can be calculated. That is, the standard deviation of multiple phase differences can be compared with a second preset threshold. If the standard deviation is greater than the second preset threshold, it indicates that moiré patterns are likely causing anomalies in the preset algorithm in the spatial domain.
[0039] In one possible implementation, inputting the first preview image and the first focus frame into the first model, so that the first model outputs the first target phase difference corresponding to the first focus frame, includes:
[0040] In the event of an abnormality in the phase difference of the third target, the first preview image and the first focus frame are input to the first model so that the first model outputs the phase difference of the first target corresponding to the first focus frame.
[0041] In this implementation, anomaly detection is performed on the phase difference calculated by the preset algorithm, and the result determines whether the first model needs to be used to infer the accurate phase difference. This combines the preset algorithm, or traditional autofocus method, with the first model trained on the moiré-enhanced image. On one hand, since the first model can handle scenes with moiré patterns in the image, accurate phase differences can be obtained in different scenarios, achieving precise focusing. On the other hand, when the preset algorithm's result is normal, directly using the phase difference calculated by the preset algorithm results in faster processing speed and lower energy consumption.
[0042] In one possible implementation, determining the first focus frame in the first preview image includes:
[0043] Detect whether the first preview image contains a human face;
[0044] If the first preview image contains a face, the first focus frame is determined based on the face region in the first preview image;
[0045] If the first preview image does not contain a human face, the first focus frame is determined based on a preset area in the first preview image.
[0046] In this implementation, the type of the first focus box in the first preview image is further distinguished. When the first preview image is determined to contain a face, the first focus box is determined by the area where the face is located, also known as the face focus box. Otherwise, the first focus box can be determined based on a preset area in the first preview image, such as the center focus box. After distinguishing different types of first focus boxes, more detailed processing can be performed. Understandably, since the first model, in addition to learning from images with moiré enhancement, also focuses on learning the focus of the face area during training, using the first model to infer the phase difference for the face focus box can achieve better results than the preset algorithm. For the center focus box, the preset algorithm can be used first, making the overall focusing process faster and the focusing effect better, thus effectively improving the user experience.
[0047] In one possible implementation, if the first preview image does not contain a face and the third phase difference is abnormal, the method further includes:
[0048] Based on the first laser ranging result corresponding to the first preview image and the second laser ranging results corresponding to each of the multiple second preview images, the position of the first motor in the terminal device is adjusted.
[0049] In one possible implementation, adjusting the position of the first motor in the terminal device based on the first laser ranging result corresponding to the first preview image and the second laser ranging results corresponding to each of the multiple second preview images includes:
[0050] Based on the first laser ranging result and the plurality of second laser ranging results, determine the fluctuation parameters of the laser ranging result;
[0051] If the fluctuation parameter is less than the third preset threshold, then the target time of flight (TOF) corresponding to the first focus frame in the first preview image is determined based on the first laser ranging result.
[0052] The position of the first motor in the terminal device is adjusted according to the target TOF.
[0053] In this implementation, when the first preview image does not contain a face and the preset algorithm's result is abnormal, focusing is achieved based on laser ranging. Laser ranging can quickly and accurately determine the distance between the terminal device and the object being photographed, which can be understood as the object distance. Then, by analyzing the relationship between the object distance, image distance, and focal length, relevant parameters in the accurate-focus state can be calculated, such as the image distance. The first motor can then be used to adjust and complete the focusing process, resulting in a clear image. It is understandable that the laser ranging process is independent of the presence of moiré patterns and is therefore unaffected by them. However, considering that laser ranging may also be unstable, multiple consecutive laser ranging results can be used to determine whether the obtained image distance is stable and effective. Even if laser ranging is unstable, the first model is still used to infer the accurate phase difference. In this way, the equally simple and fast laser ranging method can be used to handle scenarios where the preset algorithm malfunctions, because laser ranging is unaffected by moiré patterns in the image. On the other hand, potential problems with laser rangefinding were also considered, so the first model was used as a guarantee for focusing performance, balancing both focusing speed and focusing effect overall.
[0054] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory is used to store code instructions and the processor is used to run the code instructions to perform the methods described in the first aspect to the second aspect or any possible implementation of the first aspect to the second aspect.
[0055] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first to second aspects or any possible implementation of the first to second aspects.
[0056] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first to second aspects or any possible implementation of the first to second aspects.
[0057] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface, wherein the communication interface and at least one processor are interconnected via a circuit, and the at least one processor is used to run computer programs or instructions to perform the methods described in the first to second aspects or any possible implementations of the first to second aspects. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.
[0058] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).
[0059] It should be understood that the third to sixth aspects of this application correspond to the technical solutions of the first to second aspects of this application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0060] Figure 1 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 1 ;
[0061] Figure 2 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 2 ;
[0062] Figure 3 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 3 ;
[0063] Figure 4 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 1 ;
[0064] Figure 5 A schematic diagram of moiré pattern formation provided in the embodiments of this application;
[0065] Figure 6 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 1 ;
[0066] Figure 7 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 2 ;
[0067] Figure 8 This is a schematic diagram of the hardware structure of the terminal device provided in the embodiments of this application;
[0068] Figure 9 This is a schematic diagram of the software structure of the terminal device provided in the embodiments of this application;
[0069] Figure 10 A schematic flowchart illustrating the model training method provided in this application embodiment;
[0070] Figure 11 A schematic diagram illustrating the training data augmentation process provided in the embodiments of this application;
[0071] Figure 12 A schematic diagram illustrating the implementation of moiré data enhancement provided in an embodiment of this application;
[0072] Figure 13 A schematic diagram illustrating the implementation of model training provided in an embodiment of this application;
[0073] Figure 14 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 2 ;
[0074] Figure 15 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 1 ;
[0075] Figure 16 A schematic diagram illustrating the implementation of focusing window acquisition provided in an embodiment of this application;
[0076] Figure 17 A schematic diagram illustrating the implementation of model inference provided in an embodiment of this application;
[0077] Figure 18 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 2 ;
[0078] Figure 19 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 3 ;
[0079] Figure 20 A schematic diagram illustrating the implementation of time-domain phase difference anomaly detection provided in an embodiment of this application;
[0080] Figure 21 A schematic diagram illustrating the implementation of spatial phase difference anomaly detection provided in an embodiment of this application;
[0081] Figure 22 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 4 ;
[0082] Figure 23 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 3 ;
[0083] Figure 24 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 5 ;
[0084] Figure 25 A schematic diagram illustrating the implementation of laser ranging in an embodiment of this application;
[0085] Figure 26 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 3 ;
[0086] Figure 27 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0087] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0088] 1. Focusing / Focusing
[0089] In optics, focusing refers to adjusting the focal point of an optical system so that light rays from an object can form a clear image on a photosensitive element. When the object is within the effective focal length range of the lens, by adjusting the distance between the lens and the photosensitive surface, the light rays reflected or emitted by the object can be accurately converged on the photosensitive surface to form an image.
[0090] If the distance is adjusted properly, the image formed on the photosensitive surface will be very clear, which can be called being in focus; conversely, if the distance is not appropriate, the image will appear blurry.
[0091] 2. PDAF
[0092] Phase detection autofocus (PDAF) is a fast autofocus technology used in cameras and camcorders. This technology works by comparing the phase of light entering the lens to quickly determine the distance the lens should move to achieve accurate focusing.
[0093] Specifically, PDAF technology uses dedicated focusing pixels on the sensor. These pixels are not typically used directly for imaging but are specifically designed to capture light information and measure the phase difference between two light rays passing through different areas of the lens. If the two light rays are in phase, the object is in focus; if there is a phase difference, this difference is calculated to determine the direction and distance the lens needs to move, achieving precise focusing.
[0094] 3. Moiré pattern
[0095] Moiré pattern is an optical phenomenon where two patterns or lines with the same or similar spacing overlap, causing interference to form new, larger-spacing stripes or patterns. This phenomenon can be observed in many situations, such as when two layers of translucent mesh material overlap, or when photographing objects with fine textures in digital photography.
[0096] In the field of digital photography, moiré patterns specifically refer to high-frequency interference stripes generated by the interaction between the photosensitive element in digital cameras, scanners, and other devices and the subject being photographed. This is especially true when the object being photographed has a regular, fine pattern, such as fabric, textiles, or electronic screens. Because the spatial frequency of the photosensitive element's pixels is close to the spatial frequency of the stripes in the image, they interfere with each other, resulting in moiré patterns.
[0097] 4. Imaging sensor
[0098] An imaging sensor is a device that converts optical images into electrical signals, which can then be processed and analyzed to ultimately form a digital image. Imaging sensors are widely used in digital cameras, smartphones, medical imaging equipment, security monitoring systems, and other fields.
[0099] There are two main types of imaging sensors: charge-coupled devices (CCDs) and complementary metal-oxide-semiconductors (CMOSs). Both types of imaging sensors can convert light signals into electrical signals and are composed of many small sensor units, each corresponding to a pixel in the image.
[0100] 5. Other terms
[0101] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0102] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0103] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0104] 6. Electronic equipment
[0105] The electronic devices in this application embodiment may include handheld devices with shooting capabilities, vehicle-mounted devices, etc. For example, some electronic devices are: mobile phones, tablet computers, PDAs, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, personal digital assistants (PDAs), handheld devices with shooting capabilities, computing devices, terminal devices in 5G networks, or terminal devices in future evolved public land mobile networks (PLMNs), etc., and this application embodiment is not limited to these.
[0106] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0107] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0108] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0109] Based on the above introduction, the relevant technologies involved in this application will be further described in detail below.
[0110] With the development of terminal technology and the increase in user demand, most smartphones now have autofocus functionality when shooting. Below, we'll discuss this in conjunction with... Figures 1 to 3 A brief explanation of the principle of autofocus is provided, including: Figure 1 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 1 , Figure 2 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 2 , Figure 3 This application provides a schematic diagram of the implementation of phase detection autofocus in its embodiments. Figure 3 .
[0111] This is understandable. Consider how the human eye judges the distance of objects when viewing a scene. One way is because humans have two eyes, and each eye is at a different angle to an object; this difference in angle allows them to perceive the object's distance. Similarly, ... Figure 1 As shown, two additional convex lenses and corresponding sensors can be added to the camera to calculate the distance between the object and the camera. Figure 1In (a), object 101 can be considered as a point light source. The camera includes lens 102, imaging sensor 103, semi-transparent mirror 104, convex lens 109 and convex lens 111, and imaging sensors 110 and 112. Lens 102 can be understood as an abstract representation of the complex structure in the camera. The light emitted by object 101 is represented by four rays: ray 105, ray 106, ray 107, and ray 108. Rays 105 and 106 enter from above lens 102, while rays 107 and 108 enter from below lens 102. It is understood that the above-mentioned rays entering from above or below lens 102 is merely an example for illustrative purposes; the rays could also enter from the left or right side of lens 102, theoretically, as long as they are from different sides of the lens.
[0112] Assuming in Figure 1 In (a), object 101, lens 102, and imaging sensor 103 are not in a focused state. Light emitted from object 101 is focused in front of imaging sensor 103 after passing through lens 102. Semi-transparent mirror 104 is placed obliquely between lens 102 and imaging sensor 103. Some of the light rays passing through lens 102 continue to propagate along their original path, while some are reflected downwards by semi-transparent mirror 104. Part of the light rays 105 and 106 propagating above lens 102 are reflected by semi-transparent mirror 104 onto convex lens 109 and further propagated to imaging sensor 110. Similarly, part of the light rays 107 and 108 propagating below lens 102 are reflected by semi-transparent mirror 104 onto convex lens 111 and further propagated to imaging sensor 112.
[0113] Depending on whether the light emitted from object 101 converges before or after imaging sensor 103 after passing through lens 102, the positions of the light received by imaging sensor 110 and imaging sensor 112 will differ. For example... Figure 10 As shown in (a), the light emitted by object 101 is focused in front of imaging sensor 103, so the light rays propagating to imaging sensors 110 and 112 will be closer to their common center position. And... Figure 10 In (b), the light emitted by object 101 is focused behind imaging sensor 103, so the light propagating to imaging sensor 110 and imaging sensor 112 will move closer to the sides respectively.
[0114] Assuming that in the focused state, the light intensity distribution received by imaging sensors 110 and 112 is concentrated in their respective central positions, then the current focus status can be determined based on whether the actual received light is closer to the outer or inner side. Furthermore, after calibration, the relative positional difference in the light intensity distribution received by imaging sensors 110 and 112 can be used to calculate how the camera should be adjusted to achieve focus. This positional difference in the light intensity distribution received by imaging sensors 110 and 112 can be called phase difference; the specific definition of phase difference may vary in different systems.
[0115] The above describes a relatively intuitive phase-difference-based autofocus method. However, this method is not suitable for highly integrated devices, such as small mobile phones. Mobile phones and similar devices employ other similar phase-difference-based autofocus methods. Figure 2 As shown in (a), the imaging sensor 201 can be considered as composed of a number of regularly arranged sub-sensors, each corresponding to a pixel. These include a normal sub-sensor 202, a sub-sensor 204 whose left-side light is blocked, and a sub-sensor 203 whose right-side light is blocked. In some scenarios, sub-sensor 204 can also be referred to as the right pixel sensor, and sub-sensor 203 can also be referred to as the left pixel sensor.
[0116] refer to Figure 2 In (b), assuming an object 205 serves as a light source, when light propagates to the left pixel sensor 206, a light intensity distribution 207 is obtained. The left side of the light intensity distribution 207 has a weaker light intensity because most of the light is blocked, while the right side has a stronger light intensity. However, when the same light propagates to the right pixel sensor 208, the resulting light intensity distribution 209 will be significantly different from the light intensity distribution 207. In other words, by transforming some sub-sensors into the left pixel sensor 203 and the right pixel sensor 204 on the original imaging sensor 201, a similar effect can be achieved. Figure 1 The effect shown in the image.
[0117] Furthermore, such as Figure 3 As shown, suppose there is an object 301 as a light source, the camera includes a lens 304 and an imaging sensor 305, and two light rays 302 and 303, wherein light ray 302 enters from the left side of lens 304, and light ray 303 enters from the right side of lens 304. Figure 3In (a), light rays converge in front of the imaging sensor 305 after passing through lens 304. Light ray 302, which enters from the left side of the lens, propagates to the right side of the imaging sensor 305, while light ray 303, which enters from the right side of the lens, propagates to the left side of the imaging sensor 305. The light intensity distribution 306 below the imaging sensor 305 represents the light intensity distribution corresponding to the right pixel sensor, and the light intensity distribution 307 represents the light intensity distribution corresponding to the left pixel sensor. It can be seen that the light intensity distribution 307 is to the right of the light intensity distribution 306. The distance between the two is the phase difference mentioned earlier, and its unit can be understood as pixels. At this time, the phase difference can be a positive value.
[0118] Similarly, in Figure 3 In (c), after passing through lens 304, the light rays converge behind imaging sensor 305. Light ray 302, incident from the left side of the lens, propagates to the left side of imaging sensor 305, while light ray 303, incident from the right side of the lens, propagates to the right side of imaging sensor 305. At this point, light intensity distribution 307 can be seen to the left of light intensity distribution 306, and the phase difference can be negative. Figure 3 In (b), the object 301, lens 304 and imaging sensor 305 are in a state of quasi-focus, and the light intensity distribution 307 and light intensity distribution 306 overlap, with a phase difference of zero.
[0119] Understandably, in a scene, different objects are at different distances from the camera, some near and some far. A camera can only focus on objects within a certain distance range; objects beyond that range will appear blurry to varying degrees. Therefore, in real-world scenarios, it's generally necessary to consider the focus window, or in other words, which specific object in the scene should be focused on. The following section will combine... Figure 4 This article introduces a specific autofocus solution. Figure 4 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 1 :
[0120] S401. Determine if a face exists in the image. If it exists, proceed to S404; otherwise, proceed to S402.
[0121] This step typically involves a face recognition model taking the current image frame as input and outputting information related to faces in the image. This information includes the location of the face and its confidence level. If the face-related information is empty, or the confidence level is below a threshold, then it can be assumed that there is no face in the current image frame.
[0122] S402, Obtain the center focus frame.
[0123] If there is no human face in the current image frame, it is assumed by default that the object the user actually wants to photograph is located in the center of the image, and therefore the center area of the image is used as the focus frame.
[0124] S403, use the platform's focusing algorithm to calculate the phase difference.
[0125] The platform focusing algorithm in this step can be understood as the autofocus algorithm developed by camera hardware manufacturers in mobile phones, and its principle is the same as described above. Figures 1 to 3 The similarities described in the previous section will not be repeated here. The platform's autofocus algorithm can be used to obtain the phase difference, which can then be used to determine whether the image is in focus and how to make necessary adjustments.
[0126] S406. Drive the horse to the focus position according to the phase difference.
[0127] Understandably, according to the principles of optical imaging, when the object distance, image distance, and focal length satisfy the following formula, focusing can be achieved:
[0128] Formula 1
[0129] Where f represents the focal length, u is the object distance, and v is the image distance.
[0130] Understandably, the magnitude of the phase difference corresponds to the current state of the imaging system. If the phase difference is 0, it means that the image is in focus; if the phase difference is greater than 0, it means that the light from the object in the focus frame is focused in front of the imaging sensor; if the phase difference is less than 0, it means that the light from the object in the focus frame is focused behind the imaging sensor.
[0131] When an autofocus function is incorporated into a camera, calibration processing is performed, which quantifies the relationship between the phase difference and the object distance, image distance, and focal length. In some implementations, since the distance from the object to the camera is fixed (i.e., the object distance is fixed), and assuming the camera's focal length is also fixed, focusing can be achieved by changing the image distance. Specifically, based on the quantized relationship obtained from the calibration process, the image distance corresponding to the current focus is calculated using the phase difference; this is the distance between the camera lens and the imaging sensor. A motor is then used to adjust the image distance to the corresponding value, thus achieving autofocus.
[0132] S404, Obtain the face focus frame.
[0133] If a face is detected by the face recognition model in S401, the face bounding box is obtained based on the face position output by the face recognition model. Considering that the size and shape of the area occupied by a face in an image can vary, and there may be multiple faces, further processing is required after obtaining the face position information. For example, if multiple faces exist, one face can be selected as the bounding box according to certain rules, such as selecting the one closest to the center or the one occupying the largest area of the image.
[0134] S405, phase difference is obtained using AI focusing model.
[0135] In some implementations, the face focus frame also needs to be adjusted, such as by scaling, cropping, or adjusting the ratio, to suit the requirements of the focusing model. Then, the image frame and the face focus frame information are input into the focusing model, and the phase difference corresponding to the face focus frame is output to achieve autofocus.
[0136] In today's daily life and work, people often use their mobile phones to photograph electronic screens. Here are a few common examples: 1. Meeting minutes: In work meetings or lectures, participants may use their mobile phones to photograph the content on the projection screen to quickly record important information; 2. Sharing information: When seeing interesting or valuable information on social media, news websites, or other online platforms, users may choose to directly photograph the screen to share with friends, family, or colleagues; 3. Recalling the event: At concerts or large performances, audiences often use their mobile phones to photograph the large screen on the stage, not only to capture exciting moments of the performance but also as a precious record of personal memories.
[0137] However, the autofocus implementation method mentioned above for mobile phone photography has some shortcomings or problems. To better illustrate this issue, let's first consider... Figure 5 The formation of moiré patterns will be explained. Figure 5 This is a schematic diagram illustrating the formation of moiré patterns as provided in an embodiment of this application. Figure 5 In the image, image 501 is composed of many equally spaced vertical lines, and image 502 can be considered as obtained by rotating image 501 by a certain angle. For example... Figure 5 As shown, images 501 and 502 are overlaid to obtain image 503. In image 503, new horizontal stripes 504 magically appear; these stripes are moiré patterns. That is, when two patterns or lines with the same or similar spacing overlap, they interfere with each other to form new stripes or patterns with larger spacing. This phenomenon can be observed in many cases.
[0138] Due to the presence of moiré patterns, images taken by users may appear blurry. The following section discusses... Figure 6 and Figure 7 The relevant scenarios are explained, among which, Figure 6 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 1 , Figure 7 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 2 .
[0139] Scenario 1: When a user takes a picture of an electronic screen with their mobile phone, the image is blurry.
[0140] like Figure 6 As shown in (a), suppose a user is watching a movie on monitor 601, sees a great scene, and wants to share it with a friend, so they take out their phone to take a picture. The image captured at this moment is likely to look like... Figure 6 As shown in (b), different degrees of moiré patterns appeared in the phone 602 604, and the portrait on the original display was also blurry and not clear, that is, autofocus was not achieved.
[0141] This phenomenon occurs due to moiré patterns. Since typical electronic screens and imaging sensors in mobile phones can be viewed as lattice-like structures—such as the regularly arranged pixels in a liquid crystal display and the numerous sub-sensors in an imaging sensor—when light from the screen reaches the imaging sensor, it's equivalent to two patterns with similar or identical spacing overlapping. This interference creates new patterns, known as moiré patterns. The appearance of moiré patterns can be seen as a new periodic high-frequency pattern superimposed on the original image, altering the overall light intensity distribution, including the light intensity distributions of the left and right pixel sensors relied upon for autofocus. Consequently, it becomes difficult to calculate the phase difference between the light intensity distributions of the left and right pixel sensors, or in other words, the calculated phase difference may not achieve true focusing.
[0142] Scenario 2: When moiré patterns appear in the image, focus shake occurs.
[0143] like Figure 7 As shown, when a user takes a picture of an electronic screen with their mobile phone, moiré patterns appear in the image, and the image clarity constantly changes. For example, in... Figure 7 In (a), the image appears clear at this moment, but may become unclear in the next moment. Figure 7 As shown in (b), the image becomes blurry. It may then become relatively clear and blurry again, repeating this cycle. The reason for this focus jitter is that when the phone is shooting, it continuously performs autofocus. Due to the presence of moiré patterns, changes in various factors can cause significant changes in the moiré patterns, which in turn leads to changes in the light intensity distribution. Consequently, a stable phase difference cannot be obtained, and stable focus cannot be achieved.
[0144] The scenarios described above all involve the problem of blurry images caused by moiré patterns during shooting using terminal devices. Existing autofocus solutions in these scenarios cannot solve this problem. Schemes that calculate phase difference based on the light intensity distribution obtained from the left and right pixel sensors are limited by the fact that the hardware design did not initially consider addressing the moiré pattern issue, making it difficult to directly optimize and solve the problem through hardware-based algorithms. Similarly, existing focusing models used in terminal devices, such as the face focusing model mentioned earlier, may output incorrect or unstable phase differences when faced with scenes containing moiré patterns, leading to inaccurate focusing. Thus, existing technology struggles to meet users' needs for capturing clear images, negatively impacting the user experience.
[0145] Based on this, the embodiments of this application propose the following technical concept: by enhancing the moiré pattern on the data used to train the focusing model, that is, by artificially simulating moiré patterns and superimposing them onto a normal image, the scenario of moiré patterns appearing when a terminal device captures an electronic screen is simulated. Using such training data to train the focusing model, the model can learn during training how to accurately output the phase difference in the presence of moiré patterns in the image. Thus, when applying the focusing model to output the phase difference, the interference of moiré patterns can be largely avoided, and an accurate phase difference can be inferred.
[0146] In addition to improvements in the model algorithm, the method proposed in this application also incorporates laser ranging for autofocus. Specifically, when the original autofocus method malfunctions, laser ranging can be used to calculate the distance between the camera and the object using the time of flight (TOF) of the laser beam. This object distance is then used to calculate the phase difference and image distance, or the image distance can be calculated directly from the object distance. Laser ranging also avoids the focusing difficulties caused by moiré patterns and can be combined with the focusing model to obtain a better autofocus solution.
[0147] The technical solution provided in this application can be applied to terminal devices. The terminal devices will be briefly introduced below.
[0148] For example, Figure 8 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application is shown.
[0149] Figure 8This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. The terminal device 800 may include a processor 810, an external memory interface 820, an internal memory 821, a universal serial bus (USB) interface 830, a charging management module 840, a power management module 841, a battery 842, antenna 1, antenna 2, a mobile communication module 850, a wireless communication module 860, an audio module 870, a speaker 870A, a receiver 870B, a microphone 870C, a headphone jack 870D, a sensor module 880, buttons 890, a motor 891, an indicator 892, a camera 893, a display screen 894, and a subscriber identification module (SIM) card interface 895, etc. The sensor module 880 may include a pressure sensor 880A, a gyroscope sensor 880B, a barometric pressure sensor 880C, a magnetic sensor 880D, an accelerometer sensor 880E, a distance sensor 880F, a proximity light sensor 880G, a fingerprint sensor 880H, a temperature sensor 880J, a touch sensor 880K, an ambient light sensor 880L, a bone conduction sensor 880M, etc.
[0150] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the terminal device 800. In other embodiments of this application, the terminal device 800 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0151] The processor 810 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0152] The processor 810 can execute pre-defined applications. For example, in the focusing processing method proposed in the embodiments of this application, the processor 810 can be used to control the execution flow of the method and process some of its steps, such as some steps involving numerical calculations.
[0153] The terminal device 800 can perform shooting functions through an ISP, camera 893, video codec, GPU, display 894, and application processor.
[0154] The ISP (Image Signal Processor) is used to process data fed back from the camera 893. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 893.
[0155] Camera 893 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 800 may include one or N cameras 893, where N is a positive integer greater than 1.
[0156] Through the coordinated operation of the processor 810, display screen 894, and camera 893, users can preview captured images on the terminal device and take photos or videos after autofocus is complete. Researchers can also use the terminal device, in conjunction with the display screen 894 and camera 893, to collect image data from real-world scenes for analysis or model training when necessary.
[0157] The terminal device 800 implements display functions through a GPU, a display screen 894, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 894 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 810 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0158] Some terminal devices also include a neural network (NN) computing processor, also known as an NPU. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can quickly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminal devices, such as image recognition, facial recognition, autofocus, speech recognition, and text understanding.
[0159] GPUs or NPUs can be used to quickly apply models for inference. For example, a trained autofocus model can be deployed on a terminal device to quickly achieve focusing based on the GPU.
[0160] A distance sensor 880F is used to measure distance. The terminal device 800 can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal device 800 can utilize the distance sensor 880F to measure distance for rapid focusing.
[0161] Motor 891 can be used to assist in autofocusing, for example, by adjusting the image distance of the camera lens or the focal length using motor 891 integrated in the camera, in order to achieve accurate focus and capture clear images.
[0162] The software system of the terminal device 800 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of the terminal device 800.
[0163] For example, Figure 9 This is a schematic diagram of the software structure of a terminal device provided in an embodiment of this application. For example... Figure 9 As shown, the layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through interfaces. In some embodiments, the system may include an application layer, an application framework layer, the Android runtime and system libraries, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that this application uses the Android system as an example; however, the solution can also be implemented in other operating systems (such as HarmonyOS, iOS, etc.) as long as the functions implemented by each module are similar to those in the embodiments of this application.
[0164] The application layer can include a series of application packages.
[0165] like Figure 9As shown, the application layer can include applications such as camera, gallery, calendar, call, map, navigation, wireless local area networks (WLAN), Bluetooth, music, video, SMS, lock screen application, and settings application. Of course, the application layer can also include other application packages, such as third-party applications like payment applications, shopping applications, banking applications, and social applications; this application does not limit this.
[0166] Third-party applications can have functions such as facial recognition, video calls, scanning, taking photos, and recording videos.
[0167] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. These may include, for example, an activity manager, a window manager, a content provider, a view system, a resource manager, a notification manager, etc., though this embodiment does not impose any limitations on them.
[0168] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0169] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries comprise two parts: one part contains the functionalities that Java calls, and the other part consists of the Android core libraries. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0170] HAL is a wrapper around Linux kernel drivers, providing an interface to the upper layers and shielding them from the implementation details of the lower-level hardware.
[0171] HAL can include a laser sensing module, a focusing processing module, a model service module, a Wi-Fi HAL, an audio HAL, a camera service unit (Camera HAL Server), and a software code library, etc.
[0172] The laser sensing module encapsulates the control and calculation functions related to the laser sensor. For example, it can control the laser sensor to emit laser pulses, measure distance based on the time difference between the emitted and received laser pulses, and return the distance result.
[0173] The focusing processing module encapsulates the functions related to camera autofocus in the terminal device, such as calculating the phase difference based on the focus frame and calculating the position that the motor needs to be adjusted based on the phase difference.
[0174] The model service module can provide related functions of the model deployed in the terminal device. For example, it can call the inference function of the focus model, and after inputting the image frame and the focus box, it can obtain the output phase difference result after model inference, and use it for autofocus, etc.
[0175] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, audio drivers, sensor drivers, and charging drivers.
[0176] Among them, the camera driver is the driver layer of the camera device or camera device, which is mainly responsible for the interaction with the hardware module.
[0177] The technical solutions of the embodiments of this application and how the technical solutions of the embodiments of this application solve the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific examples. The following specific embodiments can be implemented independently or in combination with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0178] First, combine Figures 10 to 13 This application provides a model training method according to an embodiment, wherein... Figure 10 This is a flowchart illustrating the model training method provided in an embodiment of this application. Figure 11 This is a schematic diagram of the training data augmentation process provided in an embodiment of this application. Figure 12 This is a schematic diagram illustrating the implementation of moiré data enhancement provided in an embodiment of this application. Figure 13 This is a schematic diagram illustrating the implementation of model training provided in an embodiment of this application.
[0179] refer to Figure 10 The steps S1001 to S1004 shown in the figure are as follows:
[0180] S1001. Select multiple images to be processed from the multiple original training images contained in the first training set.
[0181] In this step, the first training set can be understood as the image data training set used to train the focus model. Each training set contains image frame data, focus box information, and actual target phase difference.
[0182] In the first training set, a portion of the images are selected as the images to be processed. The purpose is to enable the model to learn the autofocus mode in new scenes, i.e., scenes with moiré patterns, based on the modes that the model can learn from the original training set.
[0183] In some implementations, a ratio, such as 20%, can be set, and images can be randomly selected from the original training images in the first training set with a 20% probability for subsequent processing. The ratio can be determined experimentally by setting different values, such as 10%, 20%, and 30%. This means using the ratio as a hyperparameter to compare the performance of the final focusing model under different ratios, and seeing which ratio the trained focusing model can maintain the original autofocus accuracy while also handling scenarios where moiré patterns appear on electronic screens captured by the terminal device.
[0184] S1002. Perform moiré pattern overlay processing on multiple images to be processed to obtain multiple enhanced training images.
[0185] This step can be referred to. Figure 11 The following is a detailed description of a training image preprocessing workflow. S1101: Global random data augmentation, specifically including random brightness enhancement, random contrast enhancement, random occlusion (masking), and moiré enhancement. S1102: Selecting M local regions for random data augmentation, specifically including random flipping, random brightness enhancement, and random contrast enhancement. S1103: Random noise enhancement, such as adding additional Gaussian noise or Poisson noise to the image.
[0186] The various random image enhancement methods mentioned above are applied randomly to the training image data. This means different training image datasets may be enhanced using different methods. Furthermore, the same image enhancement method may not be applied exactly the same way to different training image datasets due to this randomness. In principle, all these image enhancement methods aim to improve the performance and robustness of the subsequently trained model through this image processing approach. For example, if the original images are brightly lit and contain no dark scenes, the model might not function properly or its accuracy might be compromised in low-light conditions, such as at night.
[0187] In this embodiment of the application, in order to enable the model to output a stable and effective phase difference under various conditions for autofocus after training, a new moiré data augmentation method is added. For example... Figure 12As shown, moiré patterns can be simulated using a two-dimensional sine wave function. Considering the diversity of moiré patterns in real-world scenarios, different patterns can be constructed by randomly changing the parameters of the sine wave function. For example, the position of the moiré pattern in the image can be altered by randomly changing the center parameter of the sine wave function. Furthermore, the shape of the moiré pattern can be changed by randomly changing the frequency or period parameters of the sine wave function along the X and Y axes.
[0188] In addition to using a sine wave function for simulation, other forms of mathematical functions can also be used or combined for simulation in principle, as long as the simulation result is similar to the moiré pattern seen in the real scene. This application does not impose specific limitations on the embodiments.
[0189] refer to Figure 12 In (a), the moiré pattern in image 1201 is more circular and is located in the center of the image, while the moiré pattern in image 1202 is closer to an ellipse and is tilted to the lower right.
[0190] Understandably, moiré patterns in real-world scenarios are complex and varied, and simulating them with a single sine wave function sometimes fails to fully capture their irregularity. Therefore, a method of superimposing two or more sine wave functions can be used for simulation, resulting in an image that more closely resembles real moiré patterns. For example... Figure 12 As shown in (a), moiré pattern image 1201 and moiré pattern image 1202 are superimposed to obtain moiré pattern image 1203.
[0191] Next, as Figure 12 As shown in (b), the training image 1204 to be enhanced with moiré pattern 1203 is superimposed to obtain the training image 1205 with moiré pattern enhancement. Specifically, there are different ways to superimpose the images, and the appropriate method can be selected according to the actual situation.
[0192] In some implementations, the moiré pattern image can be directly overlaid onto the original training image, which is a simple and quick method.
[0193] In some implementations, both the moiré image and the original training image have an alpha channel to represent transparency. When the two images are superimposed, the superposition can be achieved by adjusting the value of the alpha channel, as shown in the following formula:
[0194] Formula 2
[0195] in, This represents the superimposed image. It can be understood as a moiré pattern image. For the original training images, This is the transparency coefficient. In other words, the proportion of the moiré pattern image in the superimposed image can be controlled in this way. When the value is close to 1, the moiré pattern becomes more pronounced; conversely, when the value is less than 1 When the value is close to 0, the moiré pattern is not obvious.
[0196] Furthermore, to enhance the versatility of the autofocus model and move beyond focusing solely on faces, images without faces should also be included in the training dataset. This allows the trained autofocus model to focus not only on faces but also on other objects or the image center, outputting the phase difference.
[0197] S1003. For any one of the multiple augmented training images, input the augmented training image into the first model so that the first model outputs the predicted target phase difference corresponding to the augmented training image.
[0198] refer to Figure 13 The content shown in the document, the first model 1303 can be understood as an AI model used to achieve autofocus. Its specific internal structure is not specifically limited in this application embodiment. You can refer to the existing related model structure, which will not be described in detail here.
[0199] During training, the training image 1301 and the corresponding focus frame information 1302 are used as inputs to the first model 1303. After model inference, the predicted target phase difference 1304 is obtained, where the target phase difference can be understood as the phase difference corresponding to the focus frame position.
[0200] S1004. Adjust the model parameters of the first model based on the predicted target phase difference corresponding to the enhanced training image and the actual target phase difference corresponding to the enhanced training image.
[0201] like Figure 13 As shown, after the first model 1303 outputs the predicted target phase difference 1304, it compares this phase difference 1304 with the actual target phase difference 1305 to calculate the error. Then, based on the calculated error, the parameters of the first model 1303 are adjusted so that the model can gradually learn during training how to accurately output the phase difference corresponding to the focus frame when moiré patterns exist in the image. For example, the error can be calculated as the squared error between the predicted target phase difference 1304 and the actual target phase difference 1305.
[0202] The reason why adding moiré enhancement to training images improves the accuracy and robustness of the first model 1303 in the presence of moiré patterns is that we assume the label information corresponding to the training images, i.e., the actual target phase difference, is accurate. For training images with added moiré patterns, the first model 1303 adjusts its parameters through backpropagation of errors during gradual training. This allows the predicted target phase difference output by the first model 1303 to get closer and closer to the actual target phase difference, corresponding to increasingly smaller errors. Through this training process, the first model 1303 can learn how to accurately infer the phase difference in images with moiré patterns, achieving better focusing. Therefore, this method improves the generality of the first model 1303 in scenes with moiré patterns.
[0203] Based on the steps described above, the training method for this model will be further summarized and explained below. This embodiment adds moiré image enhancement processing to the original autofocus model training process. Moiré patterns are simulated using a two-dimensional sine wave function and then randomly superimposed onto the original training image. In this way, by using training data with added moiré image enhancement to train the autofocus model, the predicted phase difference output by the model gradually approaches the actual phase difference during the training process, reducing the error between the two.
[0204] The model training method proposed in this application is not complex. It mainly involves enhancing the moiré pattern in the image data used to train the model, so that after learning, the model can determine whether it is in focus even in images with moiré patterns and output the corresponding phase difference. This avoids image blurring or focus jitter caused by moiré patterns when applying the model to a terminal device, providing a better user experience.
[0205] Based on the above embodiments, the following will combine... Figures 14 to 17 The focusing processing method provided in the embodiments of this application will be described in detail. Figure 14 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 2 , Figure 15 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 1 , Figure 16 This is a schematic diagram illustrating the implementation of focus window acquisition provided in an embodiment of this application. Figure 17 This is a schematic diagram illustrating the implementation of model reasoning provided in an embodiment of this application.
[0206] The focusing method provided in this application corresponds to... Figure 15The focusing processing device 1500 shown below will be briefly described. The focusing processing device 1500 includes three modules: a focus frame confirmation module 1501, a model reasoning module 1502, and a lens control module 1503. The focus frame confirmation module 1501 can be connected to the model reasoning module 1502, and the model reasoning module 1502 can be connected to the lens control module 1503.
[0207] Next, we will combine Figure 14 The focusing processing method provided in this application embodiment will be described in detail along with the focusing processing device 1500 described above:
[0208] S1401. Determine the first focus frame in the first preview image.
[0209] This step can be done by Figure 15 The focus confirmation module in the system performs this step. The first preview image in this step can be understood as the image displayed on the screen of the terminal device when the user takes a photo or video.
[0210] For an explanation of the first focus frame, please refer to [link / reference]. Figure 16 .exist Figure 16 In (a), when the terminal device focuses, the focus frame is the center focus frame 1601, which corresponds to the center position of the image. The center focus frame 1601 can be used as the default focus frame for autofocus processing in most cases. Figure 16 In (b), when the terminal device is focusing, the focus frame is the face focus frame 1602. Generally, a face recognition model can be used to determine whether a face exists in the image and the area where the face is located. Compared to the simple center focus frame 1601, the face focus frame 1602 is more intelligent in some situations because when a face appears in the first preview image, the user usually wants to focus on the person, which can bring a better user experience.
[0211] S1402. Input the first preview image and the first focus frame into the first model so that the first model outputs the first target phase difference corresponding to the first focus frame.
[0212] This step can be done by Figure 10 The model inference module 1502 is used for execution, mainly illustrating the process of the first model in practical application, corresponding to the previous model training method. (Refer to...) Figure 17The process involves the following steps: First preview image 1701 is input into second model 1705, which can be understood as a focus box acquisition model, or a face recognition model. Second model 1705 outputs focus box information 1702, including the position and size of the focus box. It's understood that if the first preview image 1701 does not contain a focus box recognized by second model 1705, a default center focus box will be used. Then, the first preview image 1701 and focus box information 1702 are used as input to first model 1703. After inference by first model 1703, the predicted target phase difference 1704, also known as the first target phase difference, is output.
[0213] It is understandable that the model training method described above requires moiré enhancement processing on the training images, while the focusing processing method provided in this application does not require similar moiré enhancement processing. Figure 17 The first preview image 1701 with moiré patterns is only for example. The first model 1703 can process images with or without moiré patterns.
[0214] S1403. Adjust the position of the first motor in the terminal device according to the phase difference of the first target.
[0215] This step can be done by Figure 10 It is executed by the lens control module in the system.
[0216] Understandably, the phase difference value represents the current state of the camera with respect to the first focus frame in the image. The closer the phase difference is to 0, the closer it is to the focus state; conversely, the further it is from the focus state, the further it is from the focus state. The phase difference of the first target has a certain quantitative relationship with the object distance, image distance, and the focal length of the camera lens. This relationship can be obtained through a calibration process for specific hardware devices. For example, data on the phase difference, image distance, and focal length of the camera lens can be collected under different focus states. Then, by combining optical formulas and parameter fitting, the specific relationship between these physical quantities can be obtained. Thus, when the first model outputs the phase difference of the first target, this phase difference can be substituted into the aforementioned relationship to calculate the focal length or image distance at focus.
[0217] In some implementations, the lens focal length is fixed. The image distance can then be calculated using the phase difference of the first target. The distance between the lens and the imaging sensor is then adjusted by a motor to equal the image distance, achieving precise autofocus. During the adjustment process, the first motor provides feedback on the adjusted position. Generally, the first motor may not be able to adjust to the designated position perfectly in one go. Therefore, an error threshold can be set for the first motor. If the error between the current position and the designated focus position is less than this error threshold after the first motor has adjusted its position, the adjustment is considered complete; otherwise, further adjustments are made.
[0218] In summary, the focusing processing method proposed in the embodiments of this application will be summarized below. Figure 10 The corresponding model training method yields an autofocus model, also known as the first model, which is applied to the focusing process. After utilizing this autofocus model, regardless of whether moiré patterns exist in the input image, or whether a face is present in the image, whether face focusing or center focusing is performed, accurate phase difference can be output. Thus, the autofocus model trained based on moiré image enhancement has good versatility, capable of outputting accurate phase difference in various situations, ensuring that the user's captured images maintain good sharpness.
[0219] Based on the above embodiments, the following will further combine... Figures 18 to 21 Another focusing method provided in the embodiments of this application will be described in detail. Figure 18 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 2 , Figure 19 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 3 , Figure 20 This is a schematic diagram illustrating the implementation of time-domain phase difference anomaly detection provided in an embodiment of this application. Figure 21 This is a schematic diagram illustrating the implementation of spatial phase difference anomaly detection provided in an embodiment of this application. Figure 19 The focusing process shown can be viewed as Figure 14 A further form of the focusing process demonstrated.
[0220] The focusing method provided in this application corresponds to... Figure 18The focusing processing device 1800 shown below will be briefly described. The focusing processing device 1800 includes five modules: a focus frame confirmation module 1501, a model inference module 1502, a lens control module 1503, a preset algorithm module 1804, and an anomaly detection module 1805. The focus frame confirmation module 1501 can be connected to both the model inference module 1502 and the preset algorithm module 1804. The preset algorithm module 1804 is connected to the anomaly detection module 1805. The anomaly detection module 1805 can be connected to both the model inference module 1502 and the lens control module 1503. The model inference module 1502 can be connected to the lens control module 1503.
[0221] Next, we will combine Figure 19 The focusing processing method provided in this application embodiment will be described in detail along with the focusing processing device 1800 described above:
[0222] S1901, Get the current focus frame.
[0223] This step can be referenced from the previous text. Figure 10 In the corresponding embodiments, S1401 involves determining whether a face exists in the preview image and determining the corresponding focus frame, which will not be elaborated here.
[0224] S1902. Calculate the phase difference of the focus frame using a preset algorithm.
[0225] The preset algorithm in this step is the traditional autofocus method based on the left and right pixel sensors, as described earlier. Using this preset algorithm, the phase difference corresponding to the focus frame can be quickly calculated. In this embodiment, the phase difference of the current image frame calculated by the preset algorithm, or the traditional autofocus method, can be referred to as the third target phase difference.
[0226] This step can be done by Figure 18 The preset algorithm module 1804 in the middle is used to execute it.
[0227] S1903. Check if the phase difference obtained by the preset algorithm is abnormal. If it is abnormal, execute S1904; otherwise, execute S1905.
[0228] As explained above, the phase difference obtained by the preset algorithm is not necessarily accurate, especially when moiré patterns are present in the image. Under normal circumstances, when there are no moiré patterns in the image, the phase difference corresponding to the focus frame is generally a small absolute value. However, when moiré patterns are present, the phase difference calculated using the preset algorithm is likely to be a relatively large absolute value, far from zero.
[0229] Based on the above characteristics, a corresponding method can be designed to detect whether the phase difference of a preset algorithm is abnormal, such as... Figure 20 As shown, a specific method for detecting anomalies in phase difference in the time domain is illustrated. Figure 20 In (a), assuming the focus frame in image 2001 is 2002, a corresponding phase difference can be calculated according to a preset algorithm. During the shooting process, the terminal device continuously acquires the current preview image using a certain strategy, and can also calculate the phase difference for each frame. Sort the frames by time to obtain a phase difference sequence. In this embodiment, the current image frame, i.e., the image frame before the first preview image, can be called the second preview image. The phase difference calculated using a preset algorithm on multiple second preview images can also be called the second target phase difference. Sort the multiple second target phase differences by time to obtain a sequence called the first sequence.
[0230] exist Figure 20 In (b), such a sequence is exemplified, where the horizontal axis represents image frames, such as frame 1, frame 2, etc., and the vertical axis represents the phase difference. Connecting the phase differences corresponding to different image frames yields a curve. This curve can then be used to determine whether the phase difference calculated by the preset algorithm is abnormal.
[0231] In some implementations, an absolute value threshold A for the phase difference can be set. When the absolute value of a phase difference in the aforementioned curve image exceeds this set threshold A, it can be recorded as an anomaly, depending on whether the value is positive or negative. Then, a threshold M for the number of anomalies can be set. When the number of anomalies within N consecutive frames exceeds the threshold M, it can be considered an anomaly, and the phase difference of the current image frame becomes unusable.
[0232] In some implementations, considering that moiré patterns generally cause focus jitter, corresponding to significant fluctuations in phase difference across image frames, anomalies can be determined by counting the number of "peaks" and "troughs" in the curve. For example, when the phase difference of a given image frame is greater than that of the previous image frame, and the difference is greater than a pre-set threshold B, a "peak" is considered to have occurred, and the number of "peaks" can be called the first quantity. Conversely, when the phase difference of a given image frame is less than that of the previous image frame, and the difference is less than a pre-set threshold C, a "trough" is considered to have occurred, and the number of "troughs" can be called the second quantity. Similarly, a threshold K for the number of "peaks" and "troughs" can be set, also called the first preset threshold. When the number of "peaks" and "troughs" within N consecutive frames exceeds the threshold K, an anomaly can be considered.
[0233] It is understood that the above-described methods are two ways to detect phase differences in the time domain. However, there are many other ways, such as combining the two methods mentioned above. This application does not impose any specific limitations on this.
[0234] Besides detecting phase difference anomalies in the time domain, it can also be done in the spatial domain. (Reference) Figure 21 In (a), a focus frame can be divided into multiple sub-frames, or sub-windows. For example, focus frame 2106 is divided into 9 sub-windows of 3×3, namely sub-window 1, sub-window 2, etc. Furthermore, each sub-window can calculate a phase difference using traditional autofocus methods, such as... Figure 21 In (b), the phase difference table 2105 is a 3×3 table containing 9 phase differences. Considering that these 9 sub-windows all belong to the same focus window and are very close in spatial position, the phase differences corresponding to each sub-window should also be close. Therefore, by comparing whether the phase differences in multiple sub-windows are close to the same value, it can be determined whether the sub-windows obtained by the traditional autofocus algorithm are abnormal.
[0235] refer to Figure 21 In (b), a specific spatial anomaly detection process for phase difference is as follows: S2101, calculate the standard deviation of the phase difference corresponding to multiple sub-windows within a focus frame. S2102, compare whether the calculated standard deviation is greater than a preset threshold, or a second preset threshold. If the standard deviation is greater than the preset threshold, proceed to S2103; otherwise, proceed to S2104. S2103, determine that the phase difference calculated by the preset algorithm is abnormal in the spatial domain. In this case, the phase difference of the focus frame should be recalculated using other methods, such as inferring the phase difference through an autofocus model. S2104, determine that the phase difference calculated by the preset algorithm is normal in the spatial domain. That is, the phase difference calculated by the preset algorithm on the focus frame of the current image frame is reliable and can be used to adjust the camera to the focus state.
[0236] Furthermore, the phase difference calculated above, whether in the time domain or the spatial domain, can be considered abnormal if either one is abnormal. This means that moiré patterns are likely to appear in the current image, which in turn causes traditional autofocus methods to fail.
[0237] In some implementations, when a focus frame contains multiple sub-windows, corresponding to multiple phase differences, the decision of which phase difference to use for focusing is generally based on proximity. This means selecting the area closest to the device or camera among the sub-windows for focusing. This approach is adopted because users typically focus on objects that are close at hand during actual shooting. Furthermore, the magnitude of the phase difference reflects the object distance, or has a monotonic relationship with it; for example, the smaller the object distance, the smaller the phase difference. It's important to note that a smaller phase difference does not mean it approaches zero, but rather includes comparisons at negative values. For instance, if a phase difference of 0 indicates that the focus is currently accurate, moving the object in the focus frame closer to the camera will reduce the phase difference to a negative value. Therefore, the minimum phase difference among multiple phase differences can be directly used for focusing.
[0238] This step can be done by Figure 18 The anomaly detection module 1805 in the middle is used to execute it.
[0239] S1905. Adjust the position of the first motor in the terminal device according to the target phase difference.
[0240] This step is similar to the one described above. Figure 14 The steps S1403 in the corresponding embodiments are basically the same, and will not be described again here.
[0241] S1904. Input the preview image and focus frame information into the first model so that the first model outputs the corresponding target phase difference.
[0242] When it is determined in S1903 that there is an discrepancy between the phase difference calculated by the traditional autofocus method and the actual phase difference, a relatively accurate target phase difference can be obtained by inferring using the first model. More details can be found in the preceding text. Figure 14 The corresponding embodiment of S1402 will not be described again here.
[0243] Based on the process described above, the focusing processing method proposed in this application embodiment will be summarized here. In this application embodiment, by adding a phase difference anomaly detection module, the traditional autofocus method, or preset algorithm, is combined with the first model trained on the moiré enhanced image to form a complete focusing processing method. It is understandable that although the traditional autofocus method is difficult to handle the focusing failure or focusing jitter caused by moiré, it can still focus normally in most cases when there are no other scenes with similar moiré, and has a faster processing speed and lower power consumption. Therefore, by performing anomaly detection on the phase difference of the traditional autofocus method, and then determining whether to directly use the phase difference calculated by the traditional autofocus method or use the first model for inference based on the detection result, it is ensured that accurate focusing and sharpness can still be maintained even when moiré appears in the shooting scene. On the other hand, it also reduces power consumption, speeds up processing, and achieves a faster state of focus, resulting in a better user experience.
[0244] Based on the above embodiments, and in conjunction with... Figure 22 Another focusing method provided in the embodiments of this application will be described in detail. Figure 22 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 4 , Figure 22 The focusing process shown can be viewed as Figure 19 A further form of the focusing process demonstrated.
[0245] exist Figure 22 In the focusing process, some steps are similar to... Figure 19 If there are any inconsistencies, a brief description will be provided; if there are any inconsistencies, a detailed explanation will be given.
[0246] S1901, Get the current focus frame.
[0247] This step is the same as the previous one. Figure 19 The steps S1901 in the corresponding embodiments are basically the same, and will not be described again here.
[0248] S2206. Determine if it is a face focus frame. If it is a face focus frame, execute S1904; otherwise, execute S1902.
[0249] Understandably, the focus frames acquired in S1901 can be divided into two categories: one is the face focus frame containing the face, and the other can be considered a fixed-position focus frame, such as the center focus frame. Different focusing methods can produce different results for these two types of focus frames in certain situations.
[0250] For face focusing, the first model, in its design and training process, not only specifically learns about scenes where moiré patterns may occur, but also focuses on training face focusing. In other words, the first model can output accurate phase difference for faces appearing in different positions, sizes, and angles in the preview image, thus achieving precise focusing.
[0251] Therefore, in this step, different types of focus frames are processed separately, which can achieve better autofocus performance overall.
[0252] S1902. Calculate the phase difference of the focus frame using a preset algorithm.
[0253] This step is the same as the previous one. Figure 19 The steps S1902 in the corresponding embodiments are basically the same, and will not be described again here.
[0254] S1903. Check if the phase difference obtained by the preset algorithm is abnormal. If it is abnormal, execute S1904; otherwise, execute S1905.
[0255] This step is the same as the previous one. Figure 19 The steps S1903 in the corresponding embodiments are basically the same, and will not be described again here.
[0256] S1905. Adjust the position of the first motor in the terminal device according to the target phase difference.
[0257] This step is the same as the previous one. Figure 19 The steps S1905 in the corresponding embodiments are basically the same, and will not be described again here.
[0258] S1904. Input the preview image and focus frame information into the first model so that the first model outputs the corresponding target phase difference.
[0259] This step is executed if the focus box of the current image frame is a face focus box, or if the phase difference obtained by the traditional autofocus algorithm for the current image frame is abnormal. This involves using the first model to infer and calculate the target phase difference. More details on this step can be found in the previous text. Figure 19 The corresponding embodiment S1904 will not be described again here.
[0260] In summary, the focusing processing method proposed in this application focuses on differentiating the specific processing method based on whether the focusing frame of the current image frame belongs to a face focusing frame or a fixed focusing frame, such as a center focusing frame. Considering that the first model has better performance in face focusing compared to traditional autofocus methods, the face focusing scenario is directly fed to the first model for processing. In this way, by subdividing the focusing scenario, different focusing methods are prioritized for different scenarios to further optimize the focusing processing method, fully leveraging the role of the first model and improving the overall focusing processing effect.
[0261] Based on the above embodiments, the following will further combine... Figures 23 to 26 Another focusing method provided in the embodiments of this application will be described in detail. Figure 23 A schematic diagram of the focusing processing device provided in the embodiments of this application. Figure 3 , Figure 24 A schematic diagram of the focusing process of the terminal device provided in the embodiments of this application. Figure 5 , Figure 25 This is a schematic diagram illustrating the implementation of laser ranging provided in an embodiment of this application. Figure 26 A schematic diagram of the terminal device for capturing images of the electronic screen provided in the embodiments of this application. Figure 3 . Figure 24 The focusing process shown can be viewed as Figure 22 A further form of the focusing process demonstrated.
[0262] The focusing method provided in this application corresponds to... Figure 23 The focusing processing device 2300 shown below will be briefly described. The focusing processing device 2300 includes six modules: a focus frame confirmation module 1501, a model inference module 1502, a lens control module 1503, a preset algorithm module 1804, an anomaly detection module 1805, and a laser rangefinder module 2306. The focus frame confirmation module 1501 can be connected to both the model inference module 1502 and the preset algorithm module 1804. The preset algorithm module 1804 is connected to the anomaly detection module 1805. The anomaly detection module 1805 can be connected to both the laser rangefinder module 2306 and the lens control module 1503. The laser rangefinder module 2306 can be connected to both the model inference module 1502 and the lens control module 1503. The model inference module 1502 can be connected to the lens control module 1503.
[0263] Next, we will combine Figure 24 The focusing processing device 2300 described above will be used as a basis for a detailed description of the focusing processing method provided in the embodiments of this application:
[0264] S1901, Get the current focus frame.
[0265] This step is the same as the previous one. Figure 19 The steps S1901 in the corresponding embodiments are basically the same, and will not be described again here.
[0266] S2206. Determine if it is a face focus frame. If it is a face focus frame, execute S1904; otherwise, execute S1902.
[0267] This step is the same as the previous one. Figure 22 The steps S2206 in the corresponding embodiments are basically the same, and will not be described again here.
[0268] S1902. Calculate the phase difference of the focus frame using a preset algorithm.
[0269] This step is the same as the previous one. Figure 19 The steps S1902 in the corresponding embodiments are basically the same, and will not be described again here.
[0270] S1903. Check if the phase difference obtained by the preset algorithm is abnormal. If it is abnormal, execute S2407; otherwise, execute S1905.
[0271] Unlike the process described in the previous embodiment, after detecting an anomaly in the phase difference obtained by the preset algorithm, instead of directly using the first model for processing, laser ranging is preferentially used for focusing. More details on this step can be found in the previous text. Figure 19 The corresponding embodiment S1903 will not be described again here.
[0272] S2407, Using laser for distance measurement.
[0273] This step can be done by Figure 23 The laser ranging module 2306 in the system performs this operation. (See reference...) Figure 25 The content in Figure 25 In (a), a scenario of using laser for distance measurement is shown on terminal device 2501. Terminal device 2501 is equipped with a laser sensor that can emit laser pulses in front of the terminal device. When the laser pulse hits the electronic screen 2502, it is reflected. Then, when terminal device 2501 receives the reflected laser pulse, it can calculate the distance between terminal device 2501 and electronic screen 2502 based on the time of flight (TOF) of the laser pulse.
[0274] like Figure 25As shown in (b), this can be understood as a graph depicting the relationship between the intensity of the emitted and received laser pulses and time in a laser sensor. The horizontal axis represents time, and the vertical axis represents light intensity. When a laser pulse is emitted at a certain moment, a waveform resembling a "peak" (2503) will be obtained. Similarly, when the reflected laser pulse is received, a waveform resembling a "peak" (2504) will also be obtained. Assuming that the highest points of light intensity in waveforms 2503 and 2504 are taken as the moments of emission and reception, respectively, the difference between the two moments, ∆t, can be calculated. Furthermore, the object distance can be calculated using the following formula:
[0275] Formula 3
[0276] Where D represents the object distance and c represents the speed of light. In some cases, the speed of light propagation in the medium can be adjusted for better accuracy.
[0277] Besides using laser pulses to measure the distance between a camera and an object, laser ranging can also use a phase-based method, which will be briefly explained here. The phase in phase-based laser ranging differs from the phase difference in the left / right pixel sensor mentioned earlier; it refers to the period or frequency phase of the light wave itself. The laser sensor emits a modulated continuous laser beam, not a single pulse. This modulation is usually achieved by changing the intensity or frequency of the laser. Common modulation methods include sinusoidal modulation and square wave modulation. When the modulated laser beam illuminates the target object and reflects back to the laser sensor, the laser sensor detects the change in the intensity of the reflected light. Since light takes time to travel to and from the target object, this results in a phase difference between the emitted and received light. By measuring this phase difference, the travel time of the light during the round trip can be calculated, and thus the distance can be calculated. A specific calculation formula is as follows:
[0278] Formula 4
[0279] Where D represents the object distance, and λ is the modulation wavelength of the laser. This represents the phase difference of the light waves.
[0280] Since focusing is based on laser ranging, which directly measures the distance between the terminal device and the object, it is not affected by moiré patterns.
[0281] S2408. Is the laser ranging result stable and valid? If stable and valid, proceed to S2409; otherwise, proceed to S1904.
[0282] Although laser ranging is theoretically unaffected by moiré patterns, it can still be influenced by other environmental factors, leading to unstable results. In this step, the specific method for determining the stability and validity of the laser ranging result can be similar to the method described earlier for determining anomalies in the target phase difference in the time domain. This step can be performed by... Figure 23 The laser ranging module 2306 in the image performs this operation. Understandably, the terminal device continuously focuses during the shooting process, and therefore, laser ranging is also performed cyclically. For each image frame requiring laser ranging, the corresponding object distance can be measured according to the method described in S2407 above. This allows for a series of distance values obtained from laser ranging within N consecutive frames. If the laser ranging is stable, the distance values within N consecutive frames should be the same, or have only a very small, negligible deviation.
[0283] In this embodiment, the laser ranging result corresponding to the current image frame, or the first preview image, can be referred to as the first laser ranging result, while the laser ranging results corresponding to previous image frames, or multiple second preview images, can be referred to as the second laser ranging result. A fluctuation parameter can be calculated by combining the first and second laser ranging results. This fluctuation parameter is then compared with a third preset threshold. If the fluctuation parameter is less than the third preset threshold, the laser ranging result is considered stable. For example, if significant outliers appear in the distance values obtained by laser ranging within N consecutive frames, i.e., abnormal values deviating from other distance values, the number of such outliers can be used as the fluctuation parameter. Alternatively, the overall standard deviation can be used as the fluctuation parameter. If the fluctuation parameter is large, for example, greater than a set threshold, then the laser ranging result of the current frame can be considered unstable and cannot be used as a focusing criterion.
[0284] S2409. Calculate the phase difference based on the laser ranging results.
[0285] Understandably, laser ranging yields the distance between the laser sensor and the object being photographed, which is equivalent to the object distance. Therefore, after determining the object distance for focusing, the image distance or focal length can be calculated. For example, assuming the focal length of the camera in the terminal device is fixed, but the image distance is adjustable, the image distance in the focused state can be calculated based on imaging principles. This can be understood as the distance between the imaging sensor and the camera lens. Furthermore, in some implementations, the image distance used for focusing can be directly transmitted to the first motor in the terminal device for position adjustment. Alternatively, in some implementations, the target phase difference can be calculated using the quantized relationship between the target phase difference, object distance, image distance, and focal length obtained through calibration processing, and then transmitted to the first motor in the terminal device. This step can be performed by... Figure 23 The laser ranging module 2306 in the middle is used to perform this.
[0286] S1905. Adjust the position of the first motor in the terminal device according to the target phase difference.
[0287] After the focusing process provided in the embodiments of this application, such as Figure 26 As shown, even if a user takes a picture of the electronic screen with a terminal device and moiré patterns appear in the preview image, the image can be made clear and in focus after a short processing time, instead of being continuously blurry or experiencing focus jitter that fluctuates between being relatively clear and blurry.
[0288] This step is the same as the previous one. Figure 19 The steps S1905 in the corresponding embodiments are basically the same, and will not be described again here.
[0289] S1904. Input the preview image and focus frame information into the first model so that the first model outputs the corresponding target phase difference.
[0290] In this embodiment, if the focus frame of the current image frame is a face focus frame, or if the phase difference obtained by the traditional autofocus algorithm for the current image frame is abnormal and the laser ranging is unstable, this step will be executed to infer and calculate the target phase difference using the first model. More details about this step can be found in the preceding text. Figure 19 The corresponding embodiment S1904 will not be described again here.
[0291] Based on the above explanation, the focusing method proposed in this application embodiment will be summarized here. When the target phase difference obtained using the traditional autofocus method is determined to be abnormal, laser ranging is preferentially used for autofocus. The image distance or phase difference is calculated based on the object distance obtained from laser ranging, and the motor is adjusted to the corresponding position to achieve accurate autofocus. Even when laser ranging is unstable, a first model trained with moiré-enhanced images is still used to infer the phase difference for accurate focusing. In the focusing method proposed in this application embodiment, the first model can act as a fallback, ensuring autofocus can be achieved in various scenarios. Combining the traditional autofocus method with the laser ranging method, in some common scenarios, accurate focusing can be achieved while shortening the autofocus time, reducing the load and energy consumption of the terminal device. Thus, when users take photos using the terminal device, such as when photographing an electronic screen, where the traditional autofocus method fails due to moiré patterns and the preview image is blurry, laser ranging can be used for further focusing. With stable laser ranging, a clear preview image can be quickly seen, improving the user experience.
[0292] It should be noted that the module names involved in the embodiments of this application can all be defined as other names, as long as they can achieve the function of each module, and no specific restrictions are placed on the module names.
[0293] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0294] The model training method and focus processing method of the embodiments of this application have been described above. The apparatus for performing the above methods is described below. Those skilled in the art will understand that the methods and apparatus can be combined with and referenced in each other, and the related apparatus can perform the steps in the above model training method and focus processing method.
[0295] The model training and focusing methods can be applied to electronic devices with shooting capabilities. Electronic devices include terminal devices; the specific form factors of these terminal devices can be found in the aforementioned descriptions and will not be repeated here.
[0296] In one implementation, this application provides an electronic device. Figure 27 This is a schematic diagram of the hardware structure of an electronic device.
[0297] like Figure 27 As shown, the electronic device 270 includes: a processor 2701 and a memory 2702; the memory 2702 stores computer execution instructions; the processor 2701 executes the computer execution instructions stored in the memory 2702, causing the electronic device 270 to perform the above-described method.
[0298] When the memory 2702 is set up independently, the electronic device also includes a bus 2703 for connecting the memory 2702 and the processor 2701.
[0299] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.
[0300] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0301] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0302] This application provides a computer program product, which includes a computer program that, when run, causes a computer to perform the above-described method.
[0303] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0304] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A model training method, characterized in that, include: From the multiple original training images contained in the first training set, select multiple images to be processed; The multiple images to be processed are subjected to moiré overlay processing to obtain multiple enhanced training images; For any one of the multiple augmented training images, the augmented training image is input into the first model so that the first model outputs the predicted target phase difference corresponding to the augmented training image; The model parameters of the first model are adjusted based on the predicted target phase difference corresponding to the enhanced training image and the actual target phase difference corresponding to the enhanced training image.
2. The method according to claim 1, characterized in that, The process of superimposing moiré patterns on the multiple images to be processed to obtain multiple enhanced training images includes: For any one of the multiple images to be processed, generate a moiré pattern image corresponding to the image to be processed; The moiré pattern image and the image to be processed are superimposed to obtain the enhanced training image corresponding to the image to be processed.
3. The method according to claim 2, characterized in that, The generation of the moiré image corresponding to the image to be processed includes: For the image to be processed, a first sine wave and a second sine wave are generated; The first sine wave and the second sine wave are superimposed to generate a moiré pattern image corresponding to the image to be processed.
4. A focusing processing method, characterized in that, include: Determine the first focus frame in the first preview image; The first preview image and the first focus frame are input into the first model so that the first model outputs the first target phase difference corresponding to the first focus frame, wherein the first model is trained by the method according to any one of claims 1-3; The position of the first motor in the terminal device is adjusted according to the first target phase difference.
5. The method according to claim 4, characterized in that, Before inputting the first preview image and the position of the focus frame into the first model, the method further includes: Multiple second preview images are processed according to a preset algorithm to obtain the second target phase difference corresponding to the second focus frame in each of the multiple second preview images. The acquisition time of the multiple second preview images is before the acquisition time of the first preview image. The first preview image is processed according to the preset algorithm to obtain the phase difference of the third target corresponding to the first focus frame in the first preview image; Based on the second target phase difference corresponding to each of the multiple second preview images and the third target phase difference corresponding to the first preview image, detect whether the third target phase difference is abnormal; and / or, based on the sub-target phase difference corresponding to each of the multiple sub-regions included in the first focus frame, detect whether the third target phase difference is abnormal.
6. The method according to claim 5, characterized in that, The step of detecting whether the third target phase difference is abnormal based on the second target phase difference corresponding to each of the multiple second preview images and the third target phase difference corresponding to the first preview image includes: The phase differences of the second target corresponding to each of the multiple second preview images and the phase differences of the third target corresponding to the first preview image are sorted according to the image acquisition time to obtain a first sequence; In the first sequence, a first number of target phase differences located at wave crests and a second number of target phase differences located at wave troughs are determined; The sum of the first quantity and the second quantity is compared with a first preset threshold to detect whether there is an anomaly in the third target phase difference.
7. The method according to claim 5 or 6, characterized in that, The step of detecting whether the phase difference of the third target is abnormal based on the phase difference of the sub-targets corresponding to the multiple sub-regions included in the first focusing frame includes: For the phase difference of each sub-target corresponding to the multiple sub-regions included in the first focusing frame, calculate the standard deviation; The standard deviation is compared with a second preset threshold to detect whether the third target phase difference is abnormal.
8. The method according to any one of claims 5-7, characterized in that, The step of inputting the first preview image and the first focus frame into the first model, so that the first model outputs the first target phase difference corresponding to the first focus frame, includes: In the event of an abnormality in the phase difference of the third target, the first preview image and the first focus frame are input to the first model so that the first model outputs the phase difference of the first target corresponding to the first focus frame.
9. The method according to any one of claims 5-8, characterized in that, Determining the first focus frame in the first preview image includes: Detect whether the first preview image contains a human face; If the first preview image contains a face, the first focus frame is determined based on the face region in the first preview image; If the first preview image does not contain a human face, the first focus frame is determined based on a preset area in the first preview image.
10. The method according to claim 9, characterized in that, If the first preview image does not contain a human face and the third phase difference is abnormal, the method further includes: Based on the first laser ranging result corresponding to the first preview image and the second laser ranging results corresponding to each of the multiple second preview images, the position of the first motor in the terminal device is adjusted.
11. The method according to claim 10, characterized in that, The step of adjusting the position of the first motor in the terminal device based on the first laser ranging result corresponding to the first preview image and the second laser ranging results corresponding to each of the multiple second preview images includes: Based on the first laser ranging result and the plurality of second laser ranging results, determine the fluctuation parameters of the laser ranging result; If the fluctuation parameter is less than the third preset threshold, then the target time of flight (TOF) corresponding to the first focus frame in the first preview image is determined based on the first laser ranging result. The position of the first motor in the terminal device is adjusted according to the target TOF.
12. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 11.
13. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 11.
Citation Information
Cited By
A focus adjustment method, system and related apparatus
CN122293993A