Model training method, focusing processing method and electronic device
By grouping the training images and adjusting the model parameters, the problem of confidence being affected by the amount of defocus was solved, and more accurate and reliable focusing was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, the confidence level of phase detection autofocus is easily affected by the amount of defocus, resulting in inaccurate confidence and affecting the efficiency and stability of the focusing system.
By grouping multiple training images according to their defocus range, and by adjusting model parameters, the impact of defocus on confidence scores is reduced, thereby improving the accuracy of confidence scores.
This reduces the impact of defocus on confidence levels, improves the accuracy and reliability of model output, and ensures the stability and efficiency of the focusing system.
Smart Images

Figure CN120769165B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a model training method, a focusing processing method, and an electronic device. Background Technology
[0002] Phase Detection Auto Focus (PDAF) is a technology used in mobile photography to quickly achieve focusing. It directly assesses the current focus status of the lens by comparing the phase difference between the images formed on pairs of pixels at opposite positions on the left and right sides of the image sensor after light from the same object passes through the lens.
[0003] The core task of PDAF is to guide the lens to quickly adjust to the optimal focus position by accurately analyzing the phase difference (PD), thereby achieving autofocus. In existing technologies, relevant algorithms can calculate the PD from the image during the preview stage and output the confidence level corresponding to the PD. Then, based on the confidence level of the PD, a decision is made on whether to use PD-guided focusing.
[0004] However, the confidence level of PD is currently affected by the amount of defocus, which may lead to inaccurate confidence level results. Summary of the Invention
[0005] This application provides a model training method, a focusing processing method, and an electronic device to reduce the impact of defocus on confidence level and improve the accuracy and reliability of confidence level.
[0006] Firstly, embodiments of this application propose a model training method. The method includes: inputting any one of a plurality of training images into a first model to obtain the predicted phase difference of the training image output by the first model, and the predicted confidence level of the predicted phase difference. Determining a first error for each of the plurality of training images based on their respective true phase differences and predicted phase differences. Grouping the plurality of training images according to their respective defocus amounts, with each training image within the same group corresponding to a defocus amount within the same range. Finally, adjusting the model parameters of the first model based on the first errors and prediction confidence levels of the training images in the multiple groups.
[0007] In the model training method of this application, the first model is used to process images and can output the phase difference of the image and the corresponding confidence level. It can be understood that when a training image passes through the first model, the first model can output the predicted phase difference of the training image and the predicted confidence level corresponding to the predicted phase difference. Furthermore, the first model processes multiple training images individually to output the predicted phase difference and prediction confidence level for each of the multiple training images.
[0008] After the first model outputs the predicted phase difference and prediction confidence for multiple training images, the first error for each training image can be determined based on the predicted phase difference and the true phase difference of each training image. The first error indicates the degree of difference between the true phase difference and the predicted phase difference of the training images.
[0009] In one possible implementation, the first error can be calculated based on the difference between the true phase difference and the predicted phase difference. However, when the training image itself is out of focus, due to blurring and loss of detail, the difference between the predicted phase difference and the true phase difference output by the first model may be relatively large. In other words, the amount of defocus on the training image has a significant impact on the error of the training image; specifically, the greater the defocus, the greater the error of the resulting training image.
[0010] To reduce the impact of defocus on the error of the training image, this application proposes another method for calculating the first error. In this method, the initial error can be calculated first based on the true phase difference and the predicted phase difference. Then, the initial error of the training image is adjusted according to the defocus corresponding to the training image to obtain the first error of the training image, thereby reducing the impact of defocus on the first error.
[0011] Therefore, in one possible implementation, determining the first error for each of the multiple training images includes:
[0012] The initial error of each training image is determined by the difference between its true phase difference and its predicted phase difference.
[0013] Then, based on the defocus amount corresponding to each of the multiple training images, the initial errors of each of the multiple training images are adjusted to obtain the first error of each training image.
[0014] In one possible implementation, the initial error of each training image can be obtained by reducing the initial error of each training image based on the defocus amount of each training image. It can be understood that the degree of reduction in the initial error is positively correlated with the phase difference of each training image; that is, the larger the phase difference of each training image, the greater the reduction in the initial error.
[0015] In one possible implementation, the first error can be the ratio of the initial error of each training image to the defocus amount of each training image.
[0016] In this way, by utilizing the defocus amount corresponding to each training image to reduce the initial error of each training image, the impact of different defocus amounts on the error can be reduced. Specifically, the greater the defocus amount of a training image, the greater its blurriness, which may lead to a larger error between the predicted phase difference and the true phase difference. Therefore, this method can reduce the impact of defocus amount on the subsequent determination of the true confidence level for the training image, thereby improving the accuracy of model evaluation.
[0017] It is understandable that during the model training process, it is also necessary to evaluate whether the predicted phase difference and prediction confidence output by the first model are accurate, so as to adjust the first model according to the evaluation results, so that the trained first model can output relatively accurate predicted phase difference and prediction confidence.
[0018] The first error can be understood as indicating the degree of difference between the true phase difference and the predicted phase difference of the training image. In other words, the first error can be used to evaluate the accuracy of the predicted phase difference output by the first model. Simultaneously, the prediction confidence score corresponding to the predicted phase difference output by the first model can also be used to evaluate the accuracy of the predicted phase difference output by the first model; that is, the confidence score corresponding to the predicted phase difference output by the first model can be calculated based on the first error. It can be understood that the first error is negatively correlated with the confidence score of the training image; that is, the smaller the first error, the higher the confidence score of the training image. Therefore, the accuracy of the predicted phase difference output by the first model can be reflected by setting a higher confidence score for training images with smaller first errors and a lower confidence score for training images with larger first errors.
[0019] In one possible implementation, the confidence level of the training images can be determined based on the ranking of the first errors. First, the first errors of the training images can be sorted in ascending order. Then, the confidence level of the training images ranked higher in the sorting result is set to a higher value, and the confidence level of the training images ranked lower in the sorting result is set to a lower value, thus determining the confidence level of each training image.
[0020] It's understandable that when sorting the initial errors of multiple training images, the presence of defocus might cause images with larger defocus to be ranked lower, while images with smaller defocus to be ranked higher. In this case, the true confidence score of the training images will be affected by the defocus level. Therefore, grouping multiple training images can reduce the impact of defocus on the confidence score.
[0021] In one possible implementation, multiple training images can be grouped based on their respective defocus values. This includes: after obtaining the defocus values of all training images, dividing the training images into different groups according to a preset defocus value range. Within the same group, the defocus values of multiple training images belong to the same defocus value range. In other words, the defocus values of multiple training images within the same group are the same or close.
[0022] In this way, the defocus amounts of multiple training images within the same group are the same or similar, which means that the blur levels of the training images within the same group are similar. Ideally, the deviation range of the first error of each training image is similar. Therefore, in the subsequent sorting process, because the blur levels of images within the same group are similar, the impact of the defocus amount of each training image on the first error sorting is reduced, thereby reducing the impact of the defocus amount of each training image on the true confidence score.
[0023] This is understandable. Then, the first errors of multiple training images in each group are sorted to determine the true confidence level corresponding to the training images.
[0024] In one possible implementation, the first errors of the multiple training images in each group are sorted, including: sorting the multiple training images in each group in ascending order of their first errors. In the final sorting result, the smallest first error in each group is located in the first sorting position, and the largest first error is located in the last sorting position.
[0025] Then, based on the sorting results, the true confidence score corresponding to the training image with the first error located in the first position interval of the sorting results is determined as the first value, and the true confidence score corresponding to the training image with the first error located in the second position interval of the sorting results is determined as the second value. This process includes:
[0026] In one possible implementation, the ranking results of each group are divided into intervals according to a preset threshold to determine the true confidence level corresponding to the training image.
[0027] By setting a preset threshold, the sorting results of each group can be divided into two position intervals, namely the first position interval and the second position interval.
[0028] The first position interval can be the interval from the first sorting position in each group of sorting results to the sorting position of the preset threshold. The second position interval can be the interval from the next sorting position after the sorting position of the preset threshold in each group of sorting results to the last sorting position.
[0029] Then, the true confidence score corresponding to the first error located in the first position interval is marked as a first value. The true confidence score corresponding to the training image located in the second position interval is marked as a second value. For example, the true confidence score corresponding to the training image located in the first position interval is marked as 1, and the true confidence score corresponding to the training image located in the second position interval is marked as 0.
[0030] In this way, based on the ranking results of the first error within each group, the true confidence value corresponding to the training image with a smaller first error is set to a higher value, and the true confidence value corresponding to the training image with a larger first error is set to a lower value, which can determine the reliability of the predicted phase difference corresponding to each training image output by the first model.
[0031] Then, based on the prediction confidence and true confidence of each training image in the group, the second error of each training image in the group is determined.
[0032] In one possible implementation, the second error for each training image can be the difference between the true confidence of each training image and the predicted confidence of each training image output by the first model.
[0033] Finally, after obtaining the first and second errors of the training images in each group, the model parameters of the first model are adjusted according to the first and second errors of the training images in each group.
[0034] In one possible implementation, the first error can be backpropagated to the first model, and the parameters of the model can be adjusted based on the first error to make the model output a more accurate prediction of the phase difference.
[0035] In another possible implementation, the second error can be backpropagated to the first model, and the parameters of the second error model itself can be used to prompt the model to output a more accurate prediction confidence.
[0036] Furthermore, in one possible implementation, both the first error and the second error can be backpropagated to the first model, and the parameters of the model can be adjusted based on the first error and the second error to enable the model to output more accurate predicted phase difference and predicted confidence simultaneously.
[0037] In this way, by using a first error, which measures the difference between the true phase difference and the predicted phase difference corresponding to the training image, and / or a second error, which measures the difference between the true confidence and the predicted confidence corresponding to the training image, to guide the model in parameter tuning, the first model can output more accurate predicted phase difference and predicted confidence.
[0038] Secondly, embodiments of this application propose a focusing processing method. The method includes: acquiring a preview image; inputting the preview image into a first model to obtain a first phase difference of the preview image and a first confidence level corresponding to the first phase difference, wherein the first model is trained according to the above method; and finally, if the first confidence level is greater than a preset confidence level, performing focusing processing based on the first phase difference.
[0039] In the focusing processing method of this application, before shooting, multiple preview images are first acquired. These multiple preview images may be acquired at different times, under different focusing states, and from different angles.
[0040] Next, the preview image is processed by the first model, which outputs the first phase difference corresponding to the preview image and the first confidence level corresponding to the first phase difference. The first phase difference can be the predicted phase difference output by the first model. In other words, when a preview image passes through the first model, the first model outputs the first phase difference corresponding to that preview image. The first confidence level can be the predicted confidence level corresponding to the predicted phase difference output by the first model. In other words, when a preview image passes through the first model, the first model outputs the first confidence level corresponding to that preview image.
[0041] Thus, it can be understood that, based on the model training method described in the first aspect, the first model, after receiving an image, can output a relatively accurate predicted phase difference corresponding to the input image, as well as a prediction confidence level corresponding to the predicted phase difference. In other words, in focusing applications, when the preview image passes through the trained first model, the first model can output a relatively accurate first phase difference corresponding to the preview image, as well as a first confidence level corresponding to the first phase difference.
[0042] Finally, a preset confidence level is set. If the first confidence level is greater than the preset confidence level, focusing is performed based on the first phase difference. If the first confidence level is less than the preset confidence level, other more robust techniques are used to complete the image focusing process. Focusing based on the first phase difference can refer to adjusting the motor position using the first phase difference of the image, bringing the motor closer to the focal point.
[0043] It is understandable that the focusing system determines subsequent focusing processing based on the first confidence level corresponding to the image output by the first model. In this way, accurate confidence level output can avoid wasting valuable first phase difference information, thereby improving the efficiency and stability of the focusing system.
[0044] Thirdly, embodiments of this application provide a focusing device, which may be an electronic device, or a chip or chip system within an electronic device. The focusing device may include a display unit and a processing unit.
[0045] When the focusing device is an electronic device, the display unit therein can be a display screen. The display unit is used to perform the display step so that the electronic device implements a focusing method described in the first aspect or any possible implementation of the first aspect.
[0046] When the focusing device is an electronic device, the processing unit may be a processor. The focusing device may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement a focusing method described in the first aspect or any possible implementation of the first aspect.
[0047] When the focusing device is a chip or chip system within an electronic device, the processing unit can be a processor. The processing unit executes instructions stored in the storage unit to cause the electronic device to implement a focusing method described in the first aspect or any possible implementation of the first aspect. The storage unit can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).
[0048] For example, a processing unit is used to process images and achieve focusing based on a focusing system. A display unit is used to display the subject being photographed, etc.
[0049] Fourthly, embodiments of this application provide an electronic device including a processor and a memory, the memory for storing code instructions, and the processor for running the code instructions to perform the methods described in the first aspect or any possible implementation of the first aspect.
[0050] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0051] In a sixth aspect, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0052] Seventhly, this application provides a chip or chip system including at least one processor and a communication interface, wherein the communication interface and at least one processor are interconnected via a circuit, and the at least one processor is used to run a computer program or instructions to perform the methods described in the first aspect or any possible implementation thereof. The communication interface in the chip may be an input / output interface, pins, or circuits, etc.
[0053] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).
[0054] It should be understood that the third to seventh aspects of this application correspond to the technical solutions of the first and second aspects of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of a shooting scene provided in an embodiment of this application;
[0056] Figure 2 A schematic diagram of the processing flow of the PD model provided in the embodiments of this application;
[0057] Figure 3 A schematic diagram illustrating the PD value-guided focusing process provided in this application embodiment;
[0058] Figure 4 A schematic diagram of the hardware framework of the electronic device provided in the embodiments of this application;
[0059] Figure 5 A schematic diagram of the software framework of the electronic device provided in the embodiments of this application;
[0060] Figure 6 A schematic diagram illustrating PD model training provided in an embodiment of this application;
[0061] Figure 7 This is a schematic diagram illustrating the relationship between defocus amount, predicted phase difference, and confidence level provided in the embodiments of this application.
[0062] Figure 8A schematic diagram of training image grouping provided in an embodiment of this application;
[0063] Figure 9 A schematic diagram illustrating the minimum error sorting provided in the embodiments of this application;
[0064] Figure 10 A schematic diagram of the true confidence level marker provided in the embodiments of this application;
[0065] Figure 11 A schematic diagram of a PD model provided in an embodiment of this application;
[0066] Figure 12 A schematic diagram of another PD model provided in an embodiment of this application;
[0067] Figure 13 A schematic diagram illustrating the imaging application of the PD model provided in the embodiments of this application;
[0068] Figure 14 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0069] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0070] 1. Phase Detection (PD)
[0071] Phase difference is one of the key factors for achieving fast and accurate focusing. By measuring and calculating the phase difference, the camera can quickly adjust the lens position to achieve accurate focusing.
[0072] The magnitude of the phase difference directly reflects the distance between the focal point and the current sensor position. The larger the phase difference, the farther the focal point is from the current position; the smaller the phase difference, the closer the focal point is to the current position.
[0073] 2. Defocus
[0074] Defocus distance refers to the distance from the motor's current position to the focal point. A larger defocus distance indicates that the focal point is farther from the current position; a smaller defocus distance indicates that the focal point is closer to the current position. Phase difference has a linear relationship with defocus distance. When focusing, the motor is at the focal point, and the object is in sharp focus.
[0075] 3. Confidence
[0076] Confidence level is used to assess the reliability and accuracy of a model's predicted PD value. Ideally, the model's confidence level should accurately reflect the accuracy of its predictions; that is, a high-confidence prediction should highly accurately reflect the actual situation.
[0077] 4. Motor Lenspos
[0078] The motor is used to precisely control the physical position of the lens. Lenspos represents the lens's specific position on the optical axis. When the defocus is 0, the position of the motor Lenspos is the lens position when focusing. Ideally, the defocus is linearly related to the motor Lenspos.
[0079] 5. Other terms
[0080] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0081] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0082] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0083] 6. Electronic equipment
[0084] The electronic devices in this application embodiment may include handheld devices with shooting functions, vehicle-mounted devices, etc. For example, some electronic devices include: mobile phones, tablets, PDAs, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, terminal devices in 5G networks, or future evolution of public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.
[0085] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0086] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0087] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0088] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0089] To better understand the technical solution of this application, the relevant technologies involved in this application will be further described in detail below.
[0090] PDAF is a technology for achieving fast focusing in mobile photography. Its core lies in utilizing the principle of phase difference and relying on the special pixel layout on the image sensor and corresponding algorithm processing to achieve fast and accurate focusing.
[0091] For example, during PDAF execution, pixels symmetrically spaced at intervals on the image sensor can be used to cover the left and right halves of individual pixels, allowing these pixels to receive light from different directions within the lens. For instance, the left pixel receives light from the left side of the lens, and the right pixel receives light from the right side, similar to the difference in angle at which the left and right eyes perceive an object.
[0092] Then, by comparing the phase difference of the image signals received by these pixels, it is converted into a quantifiable defocus amount, and the lens position is adjusted accordingly to achieve accurate focusing.
[0093] PDAF technology has a wide range of applications in photography scenarios, which will be discussed below. Figure 1 A brief description of the shooting scene. Figure 1 This is a schematic diagram of a shooting scene provided in an embodiment of this application.
[0094] like Figure 1 As shown, the electronic device 100 is a mobile phone's camera interface, which can contain multiple controls, such as... Figure 1 In the example, based on a portrait-oriented interface, the controls at the top of the page, from left to right, are: flash on / off, AI on / off, HDR on / off, color style, and settings. The controls at the bottom of the page, from left to right, are: gallery, shutter button, and front / rear camera switch. In addition, the page includes multiple switching modes such as night scene, portrait, photo, and video. The page also includes a focus switching control. Figure 1 The multiple controls shown, containing the text 0.5, 1x, 2.5, and 5, are the focus switching controls. Clicking the focus switching controls allows you to switch the focus during shooting.
[0095] And in Figure 1 The document also shows a "More" shortcut control. Clicking the "More" shortcut control will display other operation controls on the shooting page. This embodiment does not limit the specific implementation of the operation controls included on the shooting page.
[0096] When taking a picture of subject 101, the camera captures the image by clicking the shutter button and saves it to the image library. However, before the image is displayed, the camera needs to perform a series of operations to focus. This process is as follows: Figure 1 As shown, when the camera app is accessed, the phone's camera can acquire a preview image of the subject 101. Before the camera app triggers the shooting operation, the camera can, for example, continuously acquire preview images of the captured image 101, thereby obtaining preview images at different times, where the preview images at different times can, for example, correspond to different focus states.
[0097] by Figure 1For example, suppose three preview images are obtained for the subject 101. Preview image 103 is acquired at 11:40:15, and its focus is, for example, blurry. Preview image 104 is acquired at 11:40:16, and its focus is, for example, somewhat blurry. Preview image 105 is acquired at 11:40:17, and its focus is, for example, relatively sharp.
[0098] Referring to the above Figure 1 It can be determined that preview images captured at different times can correspond to different focus states. In one implementation, for example, the terminal device will use PDAF technology to perform focus processing, so as to focus the subject during the preview stage, thereby providing a good focus state for subsequent shooting operations.
[0099] Next, we will explain in detail how PDAF achieves efficient focusing. In one implementation, PDAF technology may include a PD model, which processes the image to obtain its PD value and confidence level. The following section will combine... Figure 2 A brief introduction to the PD model, Figure 2 This is a schematic diagram of the processing flow of the PD model provided in the embodiments of this application.
[0100] exist Figure 2 The diagram illustrates the complete process of image processing, from inputting raw image data to processing it through the PD model, and finally outputting a PD value to guide focusing. The following section combines... Figure 2 The input and output process of the PD model is explained in detail.
[0101] In the PDAF process, for example, an image can be segmented into multiple basic units. In one implementation, the basic unit of segmentation can be a single pixel.
[0102] Alternatively, the basic unit of segmentation can also be a pixel block composed of multiple pixels, where a pixel block can also be understood as a focus frame. For example, in Figure 2 In the example, nine pixels within a 3×3 area can be considered as a pixel block. In actual implementation, the number of pixels contained within a pixel block can be set according to actual needs, which will not be described in this embodiment.
[0103] The following section describes the specific processing of PDAF when the basic units are pixels and pixel blocks.
[0104] Reference Figure 2For example, PDAF can be performed on pixel 200 in an image. Each pixel in the image is used as the basic unit as input to the PD model 202. Then, for each pixel, its corresponding PD value is predicted, and a confidence score associated with that predicted value is output simultaneously. It is understood that the processing method for the remaining pixels in the image is similar, and will not be described here.
[0105] Reference Figure 2 For example, PDAF can also be performed on pixel block 201 in the image. Each pixel block in the image is used as the basic unit as input to the PD model 202. Then, for each pixel block, its corresponding PD value is predicted, and a confidence score associated with that predicted value is output simultaneously. It is understood that the processing method for the remaining pixel blocks in the image is similar, and will not be described here.
[0106] Next, the focusing system needs to determine whether to use the PD value to guide the system's focusing based on the confidence level output by the PD model. The following section will combine... Figure 3 This section introduces the complete process of PDAF implementation. Figure 3 This is a schematic diagram of the PD value-guided focusing process provided in the embodiments of this application.
[0107] like Figure 3 As shown, each basic unit of the image is first input into the PD model as an independent analysis object, and then the PD model calculates and outputs the PD value and its corresponding confidence level.
[0108] Confidence level is an important indicator for evaluating the quality of a model's predictions. It quantifies the accuracy and reliability of a PD model's predictions of PD values. For example... Figure 3 As shown, when the confidence level of the model's output for a certain basic unit is high, the system tends to adopt the PD value predicted by the model and adjust the position of the lens accordingly to achieve fast and accurate focusing. When the confidence level of the model's output for a certain basic unit is low, it means that the model's prediction result for that area may not be reliable enough. At this time, the system abandons the use of the PD value output by the model and instead adopts other more robust technical means to complete the focusing process of the pixels in that basic unit.
[0109] In existing technologies, confidence levels are typically significantly influenced by the amount of defocus. Specifically, in one approach, confidence level is determined by the probability distribution and expected distribution of the defocus amount output by the model. In another approach, confidence level calculation relies on ranking the errors between the predicted and actual defocus amounts by the PD algorithm. In both approaches, confidence levels exhibit dynamic changes as the lens moves. As the lens approaches the focus position during motor movement, confidence levels increase due to the decrease in defocus; conversely, confidence levels decrease as the lens moves away from the focus position.
[0110] However, within the theoretical framework, the confidence score output of the PD model should be designed to be independent of the instantaneous position of the motor Lenspos. This design principle aims to ensure the consistency and reliability of the confidence score assessment under different focusing conditions.
[0111] Specifically, in scenes with rich textures, the system can fully rely on the PD value to guide the focusing process because the image contains sufficient detail. In this case, the ideal PD confidence level should be maintained at a high level, and this high confidence level should remain stable during the motor Lenspos adjustment of the focus position, unaffected by changes in the motor position.
[0112] Conversely, in scenes with low texture or lacking distinct features (such as backlit portraits, high-contrast scenes, scenes with repetitive textures, scenes with display content, and scenes with weak textures), the PD value may not provide sufficiently accurate focus guidance. In such cases, the ideal PD confidence level should automatically decrease to a low, constant level to reflect the limitations of PD technology in these types of scenes. Simultaneously, this low confidence level should remain constant throughout the motor-driven Lenspos focusing attempt to accurately reflect the predictive ability and reliability of PD technology in the current scene.
[0113] The output of PD confidence should be an objective reflection of the current scene characteristics and the applicability of PD technology, and should not be directly related to the specific position of the motor Lenspos, so as to ensure the stability and accuracy of the focusing system.
[0114] Therefore, it is understandable that in the related technologies described above, because the confidence level is usually significantly affected by the defocusing amount, the confidence level of the PD model may be inaccurate, potentially discarding usable PD results at the defocusing point. This could not only waste effective focusing information but also affect the overall performance and efficiency of the focusing system.
[0115] To address the aforementioned problems, this application proposes the following technical concept: A phase difference (PD) model is trained to output the phase difference (PD) and confidence score for each image. During model training, multiple images are processed using the model to output the predicted PD and predicted confidence score. The actual confidence score is then determined based on the PD error between the predicted and actual PD. Furthermore, the confidence score error of the model's prediction is determined based on the actual confidence score and the predicted confidence score shown by the model, thereby optimizing the model. Simultaneously, in determining the actual confidence score, the multiple images are grouped according to their respective phase differences, ensuring that the phase differences in each group are the same or similar. Subsequent confidence score processing is performed separately for each group, thereby reducing the impact of phase difference on the confidence score during subsequent model optimization.
[0116] The model training method and focus processing method of the embodiments of this application can be executed by an electronic device, or by a chip, chip system, or processor that supports the implementation of the model training method and focus processing method by an electronic device, or by a logic module or software that can implement all or part of the functions of an electronic device. This application does not impose specific limitations in this regard. The model training method and focus processing method of the embodiments of this application will be described in detail below using an electronic device as the execution subject.
[0117] In one implementation, the electronic device used to implement the model training method can be, for example, a server or a terminal device, and the electronic device used to implement the focus processing method can be, for example, a terminal device. The following section will first combine... Figure 4 and Figure 5 A brief introduction to the terminal equipment.
[0118] For example, Figure 4 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application.
[0119] Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0120] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device. In other embodiments of this application, the terminal device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0121] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). Different processing units may be independent devices or integrated into one or more processors. In this embodiment, processor 110 may, for example, be used to execute the various logical processes and data processes included in the xx method provided in this application.
[0122] The terminal device implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0123] Display screen 194 is used to display images, videos, etc.
[0124] The terminal device can implement the shooting function through an ISP, camera 193, video codec, GPU, display 194, and application processor. In one implementation, for example, the various units described herein can work together to capture a preview image during the shooting process and display the preview image on the display screen.
[0125] Camera 193 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal.
[0126] The software system of a terminal device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This application uses the layered architecture Android system as an example to illustrate the software structure of the terminal device.
[0127] For example, Figure 5 This is a schematic diagram of the software structure of a terminal device provided in an embodiment of this application.
[0128] like Figure 5 As shown, the layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the system may include an application layer, an application framework layer, an Android runtime and system libraries, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that this application uses the Android system as an example; however, the solution can also be implemented in other operating systems (such as HarmonyOS, iOS, etc.) as long as the functions implemented by each module are similar to those in the embodiments of this application.
[0129] The application layer can include a series of application packages.
[0130] like Figure 5 As shown, the application package may include applications such as camera, calendar, phone, map, music, settings, email, video, and social media. Of course, the application layer may also include other application packages, such as third-party applications like payment apps, shopping apps, banking apps, and social media apps; this application is not limited to these. In this embodiment, for example, a camera can be used to implement the photo-taking function. During the photo-taking process, a photo focus point (PD) can be determined based on the captured image, and the PD can guide the camera application to focus.
[0131] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0132] like Figure 5 As shown, the application framework layer may include a window manager, content provider, resource manager, view system, notification manager, etc.
[0133] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0134] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0135] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0136] The system library can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0137] The HAL layer is a wrapper around Linux kernel drivers, providing interfaces to the upper layers and shielding them from the implementation details of the lower-level hardware.
[0138] The HAL layer may include a Wi-Fi HAL, an audio HAL, a Camera HALServer unit, and software code libraries. In this embodiment, a first model may be deployed in the HAL layer. This first model is used to determine the corresponding PD (Photo Focusing) and its confidence level based on the image captured by the camera application. The PD and its confidence level can then be fed back to the camera application, which can use the PD to guide focusing if the PD confidence level is high.
[0139] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0140] The technical solutions of the embodiments of this application and how the technical solutions of the embodiments of this application solve the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific examples. The following specific embodiments can be implemented independently or in combination with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0141] The technical solution of this application involves training the PD model and applying the PD model. The following sections will introduce these two parts separately.
[0142] First, the training process of the PD model will be explained with reference to specific examples. The following section will further explain... Figures 6 to 10 To provide a detailed introduction, Figure 6 This is a schematic diagram illustrating the training of the PD model provided in an embodiment of this application. Figure 7 This diagram illustrates the relationship between defocus amount, predicted phase difference, and confidence level provided in the embodiments of this application. Figure 8 This is a schematic diagram of training image grouping provided in an embodiment of this application. Figure 9This is a schematic diagram illustrating the minimum error sorting method provided in an embodiment of this application. Figure 10 This is a schematic diagram illustrating the true confidence level labeling provided in an embodiment of this application. The training process of the PD model may include, for example, the following steps:
[0143] 1. For any one of the multiple training images, input the training image into the first model to obtain the predicted phase difference of the training image output by the first model, and the prediction confidence of the predicted phase difference.
[0144] The first model in this embodiment is used to process the image to output the phase difference of the image and the corresponding confidence level. This embodiment describes the training process for the first model; therefore, it can be understood that when a training image passes through the first model, the first model can output the predicted phase difference of the training image and the predicted confidence level corresponding to the predicted phase difference.
[0145] For example, the first model can refer to the PD model described above, or the specific implementation of the first model can be arbitrarily selected according to actual needs, as long as the first model can be used to implement the functions described above. The following explanation uses the PD model as the first model.
[0146] In this embodiment, the first model can be trained using multiple training images. During the training process, multiple training images can be input into the first model all at once to perform subsequent training. Alternatively, multiple training images can be input into the first model in batches to perform subsequent training. Or, multiple training images can be input into the first model sequentially to perform subsequent training. (See reference...) Figure 6 Multiple training images are input into the PD model 600.
[0147] However, it's important to understand that regardless of the training method used for the multiple training images, the first model processes each training image individually to output the predicted phase difference and prediction confidence for each image. For example... Figure 6 As shown, the PD model 600 outputs the predicted phase difference and prediction confidence for each of the multiple training images. Furthermore, the first model processes each of the multiple training images in a similar manner; therefore, the following explanation uses any single training image as an example.
[0148] The predicted phase difference refers to the phase difference calculated by the PD model based on the feature information of the training image after the training image has passed through it.
[0149] Prediction confidence refers to the degree of confidence the PD model has in the predicted phase difference after the training image has passed through it. A higher prediction confidence indicates greater confidence in the predicted phase difference, suggesting a higher reliability of the prediction. Conversely, a lower prediction confidence indicates less confidence in the predicted phase difference, suggesting a lower reliability of the prediction.
[0150] In one implementation, the PD model can, for example, use image features such as color features, texture features, shape features, and spatial relationship features, as well as information such as the degree of blur in the image, to analyze and calculate in order to obtain the predicted phase difference and prediction confidence of the training image.
[0151] 2. Based on the true phase difference and the predicted phase difference of each training image, determine the first error of each training image.
[0152] It is understandable that the processing method for each training image in step 2 is similar, so the following will use any training image as an example.
[0153] In one implementation, the true phase difference of the training images can be a pre-set value. Based on the above description, it can be determined that the first model can output the predicted phase difference of the training images. In this embodiment, the first error of the training images can be determined based on the true phase difference and the predicted phase difference of the training images. (Reference) Figure 6 After the PD model outputs the predicted phase difference and the true phase difference of the training image, the first error of the training image can be determined based on the predicted phase difference and the true phase difference of the corresponding training image. It can be understood that the first error is used to indicate the degree of difference between the true phase difference and the predicted phase difference of the training image.
[0154] For example, the difference between the actual phase difference and the predicted phase difference can be determined as the first error. Alternatively, the ratio of the actual phase difference to the predicted phase difference can also be determined as the first error. In actual implementation, the method for determining the first error can be determined according to actual needs, and this embodiment does not limit it in this way.
[0155] It's also understandable that when the defocus amount corresponding to the training image is large, the image may be in a defocused state. This means that the light has already undergone significant diffusion or insufficient convergence before reaching the imaging surface, potentially causing blurring and loss of detail in the training image. In this case, because the training image itself is in a defocused state, the difference between the predicted phase difference output by the first model based on the defocused training image and the true phase difference will naturally be relatively large. In other words, the defocus amount of the training image has a significant impact on the error of the training image; specifically, the greater the defocus amount, the greater the error of the corresponding training image.
[0156] To reduce the impact of defocus on the error of the training image, this application proposes another method for calculating the first error. In this method, the initial error can be calculated first based on the true phase difference and the predicted phase difference. Then, the initial error of the training image is adjusted according to the defocus corresponding to the training image to obtain the first error of the training image, thereby reducing the impact of defocus on the first error.
[0157] The initial error can be, for example, the difference between the true phase difference and the predicted phase difference, or it can be the ratio of the two. This application does not limit the calculation method of the initial error, as long as the initial error is directly calculated based on the true phase difference and the predicted phase difference, and can reflect the difference between the true phase difference and the predicted phase difference.
[0158] In one implementation, the larger the difference between the true phase difference and the predicted phase difference, the larger the initial error will be; that is, there is a positive correlation between the difference between the true phase difference and the predicted phase difference and the initial error.
[0159] For example, taking the initial error d as the difference between the true phase difference Pd and the predicted phase difference Pd' output by the PD model, the initial error d can satisfy the following formula:
[0160] d = |Pd - Pd'| Formula 1
[0161] This embodiment does not limit the calculation method of the initial error d. It can be understood that the closer the predicted phase difference output by the PD model is to the true phase difference, the smaller the value of the initial error d, which means that the predicted phase difference output by the PD model is more accurate.
[0162] Subsequently, to reduce the impact of the defocus amount of the training images on the error, the initial error d of the training images can be further adjusted based on the defocus amount to obtain the final first error for use. The purpose of adjusting the initial error d is to reduce the influence of the defocus amount on the error. For example, analogous to the normalization process, the initial error d can be reduced to eliminate the influence of the defocus amount on the error calculation.
[0163] In one implementation, for example, the initial error of the training image can be reduced based on the defocus amount of the training image to obtain a first error d1. For instance, the larger the defocus amount corresponding to the training image, the larger the range of reduction in the initial error; that is, there is a positive correlation between the range of reduction in the initial error and the defocus amount corresponding to each training image.
[0164] As an example, in one implementation, the first error d1 can satisfy the following formula:
[0165]
[0166] In Formula 2, defocus refers to the amount of defocus on the training image. Referring to Formula 2, it can be determined that, for example, the ratio of the initial error of the training image to the amount of defocus on the training image can be defined as the first error, thus eliminating the influence of defocus on the error and obtaining the first error after eliminating the influence of defocus.
[0167] However, it's understandable that in actual implementation, the method of calculating the first error based on the defocus amount of the training image and the initial error can be arbitrarily extended according to actual needs. As long as the processing logic reduces the initial error based on the defocus amount, and the degree of reduction in the initial error is proportional to the magnitude of the defocus amount, it will achieve the effect of eliminating the influence of the defocus amount on the error.
[0168] 3. Based on the defocus amount corresponding to each of the multiple training images, group the multiple training images into groups, and the defocus amounts corresponding to the multiple training images in the same group belong to the same defocus amount range.
[0169] The purpose of grouping will be explained first below:
[0170] Based on the above introduction, it can be confirmed that the PD model can output predicted phase difference and predicted confidence for the predicted phase difference. Furthermore, during the model training process, it is also necessary to measure whether the predicted phase difference and predicted confidence output by the PD model are accurate, and then adjust the PD model according to the measurement results so that the trained PD model can output relatively accurate predicted phase difference and predicted confidence.
[0171] In this embodiment, since the training images can be obtained in advance, the true phase difference can be pre-labeled on the training images. Then, the first error is obtained based on the predicted phase difference and the true phase difference. The first error is used to measure the accuracy of the predicted phase difference output by the PD model.
[0172] Furthermore, it is also necessary to measure the accuracy of the prediction confidence score output by the PD model for the predicted phase difference. The prediction confidence score indicates the reliability of the predicted phase difference, and the predicted phase difference itself is a result output by the PD model. Therefore, it is quite difficult to pre-label the true confidence score for the predicted phase difference. Thus, in this embodiment, the true confidence score of the predicted phase difference corresponding to the training image can be determined based on the first error corresponding to the training image.
[0173] As explained above, the first error reflects the accuracy of the predicted phase difference output by the PD model; specifically, the smaller the first error, the higher the accuracy of the predicted phase difference output by the PD model. Similarly, the true confidence score indicates the reliability of the predicted phase difference; specifically, the higher the reliability of the predicted phase difference, the higher the true confidence score. Since accuracy and reliability are positively correlated, the true confidence score of the predicted phase difference corresponding to the training image can be determined based on the first error corresponding to the training image.
[0174] In one implementation, for example, the first errors corresponding to multiple training images can be sorted. Then, the true confidence level of the predicted phase difference corresponding to the first-ranked training images is determined as M, and the true confidence level of the predicted phase difference corresponding to the last-ranked training images is determined as N. The specific values of M and N can be set according to actual needs. It can be understood that the true confidence level represented by M is greater than the true confidence level represented by N. That is, the higher the true confidence level of the predicted phase difference corresponding to the earlier-ranked training images, the higher the true confidence level of the predicted phase difference. In other words, the smaller the first error, the greater the true confidence level of the corresponding predicted phase difference.
[0175] However, based on the above explanation, it's understandable that the greater the defocusing amount of a training image, the greater the corresponding first error. Therefore, when sorting the first errors of multiple training images using this implementation method, a situation arises where the first errors corresponding to training images with larger defocusing amounts are always ranked lower, while the first errors corresponding to training images with smaller defocusing amounts are always ranked higher.
[0176] The resulting outcome is that the true confidence level of the predicted phase difference is lower for training images with larger defocus areas, and higher for training images with smaller defocus areas. If we then use this true confidence level to measure the accuracy of the predicted confidence level and adjust the model parameters accordingly, the final trained PD model will have lower confidence levels for the phase difference output from training images with larger defocus areas and higher confidence levels for the phase difference output from training images with smaller defocus areas during application.
[0177] However, while there may be a certain correlation between the amount of image defocus and the phase difference of an image, there is no clear correlation between the amount of image defocus and the confidence level of the image phase difference. In other words, the rule that "the greater the amount of image defocus, the lower the confidence level of the image phase difference; the smaller the amount of image defocus, the higher the confidence level of the image phase difference" is not valid.
[0178] For example, you can refer to Figure 7 To understand, such as Figure 7 As shown, suppose there are two images, image 1 and image 2. Assume the defocus amount of image 1 is 'a' and the defocus amount of image 2 is 'b', where 'a' is greater than 'b'. Based on this, assume the phase difference output by the PD model trained using the above process is 'c' for image 1 and 'd' for image 2, where 'c' is greater than 'd'. That is, the PD model can, for example, satisfy the following condition: the larger the defocus amount of the image, the larger the output phase difference.
[0179] Meanwhile, because image 1 has a large defocus, the confidence level of the PD model trained based on the above processing for the phase difference c of image 1 can be assumed to be... Figure 7 The figure is 10%. Also, because image 2 has a smaller defocus, the confidence level of the PD model trained based on the above processing for the phase difference d of image 2 can be assumed to be... Figure 7 90% shown.
[0180] However, the phase difference c output for image 1, which has a large defocus, might actually be accurate. The reason the PD model outputs a low confidence score of 10% is because image 1 has a large defocus. In other words, this 10% confidence score might be inaccurate. The reason the PD model outputs this inaccurate 10% confidence score is that during the model training phase, the defocus affects the ranking of the first error, thus causing the confidence score of the phase difference output by the trained PD model to be affected by the defocus, resulting in the confidence score not accurately reflecting the reliability of the phase difference.
[0181] To avoid the impact of defocus on error sorting, where training images with larger defocus are always sorted later and training images with smaller defocus are always sorted earlier, in this embodiment, all training images can first be divided into multiple groups or batches based on their defocus. Training images in the same group have the same or similar defocus.
[0182] Subsequently, error ranking is performed for each group. Since the defocus amount of the training images in each group is the same or similar, there will be no situation where the training image with a larger defocus amount corresponds to a smaller true confidence, or the training image with a smaller defocus amount corresponds to a larger true confidence. This can effectively eliminate the influence of defocus amount on error ranking.
[0183] The following is combined Figure 8 A brief explanation of the grouping of training images, such as... Figure 8 As shown, the figure contains eight training images and two preset defocus ranges. The eight training images are not necessarily taken in the same scene; this application does not limit the image content of the training images or the corresponding shooting scene. (See reference...) Figure 8 It can be determined that each of these 8 training images has its own defocus amount. The defocus amounts corresponding to different training images can be the same or different. The defocus amount of the training image is different from the shooting process of the training image. This embodiment does not limit the defocus amount of the training image.
[0184] and reference Figure 8 It can also be determined that, Figure 8 The two preset defocus ranges given are 290–310 and 420–430, respectively.
[0185] Next, these 8 training images can be grouped. First, the defocus amount of the 8 training images is compared with a preset defocus range to determine the defocus range to which each of the 8 training images belongs. Then, based on... Figure 8 The training images contained in the two defocus ranges shown are used to obtain the corresponding groupings for each of the two defocus ranges.
[0186] like Figure 8 As shown, the defocus amount of training image 601 is 300, the defocus amount of training image 602 is 423, the defocus amount of training image 603 is 305, the defocus amount of training image 604 is 300, the defocus amount of training image 605 is 295, the defocus amount of training image 606 is 308, the defocus amount of training image 607 is 428, and the defocus amount of training image 608 is 425.
[0187] Among them, the defocus amounts of training images 601, 603, 604, 605, and 606 fall within the defocus range of 290–310. Therefore, training images 601, 603, 604, 605, and 606 belong to the defocus range of 290–310, and accordingly, we can obtain... Figure 8 The group 609 shown may include training images 601, 603, 604, 605 and 606.
[0188] Furthermore, the defocus amounts of training images 602, 607, and 608 fall within the defocus range of 420–430. Therefore, training images 602, 607, and 608 belong to this defocus range of 420–430, and accordingly, we can obtain… Figure 8 The grouping shown in Figure 610 may include training images 602, 607, and 608. It is understood that the grouping method for all training images is similar and will not be described further here.
[0189] 4. For any group among the multiple groups, sort the first errors of each of the multiple training images in the group.
[0190] Based on the above introduction, it can be determined that the first error corresponding to the training image can be obtained by considering the defocus amount of the training image, the true phase difference, and the predicted phase difference output by the PD model. Then, for example, the first errors can be sorted to obtain the true confidence level corresponding to the training image.
[0191] The first error indicates the degree of difference between the true phase difference and the predicted phase difference of the training image. The smaller the first error corresponding to the training image, the smaller the difference between the true phase difference and the predicted phase difference, meaning the more reliable the predicted phase difference output by the PD model. Similarly, the true confidence score also reflects the reliability of the predicted phase difference output by the PD model. Specifically, the higher the confidence score, the more reliable the predicted phase difference output by the PD model.
[0192] For example, the true confidence level corresponding to the training image with the smaller first error can be set to a higher value, and the true confidence level corresponding to the training image with the larger first error can be set to a lower value, so as to accurately reflect the reliability of the predicted phase difference output by the PD model.
[0193] In one implementation, sorting can be performed separately for each group, and the implementation for each group is similar. Therefore, the following description uses any one group as an example. For instance, the first errors of the training images in the group can be sorted in ascending order to measure the magnitude of the first errors of the training images and determine the true confidence level corresponding to the training images.
[0194] The following is combined Figure 9Briefly describe the sorting of training images.
[0195] As Figure 9 shown, the figure contains two groups and two sorting sequences. The two groups are 609 and 610 respectively, and the two sorting sequences are 700 and 701.
[0196] As described in Figure 7 Step 3, group 609 contains training images 601, 603, 604, 605 and 606, and group 610 contains training images 602, 607 and 608.
[0197] Referring to Figure 9 , assume that the first errors corresponding to the 5 training images contained in group 609 are: the first error of image 601 is a, the first error of image 603 is b, the first error of image 604 is c, the first error of image 605 is d, and the first error of image 606 is e. Among them, the magnitudes of the first errors satisfy c < a < e < b < d. Assume that the first errors corresponding to the 3 training images contained in group 610 are: the first error of image 602 is f, the first error of image 607 is g, and the first error of image 608 is h. Among them, the magnitudes of the first errors satisfy g < h < f.
[0198] After that, sort according to the first error corresponding to each training image in each group. The sorting order can be from small to large, that is, the smallest first error is in the first sorting position, and the largest first error is in the last sorting position.
[0199] As Figure 9 shown, sorting the first errors corresponding to the training images in group 609 gives sorting sequence 700. The first sorting position of sequence 700 is the smallest first error c in group 609, and the last sorting position is the largest first error d in group 609. In addition, the first error a of image 601 is in the second position of sequence 700, the first error e of image 606 is in the third position of sequence 700, and the first error b of image 603 is in the fourth position of sequence 700.
[0200] Similarly, sorting the first errors corresponding to the training images in group 610 gives sorting sequence 701. The first sorting position of sequence 701 is the smallest first error g in group 610, and the last sorting position is the largest first error f in group 610. In addition, the first error h of image 608 is in the second position of sequence 701. It can be understood that the sorting method for the training images in all groups is similar and will not be introduced here.
[0201] 5. Based on the sorting results, the true confidence level corresponding to the training image of the first type is determined as the first value, and the first error of the training image of the first type is located in the first position interval in the sorting results.
[0202] 6. Based on the ranking results, determine the true confidence level corresponding to the training image of the second type as the second value, and the first error of the training image of the second type is located in the second position interval in the ranking results.
[0203] Steps 5 and 6 will be explained together below.
[0204] It is understandable that, based on the above steps, the first error can be applied to sorting multiple groups separately to obtain the sorting results corresponding to each group. The processing methods for the sorting results corresponding to each group in step 5 are similar; therefore, the following description uses the sorting result corresponding to any one group as an example.
[0205] Based on the above introduction, it can be determined that the first error reflects the accuracy of the predicted phase difference output by the PD model, and the true confidence level is used to indicate the credibility of the predicted phase difference. Accuracy and credibility are positively correlated, so it can be understood that the first error and the true confidence level are also positively correlated.
[0206] In one implementation, for example, the true confidence score corresponding to the training images with higher rankings in the first error can be set to a larger value, while the true confidence score corresponding to the training images with lower rankings in the first error can be set to a smaller value. For example, a preset threshold can be set, and the specific ranking positions of the higher and lower rankings can be determined based on this preset threshold.
[0207] For example, the preset threshold can be a pre-defined value. Based on the preset threshold, the sorting result can be divided into two position intervals, namely the first position interval and the second position interval. For example, the first position interval can be the interval from the first sorting position in the sorting result to the sorting position where the preset threshold is located, and the second position interval can be the interval from the next sorting position to the last sorting position in the sorting result where the preset threshold is located.
[0208] For ease of description, the training images in the group can be understood as a first type of training image and a second type of training image based on two position intervals. The first type of training image can be, for example, a training image where the first error is located in a first position interval, and the second type of training image can be, for example, a training image where the first error is located in a second position interval.
[0209] Finally, the true confidence scores for the first type of training images and the second type of training images can be labeled respectively. The true confidence score can be a first value or a second value, where the first value can be a higher value and the second value can be a lower value. For example, the true confidence score for the first type of training images can be determined as the first value, and the true confidence score for the second type of training images can be determined as the second value.
[0210] The first value can be 1, and the second value can be 0. Alternatively, the first value can be 0.95, and the second value can be 0.05. In this embodiment, the specific values of the first and second values are not limited, as long as the first value is greater than the second value.
[0211] The following is combined Figure 10 A brief explanation of the process for generating true confidence scores from training images.
[0212] like Figure 10 As shown in the figure, there is a sorting sequence 800. Sorting sequence 800 illustrates the process of labeling the true confidence of all training images in a group.
[0213] After obtaining a sorted sequence of groups, the sorting position corresponding to a preset threshold p% can be determined. Assuming the sorted sequence contains 100 elements, and assuming the preset threshold p% is 30%, then the sorting position corresponding to the preset threshold 30% is the position of the 30th element. For example... Figure 10 As shown, assume that the threshold p% is located at the nth sorting position in the sorting sequence of 800.
[0214] Then, the sorted sequence 800 can be divided into a first position interval and a second position interval based on the sorting position of the threshold p%. For example... Figure 10 As shown, the first position interval is the interval from the first sorting position to the nth sorting position, and the second position interval is the interval from the (n+1)th sorting position to the last sorting position. Both n and p are integers greater than or equal to 1. If n or p does not meet the integer condition, the value of n or p can be modified by rounding up or down.
[0215] Finally, the true confidence scores for all training images in the divided sequence of 800 are determined. In this ordered sequence of 800, training images located in the first position interval can be understood as the first type of training images described above, and training images located in the second position interval can be understood as the second type of training images described above.
[0216] Therefore, refer to Figure 10It can be determined that the true confidence level of multiple training images (i.e., training images of the first type) with the first error located in the first position interval is 1, and the confidence level of multiple training images (i.e., training images of the second type) with the first error located in the second position interval is 0.
[0217] 7. Based on the true confidence of each training image in the group and the predicted confidence of each training image in the group, determine the second error of each training image in the group. The second error is used to indicate the difference between the true confidence and the predicted confidence.
[0218] It is understandable that the processing method for each training image in the group in step 7 is similar, so the following will take any training image as an example.
[0219] In this embodiment, the second error of the training image can be determined based on the true confidence and predicted confidence of the training image. (See reference...) Figure 6 After obtaining the true confidence level corresponding to the training image based on the first error, the second error of the training image can be determined according to the degree of difference between the true confidence level of the training image and the predicted confidence level output by the PD model. It can be understood that the second error is used to indicate the degree of difference between the true confidence level and the predicted confidence level corresponding to the training image.
[0220] For example, the difference between the true confidence and the predicted confidence can be determined as the second error. The smaller the difference between the true confidence and the predicted confidence corresponding to the training image, the smaller the value of the second error; that is, there is a positive correlation between the difference between the true phase difference and the predicted phase difference and the initial error. Simultaneously, a smaller error means that the predicted confidence output by the PD model is more accurate. In actual implementation, the method for determining the second error can be determined according to actual needs; this embodiment does not limit this.
[0221] The method for determining the second error is similar to the method for determining the first error described above. The specific implementation of determining the second error can also be extended with reference to the content of determining the first error described in the above embodiment. This embodiment does not limit the implementation method of determining the second error.
[0222] 8. Adjust the model parameters of the first model based on the first error and / or second error of each of the multiple training images in the group.
[0223] Based on the above description, it can be determined that after the PD model outputs the predicted phase difference of the training image and the corresponding prediction confidence, the first error corresponding to the training image can be calculated based on the predicted phase difference and the true phase difference. Furthermore, the second error corresponding to the training image can be calculated based on the prediction confidence and the true confidence. Subsequently, in this embodiment, the first error and / or the second error will be backpropagated through each layer of the model, and can be used to guide the PD model in adjusting its parameters.
[0224] In one implementation, the first error of the training image can be fed back to the PD model via backpropagation. The PD model then adjusts its parameters based on this first error, resulting in a more accurate predicted phase difference output. Figure 6 As shown, the first error corresponding to the training image returns 600 to the PD model, and the model parameters are optimized.
[0225] Furthermore, in another implementation, a second error from the training images can be fed back to the PD model via backpropagation. The PD model adjusts its parameters based on this second error, resulting in a more accurate prediction confidence output. For example... Figure 6 As shown, the second error corresponding to the training image returns PD model 600, and the model parameters are optimized.
[0226] In another implementation, the first and second errors of the training images can be returned to the PD model via backpropagation. The PD model adjusts its parameters based on the first and second errors to enable it to output more accurate predicted phase difference and prediction confidence simultaneously.
[0227] Based on the above analysis, it can be understood that this application proposes two methods to reduce the impact of the defocus amount of the image on the confidence level of the first model output. In one method, considering that a larger defocus amount of the image may lead to a larger error between the predicted phase difference and the true phase difference, the initial error of each training image is adjusted based on its respective defocus amount to reduce the impact of the defocus amount on the error, thereby reducing the impact of the defocus amount on the image confidence level.
[0228] Furthermore, a larger defocus amount can lead to a greater error between the predicted and actual phase differences of an image. This can result in images with larger defocus amounts consistently having their first error ranked lower, while images with smaller defocus amounts tend to have their first error ranked higher. This can ultimately lead to images with larger defocus amounts consistently having lower true confidence scores and images with smaller defocus amounts consistently having higher confidence scores. Therefore, to address this issue, another approach involves grouping training images according to a preset defocus range before ranking the first errors. This ensures that training images within the same group have similar or identical defocus amounts, reducing the impact of defocus amount on true confidence scores when ranking the first errors within the same group.
[0229] Based on the above introduction, the following sections will introduce the possible implementation methods of the internal structure of the PD model.
[0230] The following is combined Figure 11 This article provides a detailed introduction to a PD model. Figure 11 This is a schematic diagram of a PD model provided in an embodiment of this application. The structure of the PD model can be, for example:
[0231] like Figure 11 As shown, the PD model can contain two modules: a feature extraction module and a processing module.
[0232] Feature extraction aims to extract key information or patterns from raw image data and transform them into more useful and compact feature representations for subsequent processing. (Reference) Figure 11 The feature extraction module can extract image information such as feature information and image attributes from the input image, providing useful image information for subsequent processing modules. Feature information can include low-level and high-level feature information. For example, low-level feature information can include image features such as color, texture, edges, and shape, while high-level feature information can include the shape and pose of objects in the image, contextual information, and graph structure. Image attributes can include attributes such as image blur level and image resolution.
[0233] The processing module outputs the predicted phase difference and prediction confidence of the input image. (Reference) Figure 11 The processing module can calculate and input the predicted phase difference and prediction confidence of the image by analyzing the features of the input image.
[0234] During model training, the model can be trained using the above-described model training process. The model training process is summarized below, and includes the following steps:
[0235] Multiple training images are input into the PD model either as a set of images or as a single image.
[0236] After receiving the training images, the feature extraction module processes the training images, extracts the features of the images, and transmits them to the processing module of the PD model.
[0237] After the processing module obtains the features of the image, it analyzes and calculates the features of the training image and outputs the predicted phase difference of the training image and the prediction confidence corresponding to the predicted phase difference.
[0238] Then, a first error corresponding to the training image is obtained based on the predicted phase difference and the true phase difference corresponding to the training image, and a second error is obtained based on the predicted confidence and the true confidence of the training image. The true confidence of the training image can be determined by grouping the training images based on their defocus amount, and then sorting the first errors of the training images within each group in ascending order.
[0239] Finally, the first and second errors are backpropagated to the processing module of the PD model to guide the processing module of the PD model to adjust the parameters.
[0240] Based on the above analysis, it can be understood that in this implementation, the PD model is designed to contain only two core modules. This concise design not only makes the model structure more intuitive and easy to understand, but also greatly simplifies the model training process and reduces the training difficulty.
[0241] In another implementation, the structure of the PD model is the same as described above. Figure 11 The model structures introduced are different. The following will combine... Figure 12 Another PD model will be introduced in detail. Figure 12 This is a schematic diagram of another PD model provided in an embodiment of this application. The structure of the PD model can be, for example:
[0242] like Figure 12 As shown, the PD model can contain three modules, which can be the feature extraction module, the first module, and the second module.
[0243] like Figure 11 As described in the introduction, the function of the feature extraction module is to extract image features and attributes from the input image, transforming them into a more useful and compact feature representation for subsequent processing. (Reference) Figure 12 After obtaining the input image, the feature extraction module extracts the features of the input image and then sends the image features to the first and second modules of the PD model simultaneously.
[0244] The first module of the PD model can be the module for obtaining the predicted defocus amount corresponding to the input image. After obtaining the image features transmitted by the feature extraction module, the model analyzes and calculates based on the image features, and outputs the predicted defocus amount of the input image.
[0245] The second module of the PD model can be a module for obtaining the prediction confidence corresponding to the input image. After obtaining the image features transmitted by the feature extraction module, it analyzes and calculates based on the image features, and outputs the prediction confidence of the input image.
[0246] In one implementation, the first and second modules of the PD model have similar structures.
[0247] During model training, the training process of this model is similar to... Figure 11 The model training process may vary slightly. In one implementation, the first module of the PD model is pre-trained before training the second module. When training the second module, the first module is frozen, and its parameter data is not modified.
[0248] Based on the above introduction, it can be understood that the training of the first module of the PD model can be achieved by backpropagating the first error corresponding to the training image obtained based on the predicted phase difference and the true phase difference of the training image to the first module of the PD model, and guiding the first module of the PD model to adjust its parameters.
[0249] The training process for the second module of this model is described below. This process includes:
[0250] Multiple training images are input into the PD model either as a set of images or as a single image.
[0251] After receiving the training images, the feature extraction module processes the training images, extracts the features of the images, and transmits them to the first and second modules of the PD model.
[0252] After the first and second modules acquire the image features, the first module analyzes and calculates the features of the training image and outputs the predicted phase difference of the training image. Similarly, the second module analyzes and calculates the features of the training image and outputs the prediction confidence corresponding to the predicted phase difference from the first module.
[0253] Then, the true confidence level of the training images is determined by grouping the training images based on their defocus amount and sorting the first errors corresponding to the training images within each group in ascending order. The second error corresponding to the training images is then obtained based on the predicted confidence level and the true confidence level.
[0254] Finally, the second error is backpropagated to the second module of the PD model, guiding the second module of the PD model to adjust its parameters.
[0255] Based on the above analysis, we can understand that in this implementation, the first and second modules of the PD model are trained independently. Compared to a PD model with a single processing module, due to this independence, errors or biases in the first module during training are not directly passed to the second module, potentially resulting in more accurate prediction confidence corresponding to the predicted phase difference output by the PD model. Furthermore, while this implementation involves more complex model structure and training processes compared to a single processing module PD model, it also offers greater flexibility and potential performance improvements.
[0256] Based on the above analysis, it can be determined that the role of the PD model is to output the predicted phase difference and prediction confidence of the input image. The application process of the PD model will be explained below with specific examples.
[0257] The following is combined Figure 13 The application of the PD model will be introduced in detail. Figure 13 This is a schematic diagram illustrating the imaging application of the PD model provided in this application embodiment. The application process of the PD model may include, for example, the following steps:
[0258] 1. Obtain a preview image.
[0259] A preview image is an image displayed through the viewfinder or LCD screen before the actual shooting, used to preview and confirm the focus effect. Acquiring multiple preview images allows for analysis of images taken at different times, with different focus states, and from different angles within the same scene, providing more image reference information for the PD model. Figure 13 After the device responds to the user's operation and activates the camera, the camera 1300 captures multiple preview images of the subject and inputs these preview images into the PD model. These multiple preview images may be captured in the same scene but at different times, under different focus states, and from different angles.
[0260] 2. Input the preview image into the first model to obtain the first phase difference of the preview image and the first confidence level corresponding to the first phase difference.
[0261] Based on the above introduction, it can be determined that when an image passes through the first model, the first model can output the predicted phase difference and prediction confidence corresponding to the input image. For example, the first model can refer to the PD model described above, or the specific implementation of the first model can be arbitrarily chosen according to actual needs, as long as the first model can be used to achieve the functions described above. The following explanation uses the PD model as an example of the first model.
[0262] In this embodiment, the first model can be a PD model trained based on the embodiments described above. Therefore, the first model can output a relatively accurate first phase difference and a first confidence level corresponding to the first phase difference for the preview image.
[0263] The first phase difference can be the predicted phase difference of the input image output by the PD model. That is, when an image passes through the PD model, the PD model can output the corresponding first phase difference. Correspondingly, the first confidence level can be the predicted confidence level corresponding to the predicted phase difference of the input image output by the PD model. That is, when an image passes through the PD model, the PD model can output the first confidence level of the corresponding first phase difference. For example... Figure 13 As shown, after the preview image passes through the PD model, the PD model can output the first phase difference corresponding to the preview image and the first confidence level corresponding to the first phase difference.
[0264] 3. If the first confidence level is greater than the preset confidence level, focus processing is performed based on the first phase difference.
[0265] Based on the above introduction, it can be determined that the confidence level reflects the accuracy and reliability of the first phase difference output by the PD model. It can be understood that a higher confidence level corresponding to the first phase difference output by the PD model means higher reliability, while a lower confidence level means lower reliability. Therefore, the processing method of the focusing system can be determined based on the confidence level.
[0266] In one implementation, the preset confidence level can be a pre-set value. In this embodiment, if the first confidence level is greater than the preset confidence level, it indicates that the first phase difference is highly reliable, and therefore the first phase difference can be used to guide subsequent focusing processing. Focusing processing may refer to determining the current position of the focusing motor based on the phase difference of the image and guiding the focusing motor to adjust the lens. When focusing, the position of the focusing motor is close to the focal point.
[0267] refer to Figure 13 After the PD model outputs the first phase difference of the preview image and the corresponding first confidence level, it determines whether the first confidence level of the preview image is greater than a preset confidence level. If the first confidence level is greater than the preset confidence level, it means that the focus processing can be guided by the phase difference of the preview image, and the motor is guided to adjust the lens based on the first phase difference corresponding to the preview image. If the first confidence level is less than or equal to the preset confidence level, it means that the preview image is not suitable for guiding the focus processing based on phase difference information, and other more robust technical means are used to complete the focusing process of the preview image.
[0268] Based on the above analysis, it can be determined that the PD model uses the calculated phase difference to perform focusing. Simultaneously, the PD model also outputs a confidence level associated with this phase difference to assess its reliability, thereby guiding the focusing system on whether it should focus based on this phase difference. Accurate confidence level output ensures that effective phase difference information is fully utilized, avoiding unnecessary waste, thus improving the efficiency and stability of the focusing system and achieving a more efficient focusing process.
[0269] In summary, the technical solution of this application can reduce the impact of defocusing on confidence level, avoid wasting effective focusing information, improve the accuracy and reliability of confidence level, and thus improve the overall performance and efficiency of the focusing system.
[0270] It should be noted that the module names involved in the embodiments of this application can all be defined as other names, as long as they can achieve the function of each module, and no specific restrictions are placed on the module names.
[0271] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0272] The model training method and focus processing method of the embodiments of this application have been described above. The apparatus for performing the above methods provided in the embodiments of this application is described below. Those skilled in the art will understand that the methods and apparatus can be combined with and referenced in each other, and the related apparatus provided in the embodiments of this application can perform the steps in the above model training method and focus processing method.
[0273] The model training method and focusing processing method provided in this application can be applied to electronic devices with shooting functions. Electronic devices include terminal devices, and the specific device form of the terminal device can be referred to the above-mentioned descriptions, which will not be repeated here.
[0274] In one implementation, this application provides an electronic device. Figure 14 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application.
[0275] like Figure 14As shown, the electronic device 1400 includes: a processor 1401 and a memory 1402; the memory 1402 stores computer execution instructions; the processor 1401 executes the computer execution instructions stored in the memory 1402, causing the electronic device 1400 to perform the above-described method.
[0276] When the memory 1402 is set up independently, the electronic device also includes a bus 1403 for connecting the memory 1402 and the processor 1401.
[0277] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.
[0278] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0279] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0280] This application provides a computer program product, which includes a computer program that, when run, causes a computer to perform the above-described method.
[0281] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0282] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A model training method, characterized in that, include: For any one of the multiple training images, the training image is input into the first model to obtain the predicted phase difference of the training image output by the first model, and the prediction confidence of the predicted phase difference. Based on the true phase difference of each of the multiple training images and the predicted phase difference of each of the multiple training images, the first error of each of the multiple training images is determined; Based on the defocus amount corresponding to each of the multiple training images, the multiple training images are grouped together, and the defocus amounts corresponding to the multiple training images in the same group belong to the same defocus amount range. The model parameters of the first model are adjusted based on the first error and prediction confidence of the training images in the multiple groups.
2. The method according to claim 1, characterized in that, The step of adjusting the model parameters of the first model based on the first error and prediction confidence of each training image in multiple groups includes: For any one of the multiple groups, sort the first errors of each of the multiple training images in the group; Based on the sorting results and the prediction confidence of each training image in the group, the second error of each training image in the group is determined. The model parameters of the first model are adjusted based on the first error and / or second error of each of the multiple training images in the group.
3. The method according to claim 2, characterized in that, The step of determining the second error of each of the multiple training images in the group based on the sorting result and the prediction confidence of each of the multiple training images in the group includes: Based on the sorting results, determine the true confidence level of each of the multiple training images in the group; Based on the true confidence scores corresponding to each of the multiple training images in the group and the predicted confidence scores obtained for each of the multiple training images in the group, a second error is determined for each of the multiple training images in the group. The second error is used to indicate the difference between the true confidence score and the predicted confidence score.
4. The method according to claim 3, characterized in that, Determining the true confidence level of each training image in the group based on the sorting result includes: Based on the sorting results, the true confidence level corresponding to the training image of the first type is determined to be a first value, and the first error of the training image of the first type is located in a first position interval in the sorting results; Based on the ranking results, the true confidence level corresponding to the training image of the second type is determined to be the second value, and the first error of the training image of the second type is located in the second position interval in the ranking results; The first position interval precedes the second position interval.
5. The method according to claim 4, characterized in that, The first position interval is the interval from the first sorting position to the nth sorting position, and the second position interval is the interval from the (n+1)th sorting position to the last sorting position. The nth sorting position is located at position p% in the sorting result, and n and p are both integers greater than or equal to 1.
6. The method according to any one of claims 1-5, characterized in that, The step of determining the first error of each of the multiple training images based on their true phase difference and predicted phase difference includes: Based on the true phase difference and the predicted phase difference of each of the multiple training images, the initial error of each of the multiple training images is determined, and the initial error is proportional to the difference between the true phase difference and the predicted phase difference; Based on the defocus amount corresponding to each of the multiple training images, the initial error of each of the multiple training images is reduced to obtain the first error of each of the multiple training images, wherein the degree of reduction of the initial error is proportional to the defocus amount corresponding to each of the training images.
7. The method according to any one of claims 1-5, characterized in that, The step of grouping the multiple training images according to their respective defocus values includes: For any one of the multiple preset defocus ranges, acquire multiple training images whose defocus values belong to the defocus range. Multiple training images whose defocus amount falls within the specified defocus range are grouped together.
8. A focusing processing method, characterized in that, include: Get a preview image; The preview image is input into the first model to obtain the first phase difference of the preview image and the first confidence level corresponding to the first phase difference, wherein the first model is trained according to the method of any one of claims 1 to 7; If the first confidence level is greater than the preset confidence level, focusing is performed based on the first phase difference.
9. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 8.
10. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN112102386A
Defocusing amount acquisition method and device, electronic equipment and readable storage medium
CN117714859A