METHOD AND SYSTEM FOR DEPTH-Of-FIELD REGION DETECTION AND RECOGNITION FROM A SINGLE IMAGE USING ADAPTIVELY SAMPLED LEARNING REPRESENTATION
Patent Information
- Application Number
- KR1020240044518
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2044-04-02
Smart Images

Figure 112024036459647-PAT00097_ABST
Abstract
Description
Technology Field
[0001] The following description concerns technology for recognizing the depth of field area. Background Technology
[0003] One of the most commonly addressed tasks is depth estimation from images, which involves analyzing geometric relationships within a three-dimensional (3D) scene or a scene implicitly present in an image. This analysis of relationships between objects and the environment improves object recognition accuracy and is applied in fields such as 3D modeling, physically based modeling, autonomous vehicles, video representation, and robotics. Several stereo image-based techniques have been proposed as tools for depth map estimation. However, applying stereo image-based techniques can cause blurring in the Depth-of-Field (DoF) region due to the difference between focusing and defocusing. Another type of blurring is motion blurring, which is caused by the velocity of the content in the video. DoF is one of the characteristics associated with camera focusing.
[0004] The depth of field (DoF) of an image is a critical consideration in processes such as content analysis, object detection, and Region of Interest (ROI) calculation using image or video data. Identifying DoF regions within an image is difficult because there is insufficient information to predict and approximate the focusing area in images containing DoF. To address this problem, many researchers have attempted to analyze geometric relationships by calculating the scene depth of images. However, most of these methods are applicable only to pose determination and object recognition, making them unsuitable for identifying specific ROIs contained within a single image. The problem to be solved
[0006] A training dataset can be constructed by extracting the depth of field region from a single image, and a neural network-based model for depth of field recognition can be trained using the constructed training dataset.
[0007] Using a trained model for depth of field recognition, the depth of field region can be recognized from a new image. means of solving the problem
[0009] A depth-of-field (DoF) region recognition method performed by a depth-of-field region recognition system may include the step of constructing a training dataset using depth-of-field (DoF) regions extracted from an image; and the step of training a model for depth-of-field region recognition using the constructed training dataset.
[0010] The step of constructing the above training data set includes the step of extracting a depth of field region from the image using a cross-correlation filter, and the cross-correlation filter may distinguish between features that appear blurry in the defocusing region and features that appear sharp in the focusing region from the image.
[0011] The step of constructing the above training data set includes the step of obtaining a DoF weight map from the extracted depth of field region based on the RGB color difference between the image and the Gaussian derivative image, and the image may be the original image.
[0012] The step of constructing the training data set may include the step of setting data pairs including the image and the extracted DoF weight map.
[0013] The above model for depth of field area recognition may be trained to minimize the difference between the required image and the actual DoF weight map.
[0014] The training step may include generating an input image including RGB color channels from the constructed training data set and a DoF weight map obtained from the extracted depth of field region, and dividing the generated input image into patches before inputting it into a model for depth of field region recognition.
[0015] The training step described above may include a step of calculating weight parameters for converting the DoF weight map of the input image into a Gaussian derivative image using the value obtained by multiplying the input image by the DoF weight map and the value obtained by multiplying the Gaussian derivative image by the DoF weight map.
[0016] The above training step includes a step of calculating the loss of a model for depth of field area recognition using a residual image derived through the difference between the above DoF weight map and the above input image, and the loss may be expressed through the Euclidean distance between the reconstructed map and the above DoF weight map.
[0017] The above training step may include the step of extracting only a portion of the depth of field region using a quadtree and using the extracted portion of the region as input data for a model for recognizing the depth of field region.
[0018] A depth-of-field (DoF) region recognition system may include a data set construction unit that constructs a training data set using depth-of-field (DoF) regions extracted from an image; and a model training unit that trains a model for depth-of-field region recognition using the constructed training data set. Effects of the invention
[0020] It can be used in various applications such as non-realistic rendering using areas of interest to the user, viewpoint tracking, object detection and recognition, optical character recognition, and adaptive sampling.
[0021] Focusing and defocusing areas, represented by the depth of field area, can be efficiently identified using a single image.
[0022] By using a model for recognizing the depth of field area, the depth of field area can be recognized more quickly and recognition accuracy can be improved. Brief explanation of the drawing
[0024] FIG. 1 is an example of a result of recognizing a depth of field area using a model for recognizing a depth of field area in one embodiment. FIG. 2 is an example of a case where the focal position of the DoF is different in one embodiment. FIG. 3 is an example of image detection and recognition results tested on an image with DoF applied in one embodiment. FIG. 4 is an example of the result applied to various applications in one embodiment. FIG. 5 is an example for explaining the focusing area and the defocusing area of a photograph in one embodiment. FIG. 6 is an example illustrating a weight map calculated using a DoF region in one embodiment. FIG. 7 is an example illustrating a DoF weight map obtained from an input image in one embodiment. FIG. 8 is an example illustrating the structure of a model for DoF region recognition in one embodiment. FIG. 9 is an example illustrating training results obtained using a model for DoF region recognition in one embodiment. FIG. 10 is an example illustrating that, in one embodiment, a quadtree is incorrectly partitioned due to empty space. FIG. 11 is an example of using a closed-type filter in one embodiment. FIG. 12 is an example for explaining a quadtree configuration in one embodiment. FIG. 13 is an example illustrating the operation of classifying into FD and ED based on the presence or absence of density in one embodiment. FIG. 14 is an example for illustrating the generation of a quadtree including path state values in one embodiment. FIG. 15 is an example of a quadtree node for collecting DoF patches in one embodiment. FIG. 16 is a block diagram illustrating a depth of field area recognition system in one embodiment. FIG. 17 is a flowchart illustrating a depth of field area recognition method in one embodiment. Specific details for implementing the invention
[0025] Hereinafter, embodiments will be described in detail with reference to the attached drawings.
[0027] FIG. 1 is an example of the result of recognizing the DoF using a model for depth of field area recognition in one embodiment.
[0028] Figure 1 visualizes the depth of field region extracted through the method proposed in the embodiment, and the embodiment describes a network and an application method for efficiently detecting and recognizing a blurry depth of field region in an image through camera focusing and defocusing.
[0029] In general, the DoF of an image affects object detection and recognition, rendering, ROI, and viewpoint tracking. While users can manually specify the area where the DoF applies in image processing to utilize it for image editing, the DoF of an image is one of the common characteristics representing the user's ROI. Even within the same scene, the DoF can be represented differently depending on where the focus is placed, which also affects the interpretation of the image.
[0030] FIG. 2 is an example of a case where the focal position of the DoF is different in one embodiment.
[0031] Referring to FIG. 2, the difference in images is shown depending on the change in the focused object. This may depend on the focus position of the DoF. In FIG. 2(a) and FIG. 2(b), although the scene is the same, the images can be interpreted differently due to the focus on the butterfly and the background, respectively. As such, the characteristics of the DoF can affect various application fields.
[0032] As mentioned, object interpretation varies depending on the focus position of the image, and image detection and recognition tasks are directly related to this problem.
[0033] FIG. 3 is an example of image detection and recognition results tested on an image with DoF applied in one embodiment.
[0034] Figure 3 shows the results of using YOLO for image detection and recognition in photos of four children. Figures 3(a) and 3(b) show that people were successfully detected and recognized in both images, although the focusing position and ROI are different.
[0035] In Fig. 3, all four children are recognized as people, and since there is almost no difference in accuracy, results in focus that differ from the user's intention. The most important feature in the data difference between Fig. 3(a) and Fig. 3(b) is the change in viewpoint. The viewpoint moves from left to right, but it is difficult to extract this movement when DoF information is unavailable.
[0036] It is not absolute that object 'A' will be recognized as object 'B' based on the DoF. However, various contents or different objects may exist within the same scene, and assuming that a single object is being focused by the DoF, the recognition result may vary depending on the focused object. Depending on the intensity of the DoF, it may or may not be possible to recognize it as a person. In the case of Fig. 3, most people were recognized as such, but the intensity of the DoF can lead to incorrect recognition results. The example in Fig. 3 demonstrates not merely the ability to recognize people, but the ability to decide which person to focus on. If people and animals are mixed in the scene, the DoF may recognize animals more clearly than people.
[0037] FIG. 4 is an example of the result applied to various applications in one embodiment.
[0038] If the DoF is not considered in Non-Realistic Rendering (NPR), a problem arises where object colors are severely distorted during the process of exaggeration into a cartoon style, and text recognition accuracy is degraded in Optical Character Recognition (OCR). As illustrated in Fig. 4(a), applying the NPR technique without considering the DoF results in excessive simplification of colors in a focused monkey face. Fig. 4(b), which presents the results of an OCR test, shows that text and image captions containing text are recognized. While the DoF can be considered in terms of reading and understanding text, existing approaches to applying the DoF assume that all images are sharp and do not distinguish between focusing and defocusing. In the embodiment, a method for efficiently detecting the depth of field region through a model (neural network) for DoF region recognition is proposed, and the efficiency and usefulness of the proposed method are demonstrated through experiments on the mentioned application.
[0039] In the embodiments, the proposed method will be described in more detail regarding the operations of 1) extracting DoF regions from images to build a training data set, and 2) designing an artificial neural network for training DoF regions.
[0040] A. Extracting DoF regions from images for the training dataset
[0041] We will now explain how to construct a dataset for the training phase. In general image super-resolution, high-resolution images are downsampled to generate low-resolution images, and training is performed using the loss between the two generated images. To extract the DoF region, a more specialized dataset is required, and the method for constructing such a dataset is as follows.
[0042] In the embodiment, a cross-correlation filter G calculates the DoF region of an image. The cross-correlation filter measures the level of association between two consecutive data points and is used in various fields such as image processing and computer vision. The DoF contains features that include the user's RoI, and the embodiment proposes an efficient method for identifying these features. Using the cross-correlation filter, features that appear blurry in the defocusing region and features that appear sharp in the focusing region are distinguished. To efficiently train the DoF region, Gaussian derivatives are used based on features where defocusing is close to 0 and focusing is close to 1.
[0043] DoF refers to the range of the foreground and background that is perceived as being in focus by a camera or a user. This phenomenon is also reflected in the detection or recognition of objects of interest to the user. As shown in FIG. 2, the same scene may be perceived as a butterfly or as the background depending on the focused object. In the embodiment, it was experimentally discovered that the characteristics of DoF are consistently expressed in space, and a cross-correlation filter was used to efficiently identify this. The goal is to efficiently detect the DoF region by identifying the changes occurring between the original image and the filtered image, and training a neural network based on these identified changes. In the embodiment, a cross-correlation filter was used to calculate the spatially consistent sharpness / blur characteristics of the DoF. The filter is defined as follows (see Equation 1).
[0044]
[0045] Here, H is referred to as a filter, kernel, or mask, and is the weight of each adjacent pixel, and F is the color of the adjacent pixel. The mask is modeled in various forms depending on the field of application. In the embodiment, a Gaussian type filtering technique may be used (see Equation 2).
[0046]
[0047] Here, ε is the variance. In the example, the following assumptions were made before estimating the DoF region of the image anisotropically.
[0048] 1) In the DoF region, the colors around the focused area gradually fade (see Fig. 5(a)).
[0049] 2) Areas blurred by defocusing and sharp areas are distinguished (see Fig. 5(b)).
[0050] Using the two features mentioned, the original image and Gaussian derivative image Anisotropic DoF weight map based on RGB color difference between Calculate (refer to Equation 3). Final output Is To account for potential noise that may occur when using only, the DoF weight map is obtained by refining it anisotropically.
[0051] To minimize noise in the calculated image using a cross-correlation filter and extract features of the focusing map, as follows Using Calculate.
[0052]
[0053] Here, the constant K controls the sensitivity to edges, and in the example, it is set to 2.
[0054] Equation 3 is the anisotropic DoF filter proposed in the embodiment, where , , represents the divergence, gradient, and Laplacian operators, respectively. In Equation 4, c is the diffusion coefficient, and the cross-correlation filter is calculated using mathematical formula 5.
[0055]
[0056] Here, and m represent the mask and the mask size, respectively, and were set to 15 in the example. Additionally, O represents the color difference between the blurry image and the sharp image. Also, and These are RGB colors obtained from the original image and the blurred image, and Gaussian smoothing is used as the smoothing filter.
[0057] In Equation 3, c is an edge-stopping function treated as an edge that reduces or stops diffusion when the gradient value is high. That is, the output of c(x, y, t) is When α is infinite, it becomes 0, so the diffusion rate decreases, and When is 0, it becomes 1, and the diffusion rate increases. Generally, in this process, t becomes a Gaussian kernel. As presented in the mathematical formula, the weights of the DoF region are calculated using the norm of the 3D vector, which is a representation of the RGB channel values. Referring to Fig. 6, this is an example showing a weight map calculated using the DoF region. It can be seen that the blurring effect becomes stronger as it moves from blue to red. In Fig. 6, through the above mathematical formula and Result images corresponding to this were obtained, and the weights of the defocused areas in the input images are well represented.
[0058] Anisotropic diffusion filtering is based on scale-space theory and is established and used by applying scale-space filtering such as D(x, y, t) = D0(x, y) Х G(x, y, t). When t > 0, the output image is represented as a blurry version of the input image, and as t increases, it is represented more strongly. c(x, y, t) controls the diffusion rate and is generally chosen as a function of the image gradient to preserve the edges of the image. In 1990, Pietro Perona and Jitendra Malik presented the idea of anisotropic diffusion and proposed two functions for the diffusion coefficient, one of which is Equation 4 used in the example.
[0059] Figure 7 is a DoF weight map obtained from the input image. Represents. DoF weight map It is calculated using DoF (white: focusing, black: defocusing). Referring to Figure 7, the weights of the focused area are well represented using DoF, and a dataset for network training is constructed.
[0060] B. DoF Map Training Using Convolutional Neural Networks
[0061] RGB color channels using the method described above and DoF weight map images An input image containing is generated. Before being input into the training network, each image is divided into patches. After the training data is prepared, the predicted value and ground truth Mapping function to minimize loss between The goal is to decide.
[0062] The goal of this equation is to find a function (model) f that best approximates the input image x to the DoF weighted image. The required image or approximated DoF weighted map image is and the goal is the required image and the measured DoF weight map image. The goal is to minimize the difference between them. In particular, the Mean Squared Error (MSE) loss function is used to minimize this difference. The objective function of this process is the mean squared error between the predicted image and the actual image. The goal is... The goal is to train a model f that predicts the value of and minimizes the mean squared error for the training data L (see Equation 7).
[0063]
[0064] In the examples, a super-resolution CNN-based method is used to implement the loss function L in a simple way. To apply the proposed method to the super-resolution CNN method, the DoF weight map of the input image is a Gaussian derivative image weight parameters to convert to must be calculated, and here, class are the input image and its DoF weight map, respectively, and ( : , : ) and,
[0065] and are the Gaussian derivative input images, respectively. and its DoF weight map( : , : ). also, is the target weight parameter for this step. The reason for choosing the Gaussian derivative is that, as mentioned earlier, focusing and defocusing within the image are identified as sharp and blurry forms, respectively, and the difference between them is amplified when differentiated. In the above equation, Step 1 is the input image This is the loss term required to convert to, and step 2 is the loss term that determines the weight parameters when applying the Gaussian derivative filter.
[0066] After training based on this method, the expected results were not obtained, and the generated images did not converge even when the number of training iterations was increased. The training result for 15,000 epochs was an image that converged to a grayscale image rather than a DoF weight map (see Fig. 9).
[0067] The reason this problem occurs is that after training the mapping function f(x) This is because the difference between and D is too large to obtain results that converge in the intended direction. This problem occurred in additional experiments using various images. As a solution to the problem seen in the examples, the input image is multiplied by a DoF weight map, and the resulting product is used as the network training data. , used.
[0068] In the example, a residual connection-based deep neural network method is used to improve the algorithm. Since the goal is to predict the residual map, the final loss function is calculated as follows (see Equation 8).
[0069]
[0070] Here, is the residual image ( Represents ), and x is the input image It represents. In the network training process, residual estimation, , The loss layer is calculated using the three elements of. The loss is the map reconstructed through the network and It is expressed as the Euclidean distance between them, and the reconstructed map here is the sum of the network's inputs and outputs.
[0071] This network was modeled based on a CNN, and its configuration is as follows (see Fig. 8). Residual compensation was performed by adding the feature map from the first CNN operation to the results of two convolution operations. Errors caused by convolution operations are mitigated through residual compensation. This process is repeated 10 times. Therefore, 20 convolution operations were performed. In the first cycle, only the value of the first convolution result is added once, and in subsequent cycles, the previous result value is added repeatedly. Next, the size is doubled through upscaling, and in the final step, four convolution operations are performed.
[0072] As shown in Figure 9, even with 15,000 iterations, only grayscale images far from the DoF were generated, whereas the proposed method generated an approximate DoF contour with a much smaller number of 1,400 iterations. This result indicates that the proposed method has a fast convergence speed and can extract the DoF region. Furthermore, while thousands or tens of thousands of images are typically used during the training process, the proposed method achieves an excellent learning rate with only a small dataset of 425 images due to a preprocessing method to improve convergence speed.
[0073] Additionally, we describe solver extensions used to efficiently improve the DoF extraction algorithm discussed above through adaptive sampling. Furthermore, we explain an approach that uses Quadtrees (Qt) to extract only meaningful regions from the DoF area and use them as training data. The focused region occupies a small area relative to the entire image. These small regions are subdivided into patches and used as data pairs for image-DoF weight maps to be used at the network level. The resulting reduction in the data area required for training reduces training time and memory.
[0074] The Quadtree (Qt) approach is a tree data structure in which each internal node has four child nodes. It is an algorithm that adaptively partitions a 2D rectangular space by dividing it into four quadrants and recursively subdividing it according to given criteria. Data related to leaf cells varies by application field but generally contains the "minimum unit of information of interest." In Octree, an extended version of the Quadtree (Qt) is applied to 3D space. Each internal node is divided into eight quadrants, and the cube-shaped space is recursively subdivided.
[0075] We will now explain closed-loop filtering for adaptive sampling. Before inputting data into an artificial neural network, the previously calculated A quadtree (Qt)-based DoF patch is constructed using [the filter]. The sole purpose of performing this process is to find the DoF region. If the filter F proposed in the example is not applied, the quadtree (Qt) may not be properly generated as shown in Fig. 10. Fig. 10 is This is the result of Quadtree (Qt) segmentation. In the case of monochromatic objects, classification based on focusing and defocusing is difficult. The original image shows a woman wearing brightly colored clothing and a man wearing a solid-colored T-shirt. Even with the use of a focusing filter, extracting the DoF monochromatic region is not easy (refer to the red area in Fig. 12). This pattern is It is also found in other areas and is more severe in adaptive approaches.
[0076] In Fig. 12, since the male's upper body is monochromatic, features such as color scattering or blurring are hardly detected even during defocusing. This area remains an empty region that is not subdivided by the quadtree (Qt). Therefore, since the quadtree (Qt) cannot correctly represent the DoF region, it leads to a network-level learning failure problem in subsequent processes. In the embodiment, a new filter F is proposed to mitigate this problem (see Equation 9).
[0077]
[0078] Here, , , represents the binarization, sharpening, and gamma correction filters for the image, respectively.
[0079] first, ( Gamma correction is applied. This process is performed to amplify the difference between the focusing area and the defocusing area (see Equation 10).
[0080]
[0081] Here, is an input image, and in the embodiment ... In the above equation, g represents a variable that enhances the contrast between sharp and blurry areas by applying gamma correction. Additionally, M signifies the maximum value, which was set to 255 in the example. When g is 1, brightness changes linearly, and in the example, it was set to 0.85. Although the colors obtained through this process become sharper, there is a problem of blurring at boundaries or edges. To mitigate this problem, an image sharpening approach using a Laplacian kernel, Apply (see Equation 11).
[0082]
[0083] Here, It can be represented as a mask, and in the example, a Gaussian distribution mask was applied. Finally, image binarization Using a closed filter through ( The mask image (DoF weight map) is extracted (see Fig. 11). To solve this problem A closed mask map was extracted to identify the focusing and defocusing regions by applying a filter. The spatial partitioning method using this result and the method for collecting sparse datasets will be described below.
[0084] We will now describe the collection of sparse datasets from quadtree-based DOF patches. In the examples, A quadtree (Qt) is calculated using [the method], and as shown in Fig. 11(a), the quadtree (Qt) is partitioned based on the DoF region, which is the white part of the image. To partition the quadtree (Qt) The reason for using it is not to obtain pixel information, but to extract meaningful information from space. Therefore, the training of the network process is an image patch representing the leaf nodes of a quadtree (Qt). It is performed using .
[0085] Using the method described above, the lowest nodes to be used to construct the quadtree (Qt) are generated, and the nodes are merged bottom-up to form the tree (see Fig. 12). Before combining the generated nodes into the quadtree (Qt), it is checked whether there is density (e.g., DoF weight values), and if there is density, it is compared with a threshold value (see Fig. 13) to classify the nodes into Full Density (FD) and Empty Density (ED) states. Fig. 13(a) is FIG. 13(b) is an example of patch partitioning, and FIG. 13(c) is an example of classifying FDs and EDs. The lowest nodes have specified state values, and the state value of the parent node is determined by the state value of the child nodes (see FIG. 14(a)).
[0086] Each node in the tree has data, key, and state values. The data represents the node's density value, and the key consists of x and y coordinates representing the node position and tree depth used when constructing the tree. When combining results after the network process is complete, the tree depth and position are used, and the state is described as FD, ED, or mixed.
[0087] In the example, the density of each patch is used as the data for the lowest node (see FIG. 14(a)). The depth of the lowest node is calculated using Equation 12.
[0088]
[0089] Here, d represents the depth of the current node, and and represents the total input data and the width of the current node, respectively. The parent node is created in a bottom-up manner with four nodes (see Fig. 14(b)). For the parent node, the data is the sum of the data of the child nodes. The depth decreases by 1, and the position is determined by combining the child nodes. Finally, the state value is determined by the child nodes. If the state values of all child nodes are the same, the state value of the child nodes is assigned to the parent node, and if there are FD and ED, Mix (MIX) is assigned to the state value of the parent node.
[0090] If the state values of all child nodes are identical, they are deleted. Network training is not required if the state value is ED. If it is FD, implementing a single network training is faster than for each child node. Repeating this process until the root node is reached generates a tree with state values assigned to all nodes. After the tree is completed, data and key values for all FD nodes are collected (see Fig. 15), and this dataset is used for network training. In Fig. 15, the red dashed line represents the final dataset, It is a patch.
[0091] Network training is performed using a dataset collected from a Quadtree (Qt). As previously mentioned, the values are intended solely to obtain spatial information, and D patches, which are image data corresponding to the locations of these patches, are used for training. Since the number of datasets used in this approach is smaller than the original data, the training process is efficiently optimized. Building a Quadtree (Qt) requires 1 second per image, and it took 18 hours to build the Quadtree (Qt) when using 425 data points for training. However, without using Qt, the training process took 99 hours.
[0092] FIG. 16 is a block diagram illustrating a depth of field area recognition system in one embodiment, and FIG. 17 is a flowchart illustrating a depth of field area recognition method in one embodiment.
[0093] The processor of the depth of field region recognition system (100) may include a data set building unit (1610) and a model training unit (1620). These components of the processor may be representations of different functions performed by the processor according to control commands provided by program code stored in the depth of field region recognition system. The processor and the components of the processor may control the depth of field region recognition system that performs steps (1710 to 1720) included in the depth of field region recognition method of FIG. 17. At this time, the processor and the components of the processor may be implemented to execute instructions according to the code of an operating system included in memory and the code of at least one program.
[0094] The processor can load program code stored in a file of a program for a depth of field area recognition method into memory. For example, when a program is executed in a depth of field area recognition system, the processor can control the depth of field area recognition system to load program code from a file of a program into memory under the control of an operating system. At this time, the processor may be different functional representations of the processor for executing subsequent steps (1710 to 1720) by executing commands of corresponding parts of the program code loaded into memory in each of the data set building unit (1610) and the model training unit (1620).
[0095] In step (1710), the dataset building unit (1610) can build a training dataset using the depth-of-field (DoF) region extracted from the image. The dataset building unit (1610) can extract the depth-of-field region from the image using a cross-correlation filter. The dataset building unit (1610) can obtain a DoF weight map from the extracted depth-of-field region based on the RGB color difference between the image and the Gaussian derivative image. The dataset building unit (1610) can set a data pair containing the image and the extracted DoF weight map.
[0096] In step (1720), the model training unit (1620) can train a model for depth of field region recognition using the constructed training data set. The model training unit (1620) can generate an input image including RGB color channels and a DoF weight map obtained from the extracted depth of field region from the constructed training data set, and can divide the generated input image into patches before inputting it into the model for depth of field region recognition. The model training unit (1620) can calculate weight parameters for converting the DoF weight map of the input image into a Gaussian derivative image using the value obtained by multiplying the input image by the DoF weight map and the value obtained by multiplying the Gaussian derivative image by the DoF weight map. The model training unit (1620) can calculate the loss of the model for depth of field region recognition using the residual image derived from the difference between the DoF weight map and the input image. The model training unit (1620) can use a quad tree to extract only a portion of the depth of field area and use the extracted portion as input data for the model for depth of field area recognition.
[0097] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0098] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0099] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0100] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0101] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 A depth-of-field (DoF) region recognition method performed by a depth-of-field region recognition system, comprising: a step of constructing a training dataset using a depth-of-field (DoF) region extracted from an image; and a step of training a model for depth-of-field region recognition using the constructed training dataset, wherein the step of constructing the training dataset includes a step of extracting a depth-of-field region from the image using a cross-correlation filter, obtaining a DoF weight map from the extracted depth-of-field region based on the RGB color difference between the image and a Gaussian derivative image, and setting a data pair including the image and the extracted DoF weight map, wherein the cross-correlation filter distinguishes between features that appear blurry in the defocusing region and features that appear sharp in the focusing region from the image, and wherein the image is the original image. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 A depth-of-field (DoF) region recognition method according to claim 1, wherein the model for depth-of-field region recognition is trained such that the difference between the required image and the measured DoF weight map is minimized. Claim 6 A depth-of-field (DoF) region recognition method according to claim 1, wherein the training step comprises generating an input image including RGB color channels from the constructed training data set and a DoF weight map obtained from the extracted depth-of-field region, and dividing the generated input image into patches before inputting it into a model for depth-of-field region recognition. Claim 7 A depth-of-field (DoF) region recognition method performed by a depth-of-field region recognition system, comprising: a step of constructing a training dataset using a depth-of-field (DoF) region extracted from an image; and a step of training a model for depth-of-field region recognition using the constructed training dataset, wherein the training step comprises: generating an input image including RGB color channels and a DoF weight map obtained from the extracted depth-of-field region from the constructed training dataset; dividing the generated input image into patches before inputting it into the model for depth-of-field region recognition; and calculating weight parameters for converting the DoF weight map of the input image into a Gaussian derivative image using a value obtained by multiplying the input image by the DoF weight map and a value obtained by multiplying the Gaussian derivative image by the DoF weight map. Claim 8 A depth-of-field (DoF) region recognition method according to claim 7, wherein the training step comprises the step of calculating the loss of a model for depth-of-field region recognition using a residual image derived through the difference between the DoF weight map and the input image, and wherein the loss is expressed through the Euclidean distance between the reconstructed map and the DoF weight map. Claim 9 A depth-of-field (DoF) region recognition method according to claim 8, wherein the training step comprises the step of extracting only a portion of the depth-of-field region using a quadtree and using the extracted portion of the region as input data for a model for depth-of-field region recognition. Claim 10 A depth-of-field (DoF) region recognition system comprises: a data set construction unit that constructs a training data set using a depth-of-field (DoF) region extracted from an image; and a model training unit that trains a model for depth-of-field region recognition using the constructed training data set, wherein the data set construction unit extracts a depth-of-field region from the image using a cross-correlation filter, obtains a DoF weight map from the extracted depth-of-field region based on the RGB color difference between the image and a Gaussian derivative image, and sets a data pair including the image and the extracted DoF weight map, wherein the cross-correlation filter distinguishes between features that appear blurry in the defocusing region and features that appear sharp in the focusing region from the image, and wherein the image is the original image.