AVM calibration method by use of generative artificial intelligence
Generative AI-based AVM calibration addresses the challenge of maintaining high-quality vehicle surround view images by adjusting camera parameters using marker-based data processing, ensuring accurate calibration in diverse conditions.
Patent Information
- Application Number
- PCT/KR2024/011653
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2024-08-07
- Publication Date
- 2025-08-07
AI Technical Summary
Existing AVM systems in vehicles face challenges in maintaining high-quality images due to changes in camera installation height and orientation caused by factors like passenger load, tire pressure, and part replacements, which are difficult to calibrate accurately without specialized facilities and skilled personnel.
A method using generative AI to calibrate AVM systems by processing marker-based data through a policy network, value network, and control network to adjust camera parameters, enabling fully and semi-automatic calibration in various lighting conditions.
Enables reliable AVM calibration without dedicated facilities or skilled personnel, improving calibration accuracy and reducing user intervention, even in typical lighting environments.
Smart Images

Figure KR2024011653_07082025_PF_FP_ABST
Abstract
Description
AVM Calibration Method Using Generative AI
[0001] The present invention generally relates to techniques for calibrating camera settings for an automotive AVM system.
[0002] In particular, the present invention relates to an AVM calibration technology using generative AI that searches for camera parameters for removing AVM image distortion by performing marker-based data processing on camera images of a vehicle AVM system using a generative AI model composed of a policy network, a value network, and a control network.
[0003] Around View Monitor (AVM) systems are increasingly being introduced into vehicles. AVM systems capture images of the vehicle's surroundings using cameras mounted on the vehicle, synthesize these images, and display them on an in-vehicle monitor screen. These systems also offer the ability to change the viewpoint to help drivers effectively perceive the surroundings. AVMs are also known as Surround View Monitoring (SVM).
[0004] To achieve high-quality AVM images while changing the viewpoint, precise measurements of the position and orientation of the camera installed on the vehicle, as well as the camera lens' field of view and distortion coefficient, are required. Precise measurement of these parameters requires a dedicated calibration facility with uniform lighting, as well as skilled personnel and machinery to ensure the vehicle is always positioned in the same position. On the vehicle production line, these facilities perform tolerance corrections for each vehicle. This process is also conveniently referred to as "factory calibration."
[0005] However, when a vehicle is actually used, the AVM conditions differ from those at the time of vehicle production due to factors such as the number of passengers on board, the weight of cargo loaded on the vehicle, seasonal changes in tire pressure, and parts replacement due to vehicle repairs. Consequently, the camera installation height and orientation change, which in turn degrades the quality of the AVM composite image. In general repair shops or AVM installation centers, proper AVM calibration is often difficult due to marker shape and outline distortion caused by lighting conditions. Furthermore, there is a technical barrier to identifying and resolving issues as unskilled personnel perform the calibration process.
[0006] Accordingly, there is a need for a technology that can assist unskilled personnel (vehicle drivers, external engineers) in calibrating AVM systems in typical lighting environments.
[0007] Meanwhile, prior literature related to the present invention is as follows.
[0008] (1) Republic of Korea Publication Patent No. 10-2015-0101806, “Around-view monitoring system and method using automatic recognition of grid patterns”
[0009] (2) Republic of Korea Patent Publication No. 10-2017-0026743, “Automatic correction device and method based on simplified pattern for vehicle image alignment”
[0010] (3) Republic of Korea Publication Patent No. 10-2017-0138166, “Around View Monitoring System and Tolerance Correction Method Therefor”
[0011] (4) US published patent US 2014-0051986 A1 “Systems and Methods for Registration of Multiple Vision Systems”
[0012] It is an object of the present invention to provide a technique for calibrating camera settings for a vehicle AVM system in general.
[0013] In particular, the purpose of the present invention is to provide an AVM calibration technology using generative AI that searches for camera parameters that remove AVM image distortion by processing marker-based data from camera images of a vehicle AVM system using a generative AI model composed of a policy network, a value network, and a control network.
[0014] To achieve the above purpose, the present invention proposes a method for a computer device to perform AVM calibration using a generative AI model.
[0015] The AVM calibration method using generative AI according to the present invention may be configured to include a first step of acquiring a camera image for an AVM image in a calibration field space where markers are installed; a second step of initializing camera parameters (Z) for generating an AVM image; a third step of initially configuring a marker candidate group by extracting a plurality of marker candidate images through contour-based image analysis of the camera image for each marker; a fourth step of performing a fully automatic mode AVM calibration based on marker contours for the marker candidate group using a generative AI model; a fifth step of performing a semi-automatic mode AVM calibration for the marker candidate group using the generative AI model; and a sixth step of performing neural network learning for the value network (120, 220) and the control network (130) by utilizing KL divergence for the output of the policy network (110, 210) and the output of the value network (120, 220).
[0016] According to the present invention, there is an advantage in that an AVM system can be reliably calibrated even without a dedicated calibration facility and skilled personnel.
[0017] In addition, according to the present invention, even if the fully automatic mode calibration fails, there is an advantage in that the AVM calibration can be achieved with less user intervention than in the past by performing the semi-automatic mode AVM calibration using the outline information of the marker.
[0018] Figure 1 is a conceptual diagram of a typical marker-based AVM calibration.
[0019] Figure 2 is a flowchart of a typical marker-based AVM calibration process.
[0020] Figure 3 is a flowchart of the entire AVM calibration process using generative AI according to the present invention.
[0021] Figure 4 is a flowchart of a fully automatic mode AVM calibration process in the present invention.
[0022] Figure 5 is a flowchart of a semi-automatic mode AVM calibration process in the present invention.
[0023] Figure 6 is a configuration diagram of a generative AI model for AVM calibration in the present invention.
[0024] Figure 7 is a configuration diagram of a policy network, a value network, and a control network in the present invention.
[0025] Figure 8 is a conceptual diagram showing the relationship between fully automatic mode AVM calibration and semi-automatic mode AVM calibration.
[0026] FIG. 9 is an example of view mode selection and marker selection for semi-automatic mode AVM calibration in the present invention.
[0027] FIG. 10 is an example of view mode selection and marker selection for semi-automatic mode AVM calibration in the present invention.
[0028] Hereinafter, the present invention will be described in detail with reference to the drawings.
[0029] In explaining the present invention, detailed descriptions of parts that overlap with the prior art may be omitted.
[0030] Figure 1 is a conceptual diagram of a general marker-based AVM calibration, and Figure 2 is a flowchart of a general marker-based AVM calibration process.
[0031] A vehicle is brought into a calibration workspace, and markers (21 to 24) around the vehicle are captured by multiple unit cameras (11 to 14) mounted on the vehicle, and information for correcting distortion of these marker images is obtained. For example, a process of extracting marker images and converting them into a standard shape is performed using computer vision technology. After the camera parameters obtained during this marker conversion process are stored for each camera, when the AVM system must operate later, the stored camera parameters can be applied to the images obtained from each camera to obtain good AVM images.
[0032] In the present invention, markers (21 to 24) are marks used for object recognition and coordinate recognition in the field of computer vision, and can be implemented as QR codes or image patterns. As illustrated in Fig. 1, markers (21 to 24) are placed near each corner of the vehicle.
[0033] For AVM calibration, a computer device (e.g., an AVM system) sequentially selects unit cameras (11 to 14) (S10) and acquires marker images from unit-captured images (S20). The marker images included in the unit-captured images are each slightly distorted from a standard marker shape (e.g., a square). Accordingly, camera parameters that can offset the image distortion characteristics inherent in each unit camera (11 to 14) are obtained from the marker images and stored.
[0034] To this end, marker standard coordinates according to the standard marker shape are obtained (S30), and transformation properties that map a marker image to the standard coordinates are obtained (S40). To this end, actual coordinate values are obtained for several points (e.g., corners) of the marker image appearing in the unit shooting image. Since the coordinate values before and after transformation are known for several points, transformation properties (e.g., transformation matrix) that map the actual coordinate values of the marker image to the marker standard coordinates can be obtained. Obtaining transformation properties when coordinate values before and after transformation are given is widely known in the field of machine vision, so they are not described in detail. Since this transformation property (e.g., transformation matrix) is information that allows for removing image distortion characteristics inherent in the unit camera, camera parameters for the unit camera are obtained by reflecting this transformation property (S50) and stored (S60).
[0035] The above process is performed for multiple unit cameras (11 to 14) that constitute the AVM system.
[0036] Fig. 3 is a flowchart of the entire AVM calibration process using a generative AI according to the present invention. Fig. 4 is a flowchart of the fully automatic mode AVM calibration process according to the present invention, and Fig. 5 is a flowchart of the semi-automatic mode AVM calibration process according to the present invention.
[0037] In the present invention, AVM calibration is a process of obtaining camera parameters for converting the AVM camera image into an AVM image without distortion after placing the vehicle in a calibration field and photographing standard markers marked on the floor of the calibration field with an AVM camera. In the present invention, a computer device obtains camera parameters for removing AVM image distortion based on a generative AI model. At this time, the computer device may be an AVM system control device installed in the vehicle or an external computer (e.g., a laptop computer, a dedicated computer) carried by the user.
[0038] Below, we will examine in detail the process of performing AVM calibration using a generative AI model.
[0039] Step (S100, S200): First, a computer device (e.g., AVM system) acquires a camera image for AVM video while the vehicle is placed in a calibration space where a marker is installed.
[0040] In addition, the camera parameters (Z) for AVM image generation are initialized. The initial value (Z0) of the camera parameters (Z) can be set to a standard value that is pre-saved according to the vehicle specifications (e.g., model name, year, etc.). The camera parameters (Z) generally include extrinsic parameters and intrinsic parameters. Extrinsic parameters are information representing the camera attitude information, and include information on the location where the camera is installed (e.g., spatial coordinate values x, y, z) and information on the attitude where the camera is installed (e.g., yaw, pitch, roll). Intrinsic parameters are information representing the camera lens characteristics, and it is common to use values that are pre-saved in the AVM manufacturing process, but actual measurements may be used depending on the implementation example.
[0041] Step (S300): Next, for each marker, a plurality of marker candidate images are extracted through contour-based image analysis of the camera images to initially configure marker candidate groups. For example, in the case of Fig. 1, four marker candidate groups, each consisting of, for example, 100, 105, 98, and 75 marker candidate images, are configured for four markers (21 to 24) or four cameras (11 to 14) through image analysis of the camera images.
[0042] Step (S400): Next, a fully automatic AVM calibration based on marker outlines is performed on a group of marker candidates using a generative AI model. In this process, the group of marker candidates is repeatedly filtered by the generative AI model, thereby maintaining good marker candidate images while adjusting the camera parameters (Z) in a way that removes AVM image distortion.
[0043] Figure 6 is a configuration diagram of a generative AI model (100) for AVM calibration in the present invention.
[0044] Referring to FIG. 6, the generative AI model (100) for the present invention is configured to include a policy network (110), a value network (120), and a control network (130). The generative AI model (100) of the present invention can be understood by referring to the World Model proposed by David Ha and Jurgen Schmidhuber (reference: https: / worldmodels.github.io / ).
[0045] The policy network (110) first filters the marker candidate group based on individual marker candidate images by the current camera parameters (Zk) and the MLP neural network, and then uses the intermediate camera parameters (Z k+1 ) and the individual marker probability value (r1) is calculated.
[0046] The value network (120) is an intermediate camera parameter (Z k+1 ) and the CNN neural network, the marker candidate group is secondarily filtered based on the marker candidate images (hereinafter referred to as “common marker candidate images”) that overlap the common area of adjacent cameras, and the common marker probability value (r2) is calculated.
[0047] The control network (130) controls the intermediate camera parameters (Z k+1 ) and the individual marker probability value (r1) and the common marker probability value (r2) are input into the MDN-RNN neural network to automatically adjust the camera parameters (Z k+2) is produced. The MDN-RNN neural network improves the convergence speed of AVM calibration by adjusting the camera parameters (Z).
[0048] Below, the operation of the policy network (110), value network (120), and control network (130) is described in detail.
[0049] First, the operation of the Policy Network (110) is described with reference to Fig. 7 (a).
[0050] The policy network (110) is provided with the current camera parameters (Zk) and a marker candidate group and generates individual marker candidate images (O c1 ) is calculated by the MLP neural network ('individual marker probability value'), and then the first filtering is performed on the marker candidate group based on the individual marker probability value (r1). Then, the policy network (110) provides an intermediate camera parameter (Z) that provides the optimal fitting for the marker candidate images remaining in the marker candidate group after the first filtering. k+1 ) is produced.
[0051] Specifically, each marker candidate image (O) belonging to the marker candidate group c1 ) is converted to an image based on the current camera parameters (Zk), and the converted marker candidate image (O c1 ') Extract feature vectors (e.g. size, shape, etc.) for each.
[0052] Multiple marker candidate images (O c1 ) For each, the extracted feature vector is input into a multi-layer perceptron (MLP) neural network to generate the corresponding marker candidate image (O c1 ) is a marker image, the probability (r1) (0~1) is calculated. In this specification, for convenience, it is expressed as 'individual marker probability value (r1)'. The policy network (110) calculates a plurality of marker candidate images (O) belonging to a marker candidate group. c1) is calculated for the individual marker probability value (r1), and then filtering is performed on the marker candidate group based on the individual marker probability value (r1). That is, the marker candidate image (O) with the individual marker probability value (r1) lower than the threshold is filtered. c1 ) is excluded from the marker candidate group. In this specification, it is called 'first filtering' for convenience. The size of the marker candidate group is reduced by this first filtering (e.g., 70 marker candidate images).
[0053] And, for the marker candidate group reduced in size by the first filtering, the camera parameter (Z) that optimally fits the marker candidate images to the standard marker shape k+1 ) is obtained. As an example, the optimal fitting is obtained by obtaining marker candidate images (O) belonging to the marker candidate group. c1 ) is a process of finding camera parameters that represent the minimum error (e.g., Least Square Error) from the standard marker shape when converting. At this time, the camera parameters (Z) output by the policy network (110) k+1 ) for convenience, 'intermediate camera parameters (Z k+1 ) is called.
[0054] Next, the operation of the value network (120) is described with reference to Fig. 7 (b).
[0055] The value network (120) is an intermediate camera parameter (Z k+1 ) and a marker candidate group (i.e., a marker candidate group after the first filtering) and a common marker candidate image (O) of an adjacent camera. c1 , O c2 ) is calculated by the CNN neural network as the marker probability value (r2) ('common marker probability value'), and then secondary filtering is performed on the marker candidate group based on the common marker probability value (r2).
[0056] The value network (120) receives and uses marker candidate groups obtained after the first filtering in the policy network (110). For example, assume that four marker candidate groups each consisting of 100, 105, 98, and 75 marker candidate images are formed in step (S300), and that the sizes of these four marker candidate groups are reduced to 50, 61, 43, and 39, respectively, through the first filtering in the policy network (110). The value network (120) performs a process using the four marker candidate groups whose sizes are reduced in this way.
[0057] Typically, AVM systems use wide-angle cameras (e.g., fish-eye lenses), so adjacent cameras capture the same marker in a common area. Accordingly, the value network (120) generates two marker candidate images (O) that overlap in the common area of adjacent cameras in the marker candidate group. c1 , O c2 ) for the intermediate camera parameters (Z k+1 ) is applied to transform the top view, and a synthetic marker candidate image (O) is obtained by merging them. c1|c2 ) is input into a convolutional neural network (CNN) to generate a synthetic marker candidate image (O c1|c2 ) calculates the probability (r2) (0~1) that the marker image is. In this specification, for convenience, it is expressed as 'common marker probability value (r2)'.
[0058] The value network (120) is a combination of multiple marker candidate images (O) taken of the same marker for a marker candidate group. c1 , O c2 ) and calculate the common marker probability value (r2) for them, and then perform additional filtering on the marker candidate group based on this common marker probability value (r2). That is, the marker candidate image (O) with the common marker probability value (r2) lower than the threshold is filtered. c1 , O c2) is excluded from the marker candidate group. In this specification, for convenience, it is called 'secondary filtering'. This second filtering further reduces the size of the marker candidate group (e.g., 30 marker candidate images).
[0059] Next, the operation of the control network (130) is described with reference to Fig. 7 (c).
[0060] The control network (130) controls the intermediate camera parameters (Z k+1 ) and individual marker probability values (r1) and common marker probability values (r2) are provided and the auto-adjustment camera parameters (Z k+2 ) is produced.
[0061] Intermediate camera parameters (Z k+1 ) is the currently selected individual marker candidate image (O c1 ) is obtained from the individual marker probability value (r1). The individual marker candidate image (O) is currently selected. c1 ) is the probability value derived from the common marker candidate image (O) of the adjacent camera. c1 , O c2 ) is the probability value derived from.
[0062] These intermediate camera parameters (Z k+1 ) and the individual marker probability value (r1) and the common marker probability value (r2) are input into the MDN-RNN neural network to automatically adjust the camera parameters (Z k+2 ) is obtained.
[0063] A recurrent neural network (RNN) is a neural network that processes sequence data and has the ability to recognize and learn patterns inherent in temporally continuous data (time series data). The intermediate camera parameters (Z) k+1) and the individual marker probability value (r1) and the common marker probability value (r2) are input into the RNN neural network to produce the probability density function P(z). Then, the probability density function P(z) is input into the MDN neural network to automatically adjust the camera parameters (Z k+2 ) is produced.
[0064] Mixture Density Networks (MDNs) are neural networks that perform unsupervised learning and generally have the ability to perform data clustering. In particular, MDN neural networks have the characteristic of outputting parameters of a mixture of Gaussian distributions that effectively predict the latent vector of the next stage (rounding). Accordingly, the MDN-RNN (MDN + RNN) neural network is known to have the ability to find optimal parameter values while predicting variable values of the next stage (rounding). Therefore, in the present invention, by applying MDN-RNN to adjust the camera parameter (Z), the convergence speed of AVM calibration can be improved.
[0065] Next, we describe the geometry model update and AVM evaluation operations.
[0066] Automatic adjustment camera parameters (Z) obtained in the process preceding the AVM system k+2 ) is applied, which is called geometry model update. Automatic adjustment camera parameters (Z) are applied to the unit shooting images acquired by the cameras (11 to 14) mounted on the vehicle. k+2) is applied to synthesize AVM images, and a score is calculated to evaluate the generated AVM images according to preset criteria. For example, a score value for the AVM image result can be given based on how close the four marker images placed around the vehicle are restored to the standard shape. This AVM image evaluation result can be used to automatically adjust the camera parameters (Z k+2 ) corresponds to the evaluation of distortion removal performance.
[0067] The evaluation results of this AVM image can be used to determine whether to repeat steps (S410) to (S430) further. For example, if the evaluation results of the AVM image are sufficiently satisfactory, it may be determined that no further repetition is necessary. Alternatively, if the evaluation results of the AVM image do not improve or even worsen despite multiple repetitions, it may be determined that no further repetition is necessary.
[0068] In this way, it is preferable to configure steps (S410) to (S440) in the fully automatic mode AVM calibration step (S400) to be repeated a plurality of times. At this time, the number of repetitions may be preset or may be determined by (S440). Fig. 6 (b) shows the operation of the generative AI model (100) in the first round, and Fig. 6 (c) shows the operation of the generative AI model (100) in the second round. The result of the first round (i.e., the camera parameter (Z2), the marker candidate group after filtering) is used in the second round. Generally, the result of the nth round is used in the (n+1)th round. As the rounds are repeated, inappropriate marker candidate images are removed from the marker candidate group, and the camera parameter (Z) is adjusted in a direction that reduces marker image distortion.
[0069] In summary, the fully automatic mode AVM calibration process using a generative AI model includes a policy network (110) that performs camera parameter (Z) adjustment and primary filtering of a marker candidate group based on individual marker candidates by an MLP neural network (S410), a value network (120) that performs secondary filtering of a marker candidate group based on common marker candidates of adjacent cameras by a CNN neural network (S420), a control network (130) that additionally adjusts the camera parameter (Z) based on marker probability values (r1, r2) for the marker candidate group by an MDN-RNN neural network (S430), and applies the camera parameter (Z) as a result of the adjustment to update the geometry model and perform score evaluation on the AVM image (S440).
[0070] Step (S500): Next, it is determined whether additional calibration is necessary. If it is determined that additional calibration is unnecessary, the AVM calibration process according to the present invention is completed by a fully automatic mode AVM calibration (S300). If the generative AI model (100) has been sufficiently trained, the fully automatic mode AVM calibration (S300) may be sufficient. On the other hand, if it is determined that additional calibration is necessary, the following semi-automatic mode AVM calibration (S600) is performed.
[0071] For example, if the error value produced by the policy network (110) during the optimal fitting process is smaller than a preset threshold value, it can be determined that additional calibration is unnecessary in (S500).
[0072] Additionally, the user can input whether additional calibration is required through the software menu manipulation.
[0073] Step (S600): Next, semi-automatic mode AVM calibration using a generative AI model is performed.
[0074] Figure 8 is a conceptual diagram showing the relationship between the fully automatic mode AVM calibration and the semi-automatic mode AVM calibration. Referring to Figure 8, the semi-automatic mode AVM calibration (S600) uses the marker candidate group and the automatic adjustment camera parameter (Z) that are the results of the fully automatic mode AVM calibration (S400). k+2 ) is a process of performing additional AVM calibration through user intervention. Semi-automatic mode AVM calibration (S600) uses a policy network (210) and a value network (220).
[0075] The policy network (210) is provided with basic adjustment camera parameters (Zp) and a marker candidate group and performs third-order filtering on the marker candidate group. Then, the policy network (210) provides additional adjustment camera parameters (Zp) that provide optimal fitting for the marker candidate images remaining in the marker candidate group after the third-order filtering. p+1 ) is produced. The operation of the policy network (210) is as described above with reference to Fig. 7 (a). At this time, the initial value of the basic adjustment camera parameter (Zp) input to the policy network (210) is the automatic adjustment camera parameter (Z) which is the result of the fully automatic mode AVM calibration (S400). k+2 ) is preferably set to .
[0076] The value network (220) provides additional adjustment camera parameters (Z p+1 ) and a marker candidate group (i.e., a marker candidate group after the third filtering) are provided, and a fourth filtering is performed on the marker candidate group. The operation of the value network (220) is as described above with reference to FIG. 7 (b).
[0077] In the semi-automatic mode calibration (S600), the user sequentially selects multiple markers, and the third filtering of the policy network (210) and the fourth filtering of the value network (220) are performed on the markers selected by the user, and additional adjustments are made to the camera parameters (Z) (i.e., from Zp to Zp+1 (output) is performed.
[0078] In the semi-automatic mode AVM calibration (S600), the user's view mode selection and specific marker selection on the AVM calibration screen are first identified (S610, S620). FIG. 9 and FIG. 10 are exemplary diagrams of view mode selection and marker selection for the semi-automatic mode AVM calibration according to the present invention. FIG. 9 is an exemplary diagram in which the user sequentially selects markers in a state where the left / right view mode is selected, and FIG. 10 is an exemplary diagram in which the user sequentially selects markers in a state where the front / rear view mode is selected.
[0079] The left-right view mode displays images captured by the left camera (13) and right camera (14) installed in the vehicle on the screen. In Fig. 9, the Left-Front marker (21) and Left-Rear marker (23) captured by the left camera (13) are displayed on the left side of the AVM calibration screen, and the Right-Front marker (22) and Right-Rear marker (24) captured by the right camera (14) are displayed on the right side of the screen.
[0080] The front / rear view mode displays images captured by the front camera (11) and rear camera (12) installed in the vehicle on the screen. In Fig. 10, the upper part of the AVM calibration screen displays the Front-Left marker (21) and Front-Right marker (22) captured by the front camera (11), and the right side of the screen displays the Rear-Left marker (23) and Rear-Right marker (24) captured by the rear camera (12).
[0081] In this way, the user can select the left-right view mode and the front-back view mode. As in Fig. 9 or Fig. 10, the user selects one marker while a specific view mode is selected. Fig. 9 (b) shows an example where the user selects the Left-Front marker (21) while the Left-Front view mode is selected.
[0082] When a user selects a specific marker (e.g., Left-Front marker (21)), the policy network (210) calculates a marker candidate priority for the marker candidate group of the selected marker ('selected marker') (S630). To this end, the policy network (210) calculates a marker candidate image (O) for each marker candidate image belonging to the marker candidate group of the selected marker. c1 ) is converted into an image based on the basic adjustment camera parameters (Zp), and the converted marker candidate image (O c1 ') extracts feature vectors (e.g., size, shape, etc.) for each marker, and calculates individual marker probability values (r3) by the MLP neural network. Then, based on the individual marker probability values (r3), a third filtering is performed on the marker candidate group. Then, based on the individual marker probability values (r3), a marker candidate image (O) is generated for the marker candidate group after the third filtering. c1 ) to prioritize.
[0083] The AVM calibration screen displays the highest priority marker candidate image (O c1 ) are displayed on the screen. In Figures 9 and 10, they are highlighted with red lines.
[0084] Then, the policy network (210) additionally adjusts the camera parameters (Z) to optimally fit the marker candidate images to the standard marker shape for the marker candidate group reduced in size by the third filtering. p+1 ) is obtained.
[0085] Next, the value network (220) provides additional adjustment camera parameters (Z p+1) and a marker candidate group (i.e., a marker candidate group after the third filtering) are provided, and common marker candidate images (O) of adjacent cameras for the selected marker are provided. c1 , O c2 ) is calculated by the CNN neural network, and then the 4th filtering is performed on the marker candidate group of the selected marker based on the common marker probability value (r4) (S640).
[0086] Next, for the cameras (11 to 14) related to the selection markers in the AVM system, additional adjustment camera parameters (Z) obtained in the previous process p+1 ) is applied, which is called geometry model update. This additional adjustment camera parameter (Z p+1 ) is applied to synthesize an AVM image, and then a score is calculated to evaluate the generated AVM image according to preset criteria (S650). The evaluation result of this AVM image can be used to determine whether to further repeat the process of (S630) to (S640). Alternatively, the evaluation result of this AVM image can be presented on the AVM calibration image screen and used to assist the user's judgment.
[0087] The above process (S610 to s650) is performed for a series of markers as in FIGS. 9 and 10. The above process does not have to be performed for all view modes and all markers, and is optional.
[0088] Step (S700): The camera parameters (Z) acquired through the above process are stored in the AVM system. Preferably, the camera parameters (Z) are individually acquired for each of the multiple unit cameras (11 to 14) mounted on the vehicle and stored in the AVM system. In a situation where the AVM system must operate thereafter, a good AVM image can be obtained by applying the stored camera parameters (Z) to the unit-captured images acquired from each camera (11 to 14).
[0089] Step (S800): When AVM calibration data has been sufficiently accumulated, neural network learning for the generative AI model (100) is performed.
[0090] In the present invention, neural network learning of the generative AI model (100) can be performed using the evaluation scores (S440, S650) for the geometry model of the AVM system as reward values. In this case, neural network learning of the generative AI model (100) can be performed in a simulation environment.
[0091] As shown in Equation 1, the KL divergence for the outputs of the policy network (110, 210) and the value network (120, 220) is utilized for neural network training. KL divergence (Kullback-Leibler divergence) generally expresses how much the predicted value deviates from the distribution of the reference value, and is used as a measure of the degree of inefficiency of the predicted value with respect to the actual value. KL divergence is also called relative entropy.
[0092]
[0093] At this time, neural network learning of the value network (120, 220) can be performed using data that merges the output value of the CNN neural network and the simulation environment results.
[0094] In addition, neural network learning of the control network (130) can be performed by using the individual marker probability values (r1, r3) produced by the policy network (110, 210) and the common marker probability values (r2, r4) produced by the value network (120, 220) as reward values (r1, r2)(r3, r4).
[0095] The present invention applies the concept of generative AI in that, rather than providing training data for neural network learning through supervised learning, the network output is fed back as input, and the optimal search path is found through reinforcement learning. This approach allows the RNN neural network to learn and utilize this optimal search path. Rather than manually coding calibration strategies optimized for various environments, reinforcement learning is used to discover them.
[0096] Meanwhile, the present invention can be implemented in the form of computer-readable code on a non-volatile computer-readable storage medium. Various types of storage devices exist as such non-volatile storage media, including hard disks, SSDs, CD-ROMs, NAS, magnetic tape, web disks, and cloud disks. Furthermore, the present invention can be implemented in the form of a computer program stored on a medium, coupled with hardware, to execute specific procedures.
Claims
1. A method for performing AVM calibration using a generative AI model by a computer device, Step 1: Acquiring camera images for AVM imaging in a calibration field space where markers are installed; Step 2: Initializing camera parameters (Z) for AVM image generation; A third step of initially forming a marker candidate group by extracting multiple marker candidate images by contour-based image analysis of the camera image for each marker; and A fourth step of performing a fully automatic mode AVM calibration based on marker outlines for the marker candidate group using a generative AI model; It consists of, including, The fourth step above is, Step 41, in which the policy network (110) adjusts the camera parameters (Z) and performs the first filtering of the marker candidate group by a multilayer perceptron (MLP) neural network based on individual marker candidate images, Step 42, in which the value network (120) performs secondary filtering of the marker candidate group by a convolutional neural network (CNN) based on the marker candidate image (hereinafter referred to as the 'common marker candidate image') overlapping the common area of the adjacent cameras, A method for AVM calibration using generative AI, wherein the control network (130) repeatedly performs step 43 of further adjusting camera parameters (Z) based on marker probability values for marker candidate groups by an MDN-RNN neural network multiple times.
2. In claim 1, The above step 41 is, A step in which the above policy network (110) is provided with the current camera parameters (Zk) and a marker candidate group; The above policy network (110) is configured to identify each marker candidate image (O) belonging to the marker candidate group. c1 ) is a step of converting an image based on the current camera parameters (Zk); The above policy network (110) converts the above-mentioned marker candidate image (O c1 ') A step of extracting feature vectors for each; The above policy network (110) inputs the feature vector into the MLP neural network to generate the corresponding marker candidate image (O c1 ) is a marker image (r1) (hereinafter referred to as 'individual marker probability value (r1)'); A step in which the above policy network (110) performs a first filtering on the marker candidate group based on the individual marker probability value (r1); and The above policy network (110) is an intermediate stage camera parameter (Z) that optimally fits the marker candidate images to a standard marker shape for the marker candidate group after the first filtering. k+1 ) for producing; It consists of, including, The above step 42 is, The above value network (120) is the intermediate camera parameter (Z k+1 ) and a step of providing a group of marker candidates after the first filtering; The above value network (120) overlaps the marker candidate image (O) in the common area of the adjacent cameras in the marker candidate group. c1 , O c2 ) for the intermediate camera parameters (Z k+1 ) to convert; The above value network (120) is a marker candidate image (O) after the above conversion. c1 , O c2 ) obtained by merging the synthetic marker candidate image (O c1|c2 ) is input into the CNN neural network to obtain the synthetic marker candidate image (O c1|c2 ) is a marker image (r2) (hereinafter referred to as 'common marker probability value (r2)'); and A step in which the above value network (120) performs secondary filtering on the marker candidate group based on the common marker probability value (r2); It consists of, including, The above step 43 is, The above control network (130) is the intermediate camera parameter (Z k+1 ) and a step of providing the individual marker probability value (r1) and the common marker probability value (r2); The above control network (130) is the intermediate camera parameter (Z k+1 ) and a step of inputting the individual marker probability value (r1) and the common marker probability value (r2) into a recurrent neural network (RNN) to produce a probability density function P(z); and The above control network (130) inputs the probability density function P(z) into the mixed density neural network (MDN) to automatically adjust the camera parameters (Z k+2 ) for producing; AVM calibration method using generative AI comprising:
3. In claim 2, The fourth step above is, The above auto-adjustment camera parameters (Z) in the AVM system k+2 ) to update the geometry model and synthesize the AVM image, and then evaluate the AVM image according to preset criteria. An AVM calibration method using generative AI, further comprising a 45th step of determining whether to repeat steps 41 to 43 in response to the evaluation results of the AVM image.
4. In claim 2, A fifth step of performing semi-automatic mode AVM calibration for the above marker candidate group using a generative AI model; It consists of more than , The fifth step above is, Steps to identify the user's view mode selection and specific marker selection on the AVM calibration screen; The policy network (210) sets the initial value of the basic adjustment camera parameter (Zp) for the semi-automatic mode AVM calibration to the automatic adjustment camera parameter (Z k+2 ) and the steps to set it up, The above policy network (210) selects each marker candidate image (O) belonging to the marker candidate group of the selected marker (hereinafter referred to as 'selected marker'). c1 ) is converted into an image based on the basic adjustment camera parameters (Zp) and the corresponding marker candidate image (O) is converted into a marker candidate image by the MLP neural network. c1 ) is a marker image (r3) (hereinafter referred to as 'individual marker probability value (r3)'), and A step in which the above policy network (210) calculates a marker candidate priority based on the individual marker probability value (r3) and performs a third filtering on the marker candidate group of the selected marker, The above policy network (210) provides an additional adjustment camera parameter (Z) that provides optimal fitting for the marker candidate group after the third filtering. p+1 ) and the step of producing it, The value network (220) is the additional adjustment camera parameter (Z p+1 ) and a step of providing a marker candidate group after the above 3rd filtering, The above value network (220) selects a common marker candidate image (O) of an adjacent camera for the selected marker from the marker candidate group after the third filtering. c1 , O c2 ) for the additional adjustment camera parameters (Z p+1 ) and the step of converting by applying The above value network (220) is a marker candidate image (O) after the above transformation. c1 , O c2 ) obtained by merging the synthetic marker candidate image (O c1|c2 ) is input into the CNN neural network to obtain the synthetic marker candidate image (O c1|c2 ) is a marker image (r4) (hereinafter referred to as 'common marker probability value (r4)'), and An AVM calibration method using a generative AI configured by repeatedly performing a step of performing a fourth filtering on a marker candidate group of the selected marker based on the common marker probability value (r4) of the above value network (220) for a plurality of markers.
5. In claim 4, A sixth step of performing neural network learning for the value network (120, 220) and the control network (130) by utilizing the KL divergence for the output of the policy network (110, 210) and the output of the value network (120, 220); AVM calibration method using generative AI comprising:
Citation Information
Patent Citations
around view video providing method based on machine learning for vehicle
KR101939349B1
method of providing automatic calibratiion of SVM video processing based on marker homography transformation
KR101989369B1
Device, method, and vehicle for providing around view
KR1020150014311A
System and method for monitoring around view using the grid pattern automatic recognition
KR1020150101806A
Semiconductor device and method for manufacturing the same
KR1020200031587A