Method and mobile device for generating aligned image data by alignment parameters generated by image transform ai model

By using adversarial training techniques for image transformation AI models, alignment parameters are generated, solving the problem of inconsistent output of image data from cameras with different attributes on autonomous driving mobile devices, and achieving efficient and economical performance for visual tasks.

CN121010533APending Publication Date: 2025-11-25HYUNDAI MOTOR CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411949158.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2024-12-27
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing supervised learning-based deep learning models require a large amount of benchmark real image datasets, and cameras with different attributes have differences in color, scale, distortion, etc., resulting in inconsistent outputs on autonomous driving mobile devices.

Method used

By using an image transformation AI model and adversarial training techniques, alignment parameters are generated to adjust image data captured by cameras with different attributes, ensuring consistent output.

Benefits of technology

Even with a small amount of benchmark real image data, sufficient visual task performance is ensured, the economic cost of acquiring benchmark data is reduced, and consistent image transformation is achieved through the globally optimal solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010533A_ABST
    Figure CN121010533A_ABST
Patent Text Reader

Abstract

The invention relates to a method and a mobile device for generating aligned image data through alignment parameters generated by an image transformation artificial intelligence model. A method for generating aligned image data by alignment parameters generated by an image transformation AI model includes generating, by an encoder of the image transformation AI model, at least one or more alignment parameters according to first and second camera attributes associated with a first camera and a second camera, respectively. The method further includes transforming, by an image transformer of an image transformation AI model, first image data captured by the first camera to align with second image data captured by the second camera based on the at least one alignment parameter and the brightness parameter. The method further includes training an encoder and a discriminator of the image transformation AI model by adversarial training. The image transformation AI model discriminates the transformed first image data and second image data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the rights and priority of Korean Application No. 10-2024-0067699, filed on May 24, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to a method and mobile device for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence (AI) model. More specifically, this invention relates to a method and mobile device for generating aligned image data using alignment parameters generated by an image transformation AI model that similarly transforms image data captured by cameras with different attributes using adversarial training techniques. Background Technology

[0004] Supervised learning-based deep learning models utilizing benchmark real image data have been actively used to perform various visual tasks and have demonstrated high performance compared to other learning techniques.

[0005] However, in order to achieve sufficient performance, supervised learning-based deep learning models require multiple benchmark real image datasets, and the economic cost of obtaining a large amount of benchmark real image data increases accordingly.

[0006] Therefore, it is necessary to improve the efficiency of data in order to achieve sufficient performance by utilizing a small amount of benchmark real image data.

[0007] It's possible to utilize image data captured by various cameras with different attributes within a single model; however, cameras with different attributes can exhibit various differences in color, scale, distortion, etc. Therefore, a method is needed that can still provide consistent output despite these differences.

[0008] Because cameras installed in mobile devices can be in different locations and of different types, the aforementioned problems can also occur in autonomous mobile devices. The subject matter described in this background section is intended to facilitate an understanding of the background art of the invention and therefore may include topics unknown to those skilled in the art. Summary of the Invention

[0009] The present invention is technically committed to a method and mobile device for generating aligned image data by means of alignment parameters generated by an image transformation artificial intelligence (AI) model, wherein the image transformation AI model transforms image data captured by cameras with different attributes by utilizing adversarial training techniques.

[0010] The technical problems solved by this invention are not limited to those described above. Other technical problems not described herein will be clearly understood by those skilled in the art based on the following description.

[0011] This method can be performed by means of an apparatus for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence (AI) model. The method may include: generating at least one alignment parameter using an encoder of the image transformation AI model, based on first camera attributes and second camera attributes respectively associated with a first camera and a second camera. The method may also include: transforming first image data captured by the first camera to be aligned with second image data captured by the second camera using an image transformer of the image transformation AI model, based on at least one alignment parameter and a brightness parameter. The method may further include: training the encoder and discriminator of the image transformation AI model through adversarial training. The image transformation AI model distinguishes between the transformed first image data and the second image data.

[0012] The first camera attribute and the second camera attribute may include at least one of the internal parameters or distortion coefficients of each of the first camera and the second camera.

[0013] At least one alignment parameter may include at least one of a cropping parameter for removing a predetermined region or a projection matrix for projecting the first image data onto the second image data.

[0014] The transformation may include projecting first image data based on a projection matrix. The transformation may include removing a predetermined region by reflecting cropping parameters in the projected first image data. The transformation may include adjusting the brightness by applying brightness parameters to the first image data after the predetermined region removal. The transformation may include performing a resizing operation to match the size of the brightness-adjusted first image data to the size of second image data.

[0015] The transformation may include removing a predetermined region of the first image data based on cropping parameters. The transformation may include projecting the first image data onto the first image data after the predetermined region removal, reflecting a projection matrix. The transformation may include adjusting the brightness by applying a brightness parameter to the projected first image data. The transformation may include performing a resizing operation to match the size of the brightness-adjusted first image data to the size of the second image data.

[0016] The brightness parameter can be a learnable parameter based on the discriminator loss caused by the brightness parameter, and is independent of the discriminator loss caused by the encoder.

[0017] The method may also include subordinating the clipping parameters and projection matrix to the trained encoder and determining the clipping parameters and projection matrix through regression.

[0018] Adjusting the brightness can include multiplying the first element of the brightness parameter by the total pixel and adding the second element of the brightness parameter to the total pixel.

[0019] The method may further include having a discriminator determine whether the transformed first image data is real or fake, as captured by a second camera. The method may also include training the discriminator to determine fake data. Furthermore, the method may include learning an encoder and brightness parameters to enable the discriminator to determine real data.

[0020] The method may also include normalizing the input first camera attributes and second camera attributes, as well as the output clipping parameters and projection matrix, to a predetermined range by an encoder.

[0021] A mobile device may include a memory and a processor. The memory is configured to store at least one instruction, and the processor is configured to execute the at least one instruction stored in the memory based on data obtained from the memory. The processor is further configured to: generate at least one alignment parameter based on first camera attributes and second camera attributes, respectively, associated with a first camera and a second camera, using an encoder of an image transformation AI model. The processor is further configured to: transform first image data captured by a first camera to be aligned with second image data captured by a second camera, based on at least one alignment parameter and a brightness parameter, using an image transformer of the image transformation AI model. The encoder and discriminator of the image transformation AI model are trained through adversarial training. The image transformation AI model distinguishes between the transformed first image data and the second image data.

[0022] The features of the invention briefly outlined herein are merely exemplary aspects of the invention and the subsequent detailed description thereof, and are not intended to limit the scope of the invention.

[0023] The technical problems solved by this invention are not limited to those described above. Through the following description, those skilled in the art will more clearly understand other technical problems solved by this invention that are not described herein.

[0024] According to the present invention, a method and a mobile device for generating aligned image data by means of alignment parameters generated by an image transformation AI model are provided, wherein the image transformation AI model similarly transforms image data captured by cameras with different attributes using adversarial training techniques.

[0025] Furthermore, according to the present invention, alignment parameters capable of aligning images can be generated by taking into account different features and brightness parameters of the image.

[0026] Furthermore, according to the present invention, even when performing visual tasks using image data captured by different cameras, sufficient inference performance can be ensured by aligning the image data.

[0027] Furthermore, according to the present invention, by matching the data distribution of image data captured by cameras with different geometric properties, the economic cost of obtaining benchmark real data for training a deep learning model suitable for each camera can be reduced.

[0028] Furthermore, according to the present invention, by matching the data distribution of image data, consistent inference performance based on image data captured by multiple cameras with different geometric properties can be ensured even when using a relatively simple single deep learning model.

[0029] Furthermore, according to the present invention, by utilizing an image transformation AI model that automatically manipulates images, multiple images can be consistently transformed into the optimal result based on a globally optimal solution.

[0030] The effects achievable according to the present invention are not limited to those described above, and other effects not mentioned herein will be clearly understood by those skilled in the art through the following description. Attached Figure Description

[0031] Figure 1 An example of a mobile device is shown that communicates with another device to send and receive data.

[0032] Figure 2 An example of the constituent modules of a mobile device according to the present invention is shown.

[0033] Figure 3 An example of the constituent modules of a server according to the present invention is shown.

[0034] Figure 4 An example of the functional configuration included in the image transformation artificial intelligence (AI) model according to the present invention is shown.

[0035] Figure 5 This is a flowchart illustrating the process of generating aligned image data according to the present invention.

[0036] Figure 6 This is a flowchart illustrating the detailed process of transforming image data to generate aligned image data.

[0037] Figure 7 This is a flowchart illustrating the process of training an image transformation AI model according to the present invention.

[0038] Figure 8 An example of the backwards of the image transformation AI model according to the present invention is shown. Detailed Implementation

[0039] Examples of the invention have been described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily implement the invention. However, since the examples of the invention can be implemented in various different ways, the invention is not limited to the examples described herein.

[0040] In describing examples of the invention, detailed descriptions of well-known functions or structures may unnecessarily obscure the spirit of the invention; therefore, they have not been described in detail. Identical or equivalent constituent elements in the drawings are indicated by the same reference numerals, and repeated or redundant descriptions of the same elements are omitted.

[0041] In this invention, when an element is referred to as being "connected to," "linked to," or "attached to" another element, this can mean that the element is "directly connected to," "directly linked to," or "directly linked to" the other element, or it can mean that the element is connected to, linked to, or linked to another element, and another element is located between them. Furthermore, unless otherwise expressly stated, when an element "comprises" or "has" another element, this means that the element may further include the other element without excluding another component.

[0042] In this invention, unless otherwise expressly stated, the terms first, second, etc., are used only to distinguish one element from another and do not limit the order or importance of the elements. Accordingly, a first element in one example may be referred to as a second element in another example, and similarly, a second element in one example may be referred to as a first element in another example, without departing from the scope of this invention.

[0043] In this invention, elements are distinguished from each other for clear description of each feature, but this does not necessarily mean that the elements are separate. In other words, multiple elements may be integrated into a single hardware or software unit, or a single element may be distributed and formed in multiple hardware or software units. Therefore, examples of such integration or distribution are included within the scope of this invention, even if not otherwise mentioned.

[0044] In this invention, the elements described in the various examples do not necessarily represent essential elements, and some of the elements may be optional. Therefore, examples that include a subset of the elements described in the examples are also included within the scope of this invention. Furthermore, examples that include elements other than those described in the various examples are also included within the scope of this invention.

[0045] The advantages and features of the invention, as well as the manner in which they are obtained, will become apparent to those skilled in the art from the examples of the invention described in detail below with reference to the accompanying drawings. However, examples of the invention may be embodied in many different forms, and the invention should not be construed as limited to the exemplary examples set forth herein. Rather, the examples described herein are provided to make the invention more complete and to fully convey the scope of the invention to those skilled in the art.

[0046] In this invention, each phrase such as “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B or C”, “at least one of A, B and C”, and each phrase such as “at least one of A, B or C” and “at least one of A, B, C or a combination thereof” may include any one or all possible combinations of the items listed together in the corresponding phrase.

[0047] In this invention, for ease of explanation, positional relationships are described using terms such as "upper," "lower," "left," and "right" as used in this specification. Furthermore, when the accompanying drawings shown in this invention are inverted, the positional relationships described herein can be understood in reverse. When the controllers, modules, components, devices, elements, etc., of this invention are described as having a purpose or performing operations, functions, etc., they should be considered herein as "configured" to satisfy that purpose or perform that operation or function. Each controller, module, component, device, element, etc., can be implemented independently or as part of a device including a processor and memory (such as a non-volatile computer-readable medium).

[0048] Figure 1 and Figure 2 A mobile device according to the present invention is shown. The method according to the present invention for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence (AI) model is applied to the mobile device.

[0049] Figure 1 This is a schematic diagram illustrating a mobile device communicating with another device to send and receive data.

[0050] refer to Figure 1The mobile device 100 can be powered by either electricity or fossil fuels. In the case of electricity, for example, the mobile device 100 can be a purely battery-based mobile device powered solely by a high-voltage battery, or it can use a gas-based fuel cell as its energy source. Furthermore, the fuel cell can utilize various types of gases capable of generating electricity; for example, the gas could be hydrogen. However, it is not limited to this, and various gases can be used. In the case of fossil fuels, the mobile device 100 is powered by fuels such as gasoline, diesel, or liquefied gas, and can be equipped with an engine that drives the wheel drive unit 114 through fuel combustion. The engine can be included in an energy generator 112, which provides the wheel drive unit 114 with the driving torque for the wheels.

[0051] Mobile device 100 can refer to a mobile object capable of physically moving through space. Specifically, mobile device 100 can be a vehicle, which is a ground-based mobile object that travels on the ground, and can be a regular passenger car, commercial vehicle, or purpose-built vehicle (PBV). Mobile device 100 can be a four-wheeled vehicle, such as a sedan, sports utility vehicle (SUV), or van, and can also be a vehicle with five or more wheels, such as a bus, truck, container truck, or heavy vehicle. Furthermore, mobile device 100 can include air transport vehicles, such as airplanes, drones, and helicopters, and can also include, but is not limited to, transport vehicles capable of moving at sea, such as ships and submarines.

[0052] The mobile device 100 can be driven by control in autonomous driving, which can be implemented as semi-autonomous or fully autonomous driving. Fully autonomous driving can be set up for automatic movement under the complete control of the processor 120 of the mobile device 100, without requiring user intervention even in uncertain driving conditions. Semi-autonomous driving can be set up for automatic movement requiring driver intervention in specific driving conditions. When a driving condition occurs, semi-autonomous driving can be implemented, causing the processor 120 to disable autonomous driving and switch control to the user, thus allowing the user to perform manual driving. According to the levels of autonomous driving defined by the Society of Automotive Engineers (SAE), semi-autonomous driving can correspond to levels 1 through 4, and fully autonomous driving can correspond to level 5.

[0053] Furthermore, mobile device 100 can communicate with other devices 200 and 300 or another mobile device 400. For example, other devices may include server 200, Intelligent Transportation System (ITS) device 300, and various types of user devices. Server 200 supports various controls, status management, and driving functions of mobile device 100, while ITS device 300 receives information from the ITS. For example, server 200 may be an external device operated by the mobile device manufacturer or provided for autonomous driving services, and server 200 may receive connection data from mobile device 100 or send data required for autonomous driving. To support autonomous driving and various services for mobile device 100, server 200 may send various types of information and software modules for controlling mobile device 100 in response to requests and data sent from mobile device 100 and user devices. According to the present invention, server 200 may send image transformation AI models to mobile device 100.

[0054] For example, ITS device 300 can be a Roadside Unit (RSU), and ITS device 300 can assist the user in driving their vehicle or support autonomous driving of mobile device 100 by exchanging mobility identification data, driving control and situation data, environmental data around the mobile device, and map data with mobile device 100 via V2I. With mobile device 400 via V2V, mobile device 100 can support the driver in driving the vehicle or perform autonomous driving by exchanging the data listed above.

[0055] Mobile device 100 can communicate with another mobile device or another device based on cellular communication, Wireless Access in Vehicular Environment (WAVE) communication, Dedicated Short Range Communication (DSRC) or short range communication, or any other communication scheme.

[0056] For example, mobile device 100 can communicate with server 200, ITS device 300, and mobile device 400 using LTE (a cellular communication network), 5G (a communication network such as 5G), WiFi (a communication network), or WAVE (a communication network such as WAVE). As another example, DSRC used in mobile device 100 can be used for mobile device-to-mobile device communication. The communication scheme between mobile device 100, server 200, ITS device 300, mobile device 400, and user equipment is not limited to the embodiments described above.

[0057] Figure 2This is a schematic diagram showing the constituent modules of a mobile device according to the present invention.

[0058] The mobile device 100 may include: a sensor unit 102, a transceiver 106, a display 108, an actuation unit 110, an energy generator 112, a wheel drive unit 114, a load device 116, a memory 118, and a processor 120. Each of these components is not essential; additional configurations may be provided or omitted, and one configuration may be included in or combined with another configuration, allowing a single configuration to perform multiple functions.

[0059] The sensor unit 102 may be equipped with various types of detectors for sensing various states and conditions occurring in the external and internal environment of the mobile device 100 and for identifying the location information of the mobile device 100. In other words, the sensor unit 102 may be configured as a multi-sensor module including a variety of sensors to obtain sensing data detected by each sensor.

[0060] Specifically, sensor unit 102 may be equipped with a LiDAR sensor 104a, a camera 104b, a radar sensor 104c, and a positioning sensor 104d. The camera 104b serves as a video sensor, the radar sensor 104c is used to identify dynamic and static objects present around the mobile device 100, and the positioning sensor 104d is capable of obtaining the position information of the mobile device 100. Sensor unit 102 can obtain sensor data through the aforementioned sensors, including three-dimensional (3D) recognition data, perception / observation data, and positioning information. Three-dimensional (3D) recognition data corresponds to LiDAR data, and these two terms are used interchangeably below. Perception / observation data may include radar data and image data from the camera.

[0061] According to the present invention, the Lidar sensor 104a can be a type of 3D recognition sensor, and the terms "Lidar sensor" and "3D recognition sensor" are used interchangeably below. The Lidar sensor 104a can be a sensor that observes the surrounding environment and senses the three-dimensional shape of objects based on laser scanning. Specifically, the Lidar sensor 104a can obtain three-dimensional recognition data of the surrounding environment and objects by performing laser scanning around the mobile device 100. The three-dimensional recognition data can include point clouds (i.e., detection data) representing the three-dimensional shape of objects and image data representing the surrounding environment for observation. For example, the detection data can be configured to identify each object by representing the three-dimensional contour and shape of the object and the arrangement of the objects. For example, the image data can be configured to identify the object and the surrounding environment through images of the object and the surrounding environment.

[0062] Camera 104b can acquire two-dimensional image data or image data with depth information of the environment and objects surrounding the mobile device 100. According to the present invention, since camera 104b can include different geometric properties, each camera 104b can be configured to have different internal and external parameters.

[0063] Specifically, since each of the cameras 104b according to the present invention is configured with different focal lengths, different fields of view (FOV), different lens distortions, different focal points, different lens angles, different resolutions, different points of view (POV), different distortion coefficients, different camera positions (translation), and different camera orientations (rotation), each camera 104b can acquire image data with different geometric properties.

[0064] For example, radar sensor 104c can scan with electromagnetic waves of a predetermined wavelength and can detect the behavior of an object based on electromagnetic waves reflected from the object. For example, the behavior of the object may include: the presence of the object, whether the object is moving, the distance between the moving device 100 and the object, the object's speed, and its direction of movement.

[0065] In addition to the positioning sensor 104d, the sensor unit 102 may also be equipped with a gyroscope sensor, an accelerometer, a wheel sensor, a vehicle speedometer, a speed sensor, etc., to identify the vehicle's location, driving position, and speed. Furthermore, to monitor the status of users and occupants inside the mobile device 100, as well as the operational status of internal devices operable by the user of the mobile device 100, the sensor unit 102 may include an interior-facing camera 104b, a biosensor for detecting the biosignals of the driver and occupants, and various detection modules for detecting the operation and status of internal devices.

[0066] For the purpose of describing the implementation, the present invention primarily describes a plurality of sensors of sensor unit 102, but the present invention may further include sensors for detecting various conditions not listed herein.

[0067] Transceiver 106 can support communication with server 200, ITS device 300, and mobile device 400. In this invention, transceiver 106 can send image data generated or stored during operation to server 200, and can also receive image or AI model data from server 200. In this invention, mobile device 100 can use transceiver 106 to send and receive data utilized in the method according to the invention. According to an embodiment of the invention, the AI ​​model data can be a trained image transformation AI model.

[0068] Display 108 can be used as a user interface. Through processor 120, display 108 can display the operating and control status of the mobile device 100, route / traffic information, information about remaining energy, driver requests, etc. Display 108 can be configured as a touchscreen capable of sensing driver input and receiving requests from the driver to processor 120.

[0069] Users can activate or deactivate the autonomous driving function through a software interface (such as the touchscreen of display 108) or a hardware interface (located in a predetermined location within the mobile device 100). For example, a button or key for the autonomous driving function can be installed on the steering wheel, dashboard, etc. Furthermore, the interface can be configured to provide detailed options for selecting various functions available at the corresponding level of autonomous driving.

[0070] In addition, the mobile device 100 may include: an actuation unit 110, an energy generator 112, a wheel drive unit 114, and a load device 116.

[0071] The actuation unit 110 may be equipped with at least one module for implementing driving operations, and may perform at least one driving operation of longitudinal control (such as acceleration / deceleration) and lateral control (such as steering). The actuation unit 110 may be equipped not only with pedals and a steering wheel to accept user control requests, but also with various operation modules for generating driving operations based on requests from the wheel drive unit 114.

[0072] Energy generator 112 can generate and supply power and electricity for driving electric systems such as wheel drive unit 114 and load device 116. For example, when the mobile device 100 is driven by electric energy, energy generator 112 can be configured as an electric battery, or as a combination of an electric battery and a fuel cell for charging the electric battery. In the case of the combination of an electric battery and a fuel cell, energy generator 112 may include a tank that stores materials (e.g., hydrogen) for generating electricity from the fuel cell. When the mobile device 100 is driven by fossil energy, energy generator 112 can be configured as an internal combustion engine.

[0073] The wheel drive unit 114 may include multiple wheels, a drive force transmission module, a braking module, and a steering module. The drive force transmission module generates and provides drive force to the wheels or transmits drive force. The braking module decelerates the wheels, and the steering module provides lateral control of the wheels. When the mobile device 100 is powered by electricity, the drive force transmission module may be configured as a motor module that generates drive force based on electricity output from a battery. When the mobile device 100 is powered by fossil fuels, the drive force transmission module may be equipped with a transmission device and a gearbox module that transmits power from an internal combustion engine.

[0074] The load device 116 may be an auxiliary device installed on the mobility device 100, which consumes power supplied from the energy generator 112 or converted from the output of the energy generator 112 by use by occupants or users. In this invention, the load device 116 may be a type of electrical device for non-driving purposes, other than a drive electrical system (such as wheel drive unit 114). For example, the load device 114 may be various devices installed in air conditioning systems, lighting systems, seating systems, and mobility device 100.

[0075] In addition, the mobile device 100 may include a memory 118 and a processor 120.

[0076] The memory 118 can store applications and various data for controlling the mobile device 100, and can load applications or read and record data when requested by the processor 120. In this invention, for image data obtained from camera 104b or server 200, the application stored in the memory 118 can generate at least one alignment parameter based on the encoder of an image transformation AI model, according to first camera attributes and second camera attributes associated with the respective geometric attributes of different cameras (hereinafter referred to as the first camera and the second camera). The memory 118 can store an application and at least one instruction for transforming first image data captured by a camera including first camera attributes (the first camera) to be aligned with second image data captured by a camera including second camera attributes (the second camera) based on the alignment parameter and brightness parameter, using the transformer of the image transformation AI model. Furthermore, the memory 118 can have an AI model capable of performing visual tasks using the transformed (or aligned) first image data. Specifically, through an AI model that performs visual tasks, memory 118 can store an application and at least one instruction for performing visual tasks (such as semantic segmentation, object detection, and depth estimation) by utilizing the transformed first image data.

[0077] An AI model for performing vision tasks can be trained based on 3D recognition data, image data, radar data, and location data collected from mobile device 100, server 200, and mobile device 400. The AI ​​model can be a deep neural network, such as a convolutional neural network (CNN). The AI ​​model for performing vision tasks can be updated based on data collected in real time during driving.

[0078] Processor 120 can perform overall control of mobile device 100. Processor 120 can be configured to execute applications and instructions stored in memory 118. Processor 106 can be implemented as a single processing module, and some processing can be distributed across multiple processing modules; in this invention, processor 106 can generally refer to multiple processing modules. Figure 5 and Figure 6 The above-described processing of processor 120 is described in detail.

[0079] refer to Figure 3 The server 200 implements the method for training an image transformation AI model according to the present invention, wherein the image transformation AI model not only generates alignment parameters, but also generates aligned image data through the generated alignment parameters.

[0080] Figure 3 This is a schematic diagram illustrating the constituent modules of a server according to the present invention.

[0081] refer to Figure 3 Server 200 may include a communication unit 305, a processor 310, and a memory 315. Each of the constituent elements is not essential, and additional configurations can be set or omitted. A configuration may be included in or combined with another configuration, allowing a single configuration to perform multiple functions.

[0082] According to the present invention, server 200 can train an image transformation AI model, which generates aligned image data using alignment parameters generated by the image transformation AI model.

[0083] Specifically, server 200 can use alignment parameters and brightness parameters generated by the encoder of the image transformation AI model to distinguish between the transformed first image data and the second image data, and server 200 can train the image transformation AI model through adversarial training between the encoder and the discriminator. The training process is described below.

[0084] Server 200 can distribute the image transformation AI model to mobile device 100, so that mobile device 100 can use the image transformation AI model for driving control. The image transformation AI model can align image data captured by each of multiple cameras with different geometric properties.

[0085] According to the present invention, similar to the transceiver 106 of the mobile device 100, the communication unit 305 can transmit collected image data and AI model data to the mobile device 100. The communication unit 305 can be a communication interface that not only receives various data and networks (or algorithms) for training the image transformation AI model (which supports the driving and convenience functions of the mobile device 100), but also sends information and networks related to the image transformation AI model to the mobile device 100. Furthermore, the communication unit 305 can be a communication module that not only receives data generated or stored during driving from the mobile device 100, but also sends information supporting driving, such as map information, environmental information for perceiving objects around the mobile device 100, traffic information, and weather information. The communication unit 305 can also be a communication module that sends applications related to driving and convenience functions.

[0086] The memory 315 can store programs and various data used to control the server 200, and can load programs or read and record data upon request from the processor 310. According to the present invention, the memory 315 can manage image transformation AI models and image data with different geometric properties used to train the models. The image transformation AI model can be configured to include... Figure 4 The functional configurations 415, 420, and 425 shown can collect image data for training from multiple mobile devices 100 and 400 and / or a conventional database (DB) used for training data.

[0087] Processor 310 can perform overall control of server 200. Processor 310 can be configured to execute applications and instructions stored in memory 315. Specifically, processor 310 can control server 200 to train an image transformation AI model stored in memory 315 using image data for training, and distribute the trained image transformation AI model to mobile device 100.

[0088] In this invention, the image transformation AI model used in the mobile device 100 can be a trained model, and the trained image transformation AI model in the mobile device 100 can be called the image transformation AI model.

[0089] Through the training process, processor 310 can determine the values ​​of learnable parameters that constitute the functional configuration of the image transformation AI model. Figure 7 and Figure 8 Describe the learnable parameters and the adversarial training process used to determine the learnable parameters.

[0090] Furthermore, the processor 310 can update the image transformation AI model based on the operation of the image transformation AI model distributed to the mobile device 100, feedback information received from the mobile device 100, and data of the same type as the image data.

[0091] The processor 310 can be implemented as a single processing module, and since processing in some cases can be distributed among multiple processing modules, in this invention, the processor 310 can generally refer to multiple processing modules.

[0092] In the following text, the method for generating aligned image data according to the present invention is described, by... Figure 4 and Figure 5 Describe the constituent elements of an image transformation AI model.

[0093] Figure 4 This is a schematic diagram illustrating the functional configuration included in the image transformation AI model according to the present invention. Figure 4 According to the present invention, the configuration included in the image transformation AI model can actually realize the generation of alignment parameters and the generation of aligned image data through the generated alignment parameters.

[0094] Figure 5 This is a flowchart illustrating the process of generating aligned image data according to the present invention.

[0095] The processor 120 of the mobile device 100 can process data from... Figure 4 The configuration request is shown. In embodiments of the invention, mobile device 100 is primarily described as generating aligned image data by utilizing a trained image transformation AI model, and the image transformation AI model is trained in server 200. However, without departing from the description below, this process can be distributedly controlled or interchanged. For example, mobile device 100 may train the image transformation AI model, or server 200 may generate aligned image data by utilizing the trained image transformation AI model, and may perform visual tasks using the aligned image data.

[0096] In the following text, the processor 120 of mobile device 100 and the processor 310 of server 200 are simplified to mobile device 100 and server 200, respectively, or these terms may be used interchangeably.

[0097] The image transformation AI model 410 may include an encoder 415, a transformer 420, and a discriminator 425. The first camera attributes and the second camera attributes input to the encoder 415 may represent the geometry-related attributes of each of the first and second cameras, which have different geometric attributes, mounted on the mobile device 100.

[0098] Specifically, attributes related to geometry can represent the camera's intrinsic parameters or extrinsic parameters determined by external factors based on the installation location.

[0099] The geometrically related attributes input into encoder 415 can be intrinsic parameters. However, this is not a limitation; distortion coefficients can be added to specific intrinsic parameters and can also be input into encoder 415. As an example, according to the present invention, for the first and second camera attributes related to intrinsic parameters input into encoder 415, each of focal length, FOV, lens distortion, focus, lens angle, resolution, viewpoint, and distortion coefficients, or a combination of the aforementioned intrinsic parameters, can be input.

[0100] The camera's geometry-related attributes that can be input into encoder 415 are not limited to these, and values ​​for each or a combination of internal and external parameters that can be used to generate aligned image data can also be input.

[0101] The first camera attribute of the first camera and the second camera attribute of the second camera can each be a value pre-calculated through calibration. According to the present invention, although the image transformation AI model can have pre-calculated first and second camera attributes as input to the encoder 415, a separate module for calculating camera attributes can also be provided at the front end of the encoder 415.

[0102] The processor 120 generates alignment parameters by encoder 415 based on the first camera attributes and second camera attributes of the first camera and the second camera, which have different geometric attributes (S510).

[0103] Alignment parameters represent parameters used to align image data captured by different cameras with different geometric properties. As an example, alignment parameters according to the present invention may include cropping parameters and a projection matrix, wherein the cropping parameters are used to remove predetermined regions of the image data, and the projection matrix is ​​used to project onto specific image data.

[0104] As an example, cropping parameters may include parameters for removing predetermined locations (such as the top, bottom, left, right, or center regions of image data) according to a predetermined ratio or size. For example, the cropping parameters of the present invention may consist of four parameters for partially removing the top, bottom, left, and right regions of first image data (which was captured by a camera including first camera attributes).

[0105] In the process of training the image transformation AI model according to the present invention, the cropping parameter can be a value determined by regression from the encoder belonging to the trained image transformation AI model. In the present invention, the cropping parameter can be a value determined by the encoder, which outputs the desired value based on the first camera attribute and the second camera attribute as input.

[0106] According to the present invention, the projection matrix may be a matrix that transforms first image data into second image data (which is captured by a camera including second camera attributes) in a projection manner.

[0107] To align image data, in addition to the projection matrix, encoder 415 can also generate matrices for performing translation, rotation, scaling, shearing, and reflection. As an example, encoder 415 can generate affine matrices, similarity matrices, and Euclidean matrices.

[0108] Processor 120 can normalize the first and second camera attributes input through encoder 415, as well as the output cropping parameters and projection matrix, to a predetermined range. For example, the values ​​of the output cropping parameters and projection matrix can be normalized to the range [0, 1]. By normalizing, processor 120 can improve the stability of the image transformation AI model during the training process (described below) and reduce training time. Encoder 415 can be trained through adversarial training against discriminator 425.

[0109] Next, the processor 120 transforms the first image data (which was captured by a first camera including the attributes of the first camera) based on the alignment parameters and brightness parameters through the transformer of the image transformation AI model (S520).

[0110] Specifically, the processor 120 uses the converter 420 to transform the first image data to align with the second image data (which was captured by a camera that includes the attributes of the second camera).

[0111] As an example, transformer 420 can transform the first image data using cropping parameters, projection matrix, and brightness parameters. Furthermore, transformer 420 can perform a resizing operation to adjust the size of the first image data using the size (e.g., resolution) of the second image data. Figure 6Describe the above transformation process in detail.

[0112] A brightness parameter can represent a value configured by parameters that can adjust the brightness of image data. The brightness parameter prevents the discriminator 425 from easily determining whether something is real or fake based on differences in camera exposure values ​​during the training of the encoder 415 and discriminator 425. As an example, the brightness parameter can consist of a first element and a second element, where the first element is multiplied by the total number of pixels in the image data and thus adjusts the brightness of the entire pixel, and the second element is added to the total number of pixels and thus adjusts the brightness of the entire pixel.

[0113] According to the present invention, the brightness parameter can be configured as a learnable parameter. During training, the brightness parameter is independent of the loss of the discriminator 425 caused by the weights of the encoder 415, but can be learned based on the loss of the discriminator 425 caused by the brightness parameter.

[0114] According to the present invention, since the cropping parameters, projection matrix, and brightness parameters are configured to have relatively small parameters or elements, the mobile device 100 can quickly transform image data. Furthermore, when the mobile device 100 generates the aforementioned parameters capable of automatically transforming images, it can consistently transform multiple image datasets into optimal results based on a globally optimal solution.

[0115] Mobile device 100 can perform visual tasks by utilizing a deep learning model trained to work with cameras that have different geometric properties using transformed image data. As an example, mobile device 100 can utilize the deep learning model to analyze transformed first image data, which is trained on a second image dataset (captured by a camera that includes second camera attributes).

[0116] In steps S510 and S520, and during the process of performing a visual task by analyzing the transformed first image data, the discriminator 425 may be frozen in the mobile device 100.

[0117] The discriminator 425 can be designed as a combination of a convolutional neural network (CNN) and a multi-layer perceptron (MLP). The discriminator 425 can be trained through adversarial training against the encoder 415. According to the invention, the discriminator 425 can perform a binary classification based on second image data to determine whether the image is real or fake, i.e., whether the transformed first image data was captured by a camera that includes attributes of a second camera.

[0118] pass Figure 7 and Figure 8 Describe in detail the training process of the image transformation AI model.

[0119] In this article, through Figure 6 This document describes in detail the process of transforming image data using the Transformer 420 of the Image Transformation AI Model.

[0120] Figure 6 This is a flowchart illustrating the detailed process of transforming image data to generate aligned image data.

[0121] First, via transformer 420, processor 120 projects the first image data based on the projection matrix generated by encoder 415 (S610). As an example, the projection matrix can be configured as a 3×3 matrix, and the projected first image data can be generated by performing a 2D projection transformation that maps the first image data to a specific 2D space. Therefore, processor 120 can compensate for differences in image data caused by variations in the camera's internal parameters.

[0122] Then, via transformer 420, processor 120 removes the predetermined region by applying a cropping parameter to the projected first image data (S620). Transformer 420 can appropriately remove unnecessary portions of the projected first image data by utilizing the cropping parameter.

[0123] As an example, the cropping parameters can remove a portion of the top, bottom, left, and right regions of the first projected image data according to a predetermined ratio or size, and can consist of or include four parameters that determine the location or ratio of removal for each region.

[0124] The order of steps S610 and S620 in converter 420 can be changed according to user or system settings. When the order changes, the encoder 415, discriminator 425, and brightness parameters can be learned differently depending on the processing order of converter 420 because the clipping parameters and projection matrix need to be changed.

[0125] Next, via converter 420, processor 120 adjusts the brightness by applying brightness parameters to the first image data after the predetermined area has been removed (S630).

[0126] According to the present invention, the brightness parameter learning is to compensate for differences in image brightness and thus adjust the brightness based on the exposure values ​​of cameras including different geometric properties, so that the first image data after the predetermined area is removed has a brightness similar to that of the second image data.

[0127] As an example, the brightness parameter can consist of or include a first element and a second element, where the first element is multiplied by the total pixels of the image data and thus adjusts the brightness of the total pixels, and the second element is added to the total pixels and thus adjusts the brightness of the total pixels. Accordingly, the transformer 420 can match the brightness according to a linear transformation, which multiplies the total pixels of the first image data by the first element and adds the second element.

[0128] Next, the processor 120 performs a resizing operation through the converter 420, so that the size of the first image data after brightness adjustment corresponds to that of the second image data (S640).

[0129] Transformer 420 can use image data size information (such as resolution) to perform resizing operations to adjust the image data size.

[0130] According to the present invention, the converter 420 obtains the size information (e.g., resolution) of the second image data and matches the size of the brightness-adjusted first image data with the size of the second image data.

[0131] Specifically, during the resizing process, interpolation can be performed between pixel values ​​using interpolation methods, or the size can be matched by increasing or decreasing the number of horizontal and vertical pixels using scaling factors.

[0132] Through the above processing, the processor 120 can transform the first image data captured by the first camera into an image data that is aligned with the second image data captured by the second camera.

[0133] In the following text, through Figure 7 and Figure 8 Describe the process of training an image transformation AI model on server 200.

[0134] Figure 7 This is a flowchart illustrating the process of training an image transformation AI model according to the present invention. Figure 8 This is a schematic diagram illustrating the backwards of the image transformation AI model according to the present invention. Figure 7 The training process described in the text can actually be achieved through... Figure 8 The reverse propagation operation is used to execute it.

[0135] Specifically, according to the present invention, an image transformation AI model can be trained through adversarial training of encoder 415 and discriminator 425. Since discriminator 425 is trained in server 200 by distinguishing between real and fake transformed first image data and second image data, the description of S710 and S720 is omitted. The processing of S710 and S720 is actually the same as the processing in mobile device 100 until the transformed first image data is generated.

[0136] Accordingly, through Figure 7 The description mainly focuses on steps S730 and S740, which correspond to the process of training the encoder 415, discriminator 425, and brightness parameters of the image transformation AI model through adversarial training.

[0137] The processor 310 of server 200 uses discriminator 425 to determine whether the first image data is real or fake, as captured by a second camera including attributes of the second camera (S730). In other words, based on the second image data, discriminator 425 can perform binary classification to determine whether the transformed first image data is real or fake.

[0138] Next, the processor 310 learns the encoder 415, the discriminator 425, and the brightness parameters based on the loss caused by the binary classification result of the discriminator 425 (S740). Specifically, the processor 310 follows... Figure 8 The dashed arrows represent the backpropagation of the loss caused by the binary classification result of discriminator 425.

[0139] In order to deceive the discriminator 425, the encoder 415 can be trained to minimize the size of the loss function, that is, to generate alignment parameters for transforming the first image data, so that the discriminator 425 determines (or is deceived into determining) the fake transformed first image data as real.

[0140] In order to clearly determine whether the transformed first image data is real or fake, the discriminator 425 can be trained to maximize the size of the loss function, that is, to determine that the transformed first image data is fake.

[0141] In other words, similar to the training method of Generative Adversarial Networks (GANs), the encoder 415 and discriminator 425 can be trained through adversarial training.

[0142] According to the present invention, an image transformation AI model can be trained until the value of the loss function enters a set convergence range or reaches a specific value. For example, the loss function according to the present invention may include binary cross-entropy (BCE), but is not limited thereto.

[0143] The encoder 415 and brightness parameters can be trained using the backpropagation loss value of the loss function. The gradient flow during backpropagation can be along... Figure 8 The dashed arrows propagate.

[0144] If the discriminator 425 determines that the transformed first image data is false (i.e. not captured by the second camera), the processor 310 can provide the loss value of the loss function based on the difference between the transformed first image data and the second image to the encoder 415 and the brightness parameter through backpropagation.

[0145] Based on the loss value of the loss function, the processor 310 can update the weights of the encoder 415 to output alignment parameters that minimize the difference between the first image data and the second image data.

[0146] Similarly, based on the loss value, the processor 310 can train the brightness parameters. Specifically, the training of the brightness parameters can be independent of the loss of the discriminator 425 caused by the weights of the encoder 415. The image transformation AI model trained through the above processing can be distributed to the mobile device 100.

[0147] Although the method of the present invention described above is represented as a series of operations for clarity, it is not intended to limit the order in which the steps are performed. The steps described above may be performed simultaneously or in different orders as needed. To implement the method according to the present invention, the described steps may further include different or other steps, and may include steps other than some of the steps, or may include additional steps beyond some of the steps.

[0148] The various examples of this invention do not disclose an enumeration of all possible combinations and are intended to describe representative aspects of the invention. The aspects or features described in the various examples may be applied independently or in combination of two or more.

[0149] Furthermore, various examples of the present invention can be implemented by hardware, firmware, software, or a combination thereof. When the present invention is implemented in hardware, it can be implemented using an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a general-purpose processor, a controller, a microcontroller, a microprocessor, etc.

[0150] The scope of the present invention includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) for enabling operation of methods according to various examples to be executed on a device or computer, and non-volatile computer-readable media having such software or instructions stored thereon and executable on a device or computer.

Claims

1. A method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, the method comprising: The encoder of the image transformation artificial intelligence model generates at least one alignment parameter based on the first camera attributes and the second camera attributes respectively associated with the first camera and the second camera. An image transformer using an image transformation artificial intelligence model, based on at least one alignment parameter and a brightness parameter, transforms first image data captured by a first camera into image data aligned with second image data captured by a second camera. Among them, the encoder and discriminator of the image transformation artificial intelligence model are trained through adversarial training; An AI model for image transformation distinguishes between the transformed first and second image data.

2. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 1, wherein, The first camera attribute and the second camera attribute include at least one of the internal parameters or distortion coefficients of each of the first camera and the second camera.

3. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 1, wherein, At least one alignment parameter includes at least one of a cropping parameter for removing a predetermined region or a projection matrix for projecting the first image data onto the second image data.

4. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 3, wherein, Transformations include: The first image data is projected based on the projection matrix; The predetermined region is removed by reflecting the cropping parameters in the first image data after projection. Brightness is adjusted by applying brightness parameters to the first image data after the predetermined area has been removed; Perform a resizing operation to match the size of the first image data after brightness adjustment with the size of the second image data.

5. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 3, wherein, Transformations include: Remove a predetermined region from the first image data based on the cropping parameters; The first image data is projected by reflecting the projection matrix on the first image data after the predetermined region has been removed. Brightness is adjusted by applying brightness parameters to the first image data after projection. Perform a resizing operation to match the size of the first image data after brightness adjustment with the size of the second image data.

6. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 1, wherein, The brightness parameter is a parameter that can be learned based on the discriminator's loss caused by the brightness parameter, and is independent of the discriminator's loss caused by the encoder.

7. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 3, further comprising: Make the clipping parameters and projection matrix subordinate to the trained encoder; The clipping parameters and projection matrix are determined by regression.

8. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 4 or 5, wherein, Adjusting the brightness includes: Multiply the first element of the brightness parameter by the total number of pixels; Add the second element of the brightness parameter to the entire pixel.

9. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 1, further comprising: The discriminator determines whether the transformed first image data captured by the second camera is real or fake. Train the discriminator to identify false positives; Learn the encoder and brightness parameters so that the discriminator can identify it as real.

10. The method for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model according to claim 3, further comprising: The encoder normalizes the input first camera attributes and second camera attributes, as well as the output clipping parameters and projection matrix, to a predetermined range.

11. A moving device for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, the moving device comprising: A memory configured to store at least one instruction; and A processor configured to execute at least one instruction stored in memory based on data obtained from memory; The processor is further configured as follows: The encoder of the image transformation artificial intelligence model generates at least one alignment parameter based on the first camera attributes and the second camera attributes respectively associated with the first camera and the second camera. An image transformer using an image transformation artificial intelligence model, based on at least one alignment parameter and a brightness parameter, transforms first image data captured by a first camera into image data aligned with second image data captured by a second camera. The encoder and discriminator of an image transformation AI model are trained through adversarial training. An AI model for image transformation distinguishes between the transformed first and second image data.

12. The moving apparatus according to claim 11 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The first camera attribute and the second camera attribute include at least one of the internal parameters or distortion coefficients of each of the first camera and the second camera.

13. The moving apparatus according to claim 11 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, At least one alignment parameter includes at least one of a cropping parameter for removing a predetermined region or a projection matrix for projecting the first image data onto the second image data.

14. The moving apparatus according to claim 13 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The processor is further configured as follows: The first image data is projected based on the projection matrix; The predetermined region is removed by reflecting the cropping parameters in the first image data after projection. Brightness is adjusted by applying brightness parameters to the first image data after the predetermined area has been removed; By performing a resizing operation that matches the size of the first image data after brightness adjustment with the size of the second image data, the first image data is transformed to be aligned with the second image data.

15. The moving apparatus according to claim 13 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The processor is further configured as follows: Remove a predetermined region from the first image data based on the cropping parameters; The first image data is projected by reflecting the projection matrix on the first image data after the predetermined region has been removed. Brightness is adjusted by applying brightness parameters to the first image data after projection. By performing a resizing operation that matches the size of the first image data after brightness adjustment with the size of the second image data, the first image data is transformed to be aligned with the second image data.

16. The moving apparatus according to claim 11 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The brightness parameter is a parameter that can be learned based on the discriminator's loss caused by the brightness parameter, and is independent of the discriminator's loss caused by the encoder.

17. The moving apparatus according to claim 13 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The pruning parameters and projection matrix belong to the trained encoder and are determined by regression.

18. The moving apparatus according to claim 14 or 15 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The processor is further configured to, when adjusting brightness: Multiply the first element of the brightness parameter by the total number of pixels; Add the second element of the brightness parameter to the entire pixel.

19. The moving apparatus according to claim 11 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The processor is further configured to: use a discriminator to determine whether the transformed first image data is real or fake, as captured by the second camera. Train the discriminator to identify false positives. Learn the encoder and brightness parameters so that the discriminator can identify it as real.

20. The moving apparatus according to claim 13 for generating aligned image data using alignment parameters generated by an image transformation artificial intelligence model, wherein, The processor is further configured as follows: The encoder normalizes the input first and second camera attributes, as well as the output clipping parameters and projection matrix, to a predetermined range.

Citation Information

Patent Citations

  • Cover for hair dyeing

    KR1020240067699A