Electronic device, method for executing same, and non-transitory computer readable storage medium
By using two cameras with different shutter speeds on the vehicle to acquire images and generating learning data under the control of the processor, the problems of image noise and blur in low-light or dynamic environments are solved, and the acquisition and subsequent processing of clear images are achieved.
Patent Information
- Application Number
- CN202510312582.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-19
AI Technical Summary
Electronic devices mounted on vehicles have difficulty acquiring clear images of the external environment in low-light or dynamic environments, resulting in severe noise or blur.
By installing two cameras on the vehicle, one with a slow shutter speed and the other with a fast shutter speed, multiple images are acquired respectively, and the image areas corresponding to the time are identified under the control of the processor to generate learning data for denoising and deblurring.
It enables the acquisition of clear image data in low-light or dynamic environments, supports subsequent denoising and deblurring processing, and improves image quality.
Smart Images

Figure CN120676226A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device and method for acquiring learning data. Background Art
[0002] An electronic device mounted on a vehicle for acquiring an image of the external environment may have difficulty obtaining a clear image due to noise or blur caused by the environment (eg, low light, dynamic environment, etc.).
[0003] Using artificial intelligence, improved images can be obtained by de-noising or de-blurring.
[0004] The above information is provided as background technology to assist in understanding the present disclosure. No assertion or determination is made as to whether any of the foregoing may be applicable as prior art to the present disclosure. Summary of the Invention
[0005] Means used to solve problems
[0006] According to one embodiment, an electronic device may include a communication circuit. The electronic device may include a memory for storing instructions. The electronic device may include at least one processor operably connected to the communication circuit and the memory. The instructions, when executed by the processor, may cause the electronic device to acquire multiple first images from a first camera via the communication circuit. The instructions, when executed by the processor, may cause the electronic device to acquire multiple second images from a second camera having a slower shutter speed than the first camera via the communication circuit. The instructions, when executed by the processor, may cause the electronic device to identify, from the multiple first images and the multiple second images, first and second images acquired at corresponding time points. The instructions, when executed by the processor, may cause the electronic device to set a first region of interest within an object of a specified type in the identified first image. The instructions, when executed by the processor, may cause the electronic device to set a second region of interest in the identified second image at the same location as the first region of interest. The instructions, when executed by the processor, may cause the electronic device to generate learning data from the identified first image with the first region of interest set and the identified second image with the second region of interest set.
[0007] According to one embodiment, a method performed by an electronic device may be performed in an electronic device including a communication circuit. The method may include an action of acquiring a plurality of first images from a first camera via the communication circuit. The method may include an action of acquiring a plurality of second images from a second camera having a slower shutter speed than the first camera via the communication circuit. The method may include an action of identifying, from the plurality of first images and the plurality of second images, first images and second images acquired at corresponding time points. The method may include an action of setting a first region of interest within an object of a specified type in the identified first image. The method may include an action of setting a second region of interest in the same position as the first region of interest in the identified second image. The method may include an action of generating the identified first image with the first region of interest set and the identified second image with the second region of interest set as learning data.
[0008] A non-transitory computer-readable storage medium may store a program including instructions. When executed by a processor of an electronic device including a communication circuit, the instructions may cause the electronic device to acquire multiple first images from a first camera via the communication circuit. When executed by the processor, the instructions may cause the electronic device to acquire multiple second images from a second camera having a slower shutter speed than the first camera via the communication circuit. When executed by the processor, the instructions may cause the electronic device to identify, from the multiple first images and the multiple second images, first and second images acquired at corresponding time points. When executed by the processor, the instructions may cause the electronic device to set a first region of interest within an object of a specified type in the identified first image. When executed by the processor, the instructions may cause the electronic device to set a second region of interest in the identified second image at the same location as the first region of interest. When executed by the processor, the instructions may cause the electronic device to generate, as learning data, the identified first image with the first region of interest set and the identified second image with the second region of interest set.
[0009] Effects of the Invention
[0010] According to an electronic device and method thereof according to an embodiment, learning data for learning artificial intelligence to perform denoising and / or deblurring can be acquired. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 An example of a block diagram of an electronic device is shown.
[0012] Figures 2A to 2D Shows an example of a frame acquired by a camera.
[0013] Figure 3 An example of a ground truth (GT) image and a blurred image acquired by a camera is shown.
[0014] Figure 4 An example of the operation of cropping the GT image and blurring the image is shown.
[0015] Figure 5 An example of the operation of extracting a region of interest from a GT image is shown.
[0016] Figure 6 An example of an operation of setting a region of interest in a blurred image is shown.
[0017] Figure 7 An exemplary flowchart for explaining actions of an electronic device according to an embodiment is shown.
[0018] Figure 8 An example of a block diagram illustrating an autonomous driving system for a vehicle is shown according to an embodiment.
[0019] Figure 9 and Figure 10 An example of a block diagram illustrating an autonomous driving mobile object according to an embodiment is shown.
[0020] Figure 11 An example of a gateway associated with a user device according to various embodiments is shown.
[0021] Figure 12 is a diagram for explaining the operation of an electronic device for training a neural network based on a learning data set according to one embodiment.
[0022] Figure 13 is a block diagram of an electronic device according to an embodiment.
[0023] Figure 14 This is a schematic diagram of the tractor and trailer in the disconnected state.
[0024] Figure 15 It is a schematic diagram showing the connection status of the tractor and trailer. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In conjunction with the description of the accompanying drawings, similar reference numerals may be used for similar or related components.
[0026] Over the years, the trucking industry has experienced sustained growth and expanded its service offerings to address more complex supply chains. These services include last-mile deliveries, drop-trailer programs, and intermodal transportation at ports (where freight is transported to its destination using two or more different modes of transportation, such as ship and rail, or ship and air).
[0027] As such, there are many ways to transport cargo, and manufacturers of cargo transportation-related equipment have designed different types of equipment for transporting cargo according to different transportation needs.
[0028] In this specification, trucks whose main purpose is to tow freight are generally referred to as tractors.
[0029] The tractors described in this specification can be divided into conventional trucks (or bonneted trucks), cab-over trucks (or cab-over engines), and semi-conventional trucks that are between conventional trucks and cab-over trucks, depending on the position and shape of the tractor cab.
[0030] Conventional trucks have an engine and a hood located above the front axle in front of a tractor cab, and a driver sitting behind the front axle. The tractor engine is located in front of the driver, and is a shape of tractor mainly used in North America.
[0031] On the other hand, a cab-over truck has a structure where the driver sits in front of the front axle, with the cab of the tractor located at the front end of the tractor. This is a so-called "flat face" or "flat nose" tractor with a flat front surface. The tractor's engine is located below the driver, and this is a tractor shape primarily used in most countries in Europe and Asia.
[0032] Just as there are many types of tractors depending on their purpose and needs, there are also many types of trailers towed by tractors. The most representative types of trailers are full-trailers and semi-trailers. Full-trailers and semi-trailers can be distinguished by whether the trailer is equipped with both a front axle and a rear axle. Such trailers can be connected to box trucks or tractors via a coupling device.
[0033] Specifically, a full trailer is a commercial freight trailer equipped with both front and rear axles. Designed so that its total load is supported solely by the trailer, it is capable of fully supporting its own weight without relying on a tractor, and is equipped with a drawbar for coupling to a hauling unit or towing unit such as a tractor. It is primarily used in the United States, Canada, and elsewhere.
[0034] On the other hand, a semi-trailer is a freight trailer that is only equipped with a rear axle but not a front axle, and a large part of the load can be supported by a tractor connected by a hook (hitch) that is called a "fifth wheel (fifth wheel, or steering wheel)". When the semi-trailer is separated from the tractor and is in a stationary state, the landing gear (landing gear) installed at the bottom of the semi-trailer can be vertically unfolded to the ground to support the load of the trailer. The combination of a semi-trailer and a tractor is called a semi-trailer truck (semi-trailer truck) (also referred to as "semi-trailer", "tractor-trailer (tractor-trailer)", "semi-truck (semi-truck)", "big truck (Big rig)" or "semi-trailer (semi)") in the United States. The above-mentioned "steering wheel (fifth wheel (fifth wheel))" refers to a horizontal wheel attached to the tractor axle of the trailer truck for facilitating the direction change of the trailer, also known as a "fifth wheel". A fifth wheel is a device that movably couples a tractor and a semi-trailer, and generally includes a lower portion consisting of a kingpin mounted on the semi-trailer, a trunnion plate securely fastened to the tractor, and a latching device.
[0035] In the following description, based on the above-mentioned tractor / trailer terminology, for convenience of explanation, "trailer" will refer to a cargo transport vehicle connected to a tractor for a trailer, and "tractor" will refer to a towing vehicle used to move the trailer. Furthermore, in order to maximize the elimination of rights limitations based on the embodiments described in the detailed description, in the present invention, the tractor hauling / towing a "trailer" may be used interchangeably with the "towing vehicle," and the trailer towed by the tractor may be used interchangeably with the "towed vehicle."
[0036] In addition, for the sake of convenience, preferably, the “trailer” described throughout this specification should be understood to refer to a “semi-trailer”, but is not limited thereto.
[0037] In one embodiment, the electronic device 101 and the cameras 151 and 155 may be included in (or loaded on) a vehicle (not shown). The electronic device 101 may correspond to an on-board electronic control unit (ECU) or may be included in an ECU. The ECU may be referred to as an electronic control module (ECM). The electronic device 101 may be configured as independent hardware to provide functions according to embodiments of the present invention in a vehicle. The embodiment is not limited thereto, and the electronic device 101 may correspond to a device attached to a vehicle (e.g., a black box) or may be included in the device.
[0038] Reference Figure 1 According to an embodiment, the electronic device 101 may include a communication circuit 110, a processor 120, and a memory 130. The communication circuit 110, the processor 120, and the memory 130 may be electrically connected and / or operably connected to each other via an electronic component such as a communication bus 140. Hereinafter, the operably combined hardware may refer to establishing a wired or wireless direct connection or an indirect connection between the hardware, so that the second hardware in the hardware is controlled by the first hardware. Although shown in different blocks, the embodiment is not limited thereto. Figure 1 Some of the hardware in the electronic device 101 may be provided in a single integrated circuit (SIC), such as a system on a chip (SoC). The type and / or number of hardware included in the electronic device 101 is not limited to Figure 1 For example, the electronic device 101 may only include Figure 1 Some of the hardware shown.
[0039] According to one embodiment, the communication circuit 110 of the electronic device 101 may include hardware components for supporting the transmission and / or reception of electrical signals between the electronic device 101 and an external electronic device (e.g., cameras 151 and 155). The communication circuit 110 may include, for example, at least one of a modem, an antenna, and an optical / electronic (O / E) converter. The communication circuit 110 may support the transmission and / or reception of electrical signals based on various types of protocols (e.g., Ethernet, local area network (LAN), wide area network (WAN), wireless fidelity (WiFi), Bluetooth, Bluetooth low energy (BLE), ZigBee, long term evolution (LTE), 5G NR, and / or 6G).
[0040] The electronic device 101 according to an embodiment may include hardware for processing data based on one or more instructions. The hardware for processing data may include a processor 120. For example, the hardware for processing data may include an arithmetic and logic unit (ALU), a floating point unit (FPU), a field programmable gate array (FPGA), a central processing unit (CPU) and / or an application processor (AP). The processor 120 may have a structure of a single-core processor, or may have a structure of a multi-core processor such as a dual-core, quad-core, hexa-core or octa-core.
[0041] According to an embodiment, the memory 130 of the electronic device 101 may include a hardware component for storing data and / or instructions input to and / or output from the processor 120 of the electronic device 101. For example, the memory 130 may include a volatile memory such as a random-access memory (RAM) and / or a non-volatile memory such as a read-only memory (ROM). For example, the volatile memory may include at least one of a dynamic random access memory (DRAM), a static random access memory (SRAM), a cache RAM, and a pseudo-static random access memory (PSRAM). For example, the non-volatile memory may include at least one of a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a hard disk, an optical disk, a solid state drive (SSD), and an embedded multi-media card (eMMC).
[0042] Although not shown, the electronic device 101 may further include various components. For example, the electronic device 101 may further include a display for displaying a user interface.
[0043] According to one embodiment, each camera 151, 155 may include a lens assembly or an image sensor. The lens assembly may collect light emitted from a subject serving as an image capture object. The lens assembly may include one or more lenses. According to one embodiment, each camera 151, 155 may include multiple lens assemblies. For example, in each camera 151, 155, some of the multiple lens assemblies may have the same lens properties (e.g., viewing angle, focal length, autofocus, focal ratio (fnumber), or optical zoom), or at least one lens assembly may have one or more lens properties that differ from those of another lens assembly. The lens assembly may include a wide-angle lens or a telephoto lens. According to one embodiment, the flash may include one or more light-emitting diodes (e.g., red-green-blue (RGB) LEDs, white LEDs, infrared LEDs, or ultraviolet LEDs) or a xenon lamp. For example, the image sensor may obtain an image corresponding to the subject by converting light emitted or reflected from the subject and transmitted through the lens assembly into an electrical signal. According to an embodiment, the image sensor may include one image sensor selected from a plurality of image sensors having different properties (e.g., an RGB sensor, a black and white (BW) sensor, an IR sensor, or a UV sensor), a plurality of image sensors having the same properties, or a plurality of image sensors having different properties. Each image sensor included in the image sensor may be implemented using, for example, a charge coupled device (CCD) sensor or a complementary metal oxide semiconductor (CMOS) sensor.
[0044] In one embodiment, each camera 151, 155 may have different settings. In one embodiment, each camera 151, 155 may have different settings in at least one of a plurality of settings. For example, the plurality of settings may include the angle of each camera 151, 155 (e.g., pitch angle, roll angle, and / or pan angle). For example, the plurality of settings may include angle of view, resolution, and / or lens properties (e.g., angle of view, focal length, autofocus, focal ratio, ISO sensitivity, or optical zoom). For example, the plurality of settings may include shutter speed. For example, the angles of each camera 151, 155 are set (or arranged) to be the same (or substantially the same, or corresponding) as each other through a hardware structure (e.g., a level, a housing, a fixing frame). For example, the resolution and / or lens properties (e.g., viewing angle, focal length, autofocus, focal ratio, ISO sensitivity, or optical zoom) of each camera 151, 155 can be set to be the same (or substantially the same, or corresponding). For example, the shutter speed of each camera 151, 155 can be different. For example, the shutter speed of camera 151 can be faster than the shutter speed of camera 155. In one embodiment, In this example, because the shutter speed of camera 151 is faster than that of camera 155, the image captured by camera 151 may be clearer than the image captured by camera 155. In one embodiment, because the shutter speed of camera 151 is faster than that of camera 155, the image captured by camera 155 may have at least one of lower resolution, white noise, and blur (or motion blur) than the image captured by camera 151. As a result, it may be more difficult to recognize a character string in the image captured by camera 155 than in the image captured by camera 151.
[0045] In one embodiment, the image captured by camera 151 can be used as a correct answer image (ground truth image) for reference. In one embodiment, the image captured by camera 155 can be used as a blurred image (or motion blurred image) for reference.
[0046] According to one embodiment, the cameras 151 and 155 may be disposed (or arranged) toward one direction of the vehicle (not shown). For example, the cameras 151 and 155 may be disposed (or arranged) toward the front direction (front direction and / or driving direction) of the vehicle.
[0047] exist Figure 1In the embodiment, the cameras 151 and 155 are shown as being physically separated from the electronic device 101, but this is merely an example. Depending on the embodiment, some of the cameras 151 and 155 may be integrally formed with the electronic device 101. Alternatively, both cameras 151 and 155 may be integrally formed with the electronic device 101. For example, if both cameras 151 and 155 are integrally formed with the electronic device 101, the communication circuit 110 may be a circuit for interfacing between the cameras 151 and 155 and the processor 120 (e.g., a mobile industry processor interface (MIPI)).
[0048] The following describes the operation of the electronic device 101 collecting learning data through the cameras 151 and 155.
[0049] In one embodiment, the processor 120 of the electronic device 101 may instruct (or command) the cameras 151 and 155 to take pictures. In one embodiment, the processor 120 of the electronic device 101 may instruct (or command) the cameras 151 and 155 to take pictures at the same (or substantially the same, or corresponding) time points. In one embodiment, the processor 120 may instruct (or command) the cameras 151 and 155 to take pictures when the vehicle is moving. But not limited to this. In one embodiment, the processor 120 may instruct (or command) the cameras 151 and 155 to take pictures regardless of the movement of the vehicle. For example, the time points being the same (or substantially the same, or corresponding) means that the instructions (or commands) to take pictures are generated and / or transmitted within a specified offset (or frame interval depending on the frame rate).
[0050] In one embodiment, the processor 120 may acquire an image stream from each camera 151, 155. In one embodiment, the image stream may be a data stream of images acquired at a set frame rate (e.g., frames per second (FPS)). In one embodiment, the set frame rate may be 30 FPS, but is not limited thereto. In one embodiment, the set frame rate may be less than 30 FPS (e.g., 24 FPS) or greater than 30 FPS (e.g., 60 FPS).
[0051] In one embodiment, the processor 120 may store a specified number (e.g., 10) of images acquired at a set frame rate after a capture command in the memory 130 (or a buffer in the memory 130). In one embodiment, the processor 120 may store images acquired within a specified time (e.g., 0.3 seconds) after a capture command in the memory 130 (or a buffer in the memory 130). For example, the processor 120 may store a specified number (e.g., 10) of ground truth (GT) images acquired from the camera 151 within a specified time and a specified number (e.g., 10) of blurred images acquired from the camera 155 within a specified time in the memory 130 (or a buffer in the memory 130).
[0052] In one embodiment, the processor 120 can identify whether the vehicle is moving based on the GT image and the blurred image obtained from the cameras 151 and 155. In one embodiment, the processor 120 can identify whether the vehicle is moving based on the comparison between consecutive images. For example, the processor 120 can identify whether the vehicle is moving based on the comparison between consecutive GT images (or blurred images). Here, the comparison between consecutive images can be based on comparing the feature map of the image of the first frame and the feature map of the image of the second frame after the first frame. The comparison between consecutive images can be based on the differential image between the image of the first frame and the image of the second frame. Here, the second frame can be a frame obtained after the first frame according to the set frame rate. Here, the feature map can be obtained based on a neural network. But it is not limited to this.
[0053] In one embodiment, when the difference between consecutive images is less than a specified baseline difference, the processor 120 may recognize that the vehicle has stopped. In one embodiment, when the difference between consecutive images exceeds a specified baseline difference, the processor 120 may recognize that the vehicle is moving. For example, when the size of the area exceeding the baseline difference in the difference feature map between the feature map of the image of the first frame and the feature map of the image of the second frame is greater than the baseline size, the processor 120 may recognize that the vehicle is moving. For example, when the size of the area exceeding the baseline difference in the difference feature map between the feature map of the image of the first frame and the feature map of the second frame is less than the baseline size, the processor 120 may recognize that the vehicle has stopped. For example, when the size of the area exceeding the baseline difference in the differential image between the image of the first frame and the image of the second frame is greater than the baseline size, the processor 120 may recognize that the vehicle is moving. For example, when the size of the area exceeding the baseline difference in the differential image between the image of the first frame and the image of the second frame is less than the baseline size, the processor 120 may recognize that the vehicle has stopped.
[0054] In one embodiment, the processor 120 may identify duplicate images based on the GT images and blurred images obtained from the cameras 151 and 155. In one embodiment, the processor 120 may identify duplicate images based on a comparison between consecutive images. For example, the processor 120 may identify duplicate images based on a comparison between consecutive GT images (or blurred images). Here, the comparison between consecutive images may be based on comparing a feature map of an image of a first frame with a feature map of an image of a second frame following the first frame. The comparison between consecutive images may be based on a differential image between an image of the first frame and an image of the second frame. Here, the second frame may be a frame obtained after the first frame according to a set frame rate. Here, the feature map may be obtained based on a neural network. But it is not limited thereto.
[0055] In one embodiment, when the difference between the consecutive images is below a specified reference difference, the processor 120 may identify the consecutive images as duplicate images. In one embodiment, the processor 120 may delete (or remove) one of the duplicate images.
[0056] In one embodiment, the processor 120 may determine whether the GT image and the blurred image acquired from the cameras 151 and 155 have been processed. For example, the processor 120 may determine whether the GT image and the blurred image have been processed based on whether the GT image and the blurred image were acquired while the vehicle was moving. For example, when the GT image and the blurred image were acquired while the vehicle was moving, the processor 120 may determine to process the GT image and the blurred image. However, the present invention is not limited thereto.
[0057] In one embodiment, the processor 120 may align the time between a specified number of GT images and the blurred image. In one embodiment, when determining to process a GT image and a blurred image, the processor 120 may align the time between a specified number of GT images and the blurred image. For example, the processor 120 compares a reference image among images acquired by a reference camera with images acquired by other cameras, thereby aligning the time. For example, the reference camera may be camera 151, and the reference image may be an intermediate image (e.g., the fifth image) among a specified number of images. At this time, the processor 120 may identify the blurred image acquired at the same time point (or substantially the same time point, or corresponding time point) as the fifth GT image by comparing the fifth GT image with the 10 blurred images.
[0058] For example, the comparison between the GT image and the blurred image can be based on comparing the feature map of the GT image with the feature map of the blurred image. For example, the electronic device 101 can identify the GT image and the blurred image acquired at the same time point by comparing the size of the area exceeding the benchmark difference in the difference feature map between the feature map of the GT image and the feature map of the blurred image with the benchmark size. For example, when the size of the area exceeding the benchmark difference in the difference feature map between the feature map of the GT image and the feature map of the blurred image is greater than the benchmark size, the processor 120 can identify that the GT image and the blurred image are not images acquired at the same time point. For example, when the size of the area exceeding the benchmark difference in the difference feature map between the feature map of the GT image and the feature map of the blurred image is less than the benchmark size, the processor 120 can identify that the GT image and the blurred image are images acquired at the same time point. Here, the feature map can be acquired based on a neural network. But it is not limited to this.
[0059] For example, the comparison between the GT image and the blurred image can be based on the differential image between the GT image and the blurred image. For example, the electronic device 101 can identify the GT image and the blurred image acquired at the same time point by comparing the size of the area exceeding the benchmark difference in the differential image with the benchmark size. For example, when the size of the area exceeding the benchmark difference in the differential image between the GT image and the blurred image is greater than the benchmark size, the processor 120 can recognize that the GT image and the blurred image are not images acquired at the same time point. For example, when the size of the area exceeding the benchmark difference in the differential image between the GT image and the blurred image is less than the benchmark size, the processor 120 can recognize that the GT image and the blurred image are images acquired at the same time point.
[0060] In one embodiment, acquiring the GT image and the blurred image at substantially the same time point (or at time points corresponding to each other) may mean that the acquisition time difference between the GT image and the blurred image is less than the frame interval (or half of the frame interval) according to the set frame rate. The comparison between the GT image and the blurred image can refer to Figures 2A to 2D To explain.
[0061] Figures 2A to 2D Shows an example of a frame acquired by a camera.
[0062] Reference Figure 2A , cameras 151 and 155 may start shooting at the same time point (or substantially the same time point, or time points corresponding to each other) (e.g., T1). In one embodiment, cameras 151 and 155 start shooting at substantially the same time point (or corresponding time points) may mean that the time difference between cameras 151 and 155 in starting shooting is less than the frame interval (or half of the frame interval) according to the set frame rate. Thus, in Figure 2A In the embodiment, ten GT images 211 to 220 and ten blurred images 221 to 230 may be acquired at the same time point (or substantially the same time point, or corresponding time points). For example, the GT image 211 and the blurred image 221 may be acquired at time point T1. For example, the GT images 212 to 220 and the blurred images 222 to 230 may be acquired at each time point (e.g., T2 to T10).
[0063] In one embodiment, processor 120 may align the time between the ten GT images 211 to 220 and the ten blurred images 221 to 230. For example, processor 120 may compare a GT image 215 (or a reference image) in a specified order (e.g., the fifth) among the GT images 211 to 220 acquired by camera 151 (or a reference camera) with each of the ten blurred images 221 to 230 acquired by camera 155 (or another camera). The comparison between GT image 215 and each of the ten blurred images 221 to 230 may be based on comparing a feature map of GT image 215 with a feature map of each of the ten blurred images 221 to 230. The comparison between consecutive images may be based on a difference image between GT image 215 and each of the ten blurred images 221 to 230. The feature map may be acquired based on a neural network, but is not limited thereto.
[0064] In one embodiment, the processor 120 may identify a blurred image whose difference between the GT image 215 and each of the 10 blurred images 221 to 230 is less than a specified baseline difference as an image acquired at the same time point as the GT image 215. For example, the processor 120 may identify the GT image 215 and the blurred image 225 as images acquired at the same time point. Thus, the electronic device 101 may identify a set of images acquired at the same time point based on the acquisition order (or frame order) of the 10 GT images 211 to 220 and the 10 blurred images 221 to 230. For example, the electronic device 101 may identify a blurred image from the 10 blurred images 221 to 230 that has the same frame difference as the frame difference between a specific GT image among the 10 GT images 211 to 220 and the baseline image. In one embodiment, the electronic device 101 may identify the specific image and the identified blurred image as a set of images acquired at the same time point. For example, as the GT image 215 and the blurred image 225 are identified as images acquired at the same time point, the GT image 211 and the blurred image 221 are identified as images acquired at the same time point (eg, T1).
[0065] Reference Figure 2B, the camera 151 can start shooting at a time point (for example, T3) corresponding to a time 2 frames later than the camera 155. Figure 2B In the example, eight GT images 213 to 220 and ten blurred images 221 to 230 can be acquired. For example, blurred images 221 and 222 can be acquired at time points T1 and T2. For example, GT images 213 to 220 and blurred images 223 to 230 can be acquired at each time point (e.g., T3 to T10).
[0066] In one embodiment, the processor 120 may align the time between the eight GT images 213 to 220 and the ten blurred images 221 to 230. For example, the processor 120 may compare the GT image 217 of a specified order (e.g., the fifth) among the GT images 213 to 220 acquired by the camera 151 with each of the ten blurred images 221 to 230 acquired by the camera 155.
[0067] In one embodiment, the processor 120 may identify a blurred image whose difference between the GT image 217 and each of the ten blurred images 221 to 230 is less than a specified baseline difference as an image acquired at the same time point as the GT image 217. For example, the processor 120 may identify the GT image 217 and the blurred image 227 as images acquired at the same time point. Thus, the electronic device 101 may identify a set of images acquired at the same time point based on the acquisition order (or frame order) of the eight GT images 213 to 220 and the ten blurred images 221 to 230. For example, as the GT image 217 and the blurred image 227 are identified as images acquired at the same time point, the GT image 213 and the blurred image 223 are also identified as images acquired at the same time point.
[0068] Reference Figure 2C , the camera 155 can start shooting at a time point (for example, T2) corresponding to a time one frame later than the camera 151. Figure 2C In the example, ten GT images 211 to 220 and nine blurred images 222 to 230 can be acquired. For example, the GT image 211 can be acquired at time point T1. For example, the GT images 212 to 220 and the blurred images 222 to 230 can be acquired at each time point (e.g., T2 to T10).
[0069] In one embodiment, the processor 120 may align the time between the ten GT images 211 to 220 and the nine blurred images 222 to 230. For example, the processor 120 may compare the GT image 215 of a specified order (e.g., the fifth) among the GT images 211 to 220 acquired by the camera 151 with each of the nine blurred images 222 to 230 acquired by the camera 155.
[0070] In one embodiment, the processor 120 may identify a blurred image whose difference between the GT image 215 and each of the nine blurred images 222 to 230 is less than a specified baseline difference as an image acquired at the same time point as the GT image 215. For example, the processor 120 may identify the GT image 215 and the blurred image 225 as images acquired at the same time point. Thus, the electronic device 101 may identify a set of images acquired at the same time point based on the acquisition order (or frame order) of the ten GT images 211 to 220 and the nine blurred images 222 to 230. For example, as the GT image 215 and the blurred image 225 are identified as images acquired at the same time point, the GT image 212 and the blurred image 222 are also identified as images acquired at the same time point.
[0071] Reference Figure 2D , the camera 155 can start shooting at a time point (for example, T4) corresponding to a time 3 frames later than the camera 151. Figure 2D In the example, ten GT images 211 to 220 and seven blurred images 224 to 230 can be acquired. For example, the GT images 211, 212, and 213 can be acquired at time points T1, T2, and T3. For example, the GT images 214 to 220 and the blurred images 224 to 230 can be acquired at each time point (e.g., T4 to T10).
[0072] In one embodiment, the processor 120 may align the time between the ten GT images 211 to 220 and the seven blurred images 224 to 230. For example, the processor 120 may compare the GT image 215 of a specified order (e.g., the fifth) among the GT images 211 to 220 acquired by the camera 151 with each of the seven blurred images 224 to 230 acquired by the camera 155.
[0073] In one embodiment, the processor 120 may identify a blurred image whose difference between the GT image 215 and each of the seven blurred images 224 to 230 is less than a specified baseline difference as an image acquired at the same time point as the GT image 215. For example, the processor 120 may identify the GT image 215 and the blurred image 225 as images acquired at the same time point. Thus, the electronic device 101 may identify a set of images acquired at the same time point based on the acquisition order (or frame order) of the ten GT images 211 to 220 and the seven blurred images 224 to 230. For example, as the GT image 215 and the blurred image 225 are identified as images acquired at the same time point, the GT image 214 and the blurred image 224 are also identified as images acquired at the same time point.
[0074] In one embodiment, the processor 120 may generate a learning dataset based on the GT image and the blurred image acquired at the same time point (or substantially the same time point, or corresponding time points). The learning dataset may include an image based on a pair of GT images and blurred images. The learning data may include multiple learning datasets. The generation of the learning dataset may refer to Figures 3 to 6 To explain.
[0075] Figure 3 An example of a real image and a blurred image acquired by a camera is shown.
[0076] Reference Figure 3 , the GT image 211 and the blurred image 221 may have different fields of view (FOV) depending on the physical position difference between the cameras 151 and 155. For example, the vehicle license plate 310 of the GT image 211 spans the left and right sides of the central vertical line 301 of the GT image 211, while the vehicle license plate 320 of the blurred image 221 may be located to the right of the central vertical line 301 of the blurred image 221. Figure 3 In FIG, most of the area of the vehicle in which the GT image 211 and the blurred image 221 are shown is located below the horizontal center line 305, but the field of view of the cameras 151 and 155 may be different depending on the configuration position and / or angle of each camera 151 and 155. Therefore, in order to generate learning data based on the GT image 211 and the blurred image 221, it may be necessary to perform predetermined image processing on the GT image 211 and the blurred image 221. Figures 4 to 6 , image processing operations on the GT image 211 and blurred image 221 used to generate learning data will be described.
[0077] Figure 4 An example of the operation of cropping the GT image 211 and the blurred image 221 is shown.
[0078] In one embodiment, the processor 120 may crop the GT image 211. For example, the processor 120 may crop the GT image 211 to remove the peripheral area of the GT image 211. In one embodiment, the processor 120 obtains (or extracts, or generates) the cropped GT image 420 by cropping an area 410 of a specified size (e.g., 90% of the GT image 211) of the GT image 211. In one embodiment, the processor 120 may obtain the cropped GT image 420 by cropping a specified area 410 (e.g., the central area) of the GT image 211. However, this is not limiting. In one embodiment, the processor 120 may obtain the cropped GT image 420 by cropping an area including an object. In one embodiment, the processor 120 may identify an object from the GT image 211 using an object detection (OD) model. In one embodiment, the processor 120 may obtain the cropped GT image 420 by cropping the area 410 including the identified object.
[0079] In one embodiment, the processor 120 may identify another region 430 corresponding to the cropped GT image 420 from the blurred image 221 acquired at the same time point as the GT image 211. In one embodiment, the processor 120 may identify the another region 430 by comparing the blurred image 221 with the cropped GT image 420.
[0080] For example, the comparison between the cropped GT image 420 and the blurred image 221 can be based on the differential image between the cropped GT image 420 and each specific area of the blurred image 221. In one embodiment, the specific area can be the area compared with the cropped GT image 420. In one embodiment, the specific area can be the area where the cropped GT image 420 is located when it is shifted in the blurred image 221. For example, the electronic device 101 can identify the specific area associated with the differential image with the smallest difference (or the highest matching rate) in the differential image as another area 430 of the blurred image 221. However, this is not limited to this. In one embodiment, the electronic device 101 can identify another area 430 of the blurred image 221 by comparing the cropped GT image 420 with the feature map of the blurred image 221.
[0081] For example, the comparison between the cropped GT image 420 and the blurred image 221 can be based on the distance difference between the coordinates of the feature points between the cropped GT image 420 and each specific area of the blurred image 221. For example, the electronic device 101 can identify the feature points of the cropped GT image 420. For example, the feature points may include specific positions (e.g., headlights, license plates, heads) of objects (e.g., vehicles, buildings, pedestrians) included in the cropped GT image 420. For example, the electronic device 101 can identify the feature points of the blurred image 221. For example, the electronic device 101 can identify the distance between the feature points of the cropped GT image 420 and the feature points in each specific area of the blurred image 221 that correspond to each other. In one embodiment, the electronic device 101 can identify the specific area with the shortest distance between the feature points as another area 430 of the blurred image 221.
[0082] For example, the comparison between the cropped GT image 420 and the blurred image 221 can be based on the normalized difference values of the difference images (e.g., normalized between 0 and 1) and the normalized distance of the distance of each specific region. For example, the electronic device 101 can identify the specific region with the smallest weighted sum of the normalized difference value and the normalized distance as another region 430 of the blurred image 221.
[0083] In one embodiment, the processor 120 may crop (or obtain or extract) another region 430 corresponding to the cropped GT image 420 from the blurred image 221 as a cropped blurred image 440. In one embodiment, the size of the cropped blurred image 440 is the same as the size of the cropped GT image 420. In one embodiment, the type, size, and position of the object included in the cropped GT image 420 may be the same as the type, size, and position of the object included in the cropped blurred image 440.
[0084] According to an embodiment, the processor 120 may perform image correction on the cropped GT image 420. For example, the processor 120 may increase the brightness of the cropped GT image 420.
[0085] Figure 5 An example of an operation of extracting a region of interest from the GT image 211 is shown.
[0086] Reference Figure 5, the processor 120 can identify objects from the cropped GT image 420. In one embodiment, the processor 120 can identify objects of a specified type (e.g., vehicles, traffic signs, road markings) from the cropped GT image 420 using an image segmentation model and / or an object detection model. In one embodiment, the processor 120 can identify objects from the cropped GT image 420 by adjusting the size of a mask used for object identification. In one embodiment, the processor 120 identifies objects using masks in descending order.
[0087] In one embodiment, the processor 120 may extract (or obtain) the region 510 of the recognized object (i.e., the car) from the cropped GT image 420 as the GT object image 520. In one embodiment, the processor 120 may obtain the GT object image 520 by cropping the region 510 including the recognized object.
[0088] In one embodiment, the processor 120 may identify a region of interest 530 corresponding to a text region of an object (e.g., a license plate of a vehicle, traffic guidance of a traffic sign, road guidance of a road marking) from the GT object image 520. In one embodiment, the processor 120 may identify the region of interest 530 from the GT object image 520 based on an image segmentation model. Figure 5 , the lower left GT object image 520 is shown simply by enlarging the upper right GT object image 520 , and the lower left GT object image 520 and the upper right GT object image 520 may be the same image.
[0089] In one embodiment, the processor 120 may calculate (or identify or determine) the vertices 531, 533, 535, and 537 in the region of interest 530 through graph fitting (e.g., through fitting of triangles, quadrilaterals, pentagons, or polygons). In one embodiment, the processor 120 may identify the coordinates of each vertex 531, 533, 535, and 537 from the region of interest 530 through graph fitting. For example, the electronic device 101 may identify the coordinates of each vertex 531, 533, 535, and 537 from a two-dimensional virtual coordinate system with the lower left vertex 531 of the vertices 531, 533, 535, and 537 as the origin. For example, the x-axis of the two-dimensional virtual coordinate system may be formed along the lower edge of the GT object image 520. The y-axis of the two-dimensional virtual coordinate system may be formed along the left edge of the GT object image 520.
[0090] In one embodiment, the processor 120 may mark (or set, or identify, or specify) the object region 510 and the fitted region of interest 540 in the cropped GT image 420 .
[0091] In one embodiment, the processor 120 may transform the fitted region of interest 540. For example, the processor 120 may transform the fitted region of interest 540 using a specified transformation algorithm (e.g., an algorithm for rigid-body transformation, similarity transformation, linear transformation, affine transformation, and / or perspective transformation). In one embodiment, the processor 120 may transform the fitted region of interest 540 into a license plate-shaped graphic (e.g., a rectangle), but the present invention is not limited thereto.
[0092] In one embodiment, an algorithm (or action) for identifying (or recognizing) a character string displayed on a license plate may be executed (or performed) on the (transformed) fitted region of interest 540 of the cropped GT image 420. The algorithm may include an algorithm for recognizing character strings in an image (e.g., an algorithm based on optical character recognition (OCR) functionality). In one embodiment, the character string displayed on the license plate may be identified (or recognized) based on an OCR algorithm applied to the fitted region of interest 540 of the cropped GT image 420.
[0093] In one embodiment, the processor 120 may mark the character string displayed on the license plate in the fitted region of interest 540 of the cropped GT image 420 .
[0094] Figure 6 An example of an operation of setting a region of interest in the blurred image 221 is shown.
[0095] Reference Figure 6 , the processor 120 may mark (or set, or identify, or designate) the object region 610 and the region of interest 640 in the cropped blurred image 440. In one embodiment, the processor 120 may mark (or set, or identify, or designate) the object region 610 and the region of interest 640 in the cropped blurred image 440 based on the object region 510 and the region of interest 540 in the cropped GT image 420. In one embodiment, the size and position of the object region 610 displayed in the cropped blurred image 440 may be the same as the size and position of the object region 510 displayed in the cropped GT image 420. In one embodiment, the size and position of the region of interest 640 displayed in the cropped blurred image 440 may be the same as the size and position of the region of interest 540 displayed in the cropped GT image 420.
[0096] In one embodiment, the processor 120 may mark the character string displayed on the license plate in the region of interest 640 of the cropped blurred image 440 . In one embodiment, the character string may be the same as the character string identified from the fitted region of interest 540 of the cropped GT image 420 .
[0097] In one embodiment, the processor 120 may set (or identify, or obtain) the cropped GT image 420 of the area 510 and the area of interest 540 marked with (or set, or identified, or specified) the object and the cropped blurred image 440 of the area 610 and the area of interest 640 marked with (or set, or identified, or specified) the object as a learning data set.
[0098] In one embodiment, the processor 120 may repeat the actions performed on the GT image 211 and the blurred image 221 for the next image (eg, the GT image 212 and the blurred image 222 ).
[0099] As described above, the electronic device 101 can construct learning data based on GT images and blurred images acquired at the same time point (or substantially the same time point, or corresponding time points). The electronic device 101 can construct learning data based on GT images and blurred images acquired at different shutter speeds. Thus, the electronic device 101 can construct learning data based on GT images and blurred images with minimized user intervention. In addition, the electronic device 101 can construct a large amount of learning data by minimizing user intervention. In addition, the electronic device 101 can construct an artificial intelligence model for deblurring and / or denoising based on the output results of the artificial intelligence model of the blurred image learned by the GT image. Thus, the electronic device 101 can restore the character strings within the image that cannot be recognized due to at least one of resolution, white noise and blur. In addition, the electronic device 101 can not only establish a large amount of learning data for the license plate of the vehicle, but also establish a large amount of learning data for traffic signs and / or road markings.
[0100] Figure 7 An exemplary flowchart for explaining actions of an electronic device according to an embodiment is shown.
[0101] Figure 7 You can refer to Figures 1 to 6 To explain. Figure 7 The action can be performed by the electronic device 101. Figure 7 The actions of can be performed by executing instructions stored in the memory 130. Figure 7 The actions may be performed by the processor 120 executing instructions stored in the memory 130 .
[0102] Reference Figure 7In action 710, the electronic device 101 may acquire images from the plurality of cameras 151 and 155. In one embodiment, the shutter speeds of the cameras 151 and 155 may be different. For example, the shutter speed of the camera 151 may be faster than the shutter speed of the camera 155. For example, the angles of each camera 151 and 155 may be set (or arranged) to be the same (or substantially the same, or corresponding) as each other. For example, the resolution and / or lens properties (e.g., viewing angle, focal length, autofocus, focal ratio, ISO sensitivity, or optical zoom) of each camera 151 and 155 may be set to be the same (or substantially the same, or corresponding).
[0103] In one embodiment, the electronic device 101 instructs (or commands) the cameras 151 and 155 to capture images at the same (or substantially the same, or corresponding) time points, thereby acquiring images from the cameras 151 and 155. For example, the time points being the same (or substantially the same, or corresponding) means that the capture instructions (or commands) are generated and / or transmitted within a specified offset (or frame interval depending on the frame rate).
[0104] In action 720, the electronic device 101 may identify an overlapping area between the first image and the second image. In one embodiment, the electronic device 101 may identify an overlapping area between a first image from the image acquired by the camera 151 and a second image from the image acquired by the camera 155 that was acquired at the same time as the first image.
[0105] In one embodiment, the electronic device 101 may identify the overlapping area between the first image and the second image by cropping the first image. For example, the electronic device 101 may remove the peripheral area of the first image by cropping the first image. In one embodiment, the electronic device 101 may obtain (or extract, or generate) the cropped first image by cropping an area of a specified size (e.g., 90% of the first image) of the first image. In one embodiment, the electronic device 101 may obtain the cropped first image by cropping a specified area (e.g., the central area) of the first image.
[0106] In one embodiment, the electronic device 101 may identify another region corresponding to the cropped first image from a second image acquired at the same time as the first image. In one embodiment, the electronic device 101 may identify another region by comparing the second image with the cropped first image.
[0107] For example, the comparison between the cropped first image and the second image can be based on the differential image between the cropped first image and each specific area of the second image. In one embodiment, the specific area can be the area compared with the cropped first image. In one embodiment, the specific area can be the area where the cropped first image is located when it is shifted in the second image. For example, the electronic device 101 can identify the specific area associated with the differential image with the smallest difference (or the highest matching rate) in the differential image as another area of the second image. But not limited to this. In one embodiment, the electronic device 101 can identify another area of the second image by comparing the feature maps of the cropped first image and the second image.
[0108] For example, the comparison between the cropped first image and the second image can be based on the distance difference between the coordinates of the feature points between the cropped first image and each specific area of the second image. For example, the electronic device 101 can identify the feature points of the cropped first image. For example, the feature points may include specific positions (e.g., headlights, license plates, heads) of objects (e.g., vehicles, buildings, pedestrians) included in the cropped first image. For example, the electronic device 101 can identify the feature points of the second image. For example, the electronic device 101 can identify the distance between the feature points of the cropped first image and the feature points in each specific area of the second image that correspond to each other. In one embodiment, the electronic device 101 can identify the specific area with the shortest distance between the feature points as another area of the second image.
[0109] For example, the comparison between the cropped first image and the second image can be based on the normalized difference value of each difference image (e.g., normalized between 0 and 1) and the normalized distance of each specific region. For example, the electronic device 101 can identify the specific region with the lowest weighted sum of the normalized difference value and the normalized distance as another region of the second image.
[0110] In one embodiment, the electronic device 101 may crop (or obtain or extract) another region corresponding to the cropped first image from the second image as a cropped second image. In one embodiment, the size of the cropped second image may be the same as the size of the cropped first image. In one embodiment, the type, size, and position of the object included in the cropped first image may be the same as the type, size, and position of the object included in the cropped second image.
[0111] In one embodiment, the electronic device 101 may identify the cropped first image and the cropped second image as areas where the first image and the second image overlap with each other.
[0112] In action 730 , the electronic device 101 may identify a first region of interest from the first overlapping region of the first image. In one embodiment, the electronic device 101 may sequentially perform object recognition and recognition of the first region of interest within the recognized object from the first overlapping region of the first image.
[0113] In one embodiment, the electronic device 101 can identify an object from the cropped first image. In one embodiment, the electronic device 101 can identify a specified type of object (e.g., a vehicle, a traffic sign, or a road marking) from the cropped first image using an image segmentation model and / or an object detection model. In one embodiment, the electronic device 101 can identify an object from the cropped first image by adjusting the size of a mask used for object identification. In one embodiment, the electronic device 101 identifies objects using masks in descending order.
[0114] In one embodiment, the electronic device 101 may extract (or acquire) a region of the recognized object (ie, the car) from the cropped first image as an object image. In one embodiment, the electronic device 101 may acquire the object image by cropping a region including the recognized object.
[0115] In one embodiment, the electronic device 101 can identify a region of interest corresponding to a text region of an object (e.g., a license plate of a vehicle, traffic guidance of a traffic sign, road guidance of a road marking) from an object image. In one embodiment, the electronic device 101 can identify a region of interest from an object image based on an image segmentation model. Figure 5 , the lower left object image is simply shown by enlarging the upper right object image, and the lower left object image and the upper right object image may be the same image.
[0116] In one embodiment, the electronic device 101 may calculate (or identify or determine) vertices in the region of interest by graph fitting (e.g., by fitting triangles, quadrilaterals, pentagons, or polygons). In one embodiment, the electronic device 101 may identify the coordinates of each vertex in the region of interest by graph fitting. For example, the electronic device 101 may identify the coordinates of each vertex in a two-dimensional virtual coordinate system with the lower left vertex as the origin.
[0117] In one embodiment, the electronic device 101 may identify a fitted region of interest within the region of the object in the cropped first image.
[0118] In action 740 , the electronic device 101 may identify a second ROI corresponding to the first ROI from the second overlapping area of the second image. In one embodiment, the electronic device 101 may sequentially set an object area and a second ROI within the object area from the second overlapping area of the second image.
[0119] In one embodiment, the electronic device 101 may mark (or set, or identify, or specify) the area of the object and the area of interest in the cropped second image. In one embodiment, the electronic device 101 may mark (or set, or identify, or specify) the area of the object and the area of interest in the cropped second image based on the area of the object and the area of interest in the cropped first image. In one embodiment, the size and position of the area of the object displayed in the cropped second image may be the same as the size and position of the area of the object displayed in the cropped first image. In one embodiment, the size and position of the area of interest displayed in the cropped second image may be the same as the size and position of the area of interest displayed in the cropped first image.
[0120] In action 750 , the electronic device 101 may label the object recognized from the first ROI in the second ROI. In one embodiment, the electronic device 101 may label a character string of the object recognized from the first ROI in the second ROI.
[0121] In one embodiment, the electronic device 101 may transform the fitted region of interest. For example, the electronic device 101 may transform the fitted region of interest using a specified transformation algorithm (e.g., an algorithm for rigid body transformation, similarity transformation, linear transformation, affine transformation, and / or perspective transformation). In one embodiment, the electronic device 101 may transform the fitted region of interest into a license plate-shaped graphic (e.g., a rectangle), but this is not limiting.
[0122] In one embodiment, an algorithm (or action) for identifying (or recognizing) a character string displayed in a license plate may be executed (or performed) on the (transformed) fitted region of interest of the cropped first image. The algorithm may include an algorithm for recognizing character strings in an image (e.g., an algorithm based on an optical character recognition (OCR) function). In one embodiment, the character string displayed in the license plate may be identified (or recognized) based on an OCR algorithm for the fitted region of interest of the cropped first image.
[0123] In one embodiment, the electronic device 101 may mark the character string displayed in the license plate in the region of interest of the cropped second image. In one embodiment, the character string may be the same as the character string identified from the fitted region of interest of the cropped first image.
[0124] Thereafter, the electronic device 101 may set (or identify, or acquire) the cropped first image in which the object region and the region of interest are marked and the cropped second image in which the object region and the region of interest are marked as a learning dataset.
[0125] Thereafter, the electronic device 101 may repeat the actions performed on the first image and the second image for the next image (eg, the first image and the second image).
[0126] Figure 8 An example of a block diagram representing an autonomous driving system for a vehicle according to an embodiment is shown.
[0127] according to Figure 8 The autonomous driving system 800 of a vehicle may be a deep learning network including a sensor 803, an image preprocessor 805, a deep learning network 807, an artificial intelligence (AI) processor 809, a vehicle control module 811, a network interface 813, and a communication unit 815. In various embodiments, each component may be connected through various interfaces. For example, the sensor data sensed and output by the sensor 803 may be fed to the image preprocessor 805. The sensor data processed by the image preprocessor 805 may be fed to the deep learning network 807 running in the AI processor 809. The output of the deep learning network 807 run by the AI processor 809 may be fed to the vehicle control module 811. The intermediate results of the deep learning network 807 running on the AI processor 809 may be fed to the AI processor 809. In various embodiments, the network interface 813 is connected to the electronic devices in the vehicle (e.g., Figure 1 The electronic device 101 and / or camera 151, 155) communicates with the autonomous driving system 800 to pass autonomous driving path information and / or autonomous driving control commands for autonomous driving of the vehicle to the internal block components. In one embodiment, the network interface 813 can be used to send sensor data acquired by one or more sensors 803 to an external server. In some embodiments, the autonomous driving system 800 may include more or fewer components as appropriate. For example, in some embodiments, the image preprocessor 805 may be an optional component. As another example, a post-processing component (not shown) may be included in the autonomous driving system 800 to perform post-processing on the output of the deep learning network 807 before the output is provided to the vehicle control module 811.
[0128] In some embodiments, sensor 803 may include more than one sensor. In various embodiments, sensor 803 may be attached to different locations on the vehicle. Sensor 803 may face one or more different directions. For example, sensor 803 may be attached to the front, sides, rear, and / or roof of the vehicle in various directions, such as forward-facing, rear-facing, and side-facing. In some embodiments, sensor 803 may be an image sensor such as a high dynamic range camera. In some embodiments, sensor 803 includes non-visual sensors. In some embodiments, in addition to image sensors, sensor 803 may also include radar, lidar, and / or ultrasonic sensors. In some embodiments, sensor 803 is not mounted on a vehicle with vehicle control module 811. For example, sensor 803 may be included as part of a deep learning system for capturing sensor data and may be attached to the environment or road and / or mounted on surrounding vehicles.
[0129] In some embodiments, the image preprocessor 805 can be used to preprocess sensor data from the sensor 803. For example, the image preprocessor 805 can be used to preprocess sensor data, to separate the sensor data into one or more constituent elements, and / or to post-process one or more constituent elements. In some embodiments, the image preprocessor 805 can be a GPU, a CPU, an image signal processor, or a specialized image processor. In various embodiments, the image preprocessor 805 can be a tone-mapper processor for processing high dynamic range data. In some embodiments, the image preprocessor 805 can be a constituent element of the AI processor 809.
[0130] In some embodiments, the deep learning network 807 may be a deep learning network for implementing control commands for controlling an autonomous vehicle. For example, the deep learning network 807 may be an artificial neural network such as a convolutional neural network (CNN) trained using sensor data, and the output of the deep learning network 807 is provided to the vehicle control module 811.
[0131] In some embodiments, the AI processor 809 may be a hardware processor for running the deep learning network 807. In some embodiments, the AI processor 809 is a specialized AI processor for performing inference on sensor data through a convolutional neural network. In some embodiments, the AI processor 809 may be optimized for the bit depth of the sensor data. In some embodiments, the AI processor 809 may be optimized for deep learning calculations such as calculations of neural networks including convolution, dot product, vector and / or matrix calculations. In some embodiments, the AI processor 809 may be implemented by multiple graphics processing units (GPUs) capable of efficiently performing parallel processing.
[0132] In various embodiments, the AI processor 809 performs deep learning analysis on sensor data received from one or more sensors 803 during operation of the AI processor 809 and can be coupled to a memory configured to provide the AI processor with instructions for triggering the determination of machine learning results for at least partially autonomous operation of the vehicle via an input / output interface. In some embodiments, the vehicle control module 811 processes commands output from the AI processor 809 for controlling the vehicle and can be used to translate the output of the AI processor 809 into instructions for controlling various modules of the vehicle. In some embodiments, the vehicle control module 811 is used to control the vehicle for autonomous driving. In some embodiments, the vehicle control module 811 can adjust the steering and / or speed of the vehicle. For example, the vehicle control module 811 can be used to control vehicle movement, such as deceleration, acceleration, steering, lane changes, and lane keeping. In some embodiments, the vehicle control module 811 can generate control signals for controlling vehicle lighting, such as brake lights, turn signals, and headlights. In some embodiments, the vehicle control module 811 can be used to control vehicle audio-related systems, such as the vehicle's sound system, vehicle's audio warnings, vehicle's microphone system, vehicle's horn system, etc.
[0133] In some embodiments, the vehicle control module 811 is used to control notification systems, including systems for notifying passengers and / or the driver of travel events, such as approaching a predetermined destination or a potential collision. In some embodiments, the vehicle control module 811 can be used to adjust sensors, such as the vehicle's sensors 803. For example, the vehicle control module 811 can modify the orientation of the sensors 803, change the output resolution and / or format type of the sensors 803, increase or decrease the capture rate, adjust the dynamic range, or adjust the focus of the camera. In addition, the vehicle control module 811 can turn sensors on and off individually or collectively.
[0134] In some embodiments, the vehicle control module 811 can be used to change the parameters of the image preprocessor 805 by modifying the frequency range of the filter, or by adjusting the edge detection parameters for feature and / or object detection, or by adjusting the bit depth and channels. In various embodiments, the vehicle control module 811 can be used to control the autonomous driving and / or driver assistance functions of the vehicle.
[0135] In some embodiments, the network interface 813 may function as an internal interface between the components of the autonomous driving system 800 and the communication unit 815. Specifically, the network interface 813 may be a communication interface for receiving and / or transmitting data, including voice data. In various embodiments, the network interface 813 may be connected to an external server to connect a voice call through the communication unit 815, or to receive and / or transmit text messages, or to transmit sensor data, or to update the vehicle's software through the autonomous driving system, or to update the vehicle's autonomous driving system software.
[0136] In various embodiments, the communication unit 815 may include various wireless interfaces such as cellular or WiFi. For example, the network interface 813 may be used to receive updates on operating parameters and / or instructions for the sensor 803, image preprocessor 805, deep learning network 807, AI processor 809, and vehicle control module 811 from an external server connected via the communication unit 815. For example, the machine learning model of the deep learning network 807 may be updated using the communication unit 815. According to another example, the communication unit 815 may be used to update operating parameters of the image preprocessor 805, such as image processing parameters, and / or the firmware of the sensor 803.
[0137] In another embodiment, the communication unit 815 can be used to activate communications for emergency services and emergency contacts in the event of an accident or near-accident. For example, in the event of a collision, the communication unit 815 can be used to call emergency services for assistance and can be used to notify emergency services of the details of the collision and the location of the vehicle. In various embodiments, the communication unit 815 can update or obtain an estimated time of arrival and / or destination location.
[0138] According to one embodiment, Figure 8 The illustrated autonomous driving system 800 may also be comprised of the vehicle's electronic device 101. According to one embodiment, when a user deactivates the autonomous driving function while the vehicle is operating autonomously, the AI processor 809 of the autonomous driving system 800 controls the deep learning network to learn the vehicle's autonomous driving software by inputting information related to the deactivation event as training data.
[0139] Figure 9 and Figure 10 An example of a block diagram showing an autonomous driving mobile object according to an embodiment is shown. Figure 11 An example of a gateway associated with a user device according to various embodiments is shown.
[0140] Reference Figure 9 According to this embodiment, the autonomous driving mobile body 900 may include: a control device 1000, sensing modules 904a, 904b, 904c, 904d, an engine 906 and a user interface (UI) 908.
[0141] The autonomous vehicle 900 may have an autonomous driving mode or a manual mode. For example, the vehicle may be switched from the manual mode to the autonomous driving mode or vice versa based on user input received through the user interface 908 .
[0142] When the autonomous driving mobile object 900 operates in the autonomous driving mode, the autonomous driving mobile object 900 may operate under the control of the control device 1000 .
[0143] In this embodiment, the control device 1000 may include a controller 1020 including a memory 1022 and a processor 1024 , a sensor 1010 , a wireless communication device 1030 , and an object detection device 1040 .
[0144] The object detection device 1040 may perform all or part of the functions of the distance measurement device.
[0145] That is, in this embodiment, the object detection device 1040 is a device for detecting objects located outside the autonomous driving mobile body 900. The object detection device 1040 can detect objects located outside the autonomous driving mobile body 900 and generate object information corresponding to the detection result.
[0146] The object information may include information on the presence or absence of the object, position information of the object, distance information between the moving body and the object, and relative speed information between the moving body and the object.
[0147] Objects may include lane markings, other vehicles, pedestrians, traffic signals, lights, roads, structures, speed bumps, terrain features, animals, and other objects located outside of autonomous vehicle 900. Traffic signals may include traffic lights, traffic signs, and patterns or text drawn on the road surface. Light may be generated by lights installed in other vehicles, streetlights, or sunlight.
[0148] Additionally, structures can be objects located around the road and fixed to the ground. For example, structures can include streetlights, roadside trees, buildings, utility poles, traffic lights, and bridges. Terrain can include mountains and hills.
[0149] Such an object detection device 1040 may include a camera module. The controller 1020 may extract object information from an external image captured by the camera module and cause the controller 1020 to process information related thereto.
[0150] In addition, the object detection device 1040 may also include an imaging device for identifying the external environment. In addition to LIDAR, radar, GPS devices, travel distance measurement devices (odometry), other computer vision devices, ultrasonic sensors, and infrared sensors may also be used. These devices can be selected or activated simultaneously as needed to achieve more accurate sensing.
[0151] On the other hand, the distance measurement device according to an embodiment of the present invention calculates the distance between the autonomous driving mobile object 900 and an object, is connected to the control device 1000 of the autonomous driving mobile object 900, and controls the movement of the mobile object based on the calculated distance.
[0152] For example, when there is a possibility of collision between autonomous vehicle 900 and an object based on the distance between the autonomous vehicle 900 and the object, the autonomous vehicle 900 may control the brakes to slow down or stop the vehicle. For another example, when the object is moving, the autonomous vehicle 900 may control its speed to maintain a predetermined distance from the object.
[0153] The distance measurement device according to an embodiment of the present invention may be configured as a module within the control device 1000 of the autonomous driving mobile object 900. That is, the memory 1022 and the processor 1024 of the control device 1000 may implement the collision avoidance method according to the present invention in a software manner.
[0154] In addition, the sensor 1010 can obtain various sensing information by connecting the internal / external environment of the mobile object to the sensing modules 904a, 904b, 904c, and 904d. The sensor 1010 may include a posture sensor (e.g., a yaw sensor, a roll sensor, a pitch sensor, a collision sensor, a wheel sensor), a speed sensor, a tilt sensor, a weight sensing sensor, a heading sensor, a gyro sensor, a position module, a mobile forward / backward sensor, a battery sensor, a fuel sensor, a tire sensor, a steering sensor based on steering wheel rotation, a mobile internal temperature sensor, a mobile internal humidity sensor, an ultrasonic sensor, an illumination sensor, an accelerator pedal position sensor, a brake pedal position sensor, etc.
[0155] Thus, the sensor 1010 can obtain sensing signals for mobile body posture information, mobile body collision information, mobile body direction information, mobile body position information (GPS information), mobile body angle information, mobile body speed information, mobile body acceleration information, mobile body slope information, mobile body forward / backward information, battery information, fuel information, tire information, mobile body light information, mobile body internal temperature information, mobile body internal humidity information, steering wheel rotation angle, mobile body external illuminance, pressure applied to the accelerator pedal, pressure applied to the brake pedal, etc.
[0156] In addition, in addition to this, sensor 1010 can also include an accelerator pedal sensor, a pressure sensor, an engine speed sensor, an air flow sensor, an intake air temperature sensor, a water temperature sensor, a throttle position sensor, a top dead center sensor, a crank angle sensor, etc.
[0157] As described above, the sensor 1010 may generate moving body state information based on the sensing data.
[0158] The wireless communication device 1030 is configured to implement wireless communication between the autonomous driving mobile bodies 900. For example, the autonomous driving mobile body 900 can communicate with the user's mobile phone, or other wireless communication devices 1030, other mobile bodies, central devices (traffic control devices), servers, etc. The wireless communication device 1030 can send and receive wireless signals according to the connection wireless protocol. The wireless communication protocol can be Wi-Fi, Bluetooth, Long-Term Evolution (LTE), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Global Systems for Mobile Communications (GSM), but the communication protocol is not limited thereto.
[0159] In addition, in this embodiment, the autonomous driving mobile body 900 can also realize communication between mobile bodies through the wireless communication device 1030. That is, the wireless communication device 1030 can communicate with other mobile bodies on the road through vehicle-to-vehicle (V2V) communication. The autonomous driving mobile body 900 can send and receive information such as driving warnings and traffic information through inter-vehicle communication, and can also request information from other mobile bodies or receive requests from other mobile bodies. For example, the wireless communication device 1030 can perform V2V communication through a dedicated short-range communication (DSRC) device or a cellular V2V (Cellular-V2V, C-V2V) device. In addition, in addition to communication between vehicles, communication between the vehicle and other objects (such as electronic devices carried by pedestrians, etc.) can also be realized through the wireless communication device 1030.
[0160] In addition, the wireless communication device 1030 can obtain information generated from infrastructure located on the road (traffic lights, CCTV, road side units (RSU), evolved Node B (eNode B), etc.) or various mobile bodies (Mobility) including other autonomous driving (Autonomous Driving) / non-autonomous driving (Non-Autonomous Driving) vehicles through a non-terrestrial network (Non-Terrestrial Network) that is not a terrestrial network (Terrestrial Network), and use the generated information as information for performing autonomous driving of the autonomous driving mobile body 900.
[0161] For example, the wireless communication device 1030 can conduct wireless communication with the low earth orbit (LEO) satellite system, medium earth orbit (MEO) satellite system, geostationary orbit (GEO) satellite system, high altitude platform (HAP) system, etc. that constitute the non-ground network through a non-ground network dedicated antenna installed on the autonomous driving mobile body 900.
[0162] For example, the wireless communication device 1030 can perform wireless communications with various platforms constituting the 5th Generation New Radio Non-Terrestrial Network (5GNR NTN) according to a wireless access standard corresponding to the 5th Generation New Radio Non-Terrestrial Network (5GNR NTN) standard specification currently discussed in 3GPP, but is not limited thereto.
[0163] In this embodiment, the controller 1020 can consider various information such as the location, current time, available power, etc. of the autonomous driving mobile body 900, select a platform that can appropriately perform NTN communication, and control the wireless communication device 1030 to perform wireless communication with the selected platform.
[0164] In this embodiment, the controller 1020 is a unit that controls the overall operation of each unit within the autonomous driving mobile body 900. It can be configured by the manufacturer of the mobile body during manufacturing or can be additionally configured after manufacturing to perform the autonomous driving function. Alternatively, the controller 1020 configured at the time of manufacturing can be upgraded to include a configuration for continuously performing additional functions. The controller 1020 can also be called an electronic control unit (ECU).
[0165] The controller 1020 collects various data from the connected sensors 1010, object detection device 1040, wireless communication device 1030, and the like, and transmits control signals based on the collected data to the sensors 1010, engine 906, user interface 908, wireless communication device 1030, and object detection device 1040 included in other components of the mobile body. Although not shown, the control signals may also be sent to a throttle device, a braking system, a steering device, or a navigation device related to the travel of the mobile body.
[0166] In this embodiment, the controller 1020 can control the engine 906. For example, the engine 906 can be controlled to sense the speed limit of the road on which the autonomous driving mobile body 900 is traveling and control the driving speed not to exceed the speed limit, or the engine 906 can be controlled to increase the driving speed of the autonomous driving mobile body 900 within a range not exceeding the speed limit.
[0167] Furthermore, if autonomous vehicle 900 approaches or deviates from a lane line during its travel, controller 1020 determines whether this approach or deviation is due to a normal driving situation or another driving situation. Based on the determination, controller 1020 controls engine 906 to control the vehicle's travel. Specifically, autonomous vehicle 900 can detect lane lines formed on both sides of the lane in which it is traveling. In this case, controller 1020 determines whether autonomous vehicle 900 approaches or deviates from a lane line. If it determines that autonomous vehicle 900 approaches or deviates from a lane line, controller 1020 determines whether this movement is due to a normal driving situation or another driving situation. An example of a normal driving situation is a situation where the vehicle needs to change lanes. An example of another driving situation is a situation where the vehicle does not need to change lanes. If controller 1020 determines that autonomous vehicle 900 approaches or deviates from a lane line without requiring a lane change, controller 1020 controls the vehicle's travel so that autonomous vehicle 900 does not deviate from the lane line and instead travels normally in that lane.
[0168] If there are other moving objects or obstacles in front of the moving object, the engine 906 or the braking system can be controlled to slow the moving object. In addition to speed, the trajectory, driving path, and steering angle can also be controlled. Alternatively, the controller 1020 can generate the necessary control signals based on the recognition of other external environmental information such as the moving object's lane markings, driving signals, etc. to control the movement of the moving object.
[0169] In addition to generating its own control signals, the controller 1020 can also communicate with surrounding mobile bodies or a central server, and send commands for controlling peripheral devices through the received information, thereby controlling the travel of the mobile body.
[0170] Furthermore, accurate recognition of a moving object or lane marking according to this embodiment may be difficult to achieve when the position or viewing angle of the camera module changes. Therefore, to prevent this, the controller 1020 may further generate a control signal to control the execution of camera module calibration. Therefore, in this embodiment, the controller 1020 generates a calibration control signal to the camera module so that even if the camera module's mounting position changes due to vibration or impact generated by the movement of the autonomous driving moving object 900, the camera module's normal mounting position, orientation, and viewing angle can be maintained. When the difference between the pre-stored initial mounting position, orientation, and viewing angle information of the camera module and the initial mounting position, orientation, and viewing angle information of the camera module measured during driving of the autonomous driving moving object 900 exceeds a threshold, the controller 1020 may generate a control signal to execute camera module calibration.
[0171] In this embodiment, the controller 1020 may include a memory 1022 and a processor 1024. The processor 1024 may execute software stored in the memory 1022 according to control signals from the controller 1020. Specifically, the controller 1020 stores data and commands for executing the lane detection method according to the present invention in the memory 1022. The commands may be executed by the processor 1024 to implement one or more methods disclosed herein.
[0172] In this case, the memory 1022 may be stored in a non-volatile recording medium that can be executed by the processor 1024. The memory 1022 may store software and data through appropriate internal or external devices. The memory 1022 may be composed of a memory device connected to RAM, ROM, a hard disk, and an adapter (dongle).
[0173] The memory 1022 may store at least an operating system (OS), user applications, and executable commands. The memory 1022 may also store application data and arrange data structures.
[0174] Processor 1024 may be a microprocessor or suitable electronic processor, such as a controller, microcontroller, or state machine.
[0175] Processor 1024 may be implemented as a combination of computing devices, and the computing device may be a digital signal processor, a microprocessor, or a suitable combination thereof.
[0176] On the other hand, the autonomous driving vehicle 900 may also include a user interface 908 for providing user input to the control device 1000. The user interface 908 may allow the user to input information through appropriate interaction. For example, it may be implemented through a touch screen, a keyboard, action buttons, etc. The user interface 908 transmits the input or command to the controller 1020, and the controller 1020 may execute the control action of the vehicle in response to the input or command.
[0177] In addition, the user interface 908 may enable the autonomous vehicle 900 to communicate with a device external to the autonomous vehicle 900 via the wireless communication device 1030. For example, the user interface 908 may be linked to a mobile phone, tablet computer, or other computer device.
[0178] Furthermore, in this embodiment, the autonomous vehicle 900 is described as including an engine 906, but may also include other types of propulsion systems. For example, the vehicle may be powered by electricity, hydrogen, or a hybrid system combining these. Therefore, the controller 1020 may include propulsion mechanisms corresponding to the propulsion system of the autonomous vehicle 900 and provide control signals corresponding to each propulsion mechanism.
[0179] In the following, reference will be made to Figure 10 The detailed configuration of the control device 1000 according to this embodiment will be described in more detail.
[0180] The control device 1000 includes a processor 1024. The processor 1024 can be a general-purpose single-chip or multi-chip microprocessor, a dedicated microprocessor, a microcontroller, a programmable gate array, etc. The processor can also be referred to as a CPU. In addition, in this embodiment, the processor 1024 can be used as a combination of multiple processors.
[0181] Furthermore, the control device 1000 further includes a memory 1022. The memory 1022 may also be any electronic component capable of storing electronic information. In addition to a single memory, the memory 1022 may also include a combination of memories 1022.
[0182] Data 1022b and instructions 1022a for executing the distance measurement method of the distance measurement device according to the present invention may also be stored in the memory 1022. When the processor 1024 executes the instructions 1024a, all or part of the instructions 1024a and the data 1024b required for executing the command may also be loaded onto the processor 1024.
[0183] The control device 1000 may also include a transmitter 1030a, a receiver 1030b, or a transceiver 1030c to allow transmission and reception of signals. One or more antennas 1032a, 1032b may also be electrically connected to the transmitter 1030a, the receiver 1030b, or each transceiver 1030c, and additional antennas may also be included.
[0184] The control device 1000 may further include a digital signal processor (DSP) 1070. Through the DSP 1070, the mobile object can quickly process digital signals.
[0185] The control device 1000 may also include a communication interface 1080. The communication interface 1080 may also include one or more ports and / or communication modules for connecting other devices to the control device 1000. The communication interface 1080 may allow a user to interact with the control device 1000.
[0186] The various components of the control device 1000 may also be connected together via one or more buses 1090 , which may also include a power bus, a control signal bus, a status signal bus, a data bus, etc. Under the control of the processor 1024 , the various components may communicate information with each other via the bus 1090 and perform desired functions.
[0187] On the other hand, in various embodiments, the control device 1000 can be associated with a gateway to communicate with the secure cloud. Figure 11 , the control device 1000 may be associated with a gateway 1105 for providing information acquired from at least one of the components (1101 to 1104) of the vehicle 1100 to the security cloud 1106. For example, the gateway 1105 may be included in the control device 1000. As another example, the gateway 1105 may also be configured as an additional device within the vehicle 1100 that is distinguished from the control device 1000. The gateway 1105 communicatively connects the software management cloud 1109, the security cloud 1106, and the network within the vehicle 1100 protected by the in-vehicle security software 1110, which have different networks.
[0188] For example, component 1101 may be a sensor. For example, the sensor may be used to obtain information related to at least one of the state of vehicle 1100 and the state of the surrounding area of vehicle 1100. For example, component 1101 may include sensor 1010.
[0189] For example, component 1102 may be an ECU, which may be used for engine control, transmission control, airbag control, and tire pressure management.
[0190] For example, component 1103 may be an instrument cluster. For example, the instrument cluster may represent a panel located in front of the driver's seat in a dashboard. For example, the instrument cluster may be configured to display information necessary for driving to the driver (or passenger). For example, the instrument cluster may be configured to display at least one of a visual element indicating the engine's revolutions per minute (RPM), a visual element indicating the speed of vehicle 1100, a visual element indicating the remaining fuel level, a visual element indicating a gear position, and a visual element indicating information acquired through component 1101.
[0191] For example, component 1104 may be an in-vehicle communication (telematics) device. For example, the in-vehicle communication device may represent a device that provides various mobile communication services (e.g., location information, safe driving, etc.) within the vehicle 1100 by combining wireless communication technology and GPS technology. For example, the in-vehicle communication device may be used to connect the vehicle 1100 with the driver, the cloud (e.g., the safety cloud 1106) and / or the surrounding environment. For example, the in-vehicle communication device may be configured to support high bandwidth and low latency for technologies specified in the 5G NR specification (e.g., 5G NR's V2X technology, 5G NR's NTN technology). For example, the in-vehicle communication device may be configured to support autonomous driving of the vehicle 1100.
[0192] For example, gateway 1105 can be used to connect the network within vehicle 1100 with a software management cloud 1109 and a security cloud 1106, which are external networks. For example, software management cloud 1109 can be used to update or manage at least one software required for driving and managing vehicle 1100. For example, software management cloud 1109 can be linked with in-car security software 1110 installed in the vehicle. For example, in-car security software 1110 can be used to provide security functions within vehicle 1100. For example, in-car security software 1110 can use an encryption key obtained from an external authorized server to encrypt data sent and received over the in-vehicle network to encrypt the in-vehicle network. In various embodiments, the encryption key used by in-vehicle security software 1110 can be generated in accordance with vehicle identification information (such as the license plate or vehicle identification number (VIN)) or information uniquely assigned to each user (e.g., user identification information).
[0193] In various embodiments, the gateway 1105 can transmit data encrypted by the in-vehicle security software 1110 to the software management cloud 1109 and / or the security cloud 1106 based on the encryption key. The software management cloud 1109 and / or the security cloud 1106 decrypt the data encrypted by the in-vehicle security software 1110's encryption key using a decryption key, thereby identifying the vehicle or user from which the data was received. For example, the decryption key is a unique key corresponding to the encryption key, so the software management cloud 1109 and / or the security cloud 1106 can identify the data sender (e.g., the vehicle or user) based on the data decrypted by the decryption key.
[0194] For example, the gateway 1105 is configured to support in-vehicle security software 1110 and may be associated with the control device 1000. For example, the gateway 1105 may be associated with the control device 1000 to support a connection between a client device 1107 connected to the security cloud 1106 and the control device 1000. As another example, the gateway 1105 may be associated with the control device 1000 to support a connection between a third-party cloud 1108 connected to the security cloud 1106 and the control device 1000. However, this is not limiting.
[0195] In various embodiments, gateway 1105 can be used to connect a software management cloud 1109 for managing the operating software of vehicle 1100 to vehicle 1100. For example, software management cloud 1109 can monitor whether the operating software of vehicle 1100 needs to be updated, and when it is detected that the operating software of vehicle 1100 needs to be updated, provide data for updating the operating software of vehicle 1100 through gateway 1105. As another example, software management cloud 1109 can receive a user request to update the operating software of vehicle 1100 from vehicle 1100 through gateway 1105, and provide data for updating the operating software of vehicle 1100 based on the received user request. However, the present invention is not limited to this.
[0196] Figure 12 is a diagram for explaining the operation of an electronic device for training a neural network based on a learning data set according to one embodiment.
[0197] Reference Figure 12 The described actions may be performed by the above electronic devices (for example: Figure 1 Executed by electronic device 101).
[0198] Reference Figure 12In action 1202, an electronic device according to an embodiment may obtain a learning data set. The electronic device may obtain a learning data set for supervised learning. The learning data may include input data and ground truth data pairs corresponding to the input data. The measured data may represent output data to be obtained from a neural network that receives the measured data pairs as input data. The measured data may be obtained by the electronic device.
[0199] For example, when training a neural network to recognize an image, the learning data may include information related to the image and one or more subjects included in the image. The information may include a category (category or class) of a subject that can be identified by the image. The information may include the position, width, height and / or size of a visual object corresponding to the subject within the image. The learning data set identified by action 1202 may include multiple learning data pairs. In the example of training a neural network to recognize an image, the learning data set identified by the electronic device may include multiple images and measured data corresponding to each of the multiple images.
[0200] Reference Figure 12 In action 1204, the electronic device according to an embodiment may train a neural network based on a learning data set. In an embodiment of training a neural network based on supervised learning, the electronic device may input input data included in the learning data into an input layer of the neural network. Figure 13 An example of a neural network including the input layer is described. The electronic device can obtain output data of the neural network corresponding to the input data from the output layer of the neural network that receives the input data through the input layer.
[0201] In one embodiment, the training of action 1204 may be performed based on the difference between the output data and the measured data included in the learning data and corresponding to the input data. For example, the electronic device may adjust one or more parameters related to the neural network based on a gradient descent algorithm (e.g., see later). Figure 13 The electronic device may adjust the one or more parameters to reduce the difference. The electronic device may adjust the one or more parameters to reduce the difference. The electronic device may adjust the one or more parameters to reduce the difference. The electronic device may use a function defined as a function for evaluating the performance of the neural network (such as a cost function) to perform the tuning of the neural network based on the output data. The difference between the output data and the measured data may be included as an example of the cost function.
[0202] Reference Figure 12In action 1206, the electronic device according to one embodiment may identify whether valid output data is output from the neural network trained in action 1204. Valid output data may mean that the difference (or cost function) between the output data and the measured data satisfies the conditions set for using the neural network. For example, when the average value and / or the maximum value of the difference between the output data and the measured data is below a specified threshold, the electronic device may determine that valid output data is output from the neural network.
[0203] When no valid output data is output from the neural network (action 1206 is No), the electronic device may repeatedly perform training of the neural network based on action 1204. The embodiment is not limited thereto, and the electronic device may repeatedly perform actions 1202 and 1204.
[0204] When valid output data is obtained from the neural network (yes in action 1206), the electronic device according to one embodiment may use the trained neural network according to action 1208. For example, the electronic device may input other input data, which is different from the input data input to the neural network, as learning data to the neural network. The electronic device may use the output data obtained from the neural network that received the other input data as the result of inference performed on the other input data by the neural network.
[0205] Figure 13 is a block diagram of an electronic device according to an embodiment.
[0206] Figure 13 The electronic device 1300 may include the above-mentioned electronic device 101.
[0207] For example, refer to Figure 12 The action described can be Figure 13 The electronic device 1300 and / or Figure 13 Executed by processor 1310.
[0208] Reference Figure 13 , the processor 1310 of the electronic device 1300 may perform calculations related to the neural network 1330 stored in the memory 1320. The processor 1310 may include at least one of a CPU, a GPU, and an NPU. The NPU may be implemented as a chip separate from the CPU, or may be integrated into a chip such as a CPU in the form of a system on a chip. The NPU integrated into the CPU may be referred to as a neural core and / or an AI accelerator.
[0209] Reference Figure 13, the processor 1310 may identify a neural network 1330 stored in the memory 1320. The neural network 1330 may include a combination of an input layer 1332, one or more hidden layers 1334 (or intermediate layers), and an output layer 1336. The layers described above (e.g., the input layer 1332, one or more hidden layers 1334, and the output layer 1336) may include multiple nodes. The number of hidden layers 1334 may vary depending on the embodiment, and a neural network 1330 including multiple hidden layers 1334 may be referred to as a deep neural network. The act of training the deep neural network may be referred to as deep learning.
[0210] In one embodiment, when neural network 1330 has a feedforward neural network structure, a first node included in a specific layer may be connected to all second nodes included in other layers before the specific layer. Parameters stored for neural network 1330 in memory 1320 may include weights assigned to connections between the second nodes and the first node. In neural network 1330 having a feedforward neural network structure, the value of the first node may correspond to a weighted sum of the values assigned to the second node based on the weights assigned to the connections connecting the second node and the first node.
[0211] In one embodiment, when neural network 1330 has a convolutional neural network structure, a first node included in a specific layer may correspond to a weighted sum of a portion of second nodes included in other layers prior to the specific layer. A portion of the second nodes corresponding to the first node may be identified by a filter corresponding to the specific layer. Parameters stored in memory 1320 for neural network 1330 may include weighted values representing the filter. The filter may include one or more nodes in the second node to be used to calculate the weighted sum of the first node, and a weighted value corresponding to each of the one or more nodes.
[0212] According to an embodiment, the processor 1310 of the electronic device 1300 can train the neural network 1330 using the learning data set 1340 stored in the memory 1320. Based on the learning data set 1340, the processor 1310 can perform a reference Figure 12 The described actions adjust one or more parameters stored in memory 1320 for neural network 1330.
[0213] According to one embodiment, the processor 1310 of the electronic device 1300 can perform object detection, object recognition, and / or object classification using a neural network 1330 trained based on a learning data set 1340. The processor 1310 can input an image (or video) acquired through the camera 1350 into the input layer 1332 of the neural network 1330. Based on the input layer 1332 to which the image is input, the processor 1310 can sequentially acquire the values of the nodes of the layers included in the neural network 1330 and acquire a set of values (e.g., output data) of the nodes of the output layer 1336. The output data can be used as a result of reasoning about the information included in the image using the neural network 1330. The embodiment is not limited thereto, and the processor 1310 can input an image (or video) acquired from an external electronic device connected to the electronic device 1300 via the communication circuit 1360 into the neural network 1330.
[0214] In one embodiment, the neural network 1330 trained to process an image can be used to identify (object detection) an area corresponding to a subject within the image, and / or identify (object recognition and / or object classification) the category of the subject presented within the image. For example, the electronic device 1300 can use the neural network 1330 to segment the area corresponding to the subject within the image based on a rectangular form such as a bounding box. For example, the electronic device 1300 can use the neural network 1330 to identify at least one category matching the subject from a plurality of specified categories.
[0215] Figure 14 and Figure 15 A conventional truck 10 is shown.
[0216] Figure 14 The tractor 12 is shown in a state where it is not connected to the trailer 14 .
[0217] Figure 15 The tractor 12 is shown coupled to a trailer 14. In an embodiment of the present invention, the trailer 14 is selectively coupled via a steerable wheel hitch 16 carried by the tractor 12 and secured to a kingpin 18 secured to the trailer 14 in a known manner.
[0218] This manual Figure 14 The trailer 20 shown is illustrated as a “semi-trailer” form, but this is for the convenience of explanation, and it should not be understood that the embodiments of the present invention are only applicable to the “semi-trailer” form.
[0219] As described above, the electronic device 101 may include a communication circuit 110. The electronic device 101 may include a memory 130 storing instructions. The electronic device 101 may include at least one processor 120 operatively connected to the communication circuit 110 and the memory 130. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire first images 211 to 220 from the first camera 151 via the communication circuit 110. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151 via the communication circuit 110. When executed by the processor 120, the instructions may cause the electronic device 101 to identify, from among the first images 211 to 220 and the second images 221 to 230, the first image 211 and the second image 221 acquired at corresponding time points. When executed by the processor 120, the instructions may cause the electronic device 101 to set a first region of interest 540 within an object of a specified type in the first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set a second region of interest 640 in the second image 221 at the same location as the first region of interest 540. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the first image 211 in which the first region of interest 540 is set and the second image 221 in which the second region of interest 640 is set as learning data.
[0220] In one embodiment, the second camera 155 may be arranged at an angle corresponding to that of the first camera 151. In one embodiment, the lens properties of the second camera 155 may be the same as those of the first camera 151. When executed by the processor 120, the instructions may cause the electronic device 101 to instruct the first camera 151 and the second camera 155 to capture images at corresponding time points.
[0221] When executed by the processor 120, the instructions may cause the electronic device 101 to instruct the first camera 151 and the second camera 155 to capture first images 211 to 220 and second images 221 to 230 within a specified time period after capturing images at corresponding time points. When executed by the processor 120, the instructions may cause the electronic device 101 to identify an image (e.g., image 225) captured at a time point corresponding to the reference image (e.g., image 215) among the first images 211 to 220 by comparing the first image 211 to 220 with each of the second images 221 to 230. When executed by the processor 120, the instructions may cause the electronic device 101 to identify, from the second images 221 to 230, the second image 221 having the same frame difference as the frame difference between the first image 211 and the reference image (e.g., image 215).
[0222] When executed by the processor 120, the instructions may cause the electronic device 101 to crop a region 410 of a specified size from the first image 211 to generate a cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to crop another region 430 of a specified size from the second image 221 based on the cropped first image 211 to generate a cropped second image 221.
[0223] When executed by the processor 120, the instructions may cause the electronic device 101 to identify normalized differences between the cropped first image 211 and multiple regions of the second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to identify normalized distances between feature points of the cropped first image 211 and feature points of each of multiple regions of the second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to crop a region having a minimum weighted average value between the normalized differences and the normalized distances to generate the cropped second image 221.
[0224] When executed by the processor 120, the instructions may cause the electronic device 101 to recognize an object of a specified type from the cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set the first region of interest 540 within the recognized object in the cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set the position of the object in the cropped first image 211 in the cropped second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to set the second region of interest 640 in the cropped second image 221 at the same position as the first region of interest 540 in the cropped first image 211.
[0225] When executed by the processor 120, the instructions may enable the electronic device 101 to identify a character string in the first region of interest 540 based on a character string recognition algorithm. When executed by the processor 120, the instructions may enable the electronic device 101 to mark the identified character string in the first region of interest 540 and the second region of interest 640.
[0226] When executed by the processor 120, the instructions may cause the electronic device 101 to transform the first region of interest 540 into a specified graphic based on a specified transformation algorithm. When executed by the processor 120, the instructions may cause the electronic device 101 to recognize the string from the first region of interest 540 transformed into the specified graphic based on the string recognition algorithm.
[0227] When executed by the processor 120, the instructions may cause the electronic device 101 to identify whether the electronic device 101 has moved based on differences between consecutive images of the first images 211 to 220. When executed by the processor 120, the instructions may cause the electronic device 101 to delete the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has not moved. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the learning data based on the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has moved.
[0228] As described above, the method can be performed in the electronic device 101 including the communication circuit 110. The method may include acquiring first images 211 to 220 from the first camera 151 via the communication circuit 110. The method may include acquiring second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151 via the communication circuit 110. The method may include identifying, from the first images 211 to 220 and the second images 221 to 230, first images 211 and second images 221 acquired at corresponding time points. The method may include setting a first region of interest 540 within an object of a specified type in the first image 211. The method may include setting a second region of interest 640 in the second image 221 at the same location as the first region of interest 540. The method may include generating the first image 211 having the first region of interest 540 and the second image 221 having the second region of interest 640 as learning data.
[0229] The method may include, after instructing the first camera 151 and the second camera 155 to capture images at corresponding time points, acquiring first images 211 to 220 and second images 221 to 230 within a specified time period. The method may include, by comparing a reference image (e.g., image 215) among the first images 211 to 220 with each of the second images 221 to 230, identifying an image acquired at a time point corresponding to the reference image (e.g., image 215). The method may include, from the second images 221 to 230, identifying a second image 221 having the same frame difference as the frame difference between the first image 211 and the reference image (e.g., image 215).
[0230] The method may include an act of generating a cropped first image 211 by cropping a region 410 of a specified size from the first image 211. The method may include an act of cropping another region 430 of a specified size from the second image 221 based on the cropped first image 211 to generate a cropped second image 221.
[0231] The method may include an act of identifying a normalized difference between the cropped first image 211 and a plurality of regions of the second image 221. The method may include an act of identifying a normalized distance between a feature point of the cropped first image 211 and a feature point of each of a plurality of regions of the second image 221. The method may include an act of generating the cropped second image 221 by cropping a region having a lowest weighted average between the normalized difference and the normalized distance.
[0232] The method may include an act of identifying an object of a specified type from the cropped first image 211. The method may include an act of setting the first region of interest 540 within the identified object in the cropped first image 211. The method may include an act of setting the location of the object in the cropped first image 211 in the cropped second image 221. The method may include an act of setting the second region of interest 640 in the cropped second image 221 at the same location as the first region of interest 540 in the cropped first image 211.
[0233] The method may include an action of identifying whether the electronic device 101 has moved based on a difference between consecutive images among the first images 211 to 220. The method may include an action of deleting the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has not moved. The method may include an action of generating the learning data based on the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has moved.
[0234] As described above, a non-transitory computer-readable storage medium may store a program including instructions. When executed by the processor 120 of the electronic device 101 including the communication circuit 110, the instructions may cause the electronic device 101 to acquire first images 211 to 220 from the first camera 151 via the communication circuit 110. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151 via the communication circuit 110. When executed by the processor 120, the instructions may cause the electronic device 101 to identify, from the first images 211 to 220 and the second images 221 to 230, the first image 211 and the second image 221 acquired at corresponding time points. When executed by the processor 120, the instructions may cause the electronic device 101 to set a first region of interest 540 within an object of a specified type in the first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set a second region of interest 640 in the second image 221 at the same position as the first region of interest 540. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the first image 211 in which the first region of interest 540 is set and the second image 221 in which the second region of interest 640 is set as learning data.
[0235] When executed by the processor 120, the instructions may cause the electronic device 101 to crop a region 410 of a specified size from the first image 211 to generate a cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to crop another region 430 of a specified size from the second image 221 based on the cropped first image 211 to generate a cropped second image 221.
[0236] When executed by the processor 120, the instructions may cause the electronic device 101 to identify normalized differences between the cropped first image 211 and multiple regions of the second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to identify normalized distances between feature points of the cropped first image 211 and feature points of each of multiple regions of the second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the cropped second image 221 by cropping a region having a minimum weighted average value between the normalized differences and the normalized distances.
[0237] When executed by the processor 120, the instructions may cause the electronic device 101 to recognize an object of a specified type from the cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set the first region of interest 540 within the recognized object in the cropped first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set the position of the object in the cropped first image 211 in the cropped second image 221. When executed by the processor 120, the instructions may cause the electronic device 101 to set the second region of interest 640 in the cropped second image 221 at the same position as the first region of interest 540 in the cropped first image 211.
[0238] When executed by the processor 120, the instructions may cause the electronic device 101 to identify whether the electronic device 101 has moved based on differences between consecutive images of the first images 211 to 220. When executed by the processor 120, the instructions may cause the electronic device 101 to delete the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has not moved. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the learning data based on the first images 211 to 220 and the second images 221 to 230 if the electronic device 101 has moved.
[0239] As described above, the electronic device 101 may include cameras 151 and 155. The electronic device 101 may include a memory 130 storing instructions. The electronic device 101 may include at least one processor 120 operably connected to the cameras 151 and 155 and the memory 130. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire first images 211 to 220 from the first camera 151. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151. When executed by the processor 120, the instructions may cause the electronic device 101 to identify, from the first images 211 to 220 and the second images 221 to 230, a first image 211 and a second image 221 acquired at corresponding time points. When executed by the processor 120, the instructions may cause the electronic device 101 to set a first region of interest 540 within an object of a specified type in the first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set a second region of interest 640 in the second image 221 at the same position as the first region of interest 540. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the first image 211 in which the first region of interest 540 is set and the second image 221 in which the second region of interest 640 is set as learning data.
[0240] As described above, the method can be performed in the electronic device 101 including cameras 151 and 155. The method may include acquiring first images 211 to 220 from the first camera 151. The method may include acquiring second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151. The method may include identifying, from the first images 211 to 220 and the second images 221 to 230, first images 211 and second images 221 acquired at corresponding time points. The method may include setting a first region of interest 540 within an object of a specified type in the first image 211. The method may include setting a second region of interest 640 in the second image 221 at the same location as the first region of interest 540. The method may include generating the first image 211 having the first region of interest 540 and the second image 221 having the second region of interest 640 as learning data.
[0241] As described above, a non-transitory computer-readable storage medium may store a program including instructions. When executed by the processor 120 of the electronic device 101 including the cameras 151 and 155, the instructions may cause the electronic device 101 to acquire first images 211 to 220 from the first camera 151. When executed by the processor 120, the instructions may cause the electronic device 101 to acquire second images 221 to 230 from the second camera 155 having a slower shutter speed than the first camera 151. When executed by the processor 120, the instructions may cause the electronic device 101 to identify, from the first images 211 to 220 and the second images 221 to 230, the first image 211 and the second image 221 acquired at corresponding time points. When executed by the processor 120, the instructions may cause the electronic device 101 to set a first region of interest 540 within an object of a specified type in the first image 211. When executed by the processor 120, the instructions may cause the electronic device 101 to set a second region of interest 640 in the second image 221 at the same position as the first region of interest 540. When executed by the processor 120, the instructions may cause the electronic device 101 to generate the first image 211 in which the first region of interest 540 is set and the second image 221 in which the second region of interest 640 is set as learning data.
[0242] It should be understood that an embodiment of this paper and the terms used therein are not intended to limit the technical features described herein to a specific embodiment, but rather include various modifications, equivalents or substitutes of the embodiment. In conjunction with the description of the accompanying drawings, similar figure numerals may be used for similar or related constituent elements. Unless otherwise clearly indicated in the relevant context, the singular form of the noun corresponding to a certain project may include one or more of the above-mentioned projects. Herein, phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B or C", "at least one of A, B and C" and "at least one of A, B or C" may include any one of the projects listed together in the corresponding phrases or any possible combination thereof. Terms such as "first" or "second" may be simply used to distinguish one constituent element from another constituent element, and these constituent elements are not limited in other respects (such as importance or order). If a certain component (for example, a first component) is referred to as being “coupled” or “connected” to another component (for example, a second component), with or without the term “functionally” or “communicatively”, it means that the certain component can be connected to the other component directly (for example, by wire), wirelessly, or through a third component.
[0243] In the specific embodiments of the present disclosure described above, the constituent elements included in the present disclosure are expressed as singular or plural depending on the specific embodiment proposed. However, the singular or plural expression is selected for convenience of explanation and is appropriate to the proposed situation. The present disclosure is not limited to singular or plural constituent elements. Even a constituent element expressed as plural can be constituted as singular, and even a constituent element expressed as singular can be constituted as plural.
[0244] According to an embodiment, more than one constituent element or action in the aforementioned corresponding constituent element can be omitted, or more than one other constituent element or action can be added.Alternatively or additionally, a plurality of constituent elements (e.g., module or program) can be integrated into one constituent element.In this case, the constituent element that is integrated can perform one or more functions of each constituent element in a plurality of the constituent elements identically or similarly to the function performed by the corresponding constituent element in a plurality of the constituent elements before being integrated.According to various embodiments, the action performed by module, program or other constituent elements can be performed sequentially, in parallel, iteratively or heuristically, or more than one action in the action can be performed in different orders, or be omitted, or more than one other action can be added.
[0245] On the other hand, although specific embodiments have been described in the detailed description of the present disclosure, various modifications can of course be made without departing from the scope of the present disclosure.
Claims
1. An electronic device, wherein: include: Communication circuits, memory for storing instructions, and at least one processor operatively connected to the communication circuitry and the memory; When executed by the processor, the instructions cause the electronic device to perform the following actions: Acquire a plurality of first images from the first camera through the communication circuit, acquiring, through the communication circuit, a plurality of second images from a second camera having a shutter speed slower than that of the first camera; identifying, from the plurality of first images and the plurality of second images, a first image and a second image acquired at time points corresponding to each other, setting a first region of interest within an object of a specified type in the identified first image, setting a second region of interest in the identified second image at the same position as the first region of interest, and The first recognized image in which the first region of interest is set and the second recognized image in which the second region of interest is set are generated as learning data.
2. The electronic device according to claim 1, wherein The second camera is arranged to have an angle corresponding to that of the first camera, The lens properties of the second camera are the same as those of the first camera. When executed by the processor, the instructions cause the electronic device to perform the following actions: The first camera and the second camera are instructed to shoot at corresponding time points.
3. The electronic device according to claim 1, wherein When executed by the processor, the instructions cause the electronic device to perform the following actions: Instructing the first camera and the second camera to capture multiple first images and multiple second images within a specified time period after shooting at corresponding time points. identifying an image acquired at a time point corresponding to a reference image in the plurality of first images by comparing the reference image with each of the plurality of second images, A second image having the same frame difference as the frame difference between the identified first image and the reference image is identified from the plurality of second images.
4. The electronic device according to claim 1, wherein When executed by the processor, the instructions cause the electronic device to perform the following actions: cropping a region of a specified size from the identified first image to generate a cropped first image, Based on the cropped first image, another region of a specified size is cropped from the identified second image to generate a cropped second image.
5. The electronic device according to claim 4, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying normalized differences between the cropped first image and the identified plurality of regions of the second image, identifying a normalized distance between a feature point of the cropped first image and a feature point of each of a plurality of identified regions of the second image, The region where the weighted average value between the normalized difference value and the normalized distance is the lowest is cropped to generate the cropped second image. The electronic device according to claim 4 , wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying an object of a specified type from the cropped first image, setting the first region of interest within the identified object from the cropped first image, setting the position of the object of the cropped first image in the cropped second image, A second region of interest is set in the cropped second image at the same position as the first region of interest in the cropped first image.
7. The electronic device according to claim 6, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: Based on a string recognition algorithm, identifying a string from the first region of interest, The recognized character string is marked in the first region of interest and the second region of interest.
8. The electronic device according to claim 7, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: transforming the first region of interest into a specified graphic based on a specified transformation algorithm, Based on the character string recognition algorithm, the character string is recognized from the first region of interest transformed into the designated graphic.
9. The electronic device according to claim 1, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying whether the electronic device is moving based on differences between consecutive images in the plurality of first images, deleting the plurality of first images and the plurality of second images when the electronic device is not moved; The learning data based on the plurality of first images and the plurality of second images is generated while the electronic device is moving.
10. A method performed by an electronic device, the electronic device comprising a communication circuit, wherein: The method performed by the electronic device includes the following actions: Acquire a plurality of first images from the first camera through the communication circuit, acquiring, through the communication circuit, a plurality of second images from a second camera having a shutter speed slower than that of the first camera; identifying, from the plurality of first images and the plurality of second images, a first image and a second image acquired at time points corresponding to each other, setting a first region of interest within an object of a specified type in the identified first image, setting a second region of interest in the identified second image at the same position as the first region of interest, and The first recognized image in which the first region of interest is set and the second recognized image in which the second region of interest is set are generated as learning data.
11. The method according to claim 10, wherein: The following actions are included: Instructing the first camera and the second camera to capture multiple first images and multiple second images within a specified time period after shooting at corresponding time points. identifying an image acquired at a time point corresponding to a reference image in the plurality of first images by comparing the reference image with each of the plurality of second images, and A second image having the same frame difference as the frame difference between the identified first image and the reference image is identified from the plurality of second images.
12. The method according to claim 10, wherein: The following actions are included: cropping a region of a specified size from the identified first image to generate a cropped first image, and Based on the cropped first image, another region of a specified size is cropped from the identified second image to generate a cropped second image.
13. The method according to claim 12, wherein: The following actions are included: identifying normalized differences between the cropped first image and the identified plurality of regions of the second image, identifying a normalized distance between a feature point of the cropped first image and a feature point of each of a plurality of identified regions of the second image, and The region where the weighted average value between the normalized difference value and the normalized distance is the lowest is cropped to generate the cropped second image.
14. The method according to claim 12, wherein: The following actions are included: identifying an object of a specified type from the cropped first image, setting the first region of interest within the identified object from the cropped first image, setting the position of the object of the cropped first image in the cropped second image, and A second region of interest is set in the cropped second image at the same position as the first region of interest in the cropped first image.
15. The method according to claim 10, wherein: The following actions are included: identifying whether the electronic device is moving based on differences between consecutive images in the plurality of first images, In a case where the electronic device is not moved, deleting the plurality of first images and the plurality of second images, and The learning data based on the plurality of first images and the plurality of second images is generated while the electronic device is moving.
16. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium is used to store a program including instructions; The instructions, when executed by a processor of an electronic device including a communication circuit, cause the electronic device to perform the following actions: Acquire a plurality of first images from the first camera through the communication circuit, acquiring, through the communication circuit, a plurality of second images from a second camera having a shutter speed slower than that of the first camera; identifying, from the plurality of first images and the plurality of second images, a first image and a second image acquired at time points corresponding to each other, setting a first region of interest within an object of a specified type in the identified first image, setting a second region of interest in the same position as the first region of interest in the identified second image, The first recognized image in which the first region of interest is set and the second recognized image in which the second region of interest is set are generated as learning data.
17. The non-transitory computer-readable storage medium of claim 16, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: cropping a region of a specified size from the identified first image to generate a cropped first image, Based on the cropped first image, another region of a specified size is cropped from the identified second image to generate a cropped second image.
18. The non-transitory computer-readable storage medium of claim 17, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying normalized differences between the cropped first image and the identified plurality of regions of the second image, identifying a normalized distance between a feature point of the cropped first image and a feature point of each of a plurality of identified regions of the second image, The region where the weighted average value between the normalized difference value and the normalized distance is the lowest is cropped to generate the cropped second image.
19. The non-transitory computer-readable storage medium of claim 17, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying an object of a specified type from the cropped first image, setting the first region of interest within the identified object from the cropped first image, setting the position of the object of the cropped first image in the cropped second image, A second region of interest is set in the cropped second image at the same position as the first region of interest in the cropped first image.
20. The non-transitory computer-readable storage medium of claim 16, wherein: When executed by the processor, the instructions cause the electronic device to perform the following actions: identifying whether the electronic device is moving based on differences between consecutive images in the plurality of first images, deleting the plurality of first images and the plurality of second images when the electronic device is not moved; The learning data based on the plurality of first images and the plurality of second images is generated while the electronic device is moving.