Electronic device for detecting object in vehicle indoor and operating method thereof
The electronic device in vehicles addresses the issue of unattended objects and theft by detecting and classifying moving objects within the vehicle interior and monitoring driver states, enhancing safety and prevention.
Patent Information
- Application Number
- JP2024219429
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-25
AI Technical Summary
Existing vehicles lack effective systems to detect and prevent incidents such as leaving infants or animals unattended or theft while parked, and there is a need for improved driver monitoring during operation.
An electronic device is installed in vehicles to capture interior images, generate reference images, divide them into regions, and detect moving objects using grid containers, classify the objects, and notify users or emergency services if necessary, while also monitoring driver states through face feature point analysis.
Prevents incidents of unattended infants or theft by accurately detecting moving objects and classifying them, and provides driver monitoring to ensure safe driving conditions.
Smart Images

Figure 2025094945000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an electronic device for detecting a moving object in a vehicle interior and an operating method thereof.
Background Art
[0002] When a vehicle is running, the most important things are safe driving and prevention of traffic accidents. For this purpose, various auxiliary devices for controlling the attitude of the vehicle, controlling the functions of vehicle components, etc., and safety devices such as safety belts and airbags are installed in the vehicle. Also, recently, devices such as a vehicle video recorder (Car Video Recorder, Dashboard Camera, etc.) and an event data recorder (EDR) that stores event data generated in the vehicle are installed in the vehicle, and by storing video data around the vehicle and sensor data of various sensors installed in the vehicle or on electronic devices installed in or on the vehicle, various devices that can grasp the cause of a vehicle accident when the vehicle accident occurs are increasingly being installed in the vehicle. Furthermore, mobile user terminals such as smartphones and tablet PCs equipped with a communication chip that can be connected to the Internet via a cellular-based mobile telecommunication network and / or a Wi-Fi network compliant with the IEEE 802.11 standards are also used as various vehicle electronic devices that acquire video data around the vehicle and gyro-sensor, accelerometer, GPS signals, etc., and assist the driver's driving in the vehicle, perform autonomous driving, or provide route guidance to the driver through various user applications.
[0003] On the one hand, vehicle theft incidents occur frequently not only overseas such as the United States but also domestically. In addition, accidents sometimes occur where infants are left unattended in parked vehicles and die. Also, vehicles equipped with communicable electronic devices that can provide various connected services to users in conjunction with the users' smartphones via a wireless network have become popular, and recently, the proportion of vehicles equipped with cameras for monitoring not only the external environment of the vehicle, that is, the front, rear, and sides of the vehicle, but also the internal environment of the vehicle has been increasing.
Summary of the Invention
Problems to be Solved by the Invention
[0004] From the above viewpoints, the present disclosure explores objects inside a vehicle to prevent incidents of leaving infants or animals unattended or vehicle theft incidents that may occur in the vehicle while parked. If the mobility of the explored object is identified, information regarding the identified object can be notified to the user of the vehicle and / or emergency agencies.
[0005] In relation to object monitoring in the vehicle interior environment, electronic devices and operation methods according to various embodiments are provided.
[0006] According to one embodiment, an electronic device and an operation method thereof for detecting moving objects inside a vehicle using an in-vehicle interior image can be provided.
[0007] According to another embodiment, an electronic device and an operation method thereof for acquiring an in-vehicle interior image of a vehicle can be provided.
[0008] According to still another embodiment, an electronic device and an operation method thereof for generating a background image of the vehicle interior for use in detecting moving objects from an image of the vehicle interior while parked can be provided.
[0009] According to another embodiment, an electronic device for detecting a moving object from an image of a current vehicle interior using a background image in the vehicle interior and an operating method thereof can be provided.
[0010] According to another embodiment, an electronic device for detecting a moving object by dividing a current interior image of a vehicle interior and an operating method thereof can be provided.
[0011] According to another embodiment, an electronic device for acquiring information of a moving object when the moving object is detected in a vehicle interior and an operating method thereof can be provided.
[0012] On the other hand, in relation to driver state monitoring, an electronic device and an operating method according to various embodiments are provided.
[0013] According to one embodiment, an electronic device for monitoring a driver's state while the vehicle is in motion and an operating method thereof can be provided.
[0014] According to another embodiment, an electronic device for acquiring the mounting angle of a vehicle black box and an operating method thereof can be provided.
[0015] According to another embodiment, an electronic device for monitoring a driver's state from an image using a camera not located in front of the driver and an operating method thereof can be provided.
[0016] According to another embodiment, an electronic device for determining a driver's state using feature points detected from the driver's face in an interior image of a vehicle black box and an operating method thereof can be provided.
Means for Solving the Problem
[0017] A method of operating an electronic device for detecting an object in a vehicle interior according to an embodiment may include obtaining at least one interior image of the vehicle interior, generating a reference image using the at least one interior image, and detecting an object in the vehicle based on at least one divided region of the generated reference image and at least one divided region of a target image which is the interior image of the vehicle obtained after the reference image is generated.
[0018] Here, the step of obtaining the at least one interior image may include photographing the interior of the vehicle at a predetermined frame rate for a certain period of time in the parking mode of the vehicle, and the parking mode of the vehicle can be detected and confirmed by the engine state of the vehicle.
[0019] Here, the step of generating the reference image may include setting a region of interest for each of the at least one interior image and synthesizing each interior image for which the region of interest is set.
[0020] Here, the step of setting the region of interest may include setting in a preset manner such that the window region of the vehicle is minimized.
[0021] Here, the preset manner may exclude images of a first ratio from the left and right respectively based on the center point at the lower end of the reference image, and exclude images of a second ratio from the upper end of the reference image.
[0022] Here, the step of detecting the object may include dividing the target image into a plurality of regions and detecting the object for each of the divided plurality of regions.
[0023] Here, in the step of detecting the object, a plurality of grid containers can be sequentially applied to the reference image and the target image in ascending order according to the division size of the grid container to detect the object.
[0024] Here, the plurality of grid containers include a plurality of grid containers with the same division size, and the plurality of grid containers with the same division size are set such that the start positions of the grid cells included in the grid container are different from each other. The plurality of grid containers with the same division size, where the start positions of the grid cells are set to be different from each other, can be applied to the indoor image together.
[0025] Here, in the step of detecting the object, if the object is detected using the sequentially applied grid container, the next sequentially grid container can be used to not detect the object.
[0026] Here, in the step of detecting the object, when the object is detected using the sequentially applied grid container, the object can be detected using the next sequentially grid container after the object is detected.
[0027] Here, it can further include the step of transmitting notification information regarding the detection of the object to the terminal of the vehicle user.
[0028] Here, it can further include the step of obtaining information on the detected object and the step of transmitting the obtained object information to the terminal of the vehicle user.
[0029] Here, the step of obtaining information on the object can include the step of classifying the object based on machine learning and the step of obtaining class information of the object based on the classification result.
[0030] Here, the step of classifying the object may include, when the object is a person, classifying whether the person is an adult, a child, or an infant based on the size of the object in the reference image and the ratio to the person.
[0031] An electronic device for detecting an object in a vehicle according to another embodiment includes a camera that acquires at least one in-vehicle image of the vehicle interior, and generates a reference image using the at least one in-vehicle image, and based on at least one divided region of the generated reference image and at least one divided region of a target image that is the in-vehicle image of the vehicle acquired after the reference image is generated, a processor that detects an object in the vehicle.
[0032] Here, the camera photographs the interior of the vehicle at a predetermined frame rate for a certain period of time in the parking mode of the vehicle, and the processor can detect and confirm the parking mode of the vehicle by detecting the engine state of the vehicle.
[0033] Here, the processor can set a region of interest for each of the at least one in-vehicle image, and synthesize each in-vehicle image with the region of interest set therein to generate the reference image.
[0034] Here, the processor can set the region of interest in a preset manner such that the window region of the vehicle is minimized.
[0035] Here, the preset manner can exclude images with a first ratio from the left and right respectively based on the center point at the lower end of the reference image, and exclude images with a second ratio from the upper end of the reference image.
[0036] Here, the processor can divide the target image into a plurality of regions, and detect the object for each of the divided regions.
[0037] Here, the processor can detect the object by sequentially applying a plurality of grid containers to the reference image and the target image in ascending order according to the division size of the grid container.
[0038] Here, the plurality of grid containers include a plurality of grid containers with the same division size, and the plurality of grid containers with the same division size are set so that the start positions of the grid cells included in the grid container are different from each other. The plurality of grid containers with the same division size, where the start positions of the grid cells are set to be different from each other, can be applied to the indoor image together.
[0039] Here, if the object is detected using the sequentially applied grid container, the processor cannot detect the object using the next sequentially grid container.
[0040] Here, when the object is detected using the sequentially applied grid container, the processor can detect the object using the next sequentially grid container after the object is detected.
[0041] Here, it can further include a communication circuit for transmitting notification information regarding the detection of the object to the terminal of the vehicle user.
[0042] Here, the processor acquires information of the detected object, and the electronic device can further include a communication circuit unit for transmitting the acquired object information to the terminal of the vehicle user.
[0043] Here, the processor can classify the object based on machine learning and acquire class information of the object based on the classification result.
[0044] Here, when the object is a person, the processor can classify whether the person is an adult, a child, or an infant based on the size of the object in the reference image and the ratio to the person.
[0045] Moreover, an operation method of an electronic device for monitoring the state of a driver according to another embodiment includes a step of detecting feature points of the driver's face from an in-vehicle image of a vehicle captured by an in-vehicle camera provided in the electronic device, a step of converting the coordinates of the detected feature points into coordinates on a front coordinate system, and a step of determining the state of the driver using the distance between the coordinates of at least two of the converted feature points. The front coordinate system can be a coordinate system on an image in which the in-vehicle camera captures the driver's face from the front.
[0046] Here, the step of converting into coordinates on the front coordinate system can include a step of obtaining the mounting angle of the in-vehicle camera and a step of converting the coordinates of the feature points into coordinates on the front coordinate system using the mounting angle of the in-vehicle camera.
[0047] Here, the step of obtaining the mounting angle of the in-vehicle camera can include a step of detecting a straight lane from a front image captured by a front camera provided in the electronic device, a step of detecting a vanishing point of the detected straight lane, a step of comparing the detected vanishing point with the center point of the front image to obtain the mounting angle of the front camera, and a step of obtaining the mounting angle of the in-vehicle camera from the mounting angle of the front camera.
[0048] Here, the step of obtaining the mounting angle of the in-vehicle camera can include a step of obtaining the mounting angle of the in-vehicle camera using a gyro sensor, a step of weighted addition of the mounting angle of the in-vehicle camera obtained from the mounting angle of the front camera and the mounting angle of the in-vehicle camera obtained using the gyro sensor, and a step of determining the weighted added value as the final mounting angle of the in-vehicle camera.
[0049] Here, the step of detecting the feature points of the driver's face includes detecting the center point of the left eye as the first feature point, the center point of the right eye as the second center point, and the center point of the nose as the third feature point on the driver's face. The step of determining the driver's state can be determined based on the ratio of the distance between the coordinates of the first feature point and the coordinates of the third feature point and the distance between the coordinates of the second feature point and the coordinates of the third feature point.
[0050] Here, in the step of determining the driver's state, when the ratio of the distances is greater than or equal to a first reference value or less than a second reference value, the driver's state is determined to be not looking ahead. When the ratio of the distances is greater than or equal to the second reference value and less than the first reference value, the driver's state can be determined to be in a forward-looking state.
[0051] Here, at least a plurality of the detected feature points of the driver's face are detected from one of the feature parts of the driver's face, and the feature parts of the driver's face can include at least one of eyes, nose, mouth, eyebrows, and the contour line of the face.
[0052] Here, in the step of determining the driver's state, for at least one of the driver's left and right eyes, based on the ratio (C) of the distance between the feature points at both ends of the one eye and the sum of the distances between at least one upper feature point of the one eye and the corresponding lower feature point among the plurality of eye feature points detected by the one eye, it can be determined.
[0053] Here, in the step of determining the driver's state, when the ratio (C) is less than or equal to a third reference value and continues for a certain period of time or more, the driver's state can be determined to be in a drowsy state.
[0054] Here, the step of determining the driver's state can include determining the presence or absence of the driver's drowsiness and / or the presence or absence of the driver's forward gaze by a machine learning model that has learned at least a part of the detected feature points of the driver's face.
[0055] An electronic device for monitoring a driver's state according to another embodiment includes an in-vehicle camera that captures the interior of the vehicle, and detects feature points of the driver's face from an in-vehicle image of the vehicle captured by the in-vehicle camera, converts the coordinates of the detected feature points into coordinates in a front coordinate system, and includes a processor that determines the driver's state using the distance between the coordinates of at least two of the converted feature points, and the front coordinate system can be a coordinate system on an image in which the in-vehicle camera captures the driver's face from the front.
[0056] Here, the processor can obtain the mounting angle of the in-vehicle camera, and use the mounting angle of the in-vehicle camera to convert the coordinates of the feature points into coordinates in the front coordinate system.
[0057] Here, the processor detects a straight lane from a front image captured by a front camera provided in the electronic device, detects a vanishing point of the detected straight lane, compares the detected vanishing point with the center point of the front image, obtains the mounting angle of the front camera, and can obtain the mounting angle of the in-vehicle camera from the mounting angle of the front camera.
[0058] Here, the processor obtains the mounting angle of the in-vehicle camera using a gyro sensor, and respectively weights and adds the mounting angle of the in-vehicle camera obtained from the mounting angle of the front camera and the mounting angle of the in-vehicle camera obtained using the gyro sensor, and can determine the weighted sum value as the final mounting angle of the in-vehicle camera.
[0059] Here, the processor detects the center point of the left eye as the first feature point, the center point of the right eye as the second center point, and the center point of the nose as the third feature point on the driver's face, detects the feature points of the driver's face, and determines the driver's state based on the ratio of the distance between the coordinates of the first feature point and the coordinates of the third feature point and the distance between the coordinates of the second feature point and the coordinates of the third feature point.
[0060] Here, when the ratio of the distance is greater than or equal to a first reference value or less than a second reference value, the processor determines that the driver's state is not looking ahead, and when the ratio of the distance is greater than or equal to the second reference value and less than the first reference value, the processor can determine that the driver's state is a looking-ahead state.
[0061] Here, at least a plurality of the detected feature points of the driver's face are detected from one of the feature parts of the driver's face, and the feature part of the driver's face can include at least one of eyes, nose, mouth, eyebrows, and facial contour lines.
[0062] Here, for at least one of the driver's left and right eyes, the processor can determine the driver's state based on the ratio (C) of the distance between the feature points at both ends of the one eye and the sum of the distances between at least one upper feature point of the one eye and the corresponding lower feature point among the plurality of eye feature points detected by the one eye.
[0063] Here, when the ratio (C) continues to be less than or equal to a third reference value for a certain period of time or more, the processor can determine that the driver's state is a drowsy state.
[0064] Here, the processor can determine the presence or absence of the driver's drowsiness and / or the presence or absence of the driver's looking ahead by using a machine learning model that has learned at least a part of the detected feature points of the driver's face.
Advantages of the Invention
[0065] Based on an image of the vehicle interior captured by an electronic device mounted on a vehicle, moving objects (such as people and animals) inside the vehicle can be detected, and the detected information can be notified to the user terminal of the vehicle. Thereby, it is possible to prevent the accident of leaving infants or young children unattended inside the vehicle or the theft accident of the vehicle. In addition, the processor of the in-vehicle electronic device can classify moving objects inside the vehicle by using the image data acquired through the camera and the machine learning pre-learned, and even without identifying an abnormal state from outside the vehicle, the in-vehicle electronic device alone can detect an abnormal state inside the vehicle.
[0066] In addition, based on the interior image of the vehicle acquired through a camera installed inside the vehicle during the running of the vehicle, the driving state / behavior state, etc. of the driver can be monitored. Thereby, even in a vehicle not equipped with a driver monitoring system at the time of vehicle manufacture, an embodiment according to the present invention can provide a driver monitoring system, and assistance can be provided to enable the driver to drive safely.
Brief Description of the Drawings
[0067] The following drawings are created to explain a specific example of this specification. The names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, and the technical features of this specification are not limited to the specific names used in the following drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0068] Hereinafter, some embodiments of this specification will be described in detail through exemplary drawings. When adding reference numerals to the components of each drawing, it should be noted that for the same components, as far as possible, they have the same numerals even if they are shown on different drawings. Also, when explaining the embodiments, if it is determined that a specific explanation of a related known configuration or function will impede the understanding of the embodiments, the detailed explanation thereof will be omitted.
[0069] When explaining the components of this specification, terms such as first, second, A, B, (a), (b), etc. can be used. These terms are only for distinguishing the components from other components, and the essence, order, etc. of the corresponding components are not limited by these terms. Also, unless otherwise defined, all terms used here, including technical or scientific terms, have the same meaning as generally understood by those with ordinary knowledge in the technical field. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with the meaning in the context of the related technology, and should not be interpreted in an ideal or overly formal sense unless clearly defined in this application.
[0070] In this specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, "A or B" in this specification can be interpreted as "A and / or B". For example, "A, B or C" in this specification can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0071] As used herein, the slashes ( / ) and commas used in this specification can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0072] As used herein, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, as used herein, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted the same as "at least one of A and B".
[0073] Also, as used herein, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C".
[0074] FIG. 1 is a diagram for explaining a vehicle service system applicable to an embodiment of the present invention.
[0075] In this specification, a vehicle is an example of a moving body and is not limited to vehicles. The moving body according to this specification can include various movable objects such as vehicles, people, bicycles, ships, trains, etc. Hereinafter, for convenience of explanation, the case where the moving body is a vehicle will be described as an example.
[0076] Also, in this specification, a vehicle electronic device may also be called by other names such as a vehicle infrared (Infra-Red) camera, a vehicle black box, a car dash cam, or a car video recorder.
[0077] Also, in this specification, a vehicle service system can include at least one vehicle-related service system such as a vehicle black box service system, an advanced driver assistance system (ADAS), a traffic control system, an autonomous driving vehicle service system, a vehicle teleoperated driving system, an AI (Artificial Intelligence) vehicle control system, and a V2X service system.
[0078] Referring to FIG. 1, the vehicle service system (1000) includes a vehicle electronic device (100), a vehicle service providing server (200), and a user terminal device (300). The vehicle electronic device (100) can be wirelessly connected to a wired / wireless communication network and can exchange data with the vehicle service providing server (200) and the user terminal device (300) connected to the wired / wireless communication network.
[0079] The vehicle electronic device (100) can be controlled by user control input through the user terminal device (300). For example, when the user selects an executable object installed in the user terminal device (300), the vehicle electronic device (100) can execute an operation corresponding to an event generated by the user input for the executable object. Here, the executable object can be a type of application installed in the user terminal device (300) that can remotely control the vehicle electronic device (100).
[0080] FIG. 2a is a diagram for explaining the configuration of an example of a vehicle electronic device applicable to an embodiment of the present invention.
[0081] Referring to FIG. 2a, the vehicle electronic device (100) includes at least a part of a processor (110), a power management module (111), a battery (112), a display unit (113), a user input unit (114), a sensor unit (115), a photographing unit (116), a memory (120), a communication unit (130), one or more antennas (131), a speaker (140), and a microphone (141).
[0082] The processor (110) can control the overall operation of the vehicle electronic device (100) and can be configured to implement the proposed functions, procedures, and / or methods described herein. The processor (110) can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The processor can be an AP (application processor). The processor (110) can include at least one of a DSP (digital signal processor), a CPU (central processing unit), a GPU (graphics processing unit), and a modem (modulator and demodulator).
[0083] The processor (110) can control all or part of the power management module (111), battery (112), display unit (113), user input unit (114), sensor unit (115), imaging unit (116), memory (120), communication unit (130), one or more antennas (131), speaker (140), and microphone (141). In particular, when various data are received through the communication unit (130), the processor (110) can process the received data to generate a user interface and control the display unit (113) to display the generated user interface. All or part of the processor (110) can be electrically or operably coupled with or connected to other components (e.g., the power management module (111), battery (112), display unit (113), user input unit (114), sensor unit (115), imaging unit (116), memory (120), communication unit (130), one or more antennas (131), speaker (140), and microphone (141)) within the vehicle electronic device (100).
[0084] The processor (110) can execute a signal processing function for processing the video data acquired by the imaging unit (116) and a video analysis function for obtaining information regarding the on-site situation from the video. As an example, the signal processing function includes a function of compressing the video data captured by the imaging unit (116) to reduce its capacity. The video data has a form in which multiple frames are gathered with time as the axis. That is, it can be regarded as a series of photos continuously taken during a given period of time. The capacity of such video is extremely large if not compressed, and it is very inefficient to store it in the memory as it is. Therefore, the digitized video is compressed. For video compression, methods that utilize the correlation between frames, spatial correlation, and the characteristics of vision that are sensitive to low-frequency components are used. Since the original data is lost due to compression, it can be compressed at an appropriate ratio that allows the identification of traffic accident situations of the vehicle. As video compression methods, one of various video codecs such as H.264, MPEG4, H.263, H.265 / HEVC can be used, and the video data is compressed in a manner supported by the vehicle electronic device (100).
[0085] The video analysis function can be based on deep learning and can be realized by computer vision techniques. Specifically, the video analysis function includes an image segmentation function that divides an image into multiple regions or parts and inspects each separately, an object detection function that identifies specific objects within the image, an advanced object detection model that recognizes multiple objects (e.g., a soccer field, attackers, defenders, a soccer ball, etc.) existing in one image (generates a bounding box using XY coordinates and identifies all within it), a face recognition function that not only recognizes the faces of people in the image but also identifies the identities of individuals, a boundary detection function used to identify the outer boundaries of objects or landscapes to more accurately understand the content of the image, a pattern detection function that recognizes repeated shapes, colors, and other visual representations in the image, a feature matching function that classifies by comparing the similarity of images, etc.
[0086] Such a video analysis function can also be executed by a vehicle service providing server (200) instead of the processor (110) of the vehicle electronic device (100).
[0087] The power management module (111) manages the power for the processor (110) and / or the communication unit (130). The battery (112) supplies power to the power management module (111).
[0088] The display unit (113) outputs the result processed by the processor (110).
[0089] The display unit (113) can output content, data, or signals. In various embodiments, the display unit (113) can display a video signal processed by the processor (110). For example, the display unit (113) can display a capture or still image. As another example, the display unit (113) can display a video or a camera preview image. As yet another example, the display unit (113) can display a graphical user interface (GUI) so as to be able to interact with the vehicle electronic device (100). The display unit (113) can include at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, and a 3D display. The display unit (113) can also be configured as an integrated touch screen by being combined with a sensor capable of receiving touch input, etc.
[0090] The user input unit (114) receives an input used by the processor (110). The user input unit (114) can be displayed on the display unit (113). The user input unit (114) can sense touch or hovering inputs of a finger and a pen. The user input unit (114) can sense an input caused through a rotatable structure or a physical button. The user input unit (114) can include sensors for sensing various types of inputs. The input received by the user input unit (114) can have various types. For example, the input received by the user input unit (114) can include touch and release, drag and drop, long touch, force touch, physical depression, etc. The input unit (430) can provide the received input and data related to the received input to the control unit (450). In various embodiments, the user input unit (114) can include a microphone (microphone or transducer) capable of receiving a user's voice command. In various embodiments, the user input unit (114) can include an image sensor or a camera capable of receiving a user's motion.
[0091] The sensor unit (115) includes one or more sensors. The sensor unit (115) has a function of detecting an impact applied to the vehicle or detecting a case where the amount of change in acceleration is equal to or greater than a certain value. In some embodiments, the sensor unit (115) can be an image sensor such as high dynamic range cameras. In some embodiments, the sensor unit (115) includes non-visual sensors. In some embodiments, the sensor unit (115) can include, in addition to the image sensor, RADAR, LiDAR (Light Detection And Ranging), and / or ultrasonic sensors. In some embodiments, the sensor unit (115) can include an acceleration sensor, a geomagnetic sensor, etc. to detect impacts and accelerations.
[0092] In various embodiments, the sensor unit (115) can be attached at different positions of the vehicle and / or facing one or more different directions. For example, the sensor unit (115) can be attached to the front, sides, rear, and / or roof of the vehicle in directions such as forward-facing, rear-facing, side-facing, etc.
[0093] The imaging unit (116) can capture images in at least one of the situations of the vehicle during parking, stopping, and driving. Here, the captured images can include parking lot images which are images related to the parking lot. The parking lot images can include the images captured during the period from the time when the vehicle enters the parking lot to the time when the vehicle exits the parking lot. That is, the parking lot images can include the images captured from the time when the vehicle enters the parking lot to the time when the vehicle parks (e.g., the time when the vehicle engine is turned off for parking), the images captured during the parking period of the vehicle, and the images captured from the time when the vehicle finishes parking (e.g., the time when the vehicle engine is turned on for exiting) to the time when the vehicle exits the parking lot. And the captured images can include at least one of the images of the front, rear, side, and inside of the vehicle. Also, the imaging unit (116) can include an infrared (IR) camera capable of monitoring the driver's face or pupils.
[0094] This imaging unit (116) can include a lens unit and an image sensor. The lens unit can perform the function of condensing optical signals, and the optical signals transmitted through the lens unit reach the imaging area of the image sensor to form an optical image. Here, as the image sensor, a CCD (Charge Coupled Device), a CIS (Complementary Metal Oxide Semiconductor Image Sensor), a high-speed image sensor, etc. can be used to convert the optical signals into electrical signals. And the imaging unit (116) can further include all or part of a lens unit driving part, a diaphragm, a diaphragm driving part, an image sensor control part, and an image processor.
[0095] The operating modes of the vehicle electronic device (100) can include a continuous recording mode, an event recording mode, a manual recording mode, and a parking recording mode.
[0096] The continuous recording mode is a mode that is executed when the vehicle engine is started and driving begins, and can be maintained while the vehicle is in motion. In the continuous recording mode, the vehicle video recording device (100) can execute recording in a predetermined time unit (for example, 1 to 5 minutes). In the present disclosure, the continuous recording mode and the continuous mode can be used interchangeably.
[0097] The parking recording mode can be defined as a mode that operates in a parked state when the vehicle engine is turned off or the battery supply for vehicle operation is interrupted. In the parking recording mode, the vehicle electronic device (100) can operate in a parking continuous recording mode to perform continuous recording while parked. Also, in the parking recording mode, the vehicle electronic device (100) can operate in a parking event recording mode to execute recording when a shock event is detected while parked. In this case, recording for a certain period from a predetermined time before the event occurs to a predetermined time after (for example, recording from 10 seconds before to 10 seconds after the event occurs) can be executed. In this specification, the parking recording mode and the parking mode can be used interchangeably.
[0098] The event recording mode can be defined as a mode that operates when various events occur during vehicle operation.
[0099] The manual recording mode can be defined as a mode in which the user manually activates recording. In the manual recording mode, the vehicle electronic device (100) can execute recording for a time period from a predetermined time before the user's manual recording request is generated to a predetermined time after (for example, recording from 10 seconds before to 10 seconds after the event occurs).
[0100] The memory (120) is operably coupled to the processor (110) and stores various information for operating the processor (110). The memory (120) can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. When an embodiment is implemented in software, the techniques described herein can be implemented as modules (e.g., procedures, functions, etc.) that execute the functions described herein. The modules can be stored in the memory (120) and executed by the processor (110). The memory (120) can be implemented inside the processor (110). Alternatively, the memory (120) can be implemented outside the processor (110) and communicatively connected to the processor (110) through various means known in the art.
[0101] The memory (120) can be configured inside the vehicle electronic device (100), configured to be detachable through a port provided in the vehicle electronic device (100), or exist outside the vehicle electronic device (100). When the memory (120) is configured inside the vehicle electronic device (100), it can exist in the form of a hard disk drive or a flash memory. When the memory (120) is configured to be detachable from the vehicle electronic device (100), it can exist in the form of an SD card, a Micro SD card, a USB memory, etc. When the memory (120) is configured outside the vehicle electronic device (100), it can exist in a storage space in another device or a database server through the communication unit (130).
[0102] The communication unit (130) is operably coupled to the processor (110) and transmits and / or receives wireless signals. The communication unit (130) includes a transmitter and a receiver. The communication unit (130) can include a baseband circuit for processing radio frequency signals. The communication unit (130) controls one or more antennas (131) to transmit and / or receive wireless signals. The communication unit (130) enables the vehicle electronic device (100) to communicate with other devices, where the communication unit (130) can be provided as at least one combination of various known communication modules, such as a cellular mobile communication module, a short-range wireless communication module such as a Wireless LAN (Local Area Network) module, a communication module using Low-Power Wide-Area (LPWA) technology, etc. Also, the communication unit (130) can also perform a position tracking function like a GPS (Global Positioning System) tracker.
[0103] The speaker (140) outputs voice-related results processed by the processor (110). For example, the speaker (140) can output audio data indicating that a parking event has occurred. The microphone (141) receives voice-related inputs used by the processor (110). The received voice can be a sound due to an external impact or the voice of a person related to the situation inside / outside the vehicle, and can help recognize the situation at that time together with the video captured by the imaging unit (116). The voice received through the microphone (141) can be stored in the memory (120).
[0104] FIG. 2b is a diagram illustrating another example of the vehicle electronic device (100) according to an embodiment of the present invention.
[0105] Referring to FIG. 2b, the vehicle electronic device (100) can include a communication unit (150), a control unit (160) (which can also be called a "processor"), a memory unit (170), and additional elements (180).
[0106] The communication unit (150) can transmit and receive wireless signals to communicate with an external entity. In particular, it can transmit the position of the vehicle to the external entity in real time, or transmit the video data collected by the vehicle electronic device (100) to the external entity. The communication unit (150) can be operably coupled to the control unit (160). The communication unit (150) can include a baseband circuit for processing radio frequency signals. The communication unit (150) can control one or more antennas (not shown) to transmit and / or receive wireless signals. The communication unit (130) can be provided as a combination of at least one of various known communication modules, such as a cellular mobile communication module, a short-range wireless communication module such as a Wireless LAN (Local Area Network) system, and a communication module using Low-Power Wide-Area (LPWA) technology. Further, the communication unit (150) can also execute a position tracking function such as a GPS (Global Positioning System) tracker.
[0107] The control unit (160) can execute the overall operation of the vehicle electronic device (100) and control the components of the vehicle electronic device (100). The control unit (160) can be configured in the form of a single chip including various functions such as a System On Chip (SoC). Thus, the control unit (160) configured by the SoC can include a CPU (161), a GPU (162), a Neural Processing Unit (NPU) (163), an Image Signal Processor (ISP) (164), a Memory Controller (165), I / O Controllers (166), a Direct Memory Access Controller (DMA) (167), and the like. This is an example of the configuration of the control unit (160) by the SoC, and the components of the control unit (160) can be added or changed. As an example, a small amount of memory (not shown), such as the cache memory of the CPU (161) or the ROM (Read Only Memory) of the system, is included inside the control unit (160). Also, the communication unit (150) can be included as a component within the control unit (160).
[0108] The CPU (161) processes general operations and instructions related to the operation or control of the vehicle electronic device (100). The GPU (162) is responsible for operations related to graphics and performs operations related to video processing and rendering, such as providing smooth and clear graphic representations on a high-resolution display. The NPU (163) can execute operations for the learning of Artificial Intelligence / Machine Learning (AI / ML) in order to accelerate and efficiently process artificial neural network operations. Thereby, operations using Artificial Intelligence / Machine Learning (AI / ML) can be executed in the Advanced Driver-Assistance Systems (ADAS) operation.
[0109] In FIG. 2b, although the control unit (160) is illustrated as including an NPU (163) for accelerating AI operations, the GPU (162) can also process AI operations instead of the NPU (163).
[0110] The ISP (164) can signal-process the raw image data input from the camera module and convert it into a digital image. The memory controller (165) manages data exchange between the control unit (160) and an external memory (170) (such as a RAM or flash memory) of the control unit (160).
[0111] The input / output controller (166) manages interfaces with various input / output devices provided in the vehicle electronic device (100), such as a USB (Universal Serial Bus) or an SD (Secure Digital Card) slot, and the DMA controller (167) enables data transfer between memories without the intervention of the CPU (161).
[0112] The memory (170) can be composed of a non-volatile storage device such as a RAM (Random Access Memory) or a flash memory. The RAM, as a temporary data storage, temporarily stores data generated or processed during the operation of the vehicle electronic device (100). For example, the raw image data input from the camera can be temporarily stored in the RAM before or during processing.
[0113] Data can be semi-permanently stored in the non-volatile storage device. In the vehicle electronic device (100), mainly an SD card or an embedded flash memory is used. The video data processed by the ISP (Image Signal Processor) (164) can ultimately be stored in this non-volatile memory.
[0114] The additional element (180) can include sensors such as a camera module, a display, an acceleration sensor, a microphone, etc. The camera module can capture the environment around or inside the vehicle in video. Generally, the camera module includes a front camera and a rear camera. Also, a side camera can be included. Further, in relation to the embodiments of the present invention, the camera module includes an in-vehicle camera. The display can provide real-time video or stored video to the user, or can provide menu settings. The acceleration sensor can detect the movement of the vehicle and react in situations such as sudden braking or collision to specially protect the video data at that time. The microphone can record sounds inside or outside the vehicle.
[0115] The control unit (160) composed of an SoC, the memory (170), and the additional element (180) can be connected and operate as follows.
[0116] The camera module can convert light into a digital signal and transmit it to the ISP (164). The ISP (164) processes this signal to generate a digital image, and the generated digital image can be additionally processed (e.g., compressed, AI analysis, etc.) by the GPU (162) or the NPU (163). The CPU (161) can control the components through necessary instructions by summarizing such processes. Also, the video data generated and processed through the ISP (164), the GPU (162), or the NPU (163) can be temporarily stored in the RAM through the memory controller (165) or finally stored in the memory (170) such as a non-volatile memory. The I / O controller (166) communicates with an external interface (e.g., USB port) to enable the user to access the video data, and the DMA controller (167) can independently manage the data transfer between memories to reduce the load on the CPU.
[0117] Figure 3 is a diagram for explaining the configuration of a vehicle service providing server applicable to the embodiments of the present invention.
[0118] Referring to FIG. 3, the vehicle service providing server (200) includes a communication unit (202), a processor (204), and a storage unit (206). The communication unit (202) of the vehicle service providing server (200) is connected to the vehicle electronic device (100) and / or the user terminal device (300) through a wired / wireless communication network to transmit and receive data.
[0119] FIG. 4 is a diagram for explaining the configuration of the user terminal device according to an embodiment of the present invention.
[0120] Referring to FIG. 4, the user terminal device (300) includes a communication unit (302), a processor (304), a display unit (306), and a storage unit (308). The communication unit (302) is connected to the vehicle electronic device (100) and / or the vehicle service providing server (200) through a wired / wireless communication network to transmit and receive data. The processor (304) controls the overall functions of the user terminal device (300), and transmits instructions input by the user according to the embodiments herein to the vehicle service system (1000) through the communication unit (302). When a control message related to the vehicle service is received from the vehicle service providing server (200), the processor (304) controls to display it to the user through the display unit (306).
[0121] In particular, the processor (304) can be configured in the form of a System On Chip (SoC) as an example. Thereby, the processor (304) can include processors (not shown) such as a Central Processor Unit (CPU), a Graphic Processor Unit (GPU), a Neural Processor Unit (NPU), and an Image Signal Processor (ISP).
[0122] The CPU can execute the brain functions of the processor, process the operating system and various application operations, and control the operations of the components of the user terminal device (300), such as maintaining performance and power efficiency. The GPU can be responsible for operations related to 2D (Dimension) and 3D graphics and provide smooth and clear graphic representations on a high-resolution display. The NPU can execute operations for AI and machine learning to accelerate and efficiently process artificial neural network operations. The ISP is an image signal processing device that can convert the analog signal input from the image sensor into a digital image and execute various image processing operations.
[0123] FIG. 5 is a diagram for explaining the operation of an electronic device (600) that trains a neural network based on a set of training data according to an embodiment of the present invention.
[0124] Referring to FIG. 5, in operation (S502), the electronic device (600) can acquire a set of training data. The electronic device (600) can acquire a set of training data for supervised learning. The training data can include a pair of input data and ground truth data corresponding to the input data. The ground truth data can represent the output data that the neural network attempts to obtain upon receiving the input data that is a pair of the ground truth data.
[0125] For example, when training a neural network for image recognition, the training data can include an image and information about one or more subjects included in the image. The information can include the classification (category or class) of the subject that can be identified through the image. The information can include the position, width, height, and / or size of the visual object corresponding to the subject within the image. The set of training data identified through operation (502) can include a plurality of pairs of training data. In the above example of training a neural network for image recognition, the set of training data identified by the electronic device (600) can include a plurality of images and the ground truth data corresponding to each of the plurality of images.
[0126] Referring again to FIG. 5, in operation (S504), the electronic device (600) can perform training on the neural network based on the set of training data. In one embodiment where the neural network is trained based on supervised learning, the electronic device (600) can input the input data included in the training data into the input layer of the neural network. An example of a neural network including the input layer is described with reference to FIG. 6. From the output layer of the neural network that has received the input data through the input layer, the electronic device (600) can obtain the output data of the neural network corresponding to the input data.
[0127] In one embodiment, the training of the operation (S504) can be performed based on the difference between the output data and the ground truth data included in the learning data and corresponding to the input data. For example, the electronic device (600) can adjust one or more parameters associated with the neural network based on the gradient descent algorithm so that the difference decreases. The operation of the electronic device (600) to adjust the one or more parameters can be referred to as tuning the neural network. The electronic device (600) can perform the tuning of the neural network based on the output data using a function defined to evaluate the performance of the neural network, such as a cost function. The difference between the output data and the ground truth data described above can be included as an example of the cost function.
[0128] Referring again to FIG. 5, in operation (S506), the electronic device (600) according to one embodiment can identify whether valid output data is output from the neural network trained by operation (S504). The fact that the output data is valid can mean that the difference (or cost function) between the output data and the ground truth data satisfies the conditions set for the use of the neural network. For example, when the average value and / or the maximum value of the difference between the output data and the ground truth data are less than or equal to a specified threshold, the electronic device can determine that valid output data is output from the neural network.
[0129] If valid output data is not output from the neural network (S506 - No), the electronic device (600) can repeatedly perform the training of the neural network based on operation (S504). The embodiment is not limited to this, and the electronic device (600) can repeatedly perform operations (S502, S504).
[0130] When valid output data is obtained from the neural network (Yes in S506), based on the operation (S508), the electronic device (600) according to one embodiment can use the trained neural network. For example, the electronic device (600) can input other input data that is distinguished from the input data that was input to the neural network as learning data, into the neural network. The output data obtained from the neural network that has received the other input data can be used by the electronic device (600) as a result of performing an inference on the other input data based on the neural network.
[0131] FIG. 6 is a diagram for explaining the configuration of the electronic device (600) according to an embodiment of the present invention.
[0132] The electronic device (600) in FIG. 6 can be one of the vehicle electronic device (100), the vehicle service providing server (200), or the user terminal device (300) that executes the operation by AI / ML in FIG. 1.
[0133] Referring to FIG. 6, the processor (610) of the electronic device (600) can execute computations related to the neural network (630) stored in the memory (620). The processor (610) can include at least one of a CPU (central processing unit), a GPU (graphics processing unit), or an NPU (neural processing unit). The NPU can be implemented as a chip separated from the CPU, or integrated with the CPU on the same chip in the form of a SoC (system on a chip). The NPU integrated with the CPU can be referred to as a neural core and / or an AI (artificial intelligence) accelerator.
[0134] Referring to FIG. 6, the processor (610) can identify the neural network (630) stored in the memory (620). The neural network (630) can include a combination of an input layer (632), one or more hidden layers (634) (or intermediate layers), and output layers (636). The layers described above (e.g., the input layer (632), one or more hidden layers (634), and output layer (636)) can include a plurality of nodes. The number of hidden layers (634) can vary according to embodiments, and a neural network (630) including a plurality of hidden layers (634) can be referred to as a deep neural network. The operation of training the deep neural network can be referred to as deep learning.
[0135] In one embodiment, when the neural network (630) has the structure of a feed forward neural network, the first node included in a specific layer can be connected to all of the second nodes included in other layers before the specific layer. In the memory (620), the parameters stored for the neural network (630) can include the weights assigned to the connections between the second node and the first node. In the neural network (630) having the structure of a feed forward neural network, the value of the first node can correspond to the weighted sum of the values assigned to the second nodes, based on the second nodes and the weights assigned to the connections connecting the first node.
[0136] In one embodiment, when the neural network (630) has a structure of a convolutional neural network, the first node included in a specific layer can correspond to a weighted sum of a part of the second nodes included in other layers before the specific layer. A part of the second nodes corresponding to the first node can be identified by a filter corresponding to the specific layer. In the memory (620), the parameters stored for the neural network (630) can include weights representing the filter. The filter can include one or more nodes used to calculate the weighted sum of the first node from among the second nodes, and weights corresponding to each of the one or more nodes.
[0137] The processor (610) of the electronic device (600) according to one embodiment can execute training for the neural network (630) by using the learning data set (640) stored in the memory (620). Based on the learning data set (640), the processor (610) can execute the operations described with reference to FIG. 5 to adjust one or more parameters stored in the memory (620) for the neural network (630).
[0138] The processor (610) of the electronic device (600) according to an embodiment can perform object detection, object recognition, and / or object classification by using a neural network (630) trained based on a learning dataset (640). The processor (610) can input an image (or video) acquired through a camera (650) into an input layer (632) of the neural network (630). Based on the input layer (632) into which the image is input, the processor (610) can sequentially obtain the values of the nodes of the layers included in the neural network (630) and obtain a set of values of the nodes of the output layer (636) (e.g., output data). The output data can be used as a result of inferring the information included in the image by using the neural network (630). Without being limited thereto, the processor (610) can input an image (or video) acquired from an external electronic device connected to the electronic device (600) through a communication circuit (660) into the neural network (630).
[0139] In one embodiment, a neural network (630) trained to process an image can be used to identify a region corresponding to a subject in the image (object detection) and / or to identify the class of the subject represented in the image (object recognition and / or object classification). For example, the electronic device (600) can use the neural network (630) to segment a region corresponding to the subject in the image based on a rectangular form such as a bounding box. For example, the electronic device (600) can use the neural network (630) to identify at least one class matching the subject from among a plurality of specified classes.
[0140] Hereinafter, embodiments of the present invention will be described. The embodiments described below can be implemented based on the devices or components of the devices in FIGS. 1 to 6, and the functions, methods, and procedures described in FIGS. 1 to 6.
[0141] <First Embodiment>
[0142] The first embodiment of the present invention relates to an electronic device and an operating method of the electronic device that, when an in-vehicle image acquired by a camera located inside the vehicle (including a camera capable of photographing the interior of the vehicle and having a field of view of the camera lens that can photograph the interior of the vehicle) during parking of the vehicle is used to identify an intrusion from the outside to the vehicle or the presence of a child or an animal inside the vehicle, notify the identified information to the user.
[0143] The electronic device according to the first embodiment can be an electronic device mounted on a vehicle or at least a part of the electronic device, and can include components of the electronic devices (100, 600) described in FIGS. 1 to 6. For the sake of convenience of explanation, it is assumed that the electronic device according to the first embodiment is the electronic device (600) in FIG. 6. In this case, the electronic device (600) includes a camera (650) configured to acquire an in-vehicle image during parking, a processor (610) configured to generate a background image of the vehicle interior from the in-vehicle image and identify (or detect) a moving object inside the vehicle using the background image and the currently captured in-vehicle image, and a communication circuit (660) configured to transmit information related to the detected moving object to the vehicle user's terminal directly or via a server if a moving object is detected inside the vehicle as a detection result.
[0144] Hereinafter, each component of the electronic device according to the first embodiment of the present invention will be described in more detail.
[0145] The camera (650) of the electronic device (600) can acquire an in-vehicle image of the vehicle through photographing during parking.
[0146] The processor (610) of the electronic device (600) checks whether the vehicle is in the parking mode. If it is confirmed that the vehicle is in the parking mode, the camera (650) can be controlled to capture indoor images for a predetermined time (e.g., 2 minutes) or at regular time intervals. At this time, in order to reduce the power consumption of the electronic device (600), the processor (610) can control the camera (650) to set the frame rate of image capture for vehicle interior monitoring to a preset value (e.g., 5 FPS (Frames Per Second)).
[0147] On the other hand, the processor (610) of the electronic device (600) can use the vehicle engine state detection function and / or the vehicle battery operating power state detection function to check the parking mode of the vehicle. That is, the processor (610) of the electronic device (600) can check whether the vehicle is in motion or parked / stopped based on the voltage of the vehicle battery, etc.
[0148] For example, the processor (610) can set different battery voltages for distinguishing the operating state (such as the parking state, the stopped state, the driving state, etc.) of the vehicle according to the type of the power transmission system of the vehicle.
[0149] Specifically, the processor (610) can set different battery voltages for identifying the operating state of the vehicle according to whether the power transmission system of the vehicle is an internal combustion engine vehicle (ICEV) powered by an internal combustion engine or an electric vehicle (EV) supplied with operating power from a battery.
[0150] For example, when the vehicle is an ICEV, the processor (610) can set the battery voltage in the state where the engine is engaged and the state where the engine is not engaged as follows.
[0151] - State where the engine is not engaged (Ignition_off)
[0152] The state where the engine is not engaged is a state where the vehicle is powered only by the battery. At this time, the battery voltage according to the embodiment of the present invention can be set to a value between 12.5V and 12.7V (first range).
[0153] - State where the engine is engaged (Ignition_on)
[0154] The state where the engine is engaged is a state where the generator (alternator) is charging the battery. At this time, the generator can generate a voltage higher than the battery voltage to increase the battery voltage. At this time, the battery voltage according to the embodiment of the present invention can be set to a value between 13.8V and 14.4V (second range).
[0155] On the other hand, since the electric motor of the EV is a device that rotates by receiving electricity from the battery, the voltage of the battery can vary within a certain range according to the state of charge before the operating power is supplied to the electric motor. For example, in the case of a lithium-ion battery, the voltage is 4.2V when the state of charge is 100%, and the voltage can be 3.0V when the state of charge is 0%. When the operating power is supplied to the electric motor, the voltage of the battery can vary depending on the resistance and load of the motor. Specifically, the resistance of the motor changes according to the rotational speed and torque of the motor. The greater the resistance, the greater the load, and the smaller the resistance, the smaller the load. The greater the load, the lower the voltage of the battery, and the smaller the load, the higher the voltage of the battery.
[0156] For example, when an EV accelerates or climbs a slope, the resistance of the motor is large and the load is heavy, so the voltage of the battery decreases. When the electric vehicle decelerates or travels on flat ground, the resistance of the motor is small and the load is light, so the voltage of the battery can increase. Since the voltage change of the battery affects the performance and life of the battery, a Battery Management System (BMS) (not shown) can monitor and adjust the voltage of the battery. For example, in an embodiment of the present invention, the processor (610) can identify the operating state of the EV through the voltage of the battery monitored by the BMS.
[0157] Specifically, the battery voltage can be different as follows when the operating power is supplied to the electric motor in the EV and before the operating power is supplied to the electric motor (parked state, stopped state).
[0158] - When the operating power is supplied to the electric motor (VEM_On)
[0159] When the operating power is supplied to the electric motor, the electric motor starts to rotate using the power of the battery, and the electric motor can impose a burden on the battery. Therefore, the battery voltage after the operating power is supplied to the electric motor is slightly lower than the battery voltage before the operating power is supplied to the battery. Generally, the battery voltage when the operating power is supplied to the electric motor is maintained between 13.5V and 14.1V (the third range).
[0160] - Before the operating power is supplied to the electric motor (parked state, stopped state) (VEM-Off)
[0161] When the electric motor is not supplied with operating power, the battery can be charged using the power supplied from the generator (alternator) or from an external charging station. At this time, the generator / external charging station can generate a voltage higher than the battery voltage to increase the battery voltage. Generally, the battery voltage when the electric motor is not supplied with operating power is maintained between 14.1V and 14.5V (the fourth range).
[0162] Therefore, as described above, the battery voltage of the ICEV is slightly lower than a certain threshold value (about 13V) when the engine is not engaged, and slightly higher than the above-mentioned certain threshold value (about 13V) when the engine is engaged. On the other hand, the battery voltage of the EV is lower than a certain threshold value (about 14.1V) when the electric motor is supplied with operating power, and higher than a certain threshold value (about 14.1V) when the electric motor is not supplied with operating power. However, the battery voltage can vary depending on the state of the battery. For example, when the vehicle's battery becomes old or damaged, the battery voltage may fall outside the normal range. The processor (610) can generate a background image using the vehicle image taken in the parking mode identified by the measured vehicle battery voltage as described above. In particular, when generating the background image, a region of interest (ROI) is set so that the window area of the vehicle is maximally excluded from each of the plurality of indoor images, and a background image of the vehicle interior can be generated from the indoor image with the region of interest set. In addition, the processor (610) can generate a background image using the average value of the indoor images taken at regular time intervals according to a predetermined time frame rate during the parking mode. An embodiment in which a background image (of the vehicle interior) is generated will be described with reference to FIGS. 7 to 10 below. Hereinafter, the "background image" can be referred to as the "reference image (RI)".
[0163] FIG. 7 is a diagram for explaining an example of an image taken according to the first embodiment of the present invention.
[0164] Referring to FIG. 7, (a) to (d) are examples of indoor images taken within 2 minutes after entering the parking mode. Four images are illustrated in FIG. 7. The time interval for taking indoor images in the parking mode can be different depending on the setting. The time interval for taking indoor images can be set to a value of 1 frame per second (i.e., 1 FPS) or more and usually within 10 FPS. For reference, the frame rate during normal video shooting while the vehicle is moving is 30 FPS.
[0165] In one embodiment, the processor (610) can set a Region of Interest (ROI) in each of the captured indoor images (a) to (d). The region of interest means the region that is the target of interest in image processing. The region of interest may be a part of the image or the entire image. When the region of interest is defined as the entire image, the algorithm for image processing can be applied to the selected region of interest. Thereby, the calculation time and memory usage during image processing can be reduced. In one embodiment, the region of interest can be set so that the window region is excluded from the entire image to the maximum extent.
[0166] Since this embodiment is for detecting an abnormal state inside the vehicle, it is not necessary to detect the operations outside the vehicle. Therefore, the processor (610) can exclude the window region in which the state outside the vehicle is represented from the region of interest to increase the speed of image processing. However, in some cases, the processor (610) can include the window region in the region of interest. An example of how the processor (610) sets the region of interest will be described below.
[0167] FIG. 8 is a diagram for explaining an example in which the processor sets a region of interest according to the first embodiment of the present invention.
[0168] Referring to FIG. 8, in one embodiment, when setting the region of interest, the processor (610) can exclude a certain ratio (= first ratio) (α) from the left and right respectively based on the center point (a) at the lower end of the indoor image, and can also exclude a certain ratio (= second ratio) (β) from the upper end. In FIG. 8, an example is illustrated in which 10% (α = 0.1) of the image is excluded from the left and right respectively based on the center point (a) at the lower end, and 20% (β = 0.2) of the image is excluded from the upper end. When setting the above-mentioned region of interest, the ratio excluded from the background image (hereinafter, which can be referred to as the "image exclusion ratio") can be different by prior setting. In the above manner, the region of interest is illustrated as the region within the dotted line in FIG. 8.
[0169] In order for the region of interest to be set according to the pre-set image exclusion ratio, the installation position of the camera (650) in the vehicle would have to be fixed at a specific position. The above-mentioned image exclusion ratio is set assuming that the camera (650) is installed at the above-mentioned specific position. When the camera (650) is installed at other positions than the above-mentioned specific position, the captured background image is different. If the background image is different, the image exclusion ratio at the time of setting the region of interest would also have to be different. Therefore, a camera installation guide for the user (or camera installer) would have to be provided so that the position where the camera (650) is installed in the vehicle is constant. On the other hand, in one embodiment, the above-mentioned image exclusion ratio can also be set in real time from the background image using a machine learning or deep learning model. In this case, the setting position of the camera (650) may not need to be fixed at a specific position. The region of interest can be set for each of the indoor images in FIG. 7 by the method described above. Hereinafter, the image for which the region of interest is set can be referred to as the "region of interest image" or the "ROI image".
[0170] When the region of interest image is generated, the processor (610) can generate a background image using the generated region of interest image. Specifically, the processor (610) can synthesize the region of interest images during a predetermined time (e.g., 2 minutes) after entering the parking mode to generate a background image. Usually, the processor (610) can generate a background image from indoor images captured at a rate of 1 frame per second or more (i.e., 1 FPS). The processor (610) can control the electronic device (600) to operate in a low-power mode during the parking mode, and can control the frame rate of the camera (650) to be 10 FPS or less in the low-power mode. Therefore, the processor (610) can control the camera (650) to operate at a low frame rate of 10 FPS or less in the parking mode to generate a background image while minimizing the power consumption of the electronic device (600).
[0171] FIG. 9 is a diagram showing an example of a background image generated from a region of interest image according to the first embodiment of the present invention.
[0172] Referring to FIG. 9, the background image of FIG. 9 according to an embodiment has regions with an α ratio excluded from the right and left sides of the indoor image of FIG. 8, and a region with a β ratio excluded from the upper end. As a result, in the background image of FIG. 9, the window regions on the right and left sides of the driver's seat in the indoor image of FIG. 8 are almost excluded.
[0173] On the other hand, when the parking environment is not good, such as when the parking lot is dark, the quality of the indoor image or the image of the region of interest captured in the parking mode may not be good. In such a case, when the processor (610) synthesizes the image of the region of interest to generate a background image, for example, the quality of the background image generated by removing noise using a histogram equalization method can be improved.
[0174] When the background image is generated, the processor (610) can detect (or identify) moving objects in the vehicle using the generated background image. At this time, the processor (610) can divide the current interior image into one or more regions and detect moving objects for each divided region. On the other hand, when dividing the interior image, the processor (610) can apply division filters with different division sizes in ascending order according to the division size. In this way, the size of one region divided in the entire current interior image can gradually become smaller according to the division size. As one region divided in this way becomes smaller sequentially, not only large moving objects but also small moving objects can be easily detected. On the other hand, the processor (610) can perform machine learning-based classification on the moving objects, obtain specific information about the moving objects (e.g., people (adults, children, infants), dogs, etc.) according to the class information based on the classification result, and notify this to the user.
[0175] Hereinafter, with reference to FIGS. 10 to 12, an embodiment in which the processor (610) identifies and detects moving objects in the vehicle using the background image will be described.
[0176] FIG. 10 is a diagram illustrating a case where there are no moving objects in the vehicle when confirming moving objects in the vehicle based on the background image according to the first embodiment of the present invention.
[0177] Referring to FIG. 10, (a1) and (b1) are "current in-vehicle images" taken in chronological order to detect moving objects inside the vehicle after the background image is generated. (a2) and (b2) are the results of graphically processing the difference values between the current in-vehicle images of (a1) and (b1) and the background image generated in FIG. 9. If there is no difference between the current in-vehicle image and the background image, it is displayed in black, and the parts with differences are displayed in white. The fact that both (a2) and (b2) are displayed in black means that there is no difference between the current in-vehicle images of (a1) and (b1) and the background image of FIG. 8. That is, (a2) and (b2) mean that no moving object was detected in the current images of (a1) and (b1).
[0178] In this specification, "moving object detection" means detecting an object moving inside the vehicle, and can be used in the same sense as "object movement detection" in the sense of detecting object movement.
[0179] FIG. 11 is a diagram illustrating a case where a moving object exists inside the vehicle when confirming a moving object inside the vehicle based on the background image according to the first embodiment of the present invention.
[0180] Referring to FIG. 11, (a1) and (b1) are "current in-vehicle images" taken in chronological order of the vehicle interior to detect moving objects inside the vehicle after the background image is generated. (a2) and (b2) are the results of graphically processing the difference values between the current in-vehicle images of (a1) and (b1) and the background image generated in FIG. 9. The appearance of white parts in (a2) and (b2) means that there is a difference between the current in-vehicle images of (a1) and (b1) and the background image of FIG. 8. That is, (a2) and (b2) mean that a moving object exists in the current in-vehicle images of (a1) and (b1).
[0181] On the one hand, in one embodiment, the processor (610) can divide the area of the current indoor image and identify and detect moving objects for each divided area. An example of dividing the area of the current indoor image to detect moving objects will be described below. On the other hand, in an embodiment of the present invention, since the current indoor image is compared with the background image to determine whether there are moving objects in the current indoor image, the current indoor image can be referred to as the "Target Image". Also, as described above, the background image can be referred to as the "Reference Image: RI".
[0182] FIG. 12a is a diagram illustrating a case where the area of the reference image is displayed in a grid system according to the first embodiment of the present invention, and moving objects are detected by dividing the displayed grid system into a plurality of grid cells.
[0183] For the description of the grid cells in FIG. 12a, the grid container and the grid cells will be described with reference to FIG. 12b.
[0184] FIG. 12b is a diagram for explaining the concepts of the grid container and the grid cells in the first embodiment of the present invention.
[0185] In the present invention, a frame that can cover the entire one image frame is referred to as a grid container, and a plurality of cells of the same size that constitute the grid container, which is the one image frame, are each referred to as grid cells. Since the size of the grid container needs to cover the target image, the size of the grid container can be the same as the size of the image or larger than the size of the image. Also, a grid cell refers to the smallest grid in the grid included in the grid container. That is, a grid cell is the smallest rectangular area generated by four intersecting grid lines.
[0186] In FIG. 12b, it is exemplified that one grid container contains 16 grid cells. A grid cell identifier (ID) can be assigned to each grid cell, and FIG. 12b illustrates that a grid cell ID is assigned to each grid cell. On the other hand, the division size of the grid container can be represented by the number of regions into which the corresponding grid container is divided in the X-axis direction and the Y-axis direction. For example, in FIG. 12b, one grid container is divided into 4 regions in the X-axis direction and 4 regions in the Y-axis direction. In this case, the division size of the corresponding grid container is 4X4.
[0187] Referring again to FIG. 12a, (A) to (D) illustrate examples in which the current indoor image is divided by grid containers of the same size having different division sizes.
[0188] That is, (A) is a grid container in which one grid container is divided into one region in each of the X-axis and Y-axis directions, that is, a grid container with a division size of 1x1. As a result, the 1x1 grid container includes a total of one grid cell. Therefore, as a result of applying the 1x1 grid container to the image, the image is not actually divided.
[0189] (B) is a grid container in which one grid container is divided into two regions in each of the X-axis and Y-axis directions, that is, a grid container with a division size of 2x2. As a result, the grid container of (B) includes a total of four grid cells, and the image is divided into four regions.
[0190] (C) is a grid container in which one grid container is divided into four regions in each of the X-axis and Y-axis directions, that is, a grid container with a division size of 4x4. As a result, the grid container of (C) includes a total of 16 grid cells, and the image is divided into 16 regions.
[0191] (D) is a grid container with a split size of 8x8, where one grid container is divided into 8 regions in the X-axis and Y-axis directions respectively. As a result, the grid container of (D) contains a total of 64 grid cells, and the image is divided into 64 regions.
[0192] On the other hand, (C) and (E), and (D) and (F) respectively illustrate examples where the split size of the grid container is the same, but the starting positions of the first grid cells located in the upper left region of the grid container are set differently.
[0193] That is, (E) is the same as (C) where a 4x4 grid container is applied. However, the starting position of the first grid cell in the upper left of (E) is set differently from the starting position of the first grid cell in the upper left of (C). Also, (F) is the same as (D) where an 8x8 grid container is applied. However, the starting position of the first grid cell in the upper left of (F) is set differently from the starting position of the first grid cell in the upper left of (D). The specific method of setting the starting positions of the grid cells differently will be described later.
[0194] On the other hand, the reasons for using multiple grid containers with the same split size but different starting positions of the grid cells, such as (C) and (E) and (D) and (F), are as follows.
[0195] When a moving object is located on the boundary line of a grid cell, there is a possibility that the moving object may not be detected. Therefore, by using multiple grid containers with the same split size but different positions of the grid cells, the moving object can be detected better. For example, when a grid container with a split size of 4X4 or more, such as (C) and (D) above, is applied to a target image, if multiple grid containers with the same split size but different starting positions of the grid cells are applied to the target image, the detection rate of the moving object in the corresponding target image can be improved more.
[0196] In one embodiment, the processor (610) can detect moving objects by applying grid containers (A) to (F) to the current indoor image in ascending order according to the division size. That is, the processor (610) first uses a grid container with a small division size, that is, a 1X1 grid container, to detect moving objects throughout the target image, and can sequentially detect moving objects using grid containers with larger division sizes.
[0197] The processor (610) can detect moving objects using the 1 X 1 grid container of (A). That is, the processor (610) applies the 1 X 1 grid container to the target image (TI) and the background image which is the reference image (RI) respectively. Then, calculate the pixel value (Pixel_Value_RI_Fixed) within a total of 1 grid cell in the target image, and calculate the pixel value (Pixel_Value_RI_Fixed) within a total of 1 grid cell in the reference image. Then, calculate a first difference value (Differ_1) which is the difference value of the calculated pixel values. If the first difference value is greater than a preset threshold value (threshold), it can be determined that there is a moving object within 1 grid cell in the target image. The detection of moving objects using the 1 X 1 grid container of (A) can be referred to as "one-stage detection" for convenience. The above threshold value can be preset with various experimental values, and preferably the threshold value can be set to "30". The above threshold value can also be applied identically in the object detection for each stage described below.
[0198] If no moving object is detected in the one-step detection of (A), the processor (610) can detect the moving object using the 2X2 grid container of (B). That is, the processor (610) applies the 2 X 2 grid container to the background image which is the target image (TI) and the reference image (RI) respectively. Then, pixel values (Pixel_Value_RI_Fixed) are calculated for each of the total 4 grid cells in the target image, and pixel values (Pixel_Value_RI_Fixed) are calculated for each of the total 4 grid cells in the reference image. Then, a first difference value (Differ_1), which is the difference value of the pixel values calculated for each grid cell, is calculated. If the first difference value is greater than a preset threshold, it can be determined that there is a moving object in the corresponding grid cell within the target image. The detection of the moving object using the 2 X 2 grid container of (B) can be referred to as "two-step detection" for convenience.
[0199] If no moving object is detected in the two-step detection of (B), the processor (610) can perform "three-step detection" for each grid cell in the manner described in (A) and (B) using the 4X4 grid container of (C). If no moving object is detected in the three-step detection of (C), the processor (610) can perform "four-step detection" using the 8X8 grid container of (D).
[0200] On the other hand, at each of the above steps, the processor (610) can detect the moving object using a plurality of grid containers which are set such that the sizes of the grid containers are the same but the positions of the grid cells are different. For example, in the above three-step detection, the 4X4 grid container of (C) and the 4X4 grid container of (E) can be used together, and in the above four-step detection, the 8X8 grid container of (D) and the 8X8 grid container of (F) can be used together.
[0201] As described above, detecting a moving object by sequentially using grid containers with larger and larger division sizes (i.e., grid cells of smaller sizes) is because the division size for detecting the moving object can be different depending on the size of the moving object. That is, if the size of the moving object is very large, the corresponding moving object can probably be detected even if a 1 X 1 grid container with a large size (A) of one grid cell is used. On the other hand, if the size of the moving object is small, a grid container with a smaller size of one grid cell may be required. Thereby, the processor (610) can efficiently detect moving objects of various sizes by applying grid containers with various division sizes to the target image.
[0202] So far, an example has been described in which grid containers with different division sizes are applied in ascending order with respect to the division size to detect a moving object, and if the moving object is detected by the grid container of the corresponding division size, the detection of the next stage is not executed.
[0203] In another example, even if a moving object is detected at a lower stage (e.g., the second stage), the processor (610) can execute the detection of the next stage (e.g., the third and fourth stages). That is, even if a moving object with a large size is detected at a lower stage, the processor (610) can execute the detection of the next stage to detect a moving object with a small size.
[0204] In another example, the processor (610) applies grid containers with different division sizes in descending order according to the division size, and if a moving object is detected using the grid container of the corresponding size, the detection of the moving object can be interrupted.
[0205] In another embodiment, the processor (610) applies grid containers with different division sizes in descending order according to the division size, and even if a moving object is detected using a grid container of the corresponding size, the next-stage grid container can be used to detect the moving object.
[0206] When a moving object is detected inside the vehicle according to the above-described embodiment, the processor (610) can classify the detected moving object. For this purpose, the processor (610) crops the area of the grid cell where an object is determined to exist in the target image, and can classify the moving object with the image in the cropped grid cell. For example, the processor inputs the cropped image area into a machine learning model, and the machine learning model can classify the class of the moving object existing in the input image area. At this time, the class classified by the machine learning model can include objects of pre-learned types such as children, adults, pets, etc.
[0207] Thereafter, the processor (610) can notify the user terminal of the information of the detected moving object. The information of the moving object can include notification information that a moving object has been detected and / or class information of the detected moving object.
[0208] As described above, in the following, with reference to FIG. 12c, a case where the grid containers have the same division size but the start positions of the grid cells are set to be different will be described in more detail.
[0209] FIG. 12c is a diagram for explaining a case where the grid containers have the same division size but the start positions of the grid cells are set to be different according to the first embodiment of the present invention.
[0210] Referring to FIG. 12c, (a) is an arbitrary image having a certain width and height, and (b) is a 4x4 grid container having the same size as the image of (a).
[0211] (c) shows a state in which the grid container of (b) is exactly mapped to the image of (a). At this time, the position of the grid container is in a fixed state.
[0212] (d) shows a state in which the position of the grid container has moved by the direction and magnitude of the moving vector in the state of (c). In (d), the direction of the moving vector is upper left, and the magnitude of the moving vector is exemplified as 1 / 2 of the length of the diagonal of one grid cell.
[0213] Previously in FIG. 12a, (E) and (F) were described as being set such that the divided sizes of the grid containers are the same as those of (C) and (D), respectively, but the starting positions of the grid cells are different. In an embodiment of the present invention, in this way, using a method in which the grid container moves according to the direction and magnitude of the moving vector, the divided sizes of the grid containers can be the same but the starting positions of the grid cells can be set to be different. On the other hand, in an embodiment of the present invention, the direction and magnitude of the moving vector can be variously deformed by presetting.
[0214] FIG. 12d is a diagram for explaining various examples of the moving vector in the first embodiment of the present invention.
[0215] In the case of (a), the direction of the moving vector is upper right, and the magnitude of the moving vector is 1 / 2 of the length of the diagonal of one grid cell. Thereby, a state in which the grid container has moved according to the direction and magnitude of the moving vector is illustrated.
[0216] (b), the direction of the movement vector is bottom - left, and the magnitude of the movement vector is 1 / 2 of the length of the diagonal of one grid cell. Thereby, the state in which the grid container moves according to the direction and magnitude of the movement vector is illustrated.
[0217] (c), the direction of the movement vector is bottom - right, and the magnitude of the movement vector is 1 / 2 of the length of the diagonal of one grid cell. Thereby, the state in which the grid container moves according to the direction and magnitude of the movement vector is illustrated.
[0218] Combining (d) of FIG. 12c and (a), (b), (c) of FIG. 12d, in the embodiments of the present invention, the directions of the movement vectors are illustrated as upper - left, upper - right, lower - left, and lower - right, and the direction of the movement vector is illustrated as 1 / 2 of the length of the diagonal of one grid cell. However, the direction and magnitude of the movement vector can be set to different directions (e.g., right direction, left direction, upward, downward) and various magnitudes (e.g., 1 / 4, 3 / 4 of the grid cell, etc.).
[0219] The grid containers described in FIGS. 12c and 12d are set to have the same size as the size of the image. However, in this case, when the grid container moves according to the movement vector as in (d) of FIG. 12c and (a), (b), (c) of FIG. 12d, an area where the grid container cannot cover the image will occur. In order to complement such a problem, a grid container larger than the size of the image can be used.
[0220] That is, when a 4x4 grid container having the same size as the image size is used as in the cases of (d) in FIG. 12c and (a), (b), and (c) in FIG. 12d, when the grid container moves according to the motion vector, there may occur a problem that an area of the image that cannot be covered by the area of the moved grid container is generated. In an embodiment of the present invention, in order to solve such a problem, a 5x5 grid container in which grid cells are added in the horizontal and vertical directions with the above 4x4 grid container can be used. That is, one row and one column can also be added to the 4x4 grid container.
[0221] FIG. 12e is a diagram for explaining an example of using a grid container larger than an image in consideration of the case where only the grid container moves according to the motion vector according to the first embodiment of the present invention.
[0222] (a) illustrates the use of a 5x5 grid container larger than the image. That is, when no motion vector for moving the grid container is set (or when the motion vector is "0"), since there is no need to move the grid container, as described in FIGS. 12c and 12d, even if the size of the grid container is the same as the size of the image, the processor according to the embodiment of the present invention can easily detect the moving object. However, as in (a) of FIG. 12e, in the embodiment of the present invention, in consideration of the fact that the grid container is moved according to the motion vector, a grid container larger than the size of the image can also be used.
[0223] (b) illustrates a state in which the grid container in (a) has moved according to the motion vector. Since a 5x5 grid container larger than the size of the image is used, the grid container can cover all areas of the image even after the grid container is moved according to the motion vector.
[0224] An example in which the coordinates of the grid cells are set will be described below.
[0225] Figure 12f is a diagram for explaining various examples in which the coordinates of grid cells are set according to the first embodiment of the present invention.
[0226] (A) shows an example in which, as Case 1, when the starting position of the grid container is set to coincide with the upper left of the target image, the coordinates of the grid cells are set. (B) shows an example in which, as Case 2, when the grid container moves according to the movement vector, the coordinates of the grid cells not located at the edge of the grid container are set. (C) and (D) show examples in which, as Case 3, when the grid container moves according to the movement vector, the coordinates of the grid cells located at the edge of the grid container are set.
[0227] In each case of Figure 12f, the size of one grid cell is 160 X 90, and the movement vector is assumed to be (-80, -45). Each case will be described below.
[0228] - Case 1: (A) shows an example in which a 5X5 grid container larger than the target image is applied to the target image, and at this time, the starting position of the grid container is set to coincide with the upper left of the target image. In (A), the starting position and coordinates of the grid container are (0,0). For example, the grid cell with Grid Cell ID 13 has a starting coordinate of (320, 180) and is an area with a size of 160x90.
[0229] - Case 2: (B) shows a state where the position of the grid container in (A) has moved according to the movement vector. Since the movement vector is (-80, -45), in (B), the coordinates of the starting position of each grid cell are moved 80 to the left and 45 upward. As a result, in (B), the grid cell with Grid Cell ID 13 has a starting coordinate of (240, 135) and is an area with a size of 160x90. The grid cell with the above Grid Cell ID 13 is a grid cell not located at the edge of the grid container.
[0230] - Case 3: (C) is in the same state as (B) where the position of the 5x5 grid container has been moved according to the movement vector. Since the movement vector is (-80, -45), in (C), the coordinates of the starting position of each grid cell are moved 80 to the left and 45 upward. At this time, the grid cell with Grid Cell ID 11 is the grid cell located at the left end of the grid container. The grid cell with Grid Cell ID 11 should be set as an area with a starting coordinate of (-80, 135) and a size of 160x90. However, there is no image data in the negative area on the coordinates. Therefore, when the starting coordinate is set to a negative value according to the movement vector like the grid cell located at the end of the grid container as in (C), it is not appropriate for the size of the movement vector to be directly reflected and the coordinates of the corresponding grid cell to be set. Therefore, for the grid cell located at the end of the grid container like this, the coordinates of the corresponding grid cell are set according to (D).
[0231] Referring to (D), the starting coordinate of the grid cell with Grid Cell ID 11 is not (-80, 135) but (0, 135). Also, the size of the grid cell with Grid Cell ID 11 is not 160 X 90 but 80x90. Therefore, the grid cell with Grid Cell ID 11 is an area with a starting coordinate of (0, 135) and a size of 80 X 90.
[0232] In the following, the operation of the electronic device for detecting an object in the vehicle interior will be described according to an embodiment of the present invention.
[0233] FIGS. 13a and 13b are diagrams for explaining the operation of the electronic device according to the first embodiment of the present invention.
[0234] Referring to FIGS. 13a and 13b, the electronic device generates a background image using an image of the vehicle cabin. The background image can be referred to as a "reference image (RI)" (S1301).
[0235] The reference image can be generated using at least one in-vehicle image over a certain period (e.g., 2 minutes). The in-vehicle image can be captured and obtained by an in-vehicle camera of the electronic device. Also, the in-vehicle image can be captured when the vehicle is in the parking mode. The electronic device can detect the engine state of the vehicle to confirm the parking mode of the vehicle. On the other hand, the frame rate for in-vehicle image capture can be a preset value (e.g., 5 FPS) smaller than the frame rate for video capture mode (= 30 FPS).
[0236] The electronic device can identify the size values (width, height) of the reference image and set a grid container to be applied to the reference image having the identified size values (S1303). Here, the size of the grid container applied to the reference image can match the size of the reference image, or the size of the grid container applied to the reference image can be one grid cell larger than the size of the reference image in both the horizontal and vertical directions. For example, as in the example described in FIG. 12e, if the grid container is moved according to the motion vector, a grid container that is one grid cell larger than the size of the reference image in both the horizontal and vertical directions is set, and if the motion vector is not considered, a grid container having the same size as the reference image can be set.
[0237] The electronic device checks whether the detection period (T) of a preset object has arrived (S1305). If the detection period has arrived, it proceeds to the next step to acquire a target image in the vehicle interior (S1307). Here, the detection period (T) of the object can be a period for acquiring an image frame for detecting a moving object in the vehicle at predetermined intervals.
[0238] The electronic device can perform image preprocessing on the acquired target image (S1309). The image preprocessing can include, for example, gray conversion, brightness correction, noise removal by Gaussian filtering, etc.
[0239] The gray conversion is to convert a color image into a black-and-white (gray) image. Since it is easier to process a black-and-white image than a color image, gray conversion can be performed if necessary.
[0240] The noise removal is for improving the quality of the generated background image by enhancing the quality of the indoor image. For the noise removal, as an example, techniques such as histogram equalization can be utilized. For reference, histogram equalization is a technique that "equalizes" the histogram (e.g., brightness) of an image to improve the contrast of the image. When the histogram (e.g., brightness) of an image is skewed to one side, the entire image may be too dark or too bright and the detailed information of the image may be lost. However, histogram equalization can redistribute the skewed histogram to solve such problems and improve the quality of the image.
[0241] On the one hand, the electronic device sets n to 0 (S1311). Here, "n" is a division variable that determines the division size of the grid container. The electronic device divides the grid container into 2^n X 2^n grid cells according to the above n (S1313). For example, if n = 0, the electronic device divides it into 2^0 X 2^0 = 1X1 grid cell; if n = 1, the electronic device divides it into 2^1 X 2^1 = 2X2 = 4 grid cells; if n = 2, the electronic device divides it into 2^2 X 2^2 = 4X4 = 16 grid cells; if n = 3, the electronic device divides it into 2^3 X 2^3 = 8X8 = 64 grid cells.
[0242] After that, the electronic device applies a grid container with grid cells divided into 2^n X 2^n to RI (S1315), and saves the pixel value (Pixel_Value_RI_Fixed) of each grid cell in RI (S1317). Here, the above pixel value can be the RGB value (R = 28, G = 28, B = 28) or the brightness value (intensity) of the pixels in each grid cell.
[0243] Also, the electronic device applies a grid container with grid cells divided into 2^n X 2^n to the target image (TI) (S1319), and saves the pixel value (Pixel_Value_TI_Fixed) of each grid cell in TI (S1321). Here, the above pixel value can be the RGB value (R = 28, G = 28, B = 28) or the brightness value (intensity) of the pixels in each grid cell.
[0244] After that, the electronic device calculates the difference value (Differ_1) between the pixel value (Pixel_Value_RI_Fixed) of each grid cell in RI and the pixel value (Pixel_Value_TI_Fixed) of each grid cell in TI (S1323).
[0245] The electronic device compares the first difference value (Differ_1) calculated in the above S1323 stage with a predetermined threshold value (Threshold). If the first difference value (Differ_1) is greater than the threshold value (Y in S1325 stage), the electronic device can determine that there has been a change in pixel values due to the movement of an object in the vehicle cabin room. That is, the electronic device determines that there is a moving object in the vehicle cabin (S1327).
[0246] The electronic device can crop the area where the movement of the object is identified by TI (S1329). At this time, the electronic device can identify the grid cell ID corresponding to the area where the degree of change in pixel values is greater than the threshold value and crop the area of the image where it is determined that there is an object in the target image.
[0247] The electronic device can classify the moving object existing within the cropped image area (S1331). For example, the electronic device inputs the cropped image area into a machine learning model (Machine Learning Model), and the machine learning model can classify the class of the moving object existing within the input image area. At this time, the classes classified by the machine learning model can include pre-learned types of objects such as children, adults, pets, etc.
[0248] After that, the electronic device can notify the user of the object information identified in the above S1331 stage (S1333). At this time, the above object information can include information indicating that there is a moving object in the vehicle and / or the class information of the identified object.
[0249] On the other hand, when the first difference value (Differ_1) is not greater than the threshold value (Threshold) in the above S1325 step (N in the S1325 step), the electronic device determines whether n is greater than a preset maximum value (Max) (S1335). If n is greater than the maximum value, the operation ends without executing further operations. If n is not greater than the maximum value, the process proceeds to step S1337.
[0250] The electronic device can determine the moving vector of the grid container. Here, the moving vector includes the moving direction and moving distance of the grid container (S1337). The electronic device adjusts the position of the grid container according to the determined moving vector (S1339), saves the pixel value (Pixel_Value_RI_Moved) of each grid cell in RI (S1341), and saves the pixel value (Pixel_Value_TI_Moved) of each grid cell in TI (S1343). Here, the pixel value can be the RGB value or brightness value of the pixels in each grid cell.
[0251] Thereafter, the electronic device calculates the second difference value (Differ_2) between the pixel value (Pixel_Value_RI_Moved) of each grid cell in RI and the pixel value (Pixel_Value_TI_Moved) of each grid cell in TI, and compares the calculated second difference value (Differ_2) with a preset threshold value (S1345).
[0252] Based on the comparison result in the above S1345 step, if the second difference value is greater than the threshold value (Y in S1345), the electronic device can determine that a change in the pixel value has occurred due to the movement of an object in the vehicle cabin room. That is, the electronic device determines that there is a moving object in the vehicle cabin room (S1347).
[0253] The electronic device can crop the area where the movement of the object is identified by TI (S1349). At this time, the electronic device can identify the grid cell ID corresponding to the area where the degree of change in pixel values is greater than the threshold, and crop the area of the grid cell determined to have an object in the target image.
[0254] The electronic device can classify the moving object existing in the cropped image area (S1351). Specifically, for example, the electronic device inputs the cropped image area into a machine learning model, and the machine learning model can classify the class of the moving object existing in the input image area. At this time, the classes classified by the machine learning model can include objects of pre-learned types such as children, adults, pets, etc.
[0255] After that, the electronic device can notify the user of the object information identified in the above S1351 step (S1353). At this time, the object information can include information that there is a moving object in the vehicle and / or the class information of the identified object.
[0256] On the other hand, when the second difference value (Differ_2) is not greater than the threshold (Threshold) in the above S1345 step (N in the S1345 step), the electronic device determines whether n is greater than the preset maximum value (Max) (S1355). If n is greater than the maximum value, the operation ends without performing further operations. If n is not greater than the maximum value, it proceeds to step S1357, increases n by 1 (n = n + 1), and repeatedly executes the operation from step S1311.
[0257] On the one hand, the electronic device of the present invention can be an electronic device mounted in a vehicle. Generally, in the case of an electronic device mounted in a vehicle, the hardware often has lower specifications compared to general user terminals such as smartphones and laptops. Low-spec electronic devices often contain only a CPU without a GPU or NPU, unlike recent user terminals, and may not be able to classify the classes of moving objects using a machine learning model. Therefore, in this case, the electronic device can notify the user terminal in the form of an alarm or the like that a moving object has been detected.
[0258] On the other hand, an electronic device that can classify the classes of moving objects can classify the classes of moving objects using a machine learning model. The electronic device can transmit the information of the moving object obtained according to the classification result to the user terminal. The above information of the moving object can include class information of the moving object (e.g., person (adult, child, infant), dog, etc.). On the other hand, the above class information of the moving object can be in at least one form of text, image, or video information.
[0259] The alarm that the above moving object has been detected, or the information of the above moving object, can be transmitted to the user terminal using a wireless / wired communication network. Also, the above alarm or the information of the above moving object can be transmitted to the user terminal via a server or directly to the user terminal without passing through a server.
[0260] In the following, the operations of the electronic device related to the classification of the classes of moving objects using the above machine learning and the information of the moving objects will be described.
[0261] When a moving object is detected inside the vehicle, the processor (610) can enter from the parking mode to the monitoring mode and increase the frame rate for image recording. As an example, if the frame rate is 5 FPS in the parking mode, the processor (610) can change the frame rate to 30 FPS for shooting the video of the moving object confirmed in the monitoring mode.
[0262] On the other hand, based on the image of the captured moving object (or the cropped image), the processor (610) can execute the classification for the corresponding moving object by using a machine learning algorithm such as a Support Vector Machine (SVM) for example. For reference, the SVM model, as one of the supervised learning models in machine learning, is widely used for classification and regression analysis problems.
[0263] The processor (610) can obtain the class (such as person, dog, cat, etc.) information of the moving object according to the above classification result, and can transmit the obtained information to the user terminal through the communication circuit (660).
[0264] In order to execute the image classification by machine learning, it is desirable that a Neural Processing Unit (NPU) or a Graphic Processing Unit (GPU) is provided in the electronic device of the vehicle. However, since the amount of computation required for image classification is not as large as that for image detection, an NPU or a GPU may not be necessarily required. Therefore, the image classification by a simple algorithm can also be executed by a general-purpose CPU (Central Processing Unit) which is not an NPU or a GPU.
[0265] On the one hand, in the case of an electronic device inside a vehicle, it is generally a low-spec device. Therefore, even for an electronic device capable of classifying by machine learning, the number of classifications can be limited to a certain extent. As an example, the number of classes of moving objects can be limited to about 10 that can exist inside the vehicle, such as people, dogs, cats, etc. On the other hand, when the classification result of the moving object is that the class of the moving object is a person, whether the corresponding person is an adult, a child, or a baby is important information.
[0266] Therefore, when the class of the moving object is a person, the age group information of the corresponding person can be additionally classified.
[0267] In one embodiment, the processor (610) can classify the age group of the corresponding person by using the ratio between the size of an object inside the vehicle and the size of a person who is the moving object on the current indoor image where the moving object is detected. When the electronic device (600) is equipped with an NPU or a GPU, the processor (610) can increase the number of the above classes and control the operation of the NPU or the GPU to determine the age group of a person. On the other hand, when the electronic device (600) is not equipped with an NPU or a GPU, the processor (610) can limit the number of classes to a low limit or not determine the age group of a person.
[0268] FIG. 14 is a diagram for explaining an example of classifying the age group of a person who is a moving object according to the first embodiment of the present invention.
[0269] Referring to FIG. 14, (a) is an example in which moving objects that are people are identified in two front seats. In (a), based on the size and / or ratio within the entire image of the quadrilaterals (1401, 1403) including the moving objects in the front seats, each moving object can be classified as a child and an adult respectively by using machine learning.
[0270] (b) is an example in which a moving object that is a person is identified in the rear seat. Based on the size and / or ratio within the entire image of the quadrilaterals (1411, 1413) including the moving object in the rear seat in (b), each moving object can be classified into an infant and an adult respectively using machine learning.
[0271] <Second Embodiment>
[0272] The second embodiment of the present invention relates to a driver monitoring system for monitoring driver behavior such as the presence or absence of drowsy driving and lack of forward gaze by using an in-vehicle camera installed in a cabin room while the vehicle is running, and a method for realizing the same.
[0273] The second embodiment of the present invention discloses a method and a system for monitoring a driver by using a camera installed in a vehicle interior and having an angular field of view in the driver direction in order to prevent drowsy driving or the like of the driver while the vehicle is running.
[0274] In recent years, automobile manufacturers have installed various electronic devices in vehicles to improve driver convenience and safety. One of them is a driver monitoring system (Driver Monitoring System, DMS). The DMS is mainly used to improve the safety of a vehicle during driving by detecting and analyzing a driver's behavior and state in real time. The DMS operates by utilizing various sensors and cameras. Generally, a DMS dedicated camera tracks a driver's face and detects eye movements, blink frequencies, etc. Based on such information, the DMS can determine whether the driver is in a drowsy state or whether the driver is in another dangerous situation.
[0275] However, existing DMSs require the installation of a dedicated camera for the DMS during vehicle manufacturing. Therefore, it is difficult to install a DMS additionally in a vehicle that does not have a DMS.
[0276] The second embodiment of the present invention proposes a DMS using a camera installed in the vehicle interior according to the needs of the user after the vehicle is shipped. Thereby, instead of the DMS dedicated camera installed by the vehicle manufacturer before the vehicle is shipped, the DMS can be configured using a camera installed in the passenger room or cabin room of the vehicle according to the needs of the user. Thereby, even after the vehicle is shipped by the vehicle manufacturer, the DMS can be configured simply by installing a camera in the vehicle interior according to the needs of the user.
[0277] The electronic device according to the second embodiment is an electronic device installed in a vehicle or at least a part of the electronic device, and can include the components of the electronic devices (100, 600) described in FIGS. 1 to 6. For the sake of convenience of explanation, it is assumed that the electronic device according to the second embodiment is the electronic device (600) of FIG. 6. In this case, the electronic device (600) includes a camera (650) configured to capture a front image and an image of the driver while the vehicle is running, and monitors the state of the driver in the vehicle using the image, and when it is determined that the monitoring result is that the driver is not looking ahead and / or is dozing off while driving, etc., it includes a processor (610) configured to provide a warning message to the driver. On the other hand, the camera (650) can include a front camera and an in-vehicle camera, and the in-vehicle camera can have a viewing angle opposite to the viewing angle of the front camera facing the front of the vehicle.
[0278] In the second embodiment of the present invention, the electronic device (600) acquires the mounting angle of the in-vehicle camera. The mounting angle can be acquired from a front image captured by the front camera or using a gyro sensor that can be provided in the electronic device. Further, the electronic device (600) detects feature points of the driver's face in an in-vehicle image captured by the in-vehicle camera. Thereafter, the state of the driver can be determined using the relative distance between the detected feature points. That is, the electronic device (600) can determine the driver's gaze direction and / or the presence or absence of a dozing state using the relative distance between the feature points. When it is determined that the driver is not looking ahead and / or is in a dozing state, the electronic device (600) can provide a warning to the driver.
[0279] On the other hand, the in-vehicle camera is not usually located in front of the driver's face. Therefore, the distance between the feature points of the driver's face detected in the image captured by the in-vehicle camera may not match the actual distance on the driver's face. To correct such an error, the electronic device uses the mounting angle of the in-vehicle camera. That is, the electronic device converts the coordinates of the detected feature points into coordinates on a frontal coordinate system (= a coordinate system on an image in which the in-vehicle camera captures the driver's face from the front) using the mounting angle of the in-vehicle camera. Thereafter, the electronic device can determine the driver's gaze direction and / or the presence or absence of a dozing state using the distance between the feature points converted into coordinates on the frontal coordinate system.
[0280] Based on the above, the second embodiment of the present invention will be described in detail below.
[0281] FIG. 15 is a diagram for explaining an operation method of an electronic device (600) for a DMS according to the second embodiment of the present invention.
[0282] Referring to FIG. 15, the processor (610) can acquire the mounting angle of an in-vehicle camera for in-vehicle imaging provided in the electronic device (S1510).
[0283] As one method of obtaining the mounting angle of the in-vehicle camera, a front camera is used. Specifically, the processor (610) can detect a straight lane on the road from a front image captured by the front camera, determine the vanishing point of the straight lane, and estimate the mounting angle of the front camera using the determined vanishing point of the straight lane. On the other hand, since the in-vehicle camera is located on the opposite side of the front camera (i.e., rotated 180 degrees), the processor (610) can obtain the mounting angle of the in-vehicle camera using the mounting angle of the front camera.
[0284] As another method of obtaining the mounting angle of the in-vehicle camera, a gyro sensor that can be provided in the electronic device (600) is used. Further, the processor (610) can also obtain the mounting angle of the in-vehicle camera using a weighted sum of the values respectively obtained by the above two methods.
[0285] On the other hand, the processor (610) can set the area where the driver is located in the in-vehicle image captured by the in-vehicle camera as the region of interest (S1520). Thereafter, the processor (610) can detect the feature points of the driver's face within the region of interest (S1530).
[0286] The processor (610) can convert the coordinates of the detected feature points into coordinates in the frontal coordinate system using the mounting angle of the in-vehicle camera (S1540). Thereafter, the processor (610) can determine the driver's gaze direction and / or the presence or absence of drowsy driving using the relative distances between the coordinates of the feature points in the frontal coordinate system (S1550).
[0287] When the processor (610) determines that the driver's state is forward-unwatched and / or drowsy driving, the processor (610) can provide a warning message to the driver by voice or the like (S1560).
[0288] Hereinafter, the content of each of the above-described steps will be described. The operations described below can be executed by the processor (610), but for convenience, they may be expressed as being executed by the electronic device in general.
[0289] First, the operation of the step (S1510) of obtaining the mounting angle of the indoor camera will be described. As described above, the mounting angle of the indoor camera can be obtained using the front image or the gyro sensor. With reference to FIGS. 16 to 19, a method of obtaining the mounting angle of the front camera from the front image of the front camera and then obtaining the mounting angle of the indoor camera will be described.
[0290] FIG. 16 is a diagram for explaining an example in which an electronic device detects a straight lane on a road using a front image in the second embodiment of the present invention.
[0291] When simple image processing (such as edge detection) is performed on the front image captured by the front camera of the electronic device, straight lanes (1610, 1620) on the road can be detected. Straight lanes can be detected more easily in a straight road section. Therefore, the electronic device can use sensors such as GPS and / or navigation information to determine whether the road on which the current vehicle is traveling is a straight road section, and can be set to start detecting the straight lane when it is determined that the road is a straight road section. For example, when the speed of the current vehicle is 50 km / h or more, the electronic device can determine that the current traveling road is a straight road using the latitude and longitude information obtained using the GPS coordinates.
[0292] FIG. 17 is a diagram for explaining an example in which an electronic device determines the vanishing point of a straight lane according to the second embodiment of the present invention.
[0293] Referring to FIG. 17, the electronic device can connect the straight lines (1710, 1720, 1730) detected in the front image to determine the intersection points. The electronic device can accumulate such intersection points for a certain period of time, and determine the average value of the intersection points accumulated for a certain period of time as the vanishing point (1740).
[0294] FIG. 18 is a diagram for explaining a method in which an electronic device estimates the mounting angle of the front camera from the vanishing point in the second embodiment of the present invention.
[0295] Referring to FIG. 18, the electronic device can estimate the mounting angle of the front camera from the coordinate difference between the center point (1810) of the front image captured by the front camera and the vanishing point (1820) determined in FIG. 17. As in the example of FIG. 18, the center point (1810) of the front image is located lower left than the vanishing point (1820). This means that when the plane of the front image is used as a reference, the front camera is mounted tilted in the lower left direction. Thus, the electronic device can estimate the mounting angle of the front camera from the difference in the coordinate values of the center point (1810) and the vanishing point (1820) of the front image.
[0296] The mounting angle of the front camera can be composed of components of roll, pitch, and yaw. Roll, yaw, and pitch are concepts for explaining the movement of an aircraft and are used to represent the angle and direction of the aircraft.
[0297] FIG. 19 is a diagram for explaining the concepts of roll, yaw, and pitch in the second embodiment of the present invention.
[0298] (a) Roll represents the rotational movement of the aircraft around the horizontal axis (x-axis). When the aircraft makes a roll movement, it means tilting to the left / right side around the horizontal axis. A positive roll value represents a state tilted to the right side, and a negative roll value represents a state tilted to the left side.
[0299] (b) Pitch represents the rotational movement of the aircraft around the horizontal axis (y-axis). When the aircraft makes a pitch movement, it means rotating to the left / right side around the horizontal axis. A positive pitch value represents rotating clockwise, and a negative pitch value represents rotating counterclockwise.
[0300] (c) Yaw represents the rotational movement of the aircraft around the vertical axis (z-axis). When the aircraft makes a yaw movement, it means moving forward / backward or rising / falling around the vertical axis. A positive yaw value represents the aircraft rising, and a negative yaw value represents the aircraft descending.
[0301] In an embodiment of the present invention, since the electronic device is installed on the ceiling inside a moving vehicle, the movement of the electronic device can have characteristics similar to the movement of an aircraft. Therefore, the mounting angle of the front camera can be represented by roll, pitch, and yaw components. As described above, the roll, pitch, and yaw representing the mounting angle of the front camera can be determined from the difference in coordinate values between the center point (1810) and the vanishing point (1820) of the front image in FIG. 18.
[0302] As described with reference to FIGS. 16 to 19, the electronic device can estimate the mounting angle of the front camera using a front image. The indoor camera is installed at a 180-degree difference from the front camera. Therefore, the mounting angle of the indoor camera can be easily obtained from the mounting angle of the front camera.
[0303] As another method for estimating the mounting angle of the indoor camera, when the electronic device is equipped with a gyro sensor, the electronic device can also estimate the mounting angle of the front camera using the gyro sensor. However, errors may occur in the measured values of the gyro sensor when the vehicle is running.
[0304] Therefore, the electronic device can also estimate the weighted sum of the mounting angle of the indoor camera estimated using the front image and the mounting angle of the indoor camera estimated using the gyro sensor as the mounting angle of the front camera. For an electronic device that does not include a gyro sensor, the method using the front image can be used.
[0305] So far, the method for estimating the mounting angle of the indoor camera has been described. Hereinafter, with reference to FIGS. 20 to 24, an example of an operation will be described in which a region of interest where the driver is located is set (S1520) in an indoor image captured using the indoor camera of the electronic device, feature points of the driver's face are detected in the region of interest (S1530), the coordinates of the detected feature points are converted into coordinates in a frontal coordinate system (S1540), and the driver's state is monitored using the relative distance between the coordinate values (S1550).
[0306] FIG. 20 is a diagram for explaining an example in which a region of interest where a driver is located is set in an indoor image according to a second embodiment of the present invention.
[0307] Referring to FIG. 20, a region of interest (2010) where a driver is located can be determined in a vehicle interior image. The region of interest can be determined from the position of the driver's seat. The position of the driver's seat can be preset to the left or right according to the position of the driver's seat that varies by country. Therefore, the region of interest can be determined by reflecting the preset position of the driver's seat.
[0308] When the region of interest is set, the electronic device can detect feature points of the driver's face in the region of interest. The electronic device can detect feature points within a set number of face ranges in consideration of the hardware specifications or performance.
[0309] As a first method, the electronic device can detect one feature point from one feature part such as eyes, nose, and mouth on the driver's face. This can be a method suitable when the hardware specifications of the electronic device are low.
[0310] As a second method, the electronic device sets the entire driver's face as an overall feature part, sets eyes, nose, and mouth as respective sub-feature parts, and can detect a plurality of feature points from the overall feature part and each detailed feature part. The number of feature points can be changed by setting. In this method, the overall contour and detailed contour of the face can be displayed using the feature points. This can be a method suitable when the hardware specifications of the electronic device are high. However, the first method can also be used with a high-specification electronic device.
[0311] Also, the first method and the second method can be mixed. For example, the second method can be used for eyes, and the first method can be used for nose and mouth.
[0312] FIG. 21 is a diagram for explaining an example in which one feature point is detected from a feature portion of a driver's face in the second embodiment of the present invention, and FIG. 22 is a diagram for explaining an example in which a plurality of feature points are detected from one feature portion of a driver's face according to the second embodiment of the present invention.
[0313] Referring to FIG. 21, one feature portion such as the eyes (2110), nose (2120), and mouth (2130) of the driver's face can become one feature point. The electronic device sets each feature portion from the face, and as an example, the center point of each feature portion can be detected as the feature point. That is, the electronic device can set a square of a preset size for each of the eyes (2110), nose (2120), and mouth (2130), and detect the center point of the square as the feature point.
[0314] Referring to FIG. 22, (a) is an example in which a plurality of feature points are detected from the entire face of the driver and from the eyes, nose, mouth, and eyebrows respectively, and (b) shows a preset number (e.g., 68) of feature points detected in (a). The detected feature points display the eyes and eyebrows, nose, mouth, and the contour line of the face.
[0315] On the other hand, the electronic device can determine the driver's gaze direction and / or the presence or absence of drowsy driving, etc. using the distances between the feature points detected in FIG. 21 or FIG. 22. However, these feature points are detected from an image captured using the in-vehicle camera of the electronic device. The in-vehicle camera installed in the vehicle according to the embodiment of the present invention may not be installed in front of the driver's face, unlike a general DMS system. Generally, when the driver's seat is located on the left side of the vehicle, such as in the United States, South Korea, China, etc., the in-vehicle camera installed inside the vehicle may be installed in the upper right part with respect to the driver in order to secure a wider viewing angle for monitoring the entire vehicle interior (in the case of the United Kingdom, Japan, etc., the in-vehicle camera will be installed in the upper left part of the driver). The distances between the feature points detected in the image of the driver captured at this position may not match the actual distances on the corresponding driver's face.
[0316] For example, the relative distance from the center point of the driver's nose to the tip of the right ear on the coordinates of the image captured by the in-vehicle camera located on the front right side of the driver is shorter than the relative distance from the center point of the actual driver's nose to the tip of the right ear. That is, since the position of the in-vehicle camera is not completely frontal with respect to the driver's face, the coordinates of the feature points detected in the driver's image captured by the in-vehicle camera may not match the coordinates on the driver's actual face.
[0317] To correct this coordinate error, the electronic device converts the coordinates of the detected feature points into coordinates in the frontal coordinate system. The frontal coordinate system can be defined as the coordinate system of the image captured by the camera from the front of the center point of the driver's face.
[0318] In an embodiment of the present invention, the electronic device can use the previously obtained mounting angle of the in-vehicle camera to convert the coordinates of the feature points detected on the in-vehicle image into coordinates on the frontal coordinate system.
[0319] The electronic device can determine the driver's gaze direction and / or the presence or absence of the driver's drowsy driving using the relative distance between the coordinates of the feature points on the frontal coordinate system.
[0320] FIG. 23 is a diagram for explaining an example of determining the driver's gaze direction using the feature points detected in FIG. 21 according to the second embodiment of the present invention.
[0321] FIG. 23 illustrates the relative distance between feature points according to the rotation angle of the face when the characteristic parts (eyes, nose, mouth) of the face itself are detected as one feature point as shown in FIG. 21.
[0322] (a) shows the case where the driver's face is facing the front of the driver's seat, (b) shows the case where the driver's face is facing 45 degrees to the right with respect to the front of the driver's seat, and (c) shows the case where the driver's face is facing 90 degrees to the right with respect to the front of the driver's seat. The electronic device can determine the driver's gaze direction for each case as follows.
[0323] In case (a), when the distance between the x - coordinate of the center of the left eye and the x - coordinate of the center of the nose is defined as F1, and the distance between the x - coordinate of the center of the right eye and the x - coordinate of the center of the nose is defined as F2, the electronic device calculates the distance ratio of F1 and F2. Since the driver is looking straight ahead, the calculated ratio is 1:1. If the distance ratio of F1 and F2 is 1, it can be determined that the driver is looking straight ahead. However, since an error may occur in F1 or F2, the value of the distance ratio of F1 and F2 (=F2 / F1) may not be exactly 1. Therefore, when the ratio of F1 and F2 is within a preset range (e.g., 0.8 - 1.3), it can be determined that the driver is looking straight ahead.
[0324] In case (b), when the distance between the x - coordinate of the center of the left eye and the x - coordinate of the center of the nose is defined as F3, and the distance between the x - coordinate of the center of the right eye and the x - coordinate of the center of the nose is defined as F4, since the direction of the driver's face is 45 degrees to the right, F3 is a larger value than F4. For example, when the value of F4 / F3 is 3 / 1 = 3, the value of F4 / F3 is larger than the reference value (e.g., 1.3). In this case, the electronic device can determine that the driver is looking to the right, that is, not looking straight ahead. The reference value can be set considering measurement errors of F3 or F4 and / or the normal movement range of the driver's face, etc.
[0325] In case (c), since the center of the left eye is not recognized, the distance between the x - coordinate of the left eye and the x - coordinate of the center of the nose (F5: not shown in the figure) is 0, and only the distance between the x - coordinate of the center of the right eye and the x - coordinate of the center of the nose (=F6) is recognized. Therefore, the value of F6 / F5 is infinite. In this case, the electronic device can determine that the driver is looking 90 degrees to the right, that is, not looking straight ahead.
[0326] When the driver rotates the face to the left in a manner similar to Figure 23, if the ratio of (the distance between the x - coordinate of the center of the left eye and the x - coordinate of the center of the nose) and (the distance between the x - coordinate of the center of the right eye and the x - coordinate of the center of the nose) is less than a certain value (e.g., 0.8), the electronic device can determine that the driver is looking to the left, that is, not looking straight ahead.
[0327] Thus, when a single characteristic part itself, such as eyes, nose, and mouth, is set as a single characteristic point as in the example of FIG. 23, the electronic device determines a ratio value between the distance on the x coordinate between the center point of the left eye and the center point of the nose and the distance on the x coordinate between the center point of the right eye and the center point of the nose. If the determined ratio value is equal to or greater than a first value (e.g., 1.3) or less than a second value (e.g., 0.8), it can be determined that the driver is not looking straight ahead. Also, if the determined ratio value is less than the first value (e.g., 1.3) and equal to or greater than the second value (e.g., 0.8), it can be determined that the driver is looking straight ahead.
[0328] FIG. 24 is a diagram for explaining an example of determining the presence or absence of a driver's drowsy driving state using the characteristic points of FIG. 22 according to the second embodiment of the present invention.
[0329] FIG. 24 uses the characteristic points detected in FIG. 22. The characteristic points in FIG. 24 are the characteristic points of the left eye (37 - 42), the characteristic points of the right eye (43 - 48), and a part of the characteristic points of the nose (28, 29) in FIG. 22.
[0330] The electronic device can use these characteristic points to determine the presence or absence of the driver's drowsy driving. That is, the electronic device can determine the presence or absence of the driver's drowsiness based on the ratio of the distance between the characteristic points at both ends of the eyes to the sum of the distances between the characteristic points at the upper and lower ends of the driver's eyes.
[0331] For example, assume that the distance between feature point 37 and feature point 40 of the left eye is D1, the distance between feature point 38 and feature point 42 is D2, and the distance between feature point 39 and feature point 41 is D3. The degree of eye closure (C) can be expressed as (D2 + D3) / (2 x D1). The D1 is the distance between the feature points at both ends of the driver's eyes, and this value does not change. Both D2 and D3 are the distances between the feature points on the upper side of the eye and the corresponding feature points on the lower side. When the driver's eyes close, the values of D2 and D3 both decrease. Therefore, (D2 + D3) / (2 x D1) is the ratio of the distance between the feature points at both ends of the eye (= D1) and the sum of the distances between the feature points on the upper side of the eye and the corresponding feature points on the lower side (D2 + D3). Note that the 2 multiplied by D1 is for computational convenience. In this way, the degree of eye closure (C) can be determined using the distances between multiple feature points detected from the eyes.
[0332] The degree of eye closure of the right eye can also be determined in the same way. According to this method, the electronic device determines the degree of eye closure using the distances between the feature points detected by the eyes. When the state where the degree of eye closure is below the reference value continues for a certain period of time, the electronic device can determine that the user is in a drowsy driving state.
[0333] On the other hand, when the electronic device supports machine learning, the electronic device can learn at least a part of the feature points in FIG. 24 based on machine learning to determine the presence or absence of the driver's drowsy driving. That is, the electronic device can use the feature points in FIG. 23 as input data and use a model such as a support vector machine (SVM) to determine the presence or absence of the driver's drowsy driving and the driver's gaze direction. In one embodiment, the driver's state can be classified as either normal or sleepy, and the driver's gaze direction can be classified as any of forward, 90 degrees to the left, 45 degrees to the left, 45 degrees to the right, and 90 degrees to the right.
[0334] As described with reference to FIGS. 21 to 24, the electronic device can determine the presence or absence of the driver's forward gaze and / or the presence or absence of drowsy driving by using the feature points detected from the driver's face. Thereby, when it is determined that the driver is not looking forward and / or is driving drowsily, the electronic device can provide a warning guidance to the driver with a warning sound or the like.
[0335] The embodiments of the present invention have been described so far. The embodiments of the functions, methods, and procedures of the vehicle electronic device, the vehicle service providing server, and the user terminal device disclosed in this specification can be executed through software. The constituent means of each of the above functions, methods, and procedures are code segments for performing necessary operations. The program or code segment can be stored in a processor-readable medium or transmitted by a computer data signal combined with a carrier wave in a transmission medium or a network.
[0336] The computer-readable recording medium includes all types of recording devices in which data readable by a computer system is stored. Examples of the computer-readable recording devices include ROM, RAM, CD-ROM, DVD±ROM, DVD-RAM, magnetic tape, floppy disk, hard disk, optical data storage device, and the like. Further, the computer-readable recording medium can be distributed to computer devices connected by a network, and computer-readable code can be stored and executed in a distributed manner.
[0337] The embodiments described above are not limited to the foregoing embodiments and the accompanying drawings because various substitutions, modifications, and changes can be made by those having ordinary knowledge in the technical field to which the present invention pertains without departing from the technical idea of the embodiments. Also, the embodiments described in this document are not limitedly applied, and all or part of each embodiment can be selectively combined and configured so that various modifications can be made.
[0338]
Claims
1. 1. A method of operating an electronic device for detecting an object in a vehicle interior, comprising: acquiring at least one interior image for the vehicle interior; generating a reference image using the at least one indoor image; and A method for operating an electronic device, comprising: detecting an object within the vehicle based on at least one divided area of the generated reference image and at least one divided area of a target image, which is an interior image of the vehicle acquired after the reference image is generated.
2. 2. The method of claim 1, wherein the step of acquiring at least one indoor image comprises: In a parking mode of the vehicle, capturing an image of an interior of the vehicle at a predetermined frame rate for a certain period of time, A method of operating an electronic device, wherein the parking mode of the vehicle is confirmed by detecting an engine condition of the vehicle.
3. 2. The method of claim 1, wherein the step of generating the reference image comprises: establishing a region of interest for each of the at least one indoor image; and A method of operating an electronic device, comprising the step of combining each of the room images having the region of interest defined therein.
4. 4. The method according to claim 3, wherein the step of setting the region of interest comprises: A method of operating an electronic device, comprising the step of: configuring the vehicle window area in a preset manner to be minimized.
5. In claim 4, the preset method is: A method for operating an electronic device, comprising: excluding an image of a first ratio from each of the left and right sides of a center point of a lower end of the reference image; and excluding an image of a second ratio from an upper end of the reference image.
6. 2. The method of claim 1, wherein the step of detecting an object comprises: A method for operating an electronic device, comprising: dividing the target image into a plurality of regions; and detecting the object in each of the divided plurality of regions.
7. 7. The method of claim 6, wherein the step of detecting an object comprises: A method of operating an electronic device, comprising sequentially applying a plurality of grid containers to the reference image and the target image in ascending order according to division sizes of the grid containers to detect the object.
8. In claim 7, the plurality of grid containers include a plurality of grid containers having the same division size; the plurality of grid containers having the same division size are set such that start positions of grid cells contained in the grid containers are different from each other; A method for operating an electronic device, wherein multiple grid containers having the same division size, but with the starting positions of the grid cells set to be different from each other, are applied together to the indoor image.
9. 8. The method of claim 7, wherein the step of detecting an object comprises: If the object is detected using the sequentially applied grid container, the object is not detected using a next sequential grid container.
10. 8. The method of claim 7, wherein the step of detecting an object comprises: If the object is detected using the sequentially applied grid container, the next sequential grid container after the object is detected is used to detect the object.
11. In claim 1, The method of operating an electronic device further includes the step of transmitting notification information for the detection of the object to a terminal of a user of the vehicle.
12. In claim 1, obtaining information about the detected object; and The method for operating an electronic device further includes a step of transmitting information of the obtained object to a terminal of a user of the vehicle.
13. 13. The method according to claim 12, wherein the step of acquiring information about the object comprises: classifying the object based on machine learning; and A method of operating an electronic device comprising obtaining class information of the object based on a result of the classification.
14. 14. The method of claim 13, wherein the step of classifying the objects comprises: If the object is a person, classifying the person as being an adult, a child or an infant based on the size of the object in the reference image and its proportion to the person.
15. 1. An electronic device for detecting an object in a vehicle interior, comprising: a camera for acquiring at least one interior image of the vehicle interior; and An electronic device comprising a processor that generates a reference image using the at least one interior image, and detects objects within the vehicle based on at least one divided region of the generated reference image and at least one divided region of a target image, which is an interior image of the vehicle acquired after the reference image is generated.
16. In claim 15, The camera captures an image of an interior of the vehicle at a predetermined frame rate for a certain period of time in a parking mode of the vehicle, The processor is configured to determine the parking mode of the vehicle by detecting an engine condition of the vehicle.
17. 16. The method of claim 15, wherein the processor: An electronic device that defines a region of interest for each of the at least one indoor image, and combines the indoor images with the region of interest defined therein to generate the reference image.
18. 20. The method of claim 17, wherein the processor: An electronic device that sets the region of interest in a pre-defined manner such that a window area of the vehicle is minimized.
19. 19. The method according to claim 18, wherein the preset method is: The electronic device excludes an image of a first ratio from each of the left and right sides of a center point of a lower end of the reference image, and excludes an image of a second ratio from an upper end of the reference image.
20. 16. The method of claim 15, wherein the processor: An electronic device that divides the target image into a plurality of regions and detects the object in each of the divided regions.
21. 21. The processor of claim 20, An electronic device sequentially applies a plurality of grid containers to the reference image and the target image in ascending order according to division sizes of the grid containers to detect the object.
22. 22. In claim 21, the plurality of grid containers include a plurality of grid containers having the same division size; the plurality of grid containers having the same division size are set such that start positions of grid cells contained in the grid containers are different from each other; A plurality of grid containers having the same division size, but with the starting positions of the grid cells set to be different from each other, are applied together to the indoor image.
23. 22. The method of claim 21, wherein the processor: If the object is detected using the sequentially applied grid container, the electronic device does not detect the object using a next sequential grid container.
24. 22. The method of claim 21, wherein the processor: If the object is detected using the sequentially applied grid container, the electronic device detects the object using a grid container next to the one in which the object was detected.
25. In claim 15, The electronic device further includes a communication circuit for transmitting notification information in response to the detection of the object to a terminal of a user of the vehicle.
26. In claim 15, The processor obtains information about the detected object; The electronic device further includes a communication circuitry for transmitting the obtained object information to a terminal of a user of the vehicle.
27. 27. The method of claim 26, wherein the processor: An electronic device that classifies the object based on machine learning and obtains class information of the object based on a result of the classification.
28. 28. The method of claim 27, wherein the processor: If the object is a person, the electronic device classifies the person as being an adult, a child or an infant based on the size of the object in the reference image and its proportion to the person.
29. A method for operating an electronic device for monitoring a driver's condition, comprising: detecting facial feature points of a driver from an interior image of the vehicle captured by an interior camera provided in the electronic device; transforming the coordinates of the detected feature points into coordinates on a front coordinate system; and determining a state of the driver using a distance between coordinates of at least two of the transformed feature points; A method for operating an electronic device, wherein the front coordinate system is a coordinate system on an image of the driver's face photographed from the front by the interior camera.
30. 30. The method according to claim 29, wherein the step of converting to coordinates on the front coordinate system comprises: acquiring an installation angle of the indoor camera; and converting the coordinates of the feature points into coordinates on the front coordinate system using a mounting angle of the indoor camera.
31. 31. The method according to claim 30, wherein the step of acquiring the mounting angle of the indoor camera includes: detecting a straight lane from a forward image captured by a forward camera provided in the electronic device; detecting a vanishing point of the detected straight lane; obtaining an installation angle of the front camera by comparing the detected vanishing point with a center point of the front image; and obtaining an installation angle of the indoor camera from an installation angle of the front camera.
32. 32. The method according to claim 31, wherein the step of acquiring the mounting angle of the indoor camera comprises: Obtaining an installation angle of the indoor camera using a gyro sensor; The mounting angle of the indoor camera obtained from the mounting angle of the front camera; a step of weighting and adding the mounting angle of the indoor camera acquired by using the gyro sensor; determining the weighted sum as a final mounting angle of the indoor camera.
33. 30. In claim 29, the step of detecting the facial feature points of the driver includes a step of detecting a center point of a left eye as a first feature point, a center point of a right eye as a second feature point, and a center point of a nose as a third feature point on the face of the driver; A method for operating an electronic device, wherein the step of determining the driver's state is based on a ratio of a distance between the coordinates of the first feature point and the coordinates of the third feature point and a distance between the coordinates of the second feature point and the coordinates of the third feature point.
34. 34. The method of claim 33, wherein the step of determining the state of the driver comprises: If the ratio of the distances is equal to or greater than a first reference value or is less than a second reference value, the state of the driver is determined to be a state in which the driver is not paying attention to the road ahead; A method for operating an electronic device, comprising determining that the driver's state is a forward-focused state when the ratio of the distances is greater than or equal to the second reference value and less than the first reference value.
35. 30. In claim 29, At least a plurality of the detected facial feature points of the driver are detected from one of the facial feature portions of the driver, The driver's facial features include at least one of eyes, nose, mouth, eyebrows, and facial contours.
36. 36. The method of claim 35, wherein the step of determining the state of the driver comprises: A method for operating an electronic device, comprising: determining, for at least one of the driver's left eye and right eye, based on a ratio (C) of the sum of the distance between both end eye feature points of the one eye and the distance between at least one upper eye feature point of the one eye and its corresponding lower eye feature point, among a plurality of eye feature points detected for the one eye.
37. 37. The method of claim 36, wherein the step of determining the state of the driver comprises: The method for operating an electronic device further comprises determining that the driver is in a drowsy state when the ratio (C) remains equal to or lower than a third reference value for a certain period of time or more.
38. 36. The method of claim 35, wherein the step of determining the state of the driver comprises: A method for operating an electronic device, comprising a step of determining whether the driver is drowsy and / or whether the driver is looking ahead using a machine learning model that has learned at least some of the facial feature points of the detected driver.
39. In an electronic device for monitoring the driver's condition, An interior camera for photographing the interior of the vehicle; a processor that detects feature points of the face of the driver from an interior image of the vehicle captured by the interior camera, transforms coordinates of the detected feature points into coordinates on a front coordinate system, and determines a state of the driver using a distance between the coordinates of at least two of the transformed feature points; The electronic device, wherein the front coordinate system is a coordinate system on an image of the driver's face photographed from the front by the indoor camera.
40. 40. The method of claim 39, wherein the processor: An electronic device that acquires an installation angle of the indoor camera and converts coordinates of the feature points into coordinates on the front coordinate system using the installation angle of the indoor camera.
41. 41. The method of claim 40, wherein the processor: An electronic device that detects a straight lane from a forward image captured by a forward camera provided in the electronic device, detects a vanishing point of the detected straight lane, compares the detected vanishing point with a center point of the forward image to obtain an installation angle of the forward camera, and obtains an installation angle of the indoor camera from the installation angle of the forward camera.
42. 42. The method of claim 41, wherein the processor: An electronic device that acquires an installation angle of the indoor camera using a gyro sensor, weights and adds the installation angle of the indoor camera acquired from the installation angle of the front camera and the installation angle of the indoor camera acquired using the gyro sensor, and determines the weighted sum value as the final installation angle of the indoor camera.
43. 40. The method of claim 39, wherein the processor: Detecting the center point of the left eye of the driver as a first feature point, the center point of the right eye as a second feature point, and the center point of the nose as a third feature point on the driver's face, thereby detecting the feature points of the driver's face; An electronic device that determines a state of the driver based on a ratio of a distance between the coordinates of the first feature point and the coordinates of the third feature point and a distance between the coordinates of the second feature point and the coordinates of the third feature point.
44. 44. The method of claim 43, wherein the processor: If the ratio of the distances is equal to or greater than a first reference value or is less than a second reference value, the state of the driver is determined to be a state in which the driver is not paying attention to the road ahead; When the ratio of the distances is equal to or greater than the second reference value and is less than the first reference value, the electronic device determines that the driver is in a forward-focused state.
45. 40. In claim 39, At least a plurality of the detected facial feature points of the driver are detected from one of the facial feature portions of the driver, The driver's facial features include at least one of eyes, nose, mouth, eyebrows, and facial contours.
46. 46. The method of claim 45, wherein the processor: The electronic device determines the state of the driver based on a ratio (C) of the sum of the distance between both end eye feature points of the one eye and the distance between at least one upper eye feature point of the one eye and its corresponding lower eye feature point, for each of at least one of the driver's left eye and right eye, among the multiple eye feature points detected for the one eye.
47. 47. The method of claim 46, wherein the processor: When the ratio (C) remains equal to or lower than a third reference value for a certain period of time or more, the electronic device determines that the driver is in a drowsy state.
48. 46. The method of claim 45, wherein the processor: An electronic device that determines whether the driver is drowsy and / or whether the driver is looking ahead using a machine learning model that has learned at least some of the facial feature points of the detected driver.