Electronic device and control method thereof

By overlapping part of multiple continuous image frames in the processor, generating panoramic images and identifying objects, the image blur problem caused by user motion is solved, the accuracy of object recognition is improved, and the motion-stable augmented reality function is provided.

CN113168700BActive Publication Date: 2025-05-27SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980079168.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-19
Filing Date
2019-09-30
Publication Date
2025-05-27
Estimated Expiration
2039-09-30

AI Technical Summary

Technical Problem

Prior art When capturing objects using glasses devices, images are blurred due to user's head or body movement, which affects real-time image analysis of augmented reality (AR) services and object recognition and tracking.

Method used

By configuration in the processor, partial areas of a plurality of consecutive image frames stored in the memory are overlapped to generate a panoramic image and to identify areas of a predetermined shape of the largest size in the panoramic image, thereby improving the accuracy of object recognition.

Benefits of technology

It realizes improving the accuracy of object recognition when users move quickly, provides augmented reality (AR) function including motion-stabilized objects, and reduces the impact of image blur on AR services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113168700B_ABST
    Figure CN113168700B_ABST
Patent Text Reader

Abstract

An electronic device obtains a panoramic image by overlapping partial areas of image frames, and recognizes an object from the panoramic image or an area of ​​a predetermined shape of a maximum size within the panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Apparatuses and methods consistent with the present disclosure relate to electronic devices and control methods thereof, and more particularly, to electronic devices for identifying objects and control methods thereof.

[0002] The present disclosure also relates to an artificial intelligence (AI) system that uses machine learning algorithms and their applications to simulate the functions of the human brain (e.g., recognition and judgment). Background Art

[0003] Recently, artificial intelligence systems that achieve human-level artificial intelligence (AI) have been deployed in various fields. Unlike traditional rule-based intelligent systems, artificial intelligence systems are systems contained in machines that learn, judge, and iteratively improve the performance of their functions. For example, as the use of artificial intelligence systems increases, the recognition rate and understanding of user preferences may increase accordingly. Therefore, traditional rule-based intelligent systems have been gradually replaced by deep learning-based artificial intelligence systems.

[0004] Artificial intelligence technology consists of machine learning (such as deep learning) and element technology that implements machine learning.

[0005] Machine learning is an algorithmic technology that classifies and / or trains the features of input data. Element technology is a technology that uses machine learning algorithms, such as deep learning, to simulate the recognition and judgment functions of the human brain, including language understanding, visual understanding, reasoning / prediction, knowledge representation, motion control, etc.

[0006] Artificial intelligence technology can be applied to various fields, examples of which are described below. Language understanding is a technology for recognizing and applying / processing human language / characters, including natural language processing, machine translation, dialogue systems, query response, speech recognition / synthesis, etc. Visual understanding is a technology for recognizing and processing objects as if they were perceived by humans, including object recognition, object tracking, image search, human recognition, scene understanding, spatial understanding, image enhancement, etc. Reasoning prediction is a technology for judging and logically inferring and predicting information, including knowledge / probability-based reasoning, optimization prediction, preference-based planning and recommendation. Knowledge representation is a technology for automating human experience information into knowledge data, including knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control is a technology for controlling the autonomous movement of equipment or objects, for example, the movement of vehicles and the movement of robots, including motion control (navigation, collision and travel), operation control (behavior control), etc.

[0007] Recently, various types of devices, such as eyeglass devices including cameras, have been developed. However, when capturing an object using a camera provided in the eyeglass device, image blur may occur due to the movement of the user's head or body. For example, referring to Figure 1 , the eyeglass device may include a camera, and due to the user's movement, image blur may appear in the image captured by the camera.

[0008] Image blur can pose a hindrance to augmented reality (AR) services that require real-time image analysis, and object recognition and tracking can be difficult.

[0009] Therefore, stabilization methods have been developed to remove image blur. However, traditionally, parameters can be set for each image frame. In the case of fast image motion, image data may be lost due to over-cropping. Summary of the invention

[0010]

Technical issues

[0011] One aspect of the exemplary embodiments relates to an electronic device for providing an augmented reality (AR) function including a motion-stabilized object by improving object recognition performance from a plurality of consecutive frames, and a control method thereof.

[0012]

Technical Solution

[0013] According to one embodiment, an electronic device is provided, including a memory and a processor, the processor being configured to obtain a panoramic image by overlapping a partial area of ​​a first image frame stored in the memory with a partial area of ​​at least one second image frame stored in the memory based on pixel information of a first frame and pixel information of at least one second image frame, identify an area of ​​a predetermined shape of a maximum size within the panoramic image, and identify an object from the entire area of ​​the panoramic image or the area of ​​the predetermined shape.

[0014] The processor can also be configured to obtain a panoramic image by overlapping areas with minimal difference between pixel values ​​of adjacent image frames in the first image frame and at least one second image frame based on pixel information of the first image frame and pixel information of at least one second image frame.

[0015] The processor can also be configured to obtain motion values ​​between adjacent image frames in the first image frame and at least one second image frame based on pixel information of the first image frame and pixel information of at least one second image frame, and obtain a panoramic image by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the motion values.

[0016] The processor can also be configured to convert motion values ​​based on the difference between pixel values ​​in adjacent image frames and the motion values ​​between adjacent image frames, and obtain a panoramic image by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the converted motion values.

[0017] The processor can also be configured to perform image processing including at least one of rotation, position movement or size adjustment on each of the first image frame and the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame, and obtain a panoramic image by overlapping partial areas of the frames where the image processing is performed.

[0018] The electronic device may further include a display, wherein the processor is further configured to control the display to display the first image frame, and based on the object identified from the panoramic image, control the display to display the object including at least one of a graphical user interface, characters, images, videos, or 3D models on an area where the object is displayed in the first image frame according to information about the location where the object is identified and information about the image processing.

[0019] The at least one second image frame may be an image frame captured before the first image frame, wherein the processor is further configured to update the panoramic image by overlapping a partial area of ​​the panoramic image with a partial area of ​​the third image frame based on pixel information on the panoramic image and a third image frame captured after the first image frame, identify an area of ​​a predetermined shape of a maximum size within the updated panoramic image, and re-identify an object from the updated panoramic image or the area of ​​the predetermined shape within the updated panoramic image.

[0020] The third image frame may be an image frame captured after the first image frame, and the processor may be further configured to re-identify the object in the updated panoramic image, or an area in a predetermined shape within the updated panoramic image, based on a ratio of the first image frame relative to the third image frame and an overlapping area of ​​the third image frame that is less than a predetermined ratio.

[0021] The processor may also be configured to assign weighted values ​​to multiple corresponding areas based on multiple objects identified in the panoramic image, based on at least one of multiple overlapping image frames of the multiple corresponding areas of the panoramic image, or a capture time of image frames in the multiple corresponding areas, and identify at least one of the multiple objects based on the weighted values ​​of the multiple corresponding areas.

[0022] The electronic device may also include a camera including the circuit, wherein the processor is further configured to perform continuous capture by the camera to obtain a plurality of image frames.

[0023] According to an exemplary embodiment, a method for controlling an electronic device is provided, the method comprising obtaining a panoramic image by overlapping a partial area of ​​the first image frame with a partial area of ​​the at least one second image frame based on pixel information of a first image frame and pixel information of at least one second image frame among a plurality of frames, identifying an area of ​​a predetermined shape of a maximum size within the panoramic image, and identifying an object from the entire area of ​​the panoramic image or the area of ​​the predetermined shape.

[0024] The acquiring may include acquiring a panoramic image by overlapping an area having a minimum difference between pixel values ​​of adjacent image frames between the first image frame and the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame.

[0025] The obtaining may include obtaining motion values ​​between adjacent image frames in the first image frame and at least one second image frame based on pixel information of the first image frame and pixel information of at least one second image frame, and obtaining a panoramic image by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the motion value.

[0026] The obtaining may include converting motion values ​​based on differences between pixel values ​​in adjacent image frames and motion values ​​between adjacent image frames, and obtaining a panoramic image by overlapping a partial area of ​​a first image frame with a partial area of ​​at least one second image frame based on the converted motion values.

[0027] The acquisition may include performing image processing including at least one of rotation, position movement or size adjustment relative to each of the first image frame or the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame, and acquiring a panoramic image by overlapping partial areas of the frames in which the image processing is performed.

[0028] The method may further include displaying a first image frame, and based on the object identified from the panoramic image, based on information about the position of the object identified from the panoramic image and information about image processing, displaying an object including at least one of a graphical user interface, a character, an image, a video, or a 3D model on an area where the object is displayed in the first image frame.

[0029] The at least one second frame may be a frame captured before the first image frame, wherein the method further includes updating the panoramic image by overlapping a partial image of the panoramic image with a partial area of ​​the third image frame based on pixel information on the panoramic image and a third image frame captured after the first image frame, identifying an area of ​​a predetermined shape of a maximum size within the updated panoramic image, and re-identifying the object from the entire area of ​​the updated panoramic image or the area of ​​the predetermined shape within the updated panoramic image.

[0030] The third image frame may be an image frame captured after the first image frame, wherein re-identification of the object includes re-identifying the object in an updated panoramic image, or an area in a predetermined shape within the updated panoramic image, based on a ratio of the first image frame relative to the third image frame and an overlapping area of ​​the third image frame that is less than a predetermined ratio.

[0031] Recognition of the object may include assigning weighted values ​​to the multiple corresponding areas based on the multiple objects recognized in the panoramic image, based on capture time of image frames in at least one or more of the multiple overlapping image frames of the multiple corresponding areas of the panoramic image, and recognizing at least one of the multiple objects based on the weighted values ​​of the multiple corresponding areas.

[0032] The method may further include performing continuous capturing by a camera provided in the electronic device to obtain a plurality of image frames.

[0033]

Beneficial effects

[0034] According to aspects of the present disclosure, an electronic device generates a panoramic image from a plurality of consecutive frames, recognizes an object from a frame region with a minimum motion difference, and improves the accuracy of object recognition despite rapid camera movement, thereby providing a user with augmented reality (AR) including a motion-stable object. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a view showing an eyeglass device;

[0036] Figure 2A is a block diagram of an electronic device according to an embodiment;

[0037] Figure 2B is a block diagram of an electronic device according to an embodiment;

[0038] Figure 2C is a block diagram of multiple modules of an electronic device according to an embodiment;

[0039] Figure 3 is a view explaining motion stabilization according to an embodiment;

[0040] Figure 4A is a view explaining a method for generating a panoramic image according to an embodiment;

[0041] Figure 4B is a view explaining a method for generating a panoramic image according to an embodiment;

[0042] Figure 5A is a view explaining a method for generating a panoramic image for each reference frame according to an embodiment;

[0043] Figure 5Bis a view explaining a method for generating a panoramic image for each reference frame according to an embodiment;

[0044] Figure 6 is a view explaining a method for recognizing an object and displaying an object recognition result according to an embodiment;

[0045] Fig. 7A is a view explaining a re-recognition operation of an object according to an embodiment;

[0046] Figure 7B is a view explaining a re-recognition operation of an object according to an embodiment;

[0047] Figure 8 is a view explaining a method for recognizing a final object according to an embodiment;

[0048] Fig. 9 is a block diagram showing an electronic device according to an embodiment;

[0049] Fig.10 is a block diagram of a training unit according to an embodiment;

[0050] Fig.11 is a block diagram of a response unit according to an embodiment;

[0051] Fig.12 is a view showing an example in which an electronic device according to an embodiment can operate in association with an external server to train and judge data; and

[0052] Fig.13 is a flowchart of a method for controlling an electronic device according to an embodiment. DETAILED DESCRIPTION

[0053]

Invention Mode

[0054] The embodiments of the present disclosure may be modified in many ways. Therefore, the embodiments are shown in the drawings and described in detail in the detailed description. However, it should be understood that the present disclosure is not limited to specific embodiments, but includes all modifications, equivalents and substitutions without departing from the scope and spirit of the present disclosure. In addition, well-known functions or configurations are not described in detail to avoid unnecessary details that obscure the present disclosure.

[0055] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.

[0056] All terms used in this specification, including technical and scientific terms, have the same meanings as commonly understood by technicians in the relevant fields. However, these terms may change according to the intentions of technicians in this field, legal or technical interpretations, and the emergence of new technologies. In addition, the applicant may choose some terms. These terms may be interpreted according to the meanings defined herein and, unless otherwise stated, may be interpreted based on the entire content of this specification and the common technical knowledge in the field.

[0057] Singular expressions also include plural meanings as long as it does not have a different meaning in the context. In this specification, terms such as "include" and "have / contain" should be interpreted as indicating the presence of such features, numbers, operations, elements, components or combinations thereof in the specification, and do not exclude the presence or possibility of adding one or more other features, numbers, operations, elements, components or combinations thereof.

[0058] In the present disclosure, the expression "A or B", "at least one of A and / or B", or "one or more of A and / or B", etc. includes all possible combinations of the listed items.

[0059] Terms and labels, such as "first" and "second", are used to distinguish one component from another but do not limit the components.

[0060] When one element (e.g., a first component) is referred to as being "operably (or communicatively) coupled" or "connected to" another element (e.g., a second component), the element is indirectly connected or coupled to the other element, or is connected or coupled to the other element with one or more intermediate elements (e.g., a third component) interposed therebetween. However, when one element (e.g., a first component) is "directly connected" or "directly coupled" to another element (e.g., a second component), there is no other element (e.g., a third component) between the one element and the other element.

[0061] In one embodiment, a "module", "unit" or "component" performs at least one function or operation and may be implemented as hardware (e.g., a processor or an integrated circuit), software executed by a processor, or a combination thereof. In addition, multiple "modules", "multiple" units or multiple "components" may be integrated into at least one module or chip and may be implemented as at least one processor, except for the "module", "unit" or "component" that should be implemented in specific hardware.

[0062] In this specification, the term "user" may refer to a person who uses an electronic device or a device (eg, a man-made electronic device) that uses the electronic device.

[0063] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.

[0064] Figure 2A is a block diagram of an electronic device 100 according to an embodiment.

[0065] refer to Figure 2A , the electronic device 100 may include a memory 110 and a processor 120 .

[0066] The electronic device 100 according to various embodiments may be, for example, an augmented reality (AR) glasses device. The AR glasses device may be a device for providing an augmented reality function. Augmented reality may be a technology that allows a user to perceive a virtual object overlaid on a real environment through a glasses device. For example, when a virtual object is overlaid and displayed on a real environment when viewed from a user through glasses, the user may recognize the virtual object as part of the real world. Through augmented reality, a real image may be provided by overlaying a virtual object on an actual image viewed from a user, because the actual environment and the virtual screen cannot be clearly distinguished.

[0067] The electronic device 100 according to various embodiments of the present disclosure may be a smartphone, a tablet personal computer (desktop PC), a mobile phone, a video phone, an e-book reader, a laptop personal computer (laptop PC), a netbook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device, or they may be a part thereof. The wearable device may be an accessory type device, such as a watch, a ring, a bracelet, a bracelet, a necklace, a pair of glasses, contact lenses, or a head mounted device (HMD), a fabric or clothing integrated type (e.g., an electronic device), a body attachment type (e.g., a skin pad or tattoo), or a bio-implantable circuit.

[0068] In some embodiments, an example of an electronic device may be a household appliance. The household appliance may include, for example, a television, a digital video disk (DVD) player, an audio system, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a set-top box, a home automation control panel, a security control panel, a TV box (e.g., Samsung HomeSync TM ,Apple TV TM , or GoogleTV TM ), game consoles (such as Xbox TM and PlayStation TM ), electronic dictionary, electronic key, camera or electronic photo frame.

[0069] In another embodiment, the electronic device can be any of various medical devices (for example, various portable medical measuring devices, such as blood glucose meters, heart rate meters, blood pressure monitors, or thermometers), magnetic resonance angiography (MRA), magnetic resonance imaging (MRI), computed tomography (CT), cameras, ultrasound equipment, navigation equipment, global navigation satellite systems (GNSS), event data recorders (EDR), flight data recorders (FDR), automotive infotainment equipment (for example, navigation equipment, gyrocompasses, etc.), avionics devices, security equipment, a head unit of a vehicle, an industrial or home robot, an ATM (automatic teller machine) of a financial institution, a point of sale (POS) of a store, or the Internet of Things (IoT), such as light bulbs, various sensors, electricity or gas meters, sprinklers, fire alarms, thermostats, street lights, toasters, exercise equipment, hot water tanks, heaters, boilers, etc.

[0070] The electronic device 100 may be any object that recognizes an object in a plurality of image frames.

[0071] The memory 110 may store a plurality of image frames. The plurality of image frames may be a plurality of frames captured by a camera provided in the electronic device 100. When the electronic device 100 is an AR glasses device, image blur may occur due to movement of the AR glasses caused by the user's motion while capturing the frames.

[0072] Multiple frames may be images of a scene. For example, multiple frames may be images of an intersection. However, due to the movement of the user, the intersection in the multiple frames may be captured at different positions within the frames. For example, in the first frame, the center of the intersection may be the center of the frame. However, in the second frame, the center of the intersection may not be the center of the frame.

[0073] The present disclosure is not limited thereto, but the plurality of frames may be images of a plurality of scenes. The processor 120 may divide the plurality of frames into a plurality of groups based on a plurality of corresponding scenes. Each of the plurality of groups may be a scene.

[0074] However, the present disclosure is not limited thereto. The electronic device 100 may receive a plurality of frames captured by an external device, and the memory 110 may store the received plurality of frames.

[0075] The memory 110 may be implemented as a hard disk, a nonvolatile memory, and a volatile memory, as well as any type of memory that stores data.

[0076] The processor 120 may control the overall operation of the electronic device 100 .

[0077] According to an embodiment, the processor 120 may be implemented as a digital signal processor (DSP), a microprocessor or a time controller (TCON), but is not limited thereto. The processor 120 may include one or more central processing units (CPUs), microcontroller units (MCUs), microprocessing units (MPUs), controllers, application processors (APs), or communication processors (CPs), ARM processors, etc., or may be defined by corresponding terms. The processor 120 may be implemented as a system on chip (SoC), a large-scale integrated circuit (LSI) with a built-in processing algorithm, or in the form of a field programmable gate array (FPGA). The processor 120 may perform various functions by executing computer executable instructions stored in and loaded from the memory 120.

[0078] The processor 120 may obtain a panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the at least one second frame based on pixel information on a first frame and at least one second frame among the plurality of frames stored in the memory 110 .

[0079] For example, the processor 120 may obtain a panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the second frame based on the pixel information of the first frame and the pixel information of the second frame among the multiple frames stored in the memory 110. Specifically, each of the first frame and the second frame may be an image with a resolution of 1920×1080, and if the pixel value of the first square area from the pixel (21, 1) to the pixel (1920, 1080) of the first frame is consistent with the pixel value of the second square area from the pixel (1, 1) to the pixel (1900, 1080) of the first frame, the processor 120 may obtain a panoramic image such that the first square area of ​​the first frame overlaps the second square area of ​​the second frame. This is because, although the difference between the shooting time of the first frame and the shooting time of the second frame is very small during the actual shooting process, the shooting angle may be changed by the influence of the user's motion.

[0080] For convenience of explanation, this embodiment describes that two frames overlap each other, but the processor 120 may obtain a panoramic image by overlapping three or more images.

[0081] The processor 120 can obtain a panoramic image by overlapping the area where the pixel values ​​between adjacent frames have the smallest difference based on the pixel information of each of the first frame and the second frame. In the case of more than two frames, the pixel information of each frame can be considered. This is because in the actual shooting process, although the difference between the shooting points of the first frame and the second frame is very small, the pixel value may change due to the change in the amount of light. For example, in the above example, the pixel values ​​of the first square area of ​​the first frame and the second square area of ​​the second frame may be inconsistent with each other. Therefore, the processor 120 can obtain the area with the smallest pixel value difference between the overlapping areas by moving the first frame on the second frame pixel by pixel, and obtain a panoramic image by overlapping the first frame with the second frame so that the pixel value difference is minimized.

[0082] The processor 120 may obtain a motion value between adjacent frames based on pixel information of each of the first frame and the second frame (and if more than two frames are used, additional frames), and obtain a panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the second frame based on the obtained motion value. In the above example, the processor 120 may obtain a motion value (20, 0) between the first frame and the second frame based on pixel information on the first frame and the second frame. The processor 120 may obtain a panoramic image by overlapping the first frame with the second frame based on the motion value.

[0083] However, the present disclosure is not limited thereto, and the motion value may be a value obtained when capturing a frame. For example, the electronic device 100 may further include a camera including a circuit, and the processor 120 may obtain a plurality of frames by continuously shooting with the camera. The processor 120 may sense the motion of the electronic device 100 through a sensor based on the corresponding capture times of the plurality of corresponding frames. The processor 120 may obtain the motion value between adjacent frames based on the sensed motion of the electronic device 100.

[0084] The processor 120 may convert motion values ​​obtained based on the difference between pixel values ​​in adjacent frames and motion values ​​(motion vectors) between adjacent frames, and overlap a partial area of ​​the first frame with a partial area of ​​at least one second frame based on the converted motion values ​​to obtain a panoramic image. For example, the processor 120 may stabilize motion values ​​between adjacent frames of a plurality of consecutive frames and convert the motion values. The deviation between adjacent motion values ​​may be reduced by converting the motion values.

[0085] The processor 120 may stabilize the motion value based on the pixel values ​​of adjacent frames. The processor 120 may stabilize each motion value based on a motion stabilization model. The motion stabilization module may be a model obtained by artificial algorithm training to stabilize the motion value based on the difference between the pixel values ​​in adjacent frames. The processor 120 may perform a stabilization operation on each motion value and reduce the deviation between adjacent motion values. By doing so, image blur that occurs when reproducing multiple frames may be reduced by minimizing the deviation between adjacent motion values.

[0086] The method of obtaining a panoramic image using the converted motion values ​​may be the same as the method of obtaining a panoramic image using the motion values ​​before conversion, and therefore, a detailed description thereof will be omitted.

[0087] The processor 120 may identify a region of a predetermined shape of maximum size in the panoramic image. The maximum size may correspond to the maximum size of the predetermined shape engraved within the overlapping images of the panoramic image. For example, the processor 120 may identify a square region of maximum size having the same aspect ratio as multiple frames in the panoramic image. However, the present disclosure is not limited thereto. The predetermined shape may be various shapes, such as a diagonal shape, a circular shape, etc. For ease of explanation, a square region will be an example of a region of a predetermined shape.

[0088] The processor 120 may recognize the object from the entire panoramic image or a region of a predetermined shape. For example, the processor 120 may recognize the object from the panoramic image and recognize a region of a predetermined shape of the recognized region as a final recognition region.

[0089] The processor 120 can perform at least one of image processing, such as rotation, position movement, and size adjustment, on each of the first frame and the at least one second frame based on pixel information on the first frame and the at least one second frame, and obtain a panoramic image by overlapping partial areas of the frames on which the image processing is performed.

[0090] The electronic device 100 may further include a display, and the processor 120 may control the display to display a first frame thereon, and when an object is recognized from the panoramic image, the display may be controlled to display an object thereon, the object including at least one of a graphical user interface (GUI), a character, an image, a video, or a 3D model on an area where the object is displayed in the first frame based on information about a position where the object is recognized from the panoramic image and image processing information. The processor 120 may perform object recognition within the panoramic image and display the object on the recognized area of ​​the display frame. The object may be a virtual 2D / 3D content.

[0091] The at least one second frame may be a frame captured before the first frame. The processor 120 may update the panoramic image by overlapping a partial area of ​​the panoramic image with a partial area of ​​the third frame based on the panoramic image and pixel information on the third frame captured after the first frame, identify an area of ​​a determined shape having a maximum size within the updated panoramic image, and again identify an object in the updated panoramic image or an area having a predetermined shape within the updated panoramic image.

[0092] In other words, the processor 120 may generate a panoramic image based on each of the plurality of frames, and thus generate a number of panoramic images corresponding to the plurality of frames. The processor 120 may overlap a panoramic image corresponding to a frame immediately before the current frame with the current frame to generate a panoramic image corresponding to the current frame. That is, the processor 120 may obtain a panoramic image by updating the panoramic image according to changes in the current frame. When the panoramic image is updated, the processor 120 may remove the oldest frame from the panoramic image.

[0093] The third frame may be a frame captured immediately following the first frame, and if the ratio between the first frame relative to the third frame and the overlapping area of ​​the third frame is less than a predetermined ratio, the processor 120 may obtain an image of a new scene that cannot be derived from the first and second frames. Therefore, the object may be recognized again within the panoramic image updated to identify the new object or within an area of ​​a predetermined shape of the updated panoramic image. If the ratio between the first frame relative to the third frame and the overlapping area of ​​the third frame is equal to or greater than a predetermined ratio, the processor 120 may obtain an image of a scene similar to the first frame, and the object may not be recognized again in the updated panoramic image because the scene has not substantially changed. In this case, the processor 120 may utilize the object information recognized from the panoramic image before updating, i.e., without further updating. Therefore, compared to the conventional method of recognizing an object each time an input image is provided, the power consumption of the electronic device 100 may be reduced by reducing the number of object recognition operations based on the movement of the image captured by the camera.

[0094] When multiple objects are identified from a panoramic image, the processor 120 may assign a weighted value to each of the multiple regions based on at least one of a recognition confidence value obtained as a result of object recognition, the number of overlapping frames in each of the multiple regions in the panoramic image, or a capture time of frames in each of the multiple regions, and identify at least one of the multiple objects based on the weighted value assigned to each of the multiple regions.

[0095] For example, when multiple objects are identified from a panoramic image, the processor 120 may identify the number of overlapping frames at the position of each of the multiple objects, and identify the object identified in the area with the largest number of overlapping frames as the final object. When multiple objects are identified from a panoramic image, the processor 120 may identify the capture time of the overlapping frames at the position of each of the multiple objects, assign a weight value to each of the multiple objects by assigning the highest weight to the last captured frame, and identify the object to which the highest weight value is assigned as the final object.

[0096] This method for identifying an object can reduce errors occurring in object identification by filtering the object identification results obtained from the oldest frame. For example, if there is a passing car in the first frame and there is no car in the third frame obtained after the first frame, then in the object identification results obtained from the panoramic image including the first frame, the third frame and the frames obtained subsequently, the currently non-existent car can be excluded from the final object identification results.

[0097] Figure 2B is a detailed block diagram of the electronic device 100. The electronic device 100 may include a memory 110 and a processor 120. Figure 2B , the electronic device 100 may include a communication interface 130 , a display 140 , a user interface 150 , an input / output interface 160 , a camera 170 , a speaker 180 , and a microphone 190 .

[0098] The memory 110 may be implemented as an internal memory such as a ROM (e.g., an electrically erasable programmable read-only memory (EEPROM)), a RAM, or a memory separate from the processor 120. In this case, the memory 110 may be implemented in the form of a memory embedded in the electronic device 100 or a removable memory in the electronic device 100, depending on the purpose of data storage. For example, data for driving the electronic device 100 may be stored in a memory embedded in the electronic device 100, and data for extended functions of the electronic device 100 may be stored in a memory that is attachable to or detachable from the electronic device 100. The memory embedded in the electronic device 100 can be implemented with at least one of a volatile memory (e.g., dynamic RAM (DRAM), or static RAM (SDRAM), synchronous dynamic RAM (SDRAM), etc.), a non-volatile memory (e.g., one-time programmable ROM (OTPROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), shielded ROM, flash ROM, flash memory (e.g., NAND flash memory or NOR flash memory), a hard disk drive, or a solid-state drive (SSD). The memory that can be removed from the electronic device 100 can be implemented with a memory card (e.g., compact flash, secure digital (SD), micro secure digital (SD), mini secure digital (SD), extreme digital (xD), etc.), an external memory that can be connected to a USB port (e.g., USB memory), etc.

[0099] The memory 110 may store various data, such as an operating system (O / S) software module, a motion stabilization module, a panoramic image generation module, an object recognition module, etc., for implementing functions of the electronic device 100 .

[0100] The processor 120 may control the overall operation of the electronic device 100 by executing various programs stored in and loaded from the memory 110 .

[0101] The processor 120 may include a RAM 121 , a ROM 122 , a main CPU 123 , a graphic processor or a graphic processing unit (GPU) 124 , first to nth interfaces 125 - 1 to 125 - n , and a bus 126 .

[0102] The RAM 121 , the ROM 122 , the main CPU 123 , the GPU 124 , and the first to nth interfaces 125 - 1 to 125 - n may be connected through a bus 126 .

[0103] ROM 122 may store command sets and the like for system startup. If a power-on command is input and power is supplied, CPU 123 may copy an operating system stored in storage 180 to RAM 121 according to the command stored in ROM 122, execute the operating system, and perform system booting. When booting is completed, CPU 123 may copy various programs stored in storage 180 to RAM 121, execute the application programs copied to RAM 121, and perform various operations.

[0104] The main CPU 123 may access the memory 110 and perform booting by using the O / S stored in the memory 110. The main CPU 123 may perform various operations by using various programs, content data, etc. stored in the memory 110.

[0105] The first to nth interfaces 125-1 to 125-n may be connected to various constituent elements as described above. One of the interfaces may be a network interface connected to an external device through a network.

[0106] The processor 120 and / or the GPU 124 may perform graphics processing (video processing). The processor 120 and / or the GPU 124 may generate a screen including various objects such as icons, images, texts, etc. by using a computing unit and a rendering unit. The computing unit may calculate the attribute values ​​of the object according to the screen layout, such as coordinate values, shapes, sizes, colors, etc., by using the received control commands. The rendering unit may generate screens including various layouts of the object based on the attribute values ​​calculated by the computing unit. The screen generated by the rendering unit may be displayed in the display area of ​​the display 140. The processor 120 and / or the GPU 124 may perform various processing on the video data, such as decoding, amplification, noise filtering, etc.

[0107] The processor 120 may be configured to perform processing of audio data. The processor 120 may perform various image processing, such as decoding, scaling, noise filtering, frame rate conversion, resolution conversion, etc. of audio data.

[0108] The communication interface 130 can perform communication with various types of external devices according to various types of communication methods. The communication interface 130 includes a Wi-Fi module 131, a Bluetooth module 132, an infrared communication module 133, a wireless communication module 134, etc. Each communication module can be implemented in the form of at least one hardware chip.

[0109] The processor 120 may communicate with various external devices using the communication interface 130. The external device may be a display device such as a TV, a video processing device such as a set-top box, an external server, a control device such as a remote controller, an audio output device such as a Bluetooth speaker, a lighting device, a home appliance such as a smart vacuum cleaner and a smart refrigerator, a server such as an IOT home manager, and the like.

[0110] The Wi-Fi chip 131 or the Bluetooth chip 132 can perform communication using the Wi-Fi method and the Bluetooth method, respectively. When the Wi-Fi chip 131 or the Bluetooth chip 132 is used, various connection information such as SSID and session key can be first sent and received, a communication connection can be established based on the connection information, and various information can be sent and received based on this.

[0111] The infrared communication module 133 may perform communication according to the Infrared Data Association (IrDA) technology for wirelessly transmitting data at a short distance using infrared rays between temporal rays and millimeter waves.

[0112] The wireless communication module 134 may include at least one communication chip for forming communications according to various communication standards, such as IEEE, ZigBee, third generation (3G), third generation partnership project (3GPP), long term evolution (LTE), fourth generation (4G), fifth generation (5G), etc.

[0113] Furthermore, the communication interface 130 may include at least one of a LAN (Local Area Network) module, an Ethernet module, or a wired communication module that performs communication using a pair cable, a coaxial cable, or an optical fiber cable.

[0114] According to one example, the communication interface 130 may communicate with external devices such as a remote controller and an external server using the same communication module (eg, a Wi-Fi module).

[0115] According to another example, the communication interface 130 may use different communication modules (e.g., Wi-Fi modules) to communicate with external devices such as remote controllers and external servers. For example, the communication interface 130 may use at least one of an Ethernet module or a Wi-Fi module to communicate with an external server, and may use a Bluetooth module to communicate with an external device such as a remote controller. However, this is only an example, and when communicating with multiple external devices or external servers, the communication interface 130 may use at least one communication module among various communication modules.

[0116] Meanwhile, according to an embodiment, the electronic device 100 may further include a tuner and a demodulator.

[0117] The tuner may receive a Radio Frequency (RF) broadcast signal by tuning a channel selected by a user or all pre-stored channels in an RF broadcast signal received through an antenna.

[0118] The demodulation unit may receive and demodulate the digital IF signal DIF converted by the tuner and perform channel decoding.

[0119] The display 140 may be implemented as various types of displays, such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display panel (PDP), etc. The display 140 may further include a driving circuit, a backlight unit, etc., which may be implemented in the form of an a-si TFT, a low temperature polycrystalline silicon (LTPS) TFT, an organic TFT (OTFT), etc. Meanwhile, the display 140 may be implemented as a touch screen combined with a touch sensor, a flexible display, a three-dimensional display (3D display), etc.

[0120] In addition, according to an embodiment, the display 140 may include a display panel for outputting an image, and a frame for accommodating the display panel. In particular, according to an embodiment, the frame may include a touch sensor for sensing user interaction.

[0121] The user interface 150 may be implemented as a device such as a button, a touch pad, a mouse, and a keyboard, or a touch screen capable of performing the above-mentioned display function and operation input function. The button may be a button of various types, such as a mechanical button, a touch pad, a rotary button, etc., arranged in a random area outside the body of the electronic device 100, such as a front surface unit, a side surface unit, and a rear surface unit.

[0122] The input / output interface 160 may be an interface of any one of a High Definition Multimedia Interface (HDMI), a Mobile High Definition Link (MHL), a Universal Serial Bus (USB), a Display Port (DP), Thunderbolt, a Video Graphics Array (VGA) port, an RGB port, a D-Subminiature (D-SUB), and a Digital Visual Interface (DVI). In addition to the communication interface 130, the input / output interface 160 may also be used to communicate with an external device.

[0123] The input / output interface 160 may input and output at least one of an audio signal and a video signal.

[0124] According to an example embodiment, the input / output interface 160 may include a port for inputting and outputting only an audio signal and a port for inputting and outputting only a video signal, respectively, and may be implemented as a single port for inputting and outputting audio and video signals.

[0125] The electronic device 100 may be implemented as a device without a display to transmit an image signal to a separate display device.

[0126] The camera 170 may be configured to capture still images or video images under the control of the user. The camera may capture still images at a specific point, or continuously capture still images.

[0127] The speaker 180 may be configured to output various alarm sounds or voice messages in addition to various audio data processed by the input / output interface 160 .

[0128] The microphone 190 may be configured to receive a user's voice and other sounds and convert the user's voice and other sounds into audio data.

[0129] The microphone 190 can receive the user's voice in an activated state. For example, the microphone 190 can be formed integrally with the electronic device 100 upward, forward, or sideways. The microphone 190 may include a microphone for collecting the user's voice in an analog form, an amplifier circuit for amplifying the collected user's voice, an audio-to-digital (A / D) conversion circuit for sampling the amplified user's voice and converting the sampled user's voice into a digital signal, a filter circuit for removing noise components from the converted digital signal, and the like.

[0130] The electronic device 100 may receive a user voice signal from an external device including a microphone. The received user voice signal may be a digital voice signal, but may also be an analog voice signal. For example, the electronic device 100 may receive the user voice signal via a wireless communication method such as Bluetooth, Wi-Fi, etc. The external device may be implemented as a remote control device or a smart phone.

[0131] The electronic device 100 may transmit a voice signal to an external server for voice recognition of a voice signal received from an external device.

[0132] The communication module for communicating with an external device or an external server may be implemented as one or more modules. For example, the electronic device may perform communication by using a Bluetooth module and perform communication with an external server by using an Ethernet modem or a Wi-Fi module.

[0133] The electronic device 100 may receive a voice and convert the voice into a sentence. For example, the electronic device 100 may directly apply speech-to-text (STT) to a digital voice signal received through the microphone 190 and convert the digital voice signal into text information.

[0134] The electronic device 100 may send a digital voice signal received by a voice recognition server. The voice recognition server may convert the digital voice signal into text information by using speech-to-text (STT). The voice recognition server may send the text information to another server or electronic device for performing a search corresponding to the text information. In some cases, the electronic device 100 may directly perform a search for information corresponding to the text information.

[0135] Figure 2C is a block diagram of a plurality of modules of the electronic device 100. The memory 110 may store an O / S software module 111, a motion stabilization module 112, a panoramic image generation module 113, an object recognition module 114, and an adjustment module 115.

[0136] The O / S software module 111 may be a module for controlling the overall operation of the electronic device 100. For example, the O / S software module 111 may be used to turn on or off the electronic device 100 and may include operation information such as memory management in a standby state.

[0137] The motion stabilization module 112 may be a module for stabilizing motion values ​​between multiple frames. The motion stabilization module 112 may stabilize motion values ​​between multiple frames and eliminate image blur when replaying a video for image processing.

[0138] The panoramic image generation module 113 may be a module for generating a panoramic image by overlapping a plurality of frames having stable motion values ​​through the motion stabilization module 112. When the current frame changes, the panoramic image generation module 113 may update the existing panoramic image. The panoramic image generation module 113 may remove the oldest frame from the panoramic image and overlap a new frame with the panoramic image.

[0139] The object recognition module 114 may be a module for recognizing an object from the panoramic image generated by the panoramic image generation module 113. The object recognition module 114 may recognize an object from the entire panoramic image, or recognize an object from an area of ​​a predetermined shape after determining an area of ​​a predetermined shape in the panoramic image. If a meaningful object is not recognized after recognizing an object in an area of ​​a predetermined shape, the object recognition module 114 may recognize an object from the entire panoramic image. A meaningful object may be an object recognized by the object recognition module 114. A meaningful object may be an object set by a user.

[0140] The adjustment module 115 may be a module for identifying the display position of the GUI indicating an object. The frame captured by the camera 170 may be different from the image viewed by the user through the AR glasses. When the viewing angle of the camera 170 is large, the frame captured by the camera 170 may include an area larger than the area of ​​the image viewed through the user's AR glasses. The adjustment module 115 may calculate the position to be displayed on the display 140 of the AR glasses to indicate that the object of the image viewed from the GUI indicates the object of the user. For example, the adjustment module 115 can obtain the display position of the GUI indicating the object by using a predetermined ratio between the frame captured according to the viewing angle and the image viewed through the user's AR glasses.

[0141] As described above, the processor 120 may generate a panoramic image from a plurality of consecutive frames and recognize an object from the generated panoramic image.

[0142] Hereinafter, the operation of the electronic device 100 will be described in detail with reference to the accompanying drawings.

[0143] Figure 3 2 is a diagram for explaining motion stabilization according to an embodiment of the present disclosure. Figure 3 , Ct may indicate a motion value, and Ct* may indicate a converted motion value.

[0144] The plurality of motion values ​​may be transformed to reduce the deviation between motion values ​​of adjacent frames of the plurality of frames. Figure 3 Ct shown on the left, if the function indicating the motion value is low-pass filtered, the motion value can be converted, such as Figure 3 As shown at Ct* on the left, image blur can be reduced when reproducing a plurality of frames with a changed motion value, rather than when reproducing a plurality of frames before the changed motion value.

[0145] However, reference Figure 3 , the result of converting the motion value according to the conventional technology that does not reflect the characteristics of each frame is shown on the left. The motion value can be converted based on a first function indicating the degree of adjustment of the overlapping area and a second function relative to the motion value. In the conventional technology, the first function and the second function have been applied to all motion values ​​in the same way.

[0146] Figure 3 The processor 120 is shown to stabilize each motion value based on the motion stabilization model on its left side. The motion stabilization model can be a module trained and obtained by an artificial algorithm to stabilize the motion value based on the difference of pixel values ​​between adjacent frames. When the motion stabilization model is used, the first function and the second function can be identified by reflecting the characteristics of each frame. Different functions can be applied to each motion value. Different functions mean that the parameters of the function are different.

[0147] When the image changes seriously, a stabilization function that is improved over conventional cases can be exhibited. For example, according to conventional technology, if there are parts with serious image blur and parts without image blur in multiple frames, this feature is not reflected, and stabilization can be performed based on a function. Therefore, the part with serious image blur may not shake much, but the part without image blur may shake. In this regard, when a motion stabilization model is used, stabilization can be performed based on a function generated for each motion value. Therefore, the part with serious image blur may shake less, and the part without image blur may remain in a state without shaking.

[0148] Figure 4A and Figure 4B is a view explaining a method for generating a panoramic image according to an embodiment of the present disclosure.

[0149] The processor 120 may obtain a panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the at least one second frame based on pixel information on each of the first frame and the at least one second frame of the plurality of frames. Figure 4A As shown, the processor 120 can obtain a panoramic image by using a first frame and N second frames. The processor 120 can obtain a panoramic image by overlapping N+1 frames. The first frame can be a current frame, and the second frame can be a frame captured before the first frame. The first frame and the N second frames can be consecutive frames.

[0150] The processor 120 may obtain a transfer function ΔHt based on the converted motion value Ct* and the motion value Ct, and overlap the first frame with the N second frames based on the transfer function.

[0151] The processor 120 may identify a square area 410 of the largest size in the panoramic image. Conventionally, the resolution may be reduced because overlapping square areas 420 are used in the entire area of ​​a plurality of frames. However, according to the present disclosure, a panoramic image may be generated, and the resolution may be increased to a resolution exceeding that of conventional techniques because a square area 410 of the largest size is used in the panoramic image. In addition, the resolution of the square area 410 may be higher than the resolution of the frame.

[0152] refer to Figure 4B , the processor 120 may convert the first frame into a square area 410 .

[0153] Figure 5A and Figure 5B Detailed description is given of a method for generating a panoramic image for each reference frame according to an embodiment. Figure 5A and 5B , for convenience of explanation, a panoramic image is obtained by using a first frame and two second frames.

[0154] refer to Figure 5A , the processor 120 may obtain a first panoramic image by using a first frame (frame t-1) and two second frames (frames t-2 and t-3). T-1, t-2, t-3 may only be used to explain the order of the frames and do not have a temporal meaning. That is, Figure 5A This is a view explaining generation of a first panoramic image for converting a first frame (frame t-1).

[0155] The processor 120 may obtain a first panoramic image by overlapping the first frame with the two second frames based on pixel information on the first frame and the two second frames. The processor 120 may replace the first frame (frame t-1) with a square area of ​​the largest size in the first panoramic image.

[0156] refer to Figure 5B , the processor 120 may obtain a second panoramic image using the first frame (frame t) and two second frames (frames t-1 and t-2). The processor 120 may overlap the frames by analyzing the three frames and using the first panoramic image.

[0157] The processor 120 may remove the non-overlapping portion of frame t-3 from the first panoramic image and overlap frame t to obtain a second panoramic image. In this way, a panoramic image may be generated in real time.

[0158] The processor 120 may replace the first frame (frame t) with the square area of ​​the largest size in the second panoramic image.

[0159] Processor 120 may minimize image blur by repeating the same task over multiple frames.

[0160] Figure 6 is a view explaining a method for recognizing an object and displaying an object recognition result according to an embodiment.

[0161] refer to Figure 6 , stages 1 and 5 show images displayed by the display 140 of the electronic device 100. For example, when the electronic device 100 is an AR glasses device, the processor 120 may control the camera 170 and capture frames, as shown in stages 1 and 5, and control the display 140 to display the captured frames. The processor 120 may control the display 140 to display frames in stages 1 and 5, while performing operations in stages 2 to 4.

[0162] The processor 120 may perform image processing on the frame in stage 1, and overlap the image-processed frame with the previous panoramic image to generate a panoramic image in stage 2. For example, the processor 120 may rotate the frame of stage 1 by 30 degrees clockwise to additionally overlap the frame with the previous panoramic image.

[0163] In stage 3, the processor 120 may recognize an object from the panoramic image. The processor 120 may convert information about the position of the recognized object from the panoramic image based on the image processing information about the frame of stage 1. For example, the processor 120 may rotate the information about the position of the recognized object from the panoramic image by 30 degrees in the counter-clockwise direction as a frame of stage 4. The processor 120 may convert only the information about the position of the recognized object from the panoramic image based on the image processing information, but converts the information about the position of the recognized object from the panoramic image together with the image processing frame based on the image processing information of stage 1.

[0164] The processor 120 may control the display 140 to display a graphical user interface (GUI) 610 , 620 in a region of the display object of the frame of stage 1 , such as the frame of stage 5 .

[0165] As described above, object recognition performance can be improved by using a panoramic image. In particular, even if the user moves, object recognition performance can be improved by matching with the previous frame.

[0166] Fig. 7A and Figure 7B 2 is a view explaining a re-recognition operation of an object according to an embodiment.

[0167] The processor 120 may update the panoramic image by overlapping a partial area of ​​the panoramic image with a partial area of ​​the third frame based on the panoramic image and pixel information on the third frame captured after the first frame, identify a square area of ​​the maximum size in the updated panoramic image, and re-identify the object in the identified square area. The third frame may be a frame captured after the first frame.

[0168] refer to Fig. 7A If the ratio between the size of the first frame relative to the third frame 710 and the overlapping area of ​​the third frame 710 is less than a predetermined ratio, the processor 120 may re-recognize the object within the updated panoramic image.

[0169] refer to Figure 7B , if the ratio between the size of the first frame relative to the third frame 710 and the overlapping area of ​​the third frame 710 is equal to or greater than a predetermined ratio, the processor 120 may not perform the re-recognition operation. That is, when the image blur is not obvious or the change or motion between the previous frame and the current frame is small, the processor 120 may not perform the re-recognition operation of the object.

[0170] Such operation may minimize processing by processor 120 and enable real-time object recognition.

[0171] Figure 8 is a view explaining a method for recognizing a final object according to an embodiment.

[0172] Figure 8 810-1, 810-2, and 810-3 overlap each other, and a square area 820 having a maximum size in the panoramic image is shown. It is assumed that the capture time of the three frames is in the order of 810-1, 810-2, and 810-3. That is, it is assumed that the 810-3 frame overlaps the panoramic image last.

[0173] When multiple objects are identified from a panoramic image, the processor 120 can assign a weighted value to each of the multiple regions based on at least one of the number of overlapping frames of the multiple corresponding panoramic images and the capture time of the frames in the multiple corresponding regions, and identify at least one of the multiple objects based on the weighted value of each of the multiple regions.

[0174] For example, the first object 830 and the second object 840 may be assigned a higher weighted value than the third object 850 and the fourth object 860 for appearing in a greater number of overlapping frames. Based on the capture time of the frame, the fourth object 860 may be assigned a higher weighted value than the third object 850. The first object 830 and the second object 840 may have the same weighted value. However, the processor 120 may assign weighted values ​​considering the size of the object. In this case, the first object 830 may be assigned a higher weighted value than the second object 840.

[0175] In this way, the processor 120 can identify a final object among the plurality of objects.

[0176] However, the present disclosure is not limited thereto. The processor 120 may recognize a plurality of objects from the panoramic image.

[0177] Fig. 9 is a block diagram of an electronic device 900 according to an embodiment. The electronic device 900 may be a device for stabilizing a motion value through an artificial algorithm.

[0178] refer to Fig. 9 , the electronic device 900 may include a training unit 910 and a response unit 920 .

[0179] The training unit 910 may generate or train an artificial intelligence model for stabilizing motion values ​​by using the training data. The training unit 910 may generate a recognition model having a recognition standard by using the collected training data.

[0180] The response unit 920 may stabilize the plurality of motion values ​​by using predetermined data as input data of a trained artificial intelligence model.

[0181] As an example, the training unit 910 and the response unit 920 may be included in the electronic device 900. However, the present disclosure is not limited thereto, and the training unit 910 and the response unit 920 may be installed in the electronic device 100. Specifically, at least a portion of the training unit 910 and at least a portion of the response unit 920 may be implemented as a software module or at least one hardware chip installed on the electronic device 100. For example, at least one of the training unit 910 and the response unit 920 may be manufactured as a hardware chip for artificial intelligence (AI), or a part of a conventional general-purpose processor (e.g., CPU or application processor) or a graphics processor (e.g., CPU) installed on various electronic devices as described above. The hardware chip for artificial intelligence may be a processor dedicated to probability calculation, and has a higher parallel processing capability than conventional general-purpose processing to quickly process operations in the field of artificial intelligence, such as machine training. When the training unit 910 and the response unit 920 are embodied as a software module (or a program module including instructions), the computer code of the software module may be stored in a non-transitory computer-readable medium that is readable by a computer. In this case, the software module may be executed under the control of an operating system (OS) or a predetermined application. A part of the software module may be provided by an operating system (OS), and another part may be provided by a predetermined application program.

[0182] The training unit 910 and the response unit 920 may be installed in one electronic device, or may be installed in separate electronic devices. For example, one of the training unit 910 and the response unit 920 may be included in the electronic device 100, and the other may be included in the electronic device 900. The training unit 910 and the response unit 920 may provide the response unit 920 with model information established by the training unit 910 in a wired or wireless manner, and the data input to the training unit 920 may be provided to the training unit 910 as additional training data.

[0183] Fig.10 is a block diagram of the training unit 910 according to an embodiment.

[0184] refer to Fig.10 The training unit 910 according to an example embodiment may include a training data acquisition unit 910-1 and a model training unit 910-4. The training unit 910 may further include a training data preprocessor 910-2, a training data selector 910-3, and a model evaluation unit 910-5.

[0185] The training data acquisition unit 910-1 may obtain the training data required for the artificial intelligence model for stabilizing the motion value. As an example, the training data acquisition unit 910-1 may obtain a plurality of motion values ​​and a plurality of converted motion values ​​as training data. The training data may be data collected or tested by the training unit 910 or a manufacturer of the training unit 910.

[0186] The model training unit 910-4 can train the artificial intelligence model to have a standard for stabilizing the motion value by using the training data. For example, the model training unit 910-4 can train the artificial intelligence model using at least part of the training data through supervised learning. For example, the model training unit 910-4 can train itself using the training data without any further guidance, and train the artificial intelligence model through unsupervised learning to find a standard for stabilizing the motion value. In addition, the model learning unit 910-4 can train the artificial intelligence model through reinforcement learning using, for example, feedback on whether the result of providing a response according to the learning is correct. The model learning unit 910-4 can also train the artificial intelligence model using, for example, a training algorithm including an error back propagation method or a gradient descent method.

[0187] The model training unit 910 - 4 may train a criterion regarding which training data will be used to stabilize the motion value by using the input data.

[0188] When there are multiple pre-established artificial intelligence models, the model training unit 910-4 can identify the artificial intelligence model with a high correlation between the input training data and the basic training data as the artificial intelligence model to be trained. In this case, the basic training data can be pre-classified by data type, and the artificial intelligence model can be pre-established by data type.

[0189] When the artificial intelligence model is trained, the model training unit 910-4 may store the trained artificial intelligence model. The model training unit 910-4 may store the trained artificial intelligence model in a memory of the electronic device 900. The model training unit 910-4 may store the trained artificial intelligence model in a server or memory of an electronic device connected to the electronic device 900 via a wired or wireless network.

[0190] The training unit 910 may further include a training data preprocessor 910 - 2 or a training data selector 910 - 3 to improve the response result of the artificial intelligence model or save resources or time required to generate the artificial intelligence model.

[0191] The training data preprocessor 910-2 may preprocess the obtained data so that the obtained data may be used for training to stabilize the motion value. The training data preprocessor 910-2 may produce the obtained data in a predetermined format. For example, the training data preprocessor 910-2 may divide the plurality of motion values ​​into a plurality of parts.

[0192] The training data selector 910-3 may select data required for training from the data acquired from the training data acquisition unit 910-1 and the data preprocessed by the training data preprocessor 910-2. The selected training data may be provided to the model training unit 910-4. The training data selector 910-3 may select training data required for training from the data acquired or preprocessed according to a predetermined recognition criterion. The training data selector 910-3 may select training data according to a predetermined recognition criterion through training of the model training unit 910-4.

[0193] The training unit 910 may also include a model evaluation unit 910-5 for improving the response results of the artificial intelligence model.

[0194] The model evaluation unit 910-5 may input the evaluation data into the artificial intelligence model, and if the response result output from the evaluation data does not meet the predetermined standard, the model training unit 910-4 may be allowed to train again. In this case, the evaluation data may be predefined data for evaluating the artificial intelligence model.

[0195] When there are multiple trained artificial intelligence models, the model training unit 910-5 can evaluate whether each trained artificial intelligence model meets the predetermined criteria, and identify the model that meets the predetermined criteria as the final artificial intelligence model. When there are multiple models that meet the predetermined criteria, the model evaluation unit 910-5 can identify any one or a predetermined number of preset models as the final artificial intelligence model in order of higher evaluation scores.

[0196] Fig.11 is a block diagram illustrating the response unit 920 according to an embodiment.

[0197] refer to Fig.11 , the response unit 920 according to the embodiment may include an input data acquisition unit 920 - 1 and a response result provider 920 - 4 .

[0198] The response unit 920 may further include an input data preprocessor 920 - 2 , an input data selector 920 - 3 , and a model updating unit 920 - 5 .

[0199] The input data acquisition unit 920-1 may obtain data required to stabilize the motion value. The response result provider 920-4 may apply the input data obtained from the input data acquisition unit 920-1 as an input value to the trained artificial intelligence model and stabilize the motion value. The response result provider 920-4 may apply the data selected by the input data preprocessor 920-2 or the input data selector 920-3 as an input value and obtain a response result. The response result may be identified by the artificial intelligence model.

[0200] The response result provider 920-4 may apply an artificial intelligence module to stabilize the motion value acquired from the input data acquisition unit 920-1 and stabilize the motion value from the conversation.

[0201] The response unit 920 may further include an input data preprocessor 920-2 or an input data selector 920-3 to improve the response result of the artificial intelligence model or save resources or time used to provide the response result.

[0202] The input data preprocessor 920-2 may preprocess the obtained data so that the obtained data may be used to stabilize the motion value. That is, the input data preprocessor 920-2 may manufacture the data obtained from the response result provider 920-4 in a predefined format.

[0203] The input data selector 920-3 can select the data required for providing a response from the data obtained from the input data acquisition unit 920-1 and the data preprocessed by the input data preprocessor 920-2. The selected data can be provided to the response result provider 920-4. The input data selector 920-3 can select part or all of the obtained or preprocessed data according to a predetermined identification criterion for providing a response. The input data selector 920-3 can select data according to a selection criterion predetermined by the training of the model training unit 910-4.

[0204] The model updating unit 920-5 may control the artificial intelligence model to be updated based on the response result provided by the response result provider 920-4. For example, the model updating unit 920-5 may provide the response result provided by the response result provider 920-4 to the model training unit 910-4 and request the model training unit 910-4 to further train or update the artificial intelligence model.

[0205] Fig.12 is a view showing an example in which the electronic device 100 according to an embodiment may operate in association with an external server to train and judge data.

[0206] refer to Fig.12, an external server may train a standard for stabilizing a motion value from a conversation, and the electronic device 100 may stabilize the motion value based on the training result of the server.

[0207] The model training unit 910-4 of the server S may execute the following steps: Fig.10 Functionality of the training unit 910 shown. The model training unit 910-4 of the server S can train the criteria about which filter to use for stabilizing the motion value or how to stabilize the motion value by using this information.

[0208] The response result provider 920-4 of the electronic device 100 may stabilize the motion value by applying the data selected by the input data selector 920-3 to the artificial intelligence model generated by the server S. The response result provider 920-4 of the electronic device 100 may receive the artificial intelligence model generated by the server (S) from the server (S), and stabilize the motion value by using the received artificial intelligence model.

[0209] Fig.13 is a flowchart of a method for controlling an electronic device according to an embodiment.

[0210] In step S1310, a panoramic image may be obtained by overlapping a partial area of ​​a first frame with a partial area of ​​at least one second frame based on pixel information on a first frame and at least one second frame among the plurality of frames. In step S1320, a region of a predetermined shape of a maximum size may be identified within the panoramic image. In step S1330, an object may be identified from the entire panoramic image or the region of the predetermined shape.

[0211] The acquiring of step S1310 may include acquiring a panoramic image by overlapping a region having a minimum difference between pixel values ​​in adjacent frames based on pixel information on each of the first frame and the at least one second frame.

[0212] The obtaining in step S1310 may include obtaining motion values ​​between adjacent frames based on pixel information on each of the first frame and at least one second frame, and obtaining a panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the at least one second frame based on the obtained motion values.

[0213] Obtaining S1310 may include converting the obtained motion values ​​based on a difference between pixel values ​​in adjacent frames and motion values ​​between adjacent frames, and obtaining the panoramic image by overlapping a partial area of ​​the first frame with a partial area of ​​the second frame based on the converted motion values.

[0214] Acquiring S1310 may include performing at least one of image processing, such as rotation, position shifting, or resizing, with respect to the first frame and the at least one second frame based on pixel information on each of the first frame and the at least one second frame.

[0215] The display may further include displaying a first frame, and based on the object identified from the panoramic image, based on information of the position of the object identified from the panoramic image and image processing information, displaying an object including at least one of a graphical user interface (GUI), a character, an image, a video, or a 3D model on an area where the object is displayed in the first frame.

[0216] The at least one second frame may be a frame captured before or immediately before the first frame, and the method may further include updating the panoramic image by overlapping a partial area of ​​the panoramic image with a partial area of ​​the third frame based on the panoramic image and pixel information on a third frame captured after the first frame, identifying an area of ​​a predetermined shape of a maximum size within the updated panoramic image, and re-identifying an object in the updated panoramic image or a predetermined shape within the updated panoramic image.

[0217] The third frame may be a frame captured immediately after the first frame, and if a ratio between the first frame relative to the third frame and an overlapping area of ​​the third frame is less than a predetermined ratio, re-identifying the object may include re-identifying the object within the updated panoramic image or within an area of ​​a predetermined shape within the updated panoramic image.

[0218] When a plurality of objects are identified from the panoramic image, the method may include assigning a weighted value to each of the plurality of regions based on at least one of a number of overlapping frames of the corresponding plurality of regions or a capture time of frames in each of the plurality of regions, and identifying at least one of the plurality of objects based on the weighted value of each of the plurality of regions.

[0219] The method may further include performing continuous capturing by a camera provided in the electronic device to obtain a plurality of frames.

[0220] According to various embodiments of the present disclosure, the electronic device may generate a panoramic image from a plurality of consecutive frames and improve the accuracy of object recognition by recognizing the object.

[0221] According to various embodiments of the present disclosure, the application may be installed on a conventional electronic device.

[0222] The methods according to various embodiments of the present disclosure may be embodied as software and hardware with respect to conventional electronic devices.

[0223] According to various exemplary embodiments of the present disclosure, it may be performed by an embedded server provided in an electronic device, an electronic device, or a display device.

[0224] Various embodiments of the present disclosure may be implemented as software including commands stored in a machine-readable storage medium. A machine may be a device that calls commands stored in a storage medium and may operate according to the called commands, including an electronic device (e.g., electronic device (A)) according to the disclosed example embodiments. When a command is executed by a processor, the processor may use other components to perform functions corresponding to the command, directly or under the control of the processor. The command may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-temporary storage medium. "Non-temporary" means that the storage medium does not include a signal, but is tangible, but does not distinguish whether the data is semi-permanently stored or temporarily stored on the storage medium.

[0225] According to an embodiment, the method according to various embodiments disclosed herein may be provided in a computer program product. The computer program product may be traded between a seller and a buyer as a commodity. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or through an application store (e.g., PlayStore). TM ) Online distribution. In the case of online distribution, at least a part of the computer program product may be temporarily stored or temporarily created on a storage medium, such as a memory of a manufacturer's server, an application store's server, or a relay server.

[0226] Each component (e.g., module or program) according to various embodiments may be composed of a single entity or multiple entities, and some of the above-mentioned subcomponents may be omitted, or other components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into one entity to perform the same or similar functions performed by each component before integration. According to various embodiments, the operations performed by modules, programs, or other components may be performed sequentially, in parallel, repeatedly, or heuristically, or at least some operations may be performed in a different order, or omitted, or another function may be further added.

[0227] Although embodiments have been shown and described, it will be appreciated by those skilled in the art that changes may be made to these embodiments without departing from the principles and spirit of the present disclosure. Therefore, the scope of the present disclosure is not to be interpreted as being limited to the described embodiments, but is to be defined by the appended claims and their equivalents.

Claims

1. An electronic device, comprising: Memory; and The processor is configured as: Based on the pixel information of the first image frame and the pixel information of the at least one second image frame, a panoramic image is obtained by overlapping a partial area of ​​the first image frame stored in the memory with a partial area of ​​the at least one second image frame stored in the memory, identifying a region of a predetermined shape of maximum size within the panoramic image, identifying a plurality of objects in the panoramic image or in the area of ​​the predetermined shape, identifying a plurality of regions including the plurality of objects, assigning weighting values ​​to the plurality of regions based on at least one of a number of overlapping image frames of the plurality of regions or a capture time of the image frames in the plurality of regions, and At least one object is identified among the plurality of objects based on the weighted values ​​of the plurality of regions.

2. The electronic device according to claim 1, in, The processor is also configured to obtain the panoramic image by overlapping an area having a minimum difference between pixel values ​​of adjacent image frames in the first image frame and the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame.

3. The electronic device according to claim 2, in, The processor is also configured to: obtaining a motion value between the adjacent image frames in the first image frame and the at least one second image frame based on the pixel information of the first image frame and the pixel information of the at least one second image frame, and A panoramic image is obtained by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the motion value.

4. The electronic device according to claim 3, in, The processor is also configured to: converting a motion value based on a difference between pixel values ​​in adjacent image frames and said motion value between adjacent image frames, and A panoramic image is obtained by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the converted motion value.

5. The electronic device according to claim 1, in, The processor is also configured to: performing image processing including at least one of rotation, position shifting, or size adjustment on each of the first image frame and the at least one second image frame based on the pixel information of the first image frame and the pixel information of the at least one second image frame, and A panoramic image is obtained by overlapping partial areas of frames on which image processing is performed.

6. The electronic device according to claim 5, further comprising: monitor, The processor is further configured to: controlling the display to display the first image frame, and Based on the at least one object identified from the panoramic image, according to information about the position at which the at least one object is identified and information about image processing, a display is controlled to display the at least one object including at least one of a graphical user interface GUI, characters, images, videos or 3D models on an area where the at least one object is displayed in a first image frame.

7. The electronic device according to claim 1, in, The at least one second image frame is an image frame captured before the first image frame, and The processor is further configured to: updating the panoramic image by overlapping a partial area of ​​the panoramic image with a partial area of ​​the third image frame based on the panoramic image and pixel information on a third image frame captured after the first image frame, identifying a region of a predetermined shape of maximum size within the updated panoramic image, and At least one object is re-identified from the updated panoramic image or a region of a predetermined shape within the updated panoramic image.

8. The electronic device according to claim 7, in, The third image frame is an image frame captured after the first image frame, and The processor is further configured to re-identify at least one object from the updated panoramic image or in an area of ​​a predetermined shape within the updated panoramic image based on a ratio between an overlapping area of ​​the first image frame and the third image frame associated with the third image frame being less than a predetermined ratio.

9. The electronic device according to claim 1, further comprising: Camera, including circuitry, Wherein the processor is further configured to obtain a plurality of image frames by performing continuous capture via the camera.

10. A method for controlling an electronic device, the method comprising: Based on pixel information of a first image frame and pixel information of at least one second image frame among the plurality of frames, obtaining a panoramic image by overlapping a partial area of ​​the first image frame with a partial area of ​​the at least one second image frame; Identifying a region of a predetermined shape of maximum size within the panoramic image; Identifying a plurality of objects in the panoramic image or in the region of the predetermined shape; identifying a plurality of regions including the plurality of objects; assigning weighting values ​​to the plurality of regions based on at least one of a number of overlapping image frames of the plurality of regions or a capture time of the image frames in the plurality of regions; as well as At least one object is identified among the plurality of objects based on the weighted values ​​of the plurality of regions.

11. The method according to claim 10, wherein the obtaining includes obtaining the panoramic image by overlapping an area having a minimum difference between pixel values ​​of adjacent image frames in the first image frame and the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame.

12. The method according to claim 11, wherein the obtaining comprises obtaining a motion value between the adjacent image frames in the first image frame and the at least one second image frame based on pixel information of the first image frame and pixel information of the at least one second image frame, and The panoramic image is obtained by overlapping a partial area of ​​the first image frame with a partial area of ​​the at least one second image frame based on the motion value.

13. The method of claim 12, wherein the obtaining comprises converting a motion value based on a difference between pixel values ​​in adjacent image frames and the motion value between adjacent image frames, and A panoramic image is obtained by overlapping a partial area of ​​the first image frame with a partial area of ​​at least one second image frame based on the converted motion value.

14. The method according to claim 10, in, The obtaining includes performing image processing including at least one of rotation, position shift or size adjustment with respect to each of the first image frame or the at least one second image frame based on the pixel information of the first image frame and the pixel information of the at least one second image frame, and A panoramic image is obtained by overlapping partial areas of frames on which image processing is performed.

Citation Information

Patent Citations

  • Image processing device that synthesizes image

    CN103002216A

  • Method for constructing a composite image

    EP2017783A2