A method for processing video (images) acquired from a camera linked to a computing device, and a system utilizing the same.
By attaching a gimbal to a video capturing device and controlling it for dynamic rotation, the method enhances video (image) input and processing in portable devices, enabling remote object recognition and character extraction, addressing the limitations of conventional devices.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- イ·チュンヨル
- Filing Date
- 2023-02-14
- Publication Date
- 2026-04-24
AI Technical Summary
Conventional portable computing devices face limitations in video (image) input due to restricted field of view and reception sensitivity, requiring manual adjustment for optimal capture and input, which hinders dynamic and remote interaction.
A gimbal is attached to a video (image) capturing device, controlled by the computing device to rotate around axes, enabling dynamic video (image) input and processing, with object recognition and tracking, and OCR for character extraction from captured images.
Enables remote input and higher-resolution video (image) processing, allowing for object recognition, spatial information acquisition, and character extraction from a distance, overcoming the limitations of manual adjustment and restricted field of view.
Smart Images

Figure 0007851414000007 
Figure 0007851414000008 
Figure 0007851414000009
Abstract
Description
Technical Field
[0001] In the present disclosure, a method for processing video (image) acquired from an imaging device and a system using the same are disclosed. Specifically, according to the method of the present disclosure, the computing device either integrated with or linked to the imaging device acquires the entire video (image), detects one or more objects shown in the entire video (image), performs classification to calculate the category of each of the detected objects, and generates detailed classification information including the characteristics and states of the objects as analysis results for each of the objects.
Background Art
[0002] A portable computing apparatus generally refers to a device equipped with a processor, a display, a microphone, and a speaker, and some of them can be used as a portable terminal which is a kind of communication device. Conventionally, a portable terminal has a mechanism in which user commands are input via an input device that requires user contact, such as a keypad or a touch display. However, with the development of technologies related to speech recognition technology and computer vision, it has become possible to receive user commands from a distance using input devices such as a microphone and a camera and interact according to the commands.
[0003] However, since a portable computing apparatus does not have the ability to move itself, the range of input and output is restricted depending on its physical location.
[0004] One example of the physical limitations of input devices on portable computing devices is that cameras, which are contactless input devices, are mounted on the front or back of the mobile phone, and their field of view is limited. Therefore, in order to capture the desired image, the user must hold the device and directly change the composition (field of view; FOV). Another example of such physical limitations is that microphones mounted on portable computing devices have reduced reception sensitivity depending on the direction of the sound source.
[0005] Both input and output are limited by physical position. For example, a touch display, which is an output device, generally has a shape close to a plane and is attached to some of the six surfaces of a portable terminal. The user holds the mobile phone in their hand and uses it with the touch display facing their face. Another example is an infrared projector for acquiring three-dimensional shapes. Since the direction and angle of its infrared light beam are limited to a range of specific directions on the portable terminal, similar to the example above, the user needs to hold the mobile terminal in their hand and adjust the direction of the beam according to a guide. Regarding this, the following patent publications of the Republic of Korea have been proposed: Nos. 10-2011-0032244, 10-2019-0085464, 10-2019-0074011, 10-2019-0098091, 10-2019-0106943, and 10-2018-0109499. Regarding this, see (Non-Patent Literature 1) Y. Zhou et al., "Learning to Reconstruct 3D Manhattan Wireframes From a Single Image," 2019 IEEE / CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 2019, pp.7697-7706, doi:10.1109 / ICCV.2019.00779., (Non-Patent Literature 2) Shin, D., & Kim, I. (2018). Deep Image Understanding Using Multilayered Contexts. Mathematical Problems in Engineering, 2018, 1-11. https: / / doi.org / 10.1155 / 2018 / 5847460, (Non-Patent Literature 3) Mo, K., Zhu, S., Chang, AX, Yi, L., Tripathi, S., Guibas, LJ, & Su, H. (2019).PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding. In 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE. https: / / doi.org / 10.1109 / cvpr.2019.00100, (Non-Patent Document 4) Babaee, M., Li, L., & Rigoll, G. (2019). Person identification from partial gait cycle using fully convolutional neural networks. Neurocomputing, 338, 116-125., (Non-Patent Document 5) Muhammad, U.R., Svanera, M., Leonardi, R., & Benini, S. (2018). Hair detection, segmentation, and hairstyle classification in the wild. Image and Vision Computing, 71, 25-37. https: / / doi.org / 10.1016 / j.imavis.2018.02.001, (Non-Patent Document 6) Mougeot, G., Li, D., & Jia, S. (2019). A Deep Learning Approach for Dog Face Verification and Recognition. In PRICAI 2019: Trends in Artificial Intelligence (pp. 418-430). Springer International Publishing. https: / / doi.org / 10.1007 / 978-3-030-29894-4_34, (Non-Patent Document 7) Raduly, Z., Sulyok, C., Vadaszi, Z., & Zolde, A. (2018). Dog Breed Identification Using Deep Learning.In 2018 IEEE 16th International Symposium on Intelligent Systems and Informatics(SISY).IEEE.https: / / doi.org / 10.1109 / sisy.2018.8524715,(Non-Patent Document 8)Wu, Z., Yao, T., Fu, Y., & Jiang, Y.-G.(2017).Deep learning for video classification and captioning.In Frontiers of Multimedia Research (pp. 3-29). ACM. https: / / doi.org / 10.1145 / 3122865.3122867, (Non-Patent Document 9) Wu, C.-Y., Girshick, R., He, K., Feichtenhofer, C., & Krahenbuhl, P. (2020). A Multigrid Method for Efficiently Training Video Models. In 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE. https: / / doi.org / 10.1109 / cvpr42600.2020.00023, (Non-patent document 10) Ullah, A., Ahmad, J., Muhammad, K., Sajjad, M., & Baik, SW (2018). Action Recognition in Video Sequences using Deep Bi-Directional LSTM With CNN Features. IEEE Access, 6, 1155-1166. https: / / doi.org / 10.1109 / access.2017.2778011 has been proposed. [Overview of the project] [Problems that the invention aims to solve]
[0006] The purpose of this disclosure is to present a method that can overcome the limitations of conventional portable computing devices based on the aforementioned technology, specifically the limitations related to video (image) input. This method proposes a technique in which a gimbal is attached to a video (image) capturing device such as a camera that works in conjunction with the portable computing device, and the computing device controls the gimbal to rotate the video (image) capturing device around one or more axes, thereby enabling more dynamic reception and processing of video (image) input.
[0007] This disclosure aims to solve the problems of conventional technologies and present a video (image) processing method that enables remote input via video (images) by recognizing and tracking objects from video (images) captured by cameras, etc., in a portable computing device, and actively acquiring information related to those objects and their environment. In particular, it aims to grasp the relative position information of objects in a coordinate system centered on the system described in this disclosure and spatial information based on the video (image), and to confirm the position of objects in space. [Means for solving the problem]
[0008] A method for processing video (images) performed by a computing device including a processor to solve the aforementioned problems, the method may include: a step of acquiring video (images); a step of acquiring analysis information corresponding to objects contained in the video (images) using an object analysis model; and a step of acquiring characters contained in the objects from the analysis information corresponding to the objects using an OCR model.
[0009] As an alternative, the above OCR model can perform the steps of: determining whether or not characters are displayed on the surface of the object; and, if it is determined that characters are displayed on the surface of the object, performing OCR on the characters.
[0010] As an alternative, in paragraph 2, the step of performing OCR on the above characters may include: extracting at least one video (image) sample from the object on which the above characters are displayed; determining the boundary line of the area on which the characters are displayed from the at least one video (image) sample; generating a character display video (image) from the at least one video (image) sample based on at least one of the boundary line or boundary points belonging to the boundary line; and performing OCR on the character display video (image). Alternatively, the above video (image) sample is a video (image) pattern located at at least one of the above boundary lines or boundary points, and the above video (image) pattern may include at least one of the following: some characters, the boundary of the characters, a part of the characters, or the background.
[0011] As an alternative, the step of generating the character display video (image) based on at least one of the boundary lines or boundary points belonging to the boundary line from among the at least one video (image) sample may include the steps of: acquiring a video (image) pattern located at at least one of the boundary lines or boundary points from among the at least one video (image) sample as a boundary marker; and generating the character display video (image) which is a video (image) containing the boundary marker as a partial video (image) included in the video (image) of the object on which the character is displayed.
[0012] Alternatively, the boundary marker may include at least one of the following: a leading marker which is a video (image) pattern corresponding to the first character of the above text; or a trailing marker which is a video (image) pattern corresponding to the last character of the above text.
[0013] As an alternative, the step of acquiring multiple video (image) patterns located at at least one of the boundary lines or boundary points from the above at least one video (image) sample as boundary markers may include the step of using a tracking controller to acquire a video (image) of a character that includes the leading marker and whose character recognition rate is above a threshold; and the step of determining whether or not the video (image) of the character includes the trailing marker.
[0014] As an alternative, if the image of the above characters includes the trailing marker, the process may further include a step of determining the image of the above characters as the image of the above characters.
[0015] As an alternative, if the above-mentioned image of the characters does not include the above-mentioned trailing marker, the tracking controller may be used to acquire additional image of the characters that includes the boundary marker following the last marker in the above-mentioned image of the characters, and whose character recognition rate is equal to or greater than the above-mentioned threshold; the additional image of the characters is combined with the above-mentioned image of the characters to generate an image of the combined characters; and if the above-mentioned image of the combined characters includes the above-mentioned trailing marker, the combined image of the characters may be further included as the image of the characters being displayed.
[0016] As an alternative, the step of performing the above OCR on the above characters may include: a step of using a tracking controller to acquire a video (image) of the object that includes the beginning of the above characters and whose character recognition rate is above a threshold; a step of performing the above OCR on the area of the first sentence in the video (image) of the object; and a step of performing natural language understanding (NLU) on the first text, which is the primary result of the above OCR, and calculating a first semantic value number, which is a numerical value of semantic value.
[0017] Alternatively, the process may further include a step in which, if the first semantic value is greater than or equal to the threshold, the first text is determined to be the result text which is the result of the OCR.
[0018] As an alternative, the process may further include: if the first semantic value is less than the threshold, the step of performing the OCR on the area of the sentence following the area of the first sentence; performing natural language understanding on the second text, which is the primary result of the OCR, and calculating a second semantic value, which is a numerical value of semantic value; and if the second semantic value is equal to or greater than the threshold, the step of calculating the second text as the result text, which is the result of the OCR.
[0019] Alternatively, the computing device may work in conjunction with the shooting device and gimbal to acquire the video (image), and the steps of using a tracking controller to activate one or more rotation axes of the gimbal and control the direction of the shooting device, and using the tracking controller to acquire the video (image) by zooming in or zooming out of the shooting device.
[0020] A non-temporary computer-readable medium including a computer program to solve the aforementioned problems, wherein the computer program causes a computing device to execute a method for processing images, and the method may include: a step of acquiring images; a step of acquiring analysis information corresponding to objects contained in the images from the images using an object analysis model; and a step of acquiring characters contained in the objects from the analysis information corresponding to the objects using an OCR model.
[0021] In a computing device for solving the foregoing problems, it includes a processor and a communication unit. The processor can acquire a video (image), use an object analysis model to acquire analysis information corresponding to an object included in the video (image) from the video (image), and use an OCR model to acquire characters included in the object from the analysis information corresponding to the object.
Advantages of the Invention
[0022] According to an exemplary embodiment of the present disclosure, it is possible to recognize and track one or more objects using a video (image), actively acquire information on the objects and the environment. In particular, information related to the state of an object can be acquired from a distance using a video (image), objects with which an object interacts can be discriminated using a video (image), a higher-resolution detailed video (image) related to a part of an object can be acquired from a distance, and characters printed on an object or output using other means such as a display can be grasped from a distance. Therefore, there is an effect that enables remote input using a video (image) in a portable computing device.
Brief Description of the Drawings
[0023] The following drawings attached for use in the description of the embodiments of the present invention are only a part of the plurality of embodiments of the present invention, and for a person having ordinary knowledge in the technical field to which the present invention belongs (hereinafter referred to as "ordinary technician"), it is possible to obtain other drawings based on these drawings without the need for inventive efforts. [Figure 1] Based on an embodiment in the present disclosure, it is a conceptual diagram schematically showing an exemplary configuration of a computing device that executes a method for processing a video (image) by a computing device (hereinafter referred to as "video (image) processing method"). [Figure 2]A system for executing a video (image) processing method based on an embodiment in the present disclosure, which is a conceptual diagram exemplarily showing the general hardware and software architecture including a computing device, a photographing device, and a gimbal. [Figure 3] A flowchart exemplarily showing a video (image) processing method based on an embodiment in the present disclosure. [Figure 4] A block diagram exemplarily showing modules for executing each step of a video (image) processing method based on an embodiment in the present disclosure. [Figure 5] A block diagram exemplarily showing a machine learning model used in a plurality of modules for a video (image) processing method based on an embodiment in the present disclosure. [Figure 6a] A flowchart exemplarily showing a plurality of methods used to detect the plane of the floor (ground) of an object in the video (image) processing method of the present disclosure. [Figure 6b] A flowchart exemplarily showing a plurality of methods used to detect the plane of the floor (ground) of an object in the video (image) processing method of the present disclosure. [Figure 6c] A flowchart exemplarily showing a plurality of methods used to detect the plane of the floor (ground) of an object in the video (image) processing method of the present disclosure. [Figure 6d] A flowchart exemplarily showing a plurality of methods used to detect the plane of the floor (ground) of an object in the video (image) processing method of the present disclosure. [Figure 7a] A drawing exemplarily showing object segmentation obtained by a video (image) processing method based on an embodiment in the present disclosure. [Figure 7b] An exemplary drawing for explaining a plurality of steps for detecting planes of two or more floors (grounds) in a video (image) processing method based on an embodiment in the present disclosure. [Figure 7c]This is a conceptual diagram illustrating a method for generating and using a reference plane circle and a measurement plane circle in a video (image) processing method based on one embodiment described in this disclosure. [Figure 8a] This flowchart exemplifies several methods used to determine the target position in the video (image) processing method described in this disclosure. [Figure 8b] This flowchart exemplifies several methods used to determine the target position in the video (image) processing method described in this disclosure. [Figure 8c] This flowchart exemplifies several methods used to determine the target position in the video (image) processing method described in this disclosure. [Figure 9a] This flowchart exemplifies several methods used to control the orientation of the imaging device in the video (image) processing method described in this disclosure. [Figure 9b] This flowchart exemplifies several methods used to control the orientation of the imaging device in the video (image) processing method described in this disclosure. [Figure 10a] This flowchart exemplifies several methods used to perform OCR in the video (image) processing method described in this disclosure. [Figure 10b] This flowchart exemplifies several methods used to perform OCR in the video (image) processing method described in this disclosure. [Figure 10c] This flowchart exemplifies several methods used to perform OCR in the video (image) processing method described in this disclosure. [Figure 11a] This diagram is provided as an example to illustrate multiple methods for performing OCR in the video (image) processing method described in this disclosure. [Figure 11b] This diagram is provided as an example to illustrate multiple methods for performing OCR in the video (image) processing method described in this disclosure. [Figure 11c]This diagram is provided as an example to illustrate multiple methods for performing OCR in the video (image) processing method described in this disclosure. [Figure 11d] This diagram is provided as an example to illustrate multiple methods for performing OCR in the video (image) processing method described in this disclosure. [Best Mode for Carrying Out the Invention]
[0024] All prior art referenced in this disclosure is incorporated by reference as if it were all contained herein. Unless otherwise defined, all terms used in this disclosure, including technical and scientific terms, are used in a way that is commonly understood by those of ordinary skill in the art. Terms defined in general dictionaries should be interpreted as having the same meaning as in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless otherwise defined herein.
[0025] In the detailed description of the present invention described later, the accompanying drawings illustrate specific embodiments that can be implemented by the present invention in order to clearly illustrate the object, technical solution, and advantages of the present invention. These embodiments are described in detail so as to be sufficient for a person of the ordinary skill to carry out the present invention. When describing using the accompanying drawings, the same reference numeral is used for identical components, regardless of the reference numerals in the drawings, and redundant descriptions relating thereto are omitted.
[0026] The specific structural or functional descriptions relating to the examples are disclosed for illustrative purposes only and can be modified and implemented in a variety of forms. Therefore, the examples are not limited to any particular form of disclosure, and the scope of this specification includes modifications, equivalents, or substitutions contained within the technical concept.
[0027] Expressions such as "first" and "second" are used to describe various components, but these terms should be interpreted only as distinguishing one particular component from others, and not as suggesting any specific order. Therefore, the first component can be named the second component, and similarly, the second component can be named the first component.
[0028] When it is stated that one component is "connected" or "linked" to another component, it should be understood that this may mean that it is directly connected to or linked to the other component, but it may also mean that other components exist in between.
[0029] A singular expression shall be considered to include a plural expression unless the context clearly indicates otherwise. In this specification, terms such as “include,” “contain,” “have,” or “possess” should be understood to indicate the existence of the feature, number, stage, action, component, part, or combination thereof described herein, and not to imply the exclusion of the possibility of the existence or addition of one or more other features, numbers, stages, actions, components, parts, or combinations thereof.
[0030] Furthermore, while a "part" or "sub-part" of an object can mean only a portion of it and not the whole, unless explicitly stated otherwise in the context, it should be understood to include the whole of the object. A subset of a set is equivalent to a concept that includes the set itself.
[0031] In this disclosure, “Module” may mean hardware capable of performing the functions and operations described herein, computer program code capable of performing specific functions and operations, or a recording medium on which computer program code capable of performing specific functions and operations is stored. In other words, a module may mean a functional and / or structural combination of hardware for implementing the technical concepts described herein, and / or software for driving that hardware.
[0032] Strictly speaking, a "model" refers to a function configured to calculate output data from input data in the manner trained by machine learning. Such a "model" is a type of data structure or function, and can be used by the aforementioned "modules."
[0033] However, since some ordinary engineers in fields where artificial intelligence is applied tend to confuse the terms "module" and "model," this disclosure may also use "module" and "model" interchangeably in some cases, as this usage allows ordinary engineers to easily understand the concepts without confusion.
[0034] In this disclosure, the terms "training" and "learning" refer to performing machine learning through process-based computing, and are not used to refer to the mental processes equivalent to human educational activities, as any ordinary engineer would understand. In the field of statistics, the term "machine learning" is often used to refer to a series of processes that create a target function (f) that appropriately maps an input variable (X) to an output variable (Y). The process of calculating the output variable from the input variable using the target function is called "prediction," and "appropriately mapping" means that the difference between the true value and the predicted value has been reasonably reduced. However, the reason for rationally reducing the difference rather than minimizing it is that optimization can lead to the problem of so-called overfitting, that is, when actual data that deviates from the training data is applied, the prediction may not be accurate, and appropriate empirical measures may be taken to resolve this.
[0035] Furthermore, in this disclosure, "inference" refers to the process of calculating output data from input data using a machine learning model, and is used particularly when referring to a mechanical imitation of human mental processes. Similarly, in this disclosure, "analysis" by a machine is used, like inference, to refer to a mechanical imitation of human mental processes.
[0036] In this disclosure, "Manhattan space" is used to refer to the same thing as disclosed in the non-patent document Y. Zhou et al., "Learning to Reconstruct 3D Manhattan Wireframes From a Single Image," 2019 IEEE / CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 2019, pp.7697-7706, doi:10.1109 / ICCV.2019.00779.
[0037] Furthermore, the present invention encompasses all possible combinations in the multiple embodiments shown in this disclosure. While the diverse embodiments of the present invention are distinct, it should be understood that they are not necessarily mutually exclusive. For example, certain shapes, structures, and characteristics described herein can be embodied in other embodiments without departing from the spirit and scope of the invention in one embodiment. Also, the position or arrangement of individual components in each of the embodiments disclosed herein can be modified without departing from the spirit and scope of the invention. Therefore, the detailed descriptions below should not be taken as restrictive. In the drawings, similar reference numerals indicate identical or similar functions in various respects.
[0038] Unless otherwise indicated in this specification or unless explicitly contradicted by the context, items indicated as singular shall include plural cases unless otherwise defined in the context. Furthermore, when describing the present invention, if a specific description of a related known configuration or function is deemed likely to obscure the gist of the invention, such detailed description shall be omitted.
[0039] Hereinafter, several preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings, in order to enable ordinary engineers to easily implement the present invention.
[0040] Figure 1 is a conceptual diagram illustrating an exemplary configuration of a computing device that performs a video (image) processing method based on one embodiment described in this disclosure.
[0041] As shown in Figure 1, a computing device (100) according to one embodiment of this disclosure includes a communication unit (110) and a processor (120), and is capable of communicating directly or indirectly with an external computing device (not shown) through the communication unit (110).
[0042] Specifically, the computing device (100) described above can achieve desired system performance by using a combination of typical computer hardware (e.g., a computer; a device that may include a processor, memory, storage, input and output devices, and other components of existing computing devices; electronic communication devices such as routers and switches; and electronic information storage systems such as network-attached storage (NAS) and storage area networks (SAN)) and computer software (i.e., instructions that cause the computing device to function in a particular manner). The storage described above can include not only storage devices such as hard disks and USB (Universal Serial Bus) memory, but also network-connected storage devices such as cloud servers. Here, the memory can be DDR2, DDR3, DDR4, SDP, DDP, QDP, magnetic hard disks, flash memory, etc., but is not limited to these.
[0043] The communication unit (110) of the aforementioned computing device can send and receive requests and responses with other connected computing devices, such as portable terminals. For example, such requests and responses can be sent and received using the same TCP (Transmission Control Protocol) session, but are not limited to this. For example, they can also be sent and received as UDP (User Datagram Protocol) datagrams.
[0044] Specifically, the communication unit (110) can be embodied in the form of a communication module including a communication interface. For example, the communication interface can include wireless internet interfaces such as WLAN (Wireless LAN), WiFi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World interoperability for Microwave access), HSDPA (High Speed Downlink Packet Access), 4G, and 5G, as well as short-range communication interfaces such as Bluetooth (Registered Trademark), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-WideBand), ZigBee, and NFC (Near Field Communication). In addition to the above, the communication interface can refer to any interface that can communicate with the outside (for example, a wired interface).
[0045] For example, the communication unit (110) can send and receive data from other computing devices via an appropriate communication interface as described above. Furthermore, the communication unit (110) in a broader sense may include, or be able to interact with, external input devices such as keyboards, mice, touch sensors, touchscreen input units, microphones, video cameras, or LIDAR, radar, switches, buttons, joysticks, sound cards, graphics cards, printers, displays, or external output devices such as the display unit of a touchscreen, for receiving commands or instructions. To provide a suitable user interface to the user of a computing device, such as a portable terminal, and enable interaction with the user, it is known that the computing device (100) may have a built-in display device or be able to interact with an external display device via the communication unit (110). For example, such a display device may be a touchscreen capable of touch input. A touchscreen can detect objects such as fingers or stylus pens that come into contact with or near the display by capacitive, inductive, or optical means, and can determine the detected position on the display.
[0046] The above input device may also include a microphone. Microphone types may include dynamic microphones, condenser microphones, etc., and microphones with characteristics such as omnidirectional, unidirectional, and superdirectional can be used. Beamforming microphones and microphone arrays can also be used, but are not limited to these. A microphone array refers to two or more microphones used to detect the direction of a sound source.
[0047] The output device described above may include a speaker. The type of speaker used may include, but is not limited to, omnidirectional speakers, directional speakers, or super-directional speakers using ultrasound.
[0048] Furthermore, the processor (120) of the computing device may include hardware configurations such as an MPU (microprocessing unit), CPU (central processing unit), GPU (graphics processing unit), NPU (neural processing unit), ASIC, CISC, RISC, FPGA, SOC chip or TPU (tensor processing unit), cache memory, and data bus. It may also further include software configurations for an operating system and applications that perform a specific purpose. Based on one embodiment of this disclosure, the processor (120) can perform calculations for training various neural networks. The processor (120) can perform calculations for training neural networks, such as processing input data for training in deep learning (DL), extracting feature values from input data, calculating errors, and updating weights in the neural network using backpropagation. At least one of the CPU, general-purpose graphics processing unit (GPGPU), and / or TPU of the processor (110) can process the training of network functions. For example, both the CPU and GPGPU can perform network function training and data classification using network functions. Furthermore, in one embodiment of this disclosure, it is possible to use the processors of multiple computing devices together to perform network function training and data classification using network functions. In addition, the computer program executed in the computing device (100) based on one embodiment of this disclosure can be a program that can be executed by the CPU, GPGPU, or TPU.
[0049] Throughout this specification, the terms computational model, neural network, network function, and neural network can be used interchangeably. A neural network can be composed of a set of interconnected computational units, generally called nodes. Such nodes may also be referred to as neurons. A neural network consists of at least one node. The nodes (or neurons) that make up a neural network can be interconnected by one or more links.
[0050] In a neural network, one or more nodes connected by links can, relatively speaking, be in an input-output node relationship. The concepts of input and output nodes are relative; any node that is an output node to one node can be an input node to another node, and vice versa. As mentioned above, the relationship between input and output nodes can be established around links. One input node can be connected to one or more output nodes via links, and vice versa.
[0051] In a relationship between input and output nodes connected via a single link, the output node's data can be determined based on the data input to the input node. In this case, the link connecting the input and output nodes can have weights. These weights can be variable, but they can be adjusted according to the user or algorithm to perform the function required by the neural network. For example, if one or more input nodes are interconnected to one output node by their respective links, the output node can determine its value based on the values input to the input nodes connected to it and the weights set for the links corresponding to each input node.
[0052] As mentioned above, a neural network consists of one or more nodes interconnected via one or more links, forming an input-output node relationship within the network. In a neural network, the characteristics of the network can be determined by the number of nodes and links, the correlation between nodes and links, and the weight values assigned to each link. For example, if there are two neural networks with the same number of nodes and links but different link weight values, these two neural networks can be recognized as distinct.
[0053] A neural network can consist of a set of one or more nodes. A subset of nodes constituting a neural network can form a layer. Some of the nodes constituting a neural network can form a layer based on their distance from a first input node. For example, a set of nodes that are n in distance from a first input node can form an n-th layer. The distance from the first input node can be defined based on the minimum number of links that must be traversed to reach that node from the first input node. However, this definition of a layer is arbitrary for the sake of explanation, and the position of a layer in a neural network can be defined in a way different from the above explanation. For example, the layer of nodes can be defined based on their distance from the final output node.
[0054] A first input node can refer to one or more nodes in a neural network that receive data directly without going through links in relation to other nodes. Alternatively, it can refer to a node in a neural network that does not have other input nodes connected via links in relation to other nodes based on links. Similarly, a final output node can refer to one or more nodes in a neural network that do not have an output node in relation to other nodes. Furthermore, a hidden node can refer to a node that does not fall under the category of first input node or final output node, but is still part of the neural network.
[0055] In one embodiment of the present disclosure, the neural network may have the same number of nodes in the input layer as in the output layer, and the number of nodes may decrease once before increasing again as one progresses from the input layer to the hidden layer. In another embodiment of the present disclosure, the neural network may have fewer nodes in the input layer than in the output layer, and the number of nodes may decrease as one progresses from the input layer to the hidden layer. Furthermore, in yet another embodiment of the present disclosure, the neural network may have more nodes in the input layer than in the output layer, and the number of nodes may increase as one progresses from the input layer to the hidden layer. In yet another embodiment of the present disclosure, the neural network may be a combination of the above-described neural networks. A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to the input and output layers. Using deep neural networks, it is possible to understand the latent structures of data. That is, it is possible to understand the latent structures of photographs, texts, videos, audio, and music (for example, what is depicted in a photograph, what is the content and emotion of a text, what is the content and emotion of an audio recording, etc.). Deep neural networks can include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siam networks, and generative adversarial networks (GANs). The deep neural networks mentioned above are merely examples, and this disclosure is not limited to them.
[0056] Neural networks can be trained using at least one of the following methods: supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Learning a neural network can be the process of applying knowledge to the neural network to enable it to perform specific actions. Neural networks can be trained to minimize the error in their output. During neural network training, training data is repeatedly input into the network, the network's output and target error for the training data are calculated, and the error is backpropagated from the output layer to the input layer to update the weights of each node in the neural network, in order to reduce the error. In supervised learning, training data with the correct answer labeled is used (i.e., labeled training data), while in unsupervised learning, the correct answer may not be labeled for each training data point. For example, in supervised learning for data classification, the training data could be data with a category labeled for each training data point. Labeled training data is input into the neural network, and the error can be calculated by comparing the neural network's output (category) with the labels of the training data. As another example, in unsupervised learning for data classification, the error can be calculated by comparing the input training data with the output of the neural network. The calculated error is backpropagated in the neural network in the reverse direction (i.e., from the output layer to the input layer), and through backpropagation, the connection weights of each node in each layer of the neural network can be updated. The amount of change in the connection weights of each node that are updated can be determined by the learning rate. The calculation of the neural network on the input data and the backpropagation of the error can constitute a learning cycle (epoch). The learning rate can be applied in a way that changes depending on the number of iterations of the neural network's learning cycle. For example, in the early stages of learning the neural network, a high learning rate can be used to increase efficiency by enabling the neural network to quickly achieve a certain level of performance, while in the later stages of learning, a low learning rate can be used to improve accuracy.
[0057] In neural network training, it is generally possible to use a subset of the actual data (i.e., the data that the trained neural network intends to process) as training data. This can lead to a training cycle where errors on the training data decrease, but errors on the actual data increase. Overfitting is a phenomenon where errors on the actual data increase due to excessive training on the training data. For example, a neural network that has learned to recognize cats by seeing yellow cats may fail to recognize cats of other colors as cats; this is a type of overfitting. Overfitting can cause errors in machine learning algorithms. Various optimization methods can be used to prevent such overfitting. Methods that can be applied to prevent overfitting include increasing the amount of training data, regularization, dropout (deactivating some of the network nodes during the training process), and the use of a batch normalization layer.
[0058] In one embodiment of this disclosure, a computer-readable medium on which a data structure is stored is disclosed.
[0059] A data structure can refer to the organization, management, and storage of data, enabling efficient access to and modification of that data. A data structure can also refer to the organization of data to solve a specific problem (e.g., data retrieval, data storage, data modification in the shortest possible time). A data structure can also be defined as the physical or logical relationships between multiple data elements designed to support specific data processing functions. Logical relationships between multiple data elements can include linking relationships between multiple user-defined data elements. Physical relationships between multiple data elements can include actual relationships between multiple data elements physically stored in a computer-readable storage medium (e.g., a permanent storage device). Specifically, a data structure can include a collection of data, relationships between data, and functions or instructions that can be applied to data. By leveraging effectively designed data structures, computing devices can perform calculations while minimizing the use of their resources. Specifically, computing devices can improve the efficiency of calculations, reading, ingesting, deleting, comparing, exchanging, and retrieving through effectively designed data structures.
[0060] Data structures can be classified into linear and nonlinear data structures based on their form. A linear data structure can be a structure in which only one piece of data follows one piece of data. Linear data structures can include lists, stacks, queues, and deques. A list can represent a set of data that has an internal order. Lists can include linked lists. A linked list can be a data structure in which data is linked in a linear fashion, with each piece of data having a pointer. In a linked list, pointers can contain information about the link to the next or previous piece of data. Linked lists can be represented as one-way linked lists, two-way linked lists, or circular linked lists, depending on their form. A stack can be a data array structure with restricted access to data. A stack can be a linear data structure in which data can only be processed (e.g., inserted or deleted) at one end of the data structure. Data stored in a stack can be a last-in, first-out (LIFO) data structure. A queue is also a data array structure with restrictions on access to data, but the difference from a stack is that it can be a first-in, first-out (FIFO) data structure. A deck can be a data structure that allows data to be processed at both ends of the data structure.
[0061] Nonlinear data structures can be structures where one data point is followed by multiple data points. Nonlinear data structures can include graph data structures. Graph data structures can be defined using vertices and edges, where edges can include lines connecting two different vertices. Graph data structures can also include tree data structures. A tree data structure can be a data structure where there is one path connecting two different vertices among the multiple vertices contained in the tree. In other words, it can be a data structure that does not form loops within a graph data structure.
[0062] Throughout this specification, the terms computational model, neural network, network function, and neural network can be used interchangeably. Hereafter, they will be consistently referred to as neural networks. A data structure can include a neural network. A data structure including a neural network can be stored on a computer-readable medium. Furthermore, a data structure including a neural network can include data preprocessed for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node and hierarchy of the neural network, loss functions for learning the neural network, etc. A data structure including a neural network can include any of the components of the configuration disclosed above. In other words, when constructing a data structure containing a neural network, it is possible to include all of the following, or any combination thereof: data preprocessed for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node and hierarchy of the neural network, loss functions for learning the neural network, etc. Beyond the aforementioned configurations, a data structure containing a neural network can also include any other information that determines the characteristics of the neural network. Furthermore, the data structure can include, and is not limited to, any form of data used or generated during the computational process of the neural network. Computer-readable media can include computer-readable recording media and / or computer-readable transmission media. A neural network can generally be composed of a set of interconnected computational units called nodes. Such nodes may also be called neurons. A neural network is composed of at least one node.
[0063] A data structure can include data to be input to a neural network. A data structure including data to be input to a neural network can be stored on a computer-readable medium. The data to be input to a neural network can include learning data input during the learning process of the neural network and / or input data to be input to a neural network after learning is complete. The data to be input to a neural network can include pre-processed data and / or data subject to pre-processing. Pre-processing can include data processing processes for inputting data into a neural network. Therefore, a data structure can include data subject to pre-processing and data generated by pre-processing. The aforementioned data structures are merely examples, and this disclosure is not limited thereto.
[0064] The data structure may include the weights of a neural network. (In this specification, weights and parameters can be considered to have the same meaning.) The data structure including the weights of a neural network may be stored in a computer-readable medium. A neural network may include multiple weights. The weights may be variable, but can be varied according to the user or algorithm in order to perform the function required by the neural network. For example, if one or more input nodes are interconnected to one output node by their respective links, the output node can determine the value of the data output from the output node based on multiple values input to the input nodes connected to the output node and the weights set for the links corresponding to each input node. The data structures described above are illustrative examples, and this disclosure is not limited thereto.
[0065] As an example rather than a limitation, the weights may include weights that change during the learning process of the neural network and / or weights after the neural network has finished learning. The weights that change during the learning process of the neural network may include the weights at the start of the learning cycle and / or weights that change during the learning cycle. The weights after the neural network has finished learning may include weights after the learning cycle has finished. Therefore, a data structure containing the weights of a neural network may include a data structure containing weights that change during the learning process of the neural network and / or weights after the neural network has finished learning. Accordingly, the aforementioned weights and / or each combination of weights shall be included in the data structure containing the weights of a neural network. The aforementioned data structures are merely examples, and this disclosure is not limited thereto.
[0066] Data structures containing neural network weights can be stored on a computer-readable storage medium (e.g., memory, hard disk) after undergoing a serialization process. Serialization can be a process of converting data structures into a form that can be stored on the same or different computing devices and later reconfigured for use. Computing devices can serialize data structures and send and receive data over a network. Serialized data structures containing neural network weights can be reconfigured on the same or other computing devices through deserialization. Data structures containing neural network weights are not limited to serialization. Furthermore, data structures containing neural network weights can include data structures that enhance computational efficiency while minimizing the use of computing device resources (e.g., B-trees, tries, m-way search trees, AVL trees, Red-Black trees in nonlinear data structures). The foregoing are illustrative examples, and this disclosure is not limited thereto.
[0067] The data structure can include the hyperparameters of the neural network. The data structure containing the neural network's hyperparameters can be stored in a computer-readable medium. The hyperparameters can be variable depending on the user. Examples of hyperparameters include the learning rate, cost function, number of learning cycle iterations, weight initialization (e.g., setting the range of weight values to be initialized), and number of hidden units (e.g., number of hidden layers, number of nodes in hidden layers). The aforementioned data structures are illustrative and the disclosure is not limited thereto.
[0068] Figure 2 is a conceptual diagram illustrating the overall hardware and software architecture of a system that performs a video (image) processing method based on one embodiment of this disclosure, including a computing device, a camera, and a gimbal.
[0069] Using Figure 2, an overview of the configuration of the method and apparatus according to the present invention is shown. The computing device (100) may include an imaging device (200) and may be linked to an external imaging device (200) wirelessly or wired. The computing device (100) may be linked wirelessly or wired to or include a gimbal (300) that performs the function of controlling the attitude of the imaging device (200). To control the attitude of the imaging device (200), the gimbal (300) may include the imaging device (200) or be equipped with a predetermined mechanism (e.g., a suction cup) that can fix the imaging device (200).
[0070] For the aforementioned attitude control, the gimbal (300) may have one or more axes of rotation, examples of which are described in Published Japanese Patent No. 10-2019-0036323. The gimbal (300) can actively improve the input range of the imaging device (200) by controlling its attitude.
[0071] If the gimbal (300) has one axis of rotation, that axis of rotation can be the yaw axis (Y). The yaw axis allows the camera device (200) to interact with objects in space with minimal rotation.
[0072] If the gimbal (300) has two axes of rotation, these axes can be the yaw axis and the pitch axis (P). If the gimbal (300) has three axes of rotation, these axes can be the yaw axis, the pitch axis, and the roll axis (R).
[0073] The gimbal (300) may include a power supply unit (310) as a hardware component. The power supply unit (310) may receive external power via wired or wireless means, and via direct current or alternating current. The power supplied to the power supply unit (310) can be used in the gimbal (300) or the computing device (100). The power supply unit (310) can also be used to charge a battery built into the gimbal (300) or a battery built into the computing device.
[0074] Furthermore, the gimbal (300) may include at least one gimbal motor (330; not shown) as a component of its hardware. Each of the gimbal motors (330) is configured to change the direction of the camera (200) or the computing device (100) in which the camera (200) is built along the aforementioned axis of rotation, and the gimbal motors (330) may be, but are not limited to, DC motors, stepper motors, or brushless motors. The gimbal (300) may further include gears for converting the torque of the motors in addition to the gimbal motors (330).
[0075] The motors (330) of the gimbal (300) are used to orient the camera (200) or computing device (100) fixed to the gimbal toward a specific object. It is preferable, however, that each of its rotation axes is positioned parallel to the yaw, pitch, roll, and other axes of the camera (200) or computing device (100), but it is easy for an ordinary engineer to understand that there is no reason to limit this arrangement.
[0076] The gimbal (300) may further include at least one sensor (340; not shown in the illustration) as a component of its hardware. The sensor (340) can perform the function of detecting one or more of the following for a fixed part of the gimbal (300) or motor (330): position, angular position, displacement, angular displacement, velocity, angular velocity, acceleration, and angular acceleration. The types of sensors (340) mentioned above include, but are not limited to, acceleration sensors, gyroscopes, magnetic sensors such as geomagnetic sensors, Hall sensors, pressure sensors, infrared sensors, proximity sensors, motion sensors, dimming sensors, image sensors, GPS sensors, temperature sensors, humidity sensors, barometric pressure sensors, and LiDAR sensors.
[0077] Sensors (340) that could not be mounted on the computing device (100), particularly portable computing devices, due to weight and volume limitations, can be mounted on the gimbal (300), which can then be used to acquire information about the surroundings of the gimbal (300).
[0078] The specific functions and effects of the present invention, which can be achieved by the individual components schematically described using Figure 2, will be described in detail below using Figures 3 to 11d. The multiple components shown in Figure 2 are examples of implementation in a single computing device for the sake of clarity, but it is obvious that the computing device (100) that performs the method of this disclosure can also be configured with multiple devices working together. For example, a gimbal (300) can be configured as an independent computing device, and the gimbal (300) and the computing device (100), such as a portable computing device like a handheld terminal, can work together, in which case the gimbal (300) can also perform at least some of the functions performed by the handheld terminal (100). In other words, an ordinary person of the art can configure multiple devices to work together to perform the method of this disclosure in a variety of ways.
[0079] Figure 3 is a flowchart illustrating an example of a video (image) processing method based on one embodiment of this disclosure, and Figure 4 is a block diagram illustrating modules that perform each stage of the video (image) processing method based on one embodiment of this disclosure. Furthermore, Figure 5 is a block diagram illustrating machine learning models used in multiple modules for the video (image) processing method based on one embodiment of this disclosure.
[0080] As shown in Figure 3, the video (image) processing method based on this disclosure first includes a video (image) acquisition step (S1000), in which a video (image) input module (4100) embodied by a computing device (100) acquires the entire video (image) from a shooting device (200) which is included in the computing device (100) or which is linked through the communication unit (110) of the computing device (100).
[0081] Here, "the whole image" refers to an image that is contrasted with an image that is part of the whole image, such as the object image described later.
[0082] Next, the above video (image) processing method further includes a category classification step (S2000), in which an object analysis module (4200) embodied by a computing device (100) detects one or more objects in the overall video (image) and performs classification to calculate the category of each of the detected objects.
[0083] Here, the category of an object refers to the result of classifying an object into categories such as people, trees, puppies, etc.
[0084] During the category classification stage (S2000), simultaneously with the classification, it is possible to calculate the position of one or more objects from the overall video (image) as part of the 2D measurement.
[0085] The difference between the aforementioned 2D measurement and the 3D measurement described later is that 2D measurement is performed based on a 2D coordinate system shown in the image (video) without considering information related to 3D depth (or depth), while 3D measurement is performed considering not only the 2D coordinates in the image (video) but also information related to depth (or depth).
[0086] Specifically, the category classification stage (S2000) may include a stage (S2100) in which the overall video (image) is resized to a lower resolution than the original resolution of the overall video (image), and a stage (S2200) in which the resized video (image) is input into an object analysis model (M420) to calculate the category, position, and importance of each of the objects.
[0087] For example, the processor (120) of the computing device (100) can use an object analysis model (M420) to obtain analysis information from video (images) that corresponds to the objects contained in the video (images). The analysis information may include at least one of the following: classification information indicating the category of the object, location information indicating the location of the object, and / or importance information indicating the priority of the object in the video (images).
[0088] The object analysis model (M420) is a model for analyzing objects contained in video (images), and can include object classification models, localization models, object detection models, segmentation models, and more.
[0089] Based on one embodiment of this disclosure, the processor (120) can classify the types (classes) of objects in a given video (image) using an object analysis model (M420) that performs object classification. For example, if a person is shown in the given video (image), the processor (120) can use the object analysis model (M420) to obtain the output that "the type of the input video (image) is a person." The foregoing is merely illustrative, and this disclosure is not limited thereto.
[0090] Based on one embodiment of the present disclosure, the processor (120) can also output location information indicating the position of objects within a video (image) using an object analysis model (M420) that performs classification and localization. For example, when the processor (120) uses an object analysis model (M420) that performs classification and localization, it can recognize objects within a video (image) using a bounding box and output location information. The bounding box can transmit the location information of an object by outputting the coordinates of the left, right, top, and bottom of the box. The foregoing is merely illustrative, and the present disclosure is not limited thereto.
[0091] Based on one embodiment of the present disclosure, the processor (120) can sense at least one object using an object analysis model (M420) that performs object sensing. The processor (120) can sense multiple objects and extract location information by simultaneously classifying and localizing at least one object using the object analysis model (M420) that performs object sensing. The foregoing is illustrative, and the present disclosure is not limited thereto.
[0092] Based on one embodiment of the present disclosure, the processor (120) can detect objects by classifying pixels using a segmentation object analysis model (M420) and separating the boundaries of objects in an image from the background. The foregoing is illustrative and the present disclosure is not limited thereto.
[0093] It is common knowledge among engineers that the size adjustments performed in stage (S2100) are intended to reduce the computational load on the object analysis model (M420) and improve the processing speed of the object analysis module (4200).
[0094] The importance level resulting from stage (S2200) can be used as a measure to determine the priority of the above objects. Conversely, objects with higher priority can be assigned a higher importance level, and / or objects with lower priority can be assigned a lower importance level. Priority will be discussed in more detail later.
[0095] For objects whose importance is above a predetermined threshold value, the computing device (100) can extract and sample object images (images) of those objects from the overall image (image) {crop feed stage; S2500}. In stage (S2200), resized images (images) are used, and some of the data from the overall image (image) is lost. Therefore, the crop feed stage (S2500) allows the use of the image (image) information before the loss for objects with relatively high importance.
[0096] Continuing with Figure 3, the image processing method based on this disclosure further includes a detailed classification step (S3000) in which a detailed classification module (4300) embodied by a computing device (100) generates detailed classification information, including the characteristics and state of each detected object, as an analysis result for that object.
[0097] Here, the concept of an object includes spatial objects, which are objects that correspond to the "space" itself in which the entire image was captured. The detailed classification information of a spatial object can include information related to that space. Object characteristics and state
[0098] The properties of the object described above refer to the aspects of the object that are virtually unchanging over time, while the state of the object described above refers to the aspects of the object that can generally change over time. Specifically, the characteristics of the object described above may include information about partial objects that are components of or belong to the object. For example, if a person is detected as an object from the whole image, then parts of that person such as their arms, legs, and eyes, as well as the clothes and footwear they are wearing, are partial objects of that object.
[0099] To detect partial objects as described above, the detailed classification stage (S3000) may include a stage (S3920) in which a partial object is a component that forms part of or belongs to the above object, and, if the partial object is detected, an additional stage (S3940) in which analysis results of the characteristics and state of the partial object are generated as part of the detailed classification information.
[0100] Furthermore, the characteristics of the object may include at least one of the following: the main color of the object; a subordinate, which is information referring to the partial object of the object; a subject, which is information referring to the other object if the object is a partial object of another object; the size of the object; one or more materials of the object, including the main material of the object; the transparency of the object; text displayed on the surface of the object; and whether or not the object is movable under its own power.
[0101] Here, the size of the object can be the size measured by two-dimensional or three-dimensional measurement. Furthermore, the transparency of the object is a property that the object can have if it is an object that has transparent parts, such as a glass window. For example, an opaque object can have a value of 0, while an object made of a transparent material like glass can have a positive value.
[0102] On the other hand, the state of the object may include at least one of the following: the position of the object, the orientation of the object, the action of the object, the direction of the object, whether or not the object is in contact with the floor (ground), and the velocity of the object.
[0103] Here, the position of the object can be a position measured by two-dimensional or three-dimensional measurement. The orientation of the object can be inferred from the positional information of a part of the object, and the behavior of the object can be inferred from the temporally continuous orientation.
[0104] Furthermore, the orientation of the object can be inferred from the object's position information or its actions.
[0105] Whether or not the above object is in contact with the floor (ground) refers to whether or not the object is in contact with the plane of the floor (ground) of the spatial object to which it belongs. Examples of objects that have a true value include chairs, desks, utility poles, and car tires.
[0106] The detailed classification information for the above object may further include attributes that refer to information relating to the input and output of the object in the system, in addition to the above characteristics and state. The attributes of the above object may include the data input time, which includes the time when the raw data of the above object was first input, and at least one of the operational rights related to the system under this disclosure that have been granted to the above object.
[0107] In order to generate the detailed classification information of the above-mentioned objects, the detailed classification stage (S3000) may include the following steps: selecting a detailed classification model (M430), which is a set of models consisting of at least one model that has been pre-trained to conform to the above-mentioned category, in order to acquire the above-mentioned characteristics and state of each object belonging to the above-mentioned category to which the above-mentioned object belongs, for each individual object whose importance is an object of a predetermined second numerical value or higher (S3200); and inputting an individual object video (image), which is an image of the above-mentioned individual object, into the selected detailed classification model (M430), and generating an object record that includes the identifier of the above-mentioned individual object and the above-mentioned detailed classification information as an object record belonging to the above-mentioned individual object by the identifier (S3400).
[0108] Here, the detailed classification model (M430) is used to distinguish one or more objects from each other. In other words, the detailed classification information generated by the detailed classification model (M430) can assign an identifier to each of multiple objects that can distinguish each other from other objects.
[0109] Furthermore, in this context, an object record refers to a record containing information belonging to the identifier of each object. Examples of information belonging to the identifier of each object include, if the object is a person, it may include the shape of the person's face, height, gait, tattoos, hairstyle, etc. If the object is a dog, it may include the shape of the dog's head, fur growth pattern and color, breed, etc. An object record may also contain information about other objects that belong to each object, and this information may be the identifier of those other objects.
[0110] Among these, the fact that the way a person walks can be classified using artificial intelligence methodologies is described, for example, in the non-patent literature paper Babaee, M., Li, L., & Rigoll, G. (2019). Person identification from partial gait cycle using fully convolutional neural networks. Neurocomputing, 338, 116-125.
[0111] Furthermore, the possibility of classifying human hairstyles using artificial intelligence methodologies is described, for example, in the non-patent literature paper Muhammad, UR, Svanera, M., Leonardi, R., & Benini, S. (2018). Hair detection, segmentation, and hairstyle classification in the wild. Image and Vision Computing, 71, 25-37. https: / / doi.org / 10.1016 / j.imavis.2018.02.001.
[0112] However, ordinary engineers can understand that information other than that presented in the aforementioned prior literature can also be obtained through artificial intelligence methodologies.
[0113] The step (S3200) of selecting a detailed classification model (M430) for each individual object can be performed by a classification model selection module (4320) implemented by a computing device (100). After the category of an object is obtained, the classification model selection module (4320) performs the function of selecting a detailed classification model that fits that category, along with the algorithm applied to that detailed classification model.
[0114] For example, if the object category is "person," the classification model selection module (4320) can select a detailed classification model that generates detailed classification information that can identify a person, such as the shape of their face, height, gait, tattoos, and hairstyle. Similarly, if the object category is "dog," it can select a detailed classification model that generates detailed classification information that can identify a dog, such as the shape of its head, the way its fur grows and its color, and its breed.
[0115] As stated in the non-patent literature paper Mougeot, G., Li, D., & Jia, S. (2019). A Deep Learning Approach for Dog Face Verification and Recognition. In PRICAI 2019: Trends in Artificial Intelligence (pp. 418-430). Springer International Publishing. https: / / doi.org / 10.1007 / 978-3-030-29894-4_34, it is possible to classify not only the shape of a human face but also the shape of a dog's head using artificial intelligence methodologies.
[0116] Furthermore, regarding the methodology of artificial intelligence for classifying dog breeds, it is possible to refer to the non-patent literature paper: Raduly, Z., Sulyok, C., Vadaszi, Z., & Zolde, A. (2018). Dog Breed Identification Using Deep Learning. In 2018 IEEE 16th International Symposium on Intelligent Systems and Informatics (SISY). IEEE. https: / / doi.org / 10.1109 / sisy.2018.8524715.
[0117] On the other hand, since the detailed classification model (M430) is a collection of models, it can include a measurement model (M431) that calculates at least one of the following by performing at least one of the following on the object: the position of the object, whether or not the object is in contact with the floor (ground), the orientation of the object, the velocity of the object, the posture of the object, and / or the size of the object. Here, the size of the object can include at least one of the following: a one-dimensional dimension including height, width and / or depth (or length), a two-dimensional dimension including the surface area of the object, and / or a three-dimensional dimension including the volume of the object.
[0118] For example, the volume of an object can be calculated based on at least one (or more) of the following: the object's segmentation, its orientation, and / or its depth (or length) information.
[0119] The measurement model (M431) can calculate at least one of the following, if the object is a spatial object: the orientation and coordinates of the spatial object, the position of the system which is the origin of at least one of these; the Manhattan space which is the volume space containing the objects contained within the spatial object's space; the plane of the floor (ground) detected from the space; the gravity vector applied to the space; the empty volume space obtained by removing the objects contained within the space from the space; the partial objects of the spatial object; and the direction of the space. The space can be indoors or outdoors.
[0120] Here, a partial object of a spatial object refers to a partial object that constitutes the space of that spatial object, but that partial object is fixed in that space. Examples of the aforementioned partial objects of a spatial object can include glass windows, doors, walls, kitchens, roads, pedestrian bridges, etc.
[0121] The system's location refers to the location of the system in an indoor or outdoor space that serves as the origin of direction or coordinates, as defined in this disclosure.
[0122] Furthermore, the direction of the space mentioned above refers to the direction of the object which is the spatial object mentioned above, but it may be based on a part of the system or space described in this disclosure.
[0123] Depth information in 3D measurement can be predicted by an artificial intelligence methodology that derives depth information from the above video (image), or it can be provided as supplementary information to the above video (image) from other sensors (340) such as lidar, radar, and ultrasonic sensors.
[0124] Specifically, in step (S3400), the three-dimensional measurement may further include a process (S3410) to identify a length reference object which is an object that satisfies the condition that the deviation in length of at least one part of the object or a partial object which is a component of the object or a component belonging to the object is smaller than a predetermined standard, and to measure the two-dimensional length of the length reference object, and a process (S3420) to detect the plane of the floor (ground) of the object.
[0125] The purpose of this 3D measurement is to determine the relative position between the system and the object, and / or the absolute position of the object, using information from one or more objects.
[0126] For example, the height of an adult male, which is an example of a length reference object, can be used as a reference length for objects belonging to other categories, such as doors, pencils, and cups, and this can also be used to measure distance.
[0127] Another example of a length-reference object is when multiple doors are detected in a single overall video (image) or space (or spatial object). In a single space, doors of the same design can be used as length-reference objects for distance determination because they are roughly the same height.
[0128] Furthermore, partial objects can also be used as length reference objects if the deviation of their lengths is relatively small. For example, the horizontal length of a human eyeball can be used as a length reference object because its standard deviation is relatively small.
[0129] In process (S3410), when measuring the two-dimensional length of a length reference object, if the orientation information of the length reference object is available, it is possible to calculate a corrected length that reflects the tilt due to the orientation of the length reference object. The object's orientation and its calculation will be described later.
[0130] The flat surface of the floor (ground) Regarding the detection of the plane of the floor (ground) of an object in process (S3420), various methods are possible, but for example, one method is to detect the plane of the floor (ground) in contact with the object along the direction of the vector of gravity acting on the object.
[0131] Furthermore, for detecting the plane of the object's floor (ground), there are also methods that use Manhattan spatial detection, methods that generate and extend the floor (ground) plane between two or more of the above-mentioned length reference objects, and methods that generate and extend the floor (ground) plane between the length reference object before the movement and the length reference object after the movement when the length reference object moves.
[0132] Figures 6a to 6d are flowcharts illustrating several methods used to detect the plane of an object's floor (ground) in the image processing method described in this disclosure.
[0133] According to Figure 6a, a specific first embodiment of detecting the plane of the object's floor (ground) (S3420) begins with the step (S3422a) of detecting at least one object of equivalent size that belongs to a category of equivalent size, which is a category in which at least one part of an object or a partial object of that object has a size deviation smaller than a predetermined standard. Here, size can be one or more of width, height, and depth (or length).
[0134] For example, a category that satisfies the condition that the size deviation is smaller than a predetermined standard could be the category of desks. Certain types of desks have a relatively small height deviation. If the height of a certain type of desk is, for example, 70-74 cm on average, then it is possible to set the standard length of an object of equivalent size to 72 cm.
[0135] By any choice, objects of the same size as described above can also be objects relative to the floor (ground). Floor (ground)-referenced objects are objects used to detect the plane of the floor (ground), and refer to objects whose object segmentation is at least part in general contact with the floor (ground). For example, a desk and a chair are examples of such floor (ground)-referenced objects.
[0136] Following step (S3422a), the first embodiment of the process (S3420) includes step (S3424a) of detecting a floor (ground) contact point, which is a contact point with the lowest level of object segmentation of the object of the same size as described above.
[0137] Figure 7a is a diagram illustrating the object segmentation obtained by the video (image) processing method described in this disclosure.
[0138] While the method for obtaining object segmentation is generally known to engineers, one can refer to non-patent literature such as the paper Mo, K., Zhu, S., Chang, AX, Yi, L., Tripathi, S., Guibas, LJ, & Su, H. (2019). PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding. In 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE.
[0139] As shown in Figure 7a, in the object segmentation of the chair given as an example, the lowest part of the region corresponding to the chair legs (712, 714, 716, 718) can be considered to be in contact with the floor (ground), and therefore it is possible to detect the floor (ground) contact point (720) which is the point of contact with the floor (ground) at that lowest part. As another example, the lowest part of the object segmentation of the door is in contact with the floor (ground), and therefore it is possible to detect the floor (ground) contact point which is the point of contact with it.
[0140] Following step (S3424a), the first embodiment of process (S3420) includes step (S3426a) of determining the three-dimensional distance of the equivalent-sized objects based on their two-dimensional lengths. For example, a chair located further away will have at least one (or more) smaller measurement of its two-dimensional length (height), width, and / or depth (or length) compared to a chair located closer, and this can be used to determine the three-dimensional distance.
[0141] Following step (S3426a), the first embodiment of process (S3420) further includes step (S3428a) of connecting the floor (ground) contact points (720) to generate a three-dimensional reference plane (730), extending the three-dimensional reference plane to generate a floor (ground) plane (740). As shown in Figure 7a, a three-dimensional reference plane can be generated by connecting the lowest points of the chair legs, and this can be further extended to a floor (ground) plane (740).
[0142] Here, the extension from the 3D reference plane extends to the boundaries of gravity-horizontal objects, such as the corners of gravity-horizontal objects, the starting points of walls, and the bottom layers of gravity-horizontal objects, resulting in the plane of the floor (ground). Here, gravity-horizontal objects generally refer to objects positioned upright parallel to the direction of gravity, such as glass windows and walls. Specifically, gravity-horizontal objects can be objects (e.g., doors, windows, etc.) that have lines and / or surfaces positioned horizontally with respect to the gravity vector. Gravity-horizontal objects can also be objects (e.g., door frames, walls, etc.) that have one or more lines and / or surfaces in contact with the plane of the floor (ground). Gravity-horizontal objects can also have a width-to-height ratio (width divided by height) of 1 or less.
[0143] On the other hand, the second embodiment of detecting a floor (ground) plane using a large number of objects (S3420) begins with the step (S3422b) of detecting a large number of similar objects whose similarity to each other is the same as or greater than a predetermined value. For example, doors, desks, chairs, etc., of a certain design may be such similar objects.
[0144] Considering that the video (image) processing method described in this disclosure can be performed continuously and repeatedly, if the video (image) processing method has already been provided to a space in which the above-mentioned similar objects exist, in step (S3422b), it is possible to detect the above-mentioned similar objects by referring to the position and length of similar objects that have been previously detected.
[0145] Following step (S3422b), a second embodiment of the process (S3420) further includes step (S3424b) of detecting at least one bottom-level object, which is the object located at the bottom of the numerous similar objects. As in the example shown in Figure 7a, the chair legs and the bottom of the door are bottom-level objects because they are in contact with the floor (ground).
[0146] Following step (S3424b), the second embodiment of the process (S3420) further includes step (S3426b) of detecting floor (ground) contact points, which are the contact points of each of the bottom-level object segments to the bottom level of the object segmentation, and following step (S3426b), further includes step (S3428b) of connecting the floor (ground) contact points to generate a three-dimensional reference plane, and extending the three-dimensional reference plane to generate a floor (ground) plane.
[0147] As an example of a modification of the second embodiment described above, the process (S3420) includes the steps of: detecting a large number of similar objects whose similarity to each other is the same as or greater than a predetermined value (S3422b'); determining whether one of the large number of similar objects is located completely above another of the large number of similar objects (S3424b'); and, based on whether or not it is located completely above, assigning objects from the large number of similar objects that have differences in their layered floor (ground) surfaces to a set of multiple distinct layered floor (ground) surfaces. The process can include the inclusion step (S3426b') and, for each of the above-mentioned sets of multiple distinct layered floor (ground) surfaces, the step (S3428b') of generating a 3D reference plane by connecting the floor (ground) contact points, which are the points of contact with the lowest level of the object segmentation of the layered floor (ground) surface set object that belongs to each of the above-mentioned sets of layered floor (ground) surfaces, and then extending the above-mentioned 3D reference plane to generate a layered floor (ground) (ground) plane, thereby generating the above-mentioned layered floor (ground) (ground) plane as two or more floor (ground) planes.
[0148] Figure 7b is an illustrative diagram illustrating multiple steps for detecting two or more floor (ground) planes in a video (image) processing method based on one embodiment of the present disclosure.
[0149] Since stage (S3422b') is the same as stage (S3422b), referring to Figure 7b, the "complete upper stage" of stages (S3424b') and (S3426b') can be explained as follows: First, it can be assumed that objects whose similarity is above a predetermined threshold have the same height h within the error range.
[0150] In the overall image, if each object is positioned far apart from the others along the vertical axis (y-axis), it is possible to detect (i) that there is distance between the objects, or (ii) that there is a difference in the height of the floor surface (ground) supporting the objects. If the bottom left corner of the overall image is taken as the origin of the coordinate system, when there is distance between objects, the object with a lower y-axis coordinate value (700b) will be closer to the camera than the other objects (700a), and therefore its height (h b The height (h) of an object (700a) that is larger and has a relatively higher y-axis coordinate value a ) is the height (h b It must become smaller by a certain ratio relative to ).
[0151] In other words, objects whose height changes at a constant ratio depending on the y-axis coordinate value (e.g., 700b and 700a), and objects that have differences in the horizontal axis (x-axis) coordinate value but only a difference below a certain level in the y-axis coordinate value (e.g., 700b and 700d), can be determined to exist on the same floor surface (ground), and the corresponding set of floor surfaces (ground) can be generated.
[0152] For example, even if the y-axis coordinate value is greater than that of the other object (700a), the height (h) of the other object (700a) is greater. a ) from a certain ratio to height (h c ) does not decrease, height (h cIf there is an object (700c) that is the same as or higher than another object, it can be determined that it exists on a different layer floor (ground), and therefore it is possible to generate a set of different layer floor (ground) objects.
[0153] By repeating these inference assumptions, it is possible to generate the above-mentioned set of floor surfaces (ground surfaces).
[0154] Alternatively, instead of the height of each object, the height or area of the three-dimensional reference plane that each object occupies can be used.
[0155] On the other hand, in detecting the plane of the floor (ground) (S3420), the third embodiment using Manhattan space begins with the step (S3422c) of detecting at least one length-equivalent object that belongs to a length-equivalent category, which is a category that satisfies the condition that the deviation of the length of at least a part of an object or a partial object of an object that belongs to a specific category among at least one category, is smaller than a predetermined standard.
[0156] The aforementioned objects of equivalent length can include desks with relatively small height differences.
[0157] Following step (S3422c), a third embodiment of the process (S3420) further includes the steps of detecting a Manhattan space generated by the collection of objects of equal length (S3424c), detecting the floor (ground) of the Manhattan space (S3426c), and extending the floor (ground) of the Manhattan space horizontally to generate a plane of the floor (ground) (S3428c).
[0158] More specifically, the detection of the Manhattan space (S3424c) described above can be comprised of the following steps: a first step (S3424c-1) in which floor (ground) objects, which constitute the floor (ground) surface, are detected and the boundaries of the floor (ground) objects are generated; a second step (S3424c-2) in which wall objects, which are objects perpendicular to the boundaries of the floor (ground) objects, are detected; and a third step (S3424c-3) in which ceiling objects, which are objects other than the floor (ground) objects, are perpendicular to the wall objects.
[0159] In other words, here, Manhattan space refers to the space enclosed by floor (ground) objects, wall objects, and ceiling objects. Examples of wall objects perpendicular to the floor (ground) object include glass windows and doors.
[0160] Unlike the aforementioned embodiments, the fourth embodiment, which uses multiple objects with the same pattern to detect the plane of the floor (ground) (S3420), begins with the step of detecting multiple identical pattern objects, which are a large number of objects with the same pattern (S3422d).
[0161] Following step (S3422d), the fourth embodiment of the process (S3420) further includes, once the same pattern object is detected, a step (S3424d) of detecting the lower part of the same pattern object, and a step (S3426d) of measuring the relative distance between the same pattern objects based on either the blockage between the same pattern objects or the difference in length between the same pattern objects in the image.
[0162] For example, in step (S3426d), if one object hides and obscures the object segmentation of other objects, that one object can be detected as being closer to the system than the other objects. Also, if, among multiple objects with the same pattern, one object is smaller than the others, that one object can be detected as being further away from the system than the other objects.
[0163] Following step (S3426d), the fourth embodiment of the process (S3420) further includes step (S3428d) of generating a virtual plane from the floor (ground) contact points of two or more objects from the same pattern object, or from three or more points contained within one object from the same pattern object. When there are multiple generated virtual planes, it is considered that an ordinary engineer can easily understand that it is possible to generate layered virtual planes, with each layer being a layer, based on differences in length, state, and position.
[0164] Following step (S3428d), the fourth embodiment of the process (S3420) further includes step (S3429d) of extending the virtual plane to generate a floor (ground) plane.
[0165] Returning to the 3D measurement part in step (S3400), and continuing the explanation, the 3D measurement may further include the process of setting a virtual length reference line on the detected floor (ground) plane after detecting the plane of the floor (ground) (S3420) (S3430); and the process of measuring the distance between the object and the position of at least one origin of the system among the orientation and coordinates of the spatial object, or the position of the object relative to the position of the system (S3440).
[0166] On the other hand, the two-dimensional measurement in step (S3400) may include posture and direction measurement, which calculates the two-dimensional posture and direction of the object based on the relative positions of multiple partial objects of the object, and area measurement, which calculates the two-dimensional area of the object. This can be done by the detailed classification module (4300) using the measurement model (M431) of the detailed classification model (M430).
[0167] Due to the nature of 2D measurement, the area here refers to the area of the entire image or object image, without considering depth (or length). This 2D area can be measured through object segmentation classification and measurement.
[0168] Furthermore, in step (S3400), in order to generate the characteristics of the object, the detailed classification model (M430) may further include a deepened characteristics model (M432) for calculating at least one of the following: information on a partial object that is part of or a component belonging to the object; deepened classification information of the object; the main color of the object; the subordinates of the object; the main body of the object; one or more materials of the object; the transparency of the object; and whether or not the object is able to move on its own.
[0169] Here, the deepened classification information for the above object refers to information obtained by dividing the object's category into deepened classifications. For example, if the object's category is "dog," then the deepened classification information could be the breed of that dog.
[0170] Posture and behavior Furthermore, the detailed classification model (M430) may further include a posture discrimination model (M433) that calculates the posture of the object based on the processing results of the measurement model (M431) and the deepening characteristics model (M432), and may further include an action discrimination model (M434) that classifies the actions of the object based on the temporally continuous posture calculated from the posture discrimination model (M433).
[0171] For example, the processor (120) of the computing device (100) can use an attitude discrimination model (M433) to obtain attitude information related to an object from the corresponding analysis information.
[0172] Furthermore, the processor (120) of the computing device (100) can use an action discrimination model (M434) to obtain action information related to an object from the posture information related to the object.
[0173] Here, the term "action" encompasses both behaviors that do not reflect context and behaviors that do reflect context, but behavior will be discussed later.
[0174] Because the method for determining an object's posture may differ depending on its category, the posture determination model (M433) can be a category-specific posture determination model that varies depending on the category. For example, a dog posture determination model that calculates a dog's posture when it is sitting may be different from a human posture determination model that calculates a human's posture when it is sitting. Therefore, in the posture determination model (M433), it is possible to determine which posture determination method to apply to an object from among multiple posture determination methods based on the analysis information corresponding to the object. The posture determination model (M433) can change and apply the posture determination method depending on the classification information that indicates the object's category.
[0175] Similarly, since the method of determining an action may differ depending on the object's category, the action determination model (M434) can be a category-specific action determination model that differs depending on the category. The action determination model (M434) can be provided with a different method of determining an action based on classification information indicating the object's category.
[0176] The posture discrimination model (M433) can generate posture information for an object from the analysis data, including information about the object's corresponding N-dimensional (e.g., 1D, 2D, 3D, etc.) posture. The posture information for an object can include information about the object's posture over time.
[0177] The action discrimination model (M434) can obtain action information related to an object, including the movement vectors of N dimensions (e.g., 1D, 2D, 3D, etc.) corresponding to the object, from the object's orientation information. Furthermore, the action discrimination model (M434) can generate action information related to the object's context, including the action classification of N dimensions (e.g., 1D, 2D, 3D, etc.) corresponding to the object, from the object's orientation information. The object's context can include at least one of the following: the object's state or / or the purpose of the action. The action classification can be calculated by identifying the positions of partial objects included in the object, determining the object's orientation based on the positions of the partial objects, and determining the object's action classification based on the object's orientation. A partial object can include at least one of the parts of the object or components belonging to the object.
[0178] The determination of posture and action can be performed in two dimensions or in three dimensions. Therefore, the movement vector can include a two-dimensional movement vector that includes at least one of the two-dimensional direction and / or velocity of the object. Furthermore, the movement vector can include a three-dimensional movement vector of the object calculated based on at least one of the position, velocity and / or acceleration of the camera that captured the video (image).
[0179] In an embodiment in which posture and action are determined in two dimensions, in step (S3400), the detailed classification module (4300) can calculate two-dimensional motion information of the object, including the two-dimensional movement vector of the object and the two-dimensional action classification of the object, from a plurality of temporally continuous object images. Therefore, the two-dimensional action classification can be calculated by performing the steps of identifying the position of each of the partial objects included in the object (S3450a), determining the two-dimensional posture of the object based on the relative positions of each of the partial objects using a posture discrimination model (M433) (S3460a), and determining the two-dimensional action classification of the object based on the temporally continuous two-dimensional posture of the object using an action discrimination model (M434) (S3470a).
[0180] Here, the 2D movement vector indicates the 2D direction and velocity of the object, and the 2D action classification indicates the type of action determined from the object's 2D orientation.
[0181] On the other hand, in an embodiment in which posture and action are determined in three dimensions, in step (S3400), the detailed classification module (4300) can calculate three-dimensional motion information of the object, including the three-dimensional movement vector of the object and the three-dimensional action classification of the object, from a plurality of temporally continuous object images. Therefore, the three-dimensional action classification can be calculated by performing the steps of: identifying the position of each of the partial objects included in the object (S3450b); determining the three-dimensional posture of the object based on the relative positions of each of the partial objects using a posture discrimination model (M433) (S3460b); and determining the three-dimensional action classification of the object based on the temporally continuous three-dimensional posture of the object using an action discrimination model (M434) (S3470b).
[0182] Unlike the determination of two-dimensional posture and actions, the determination of three-dimensional posture and actions requires that movement in the depth (or length) direction from the imaging device also be reflected. Therefore, the computing device (100) can use the linked sensor (340) to calculate or estimate at least one (one or more) of the position, velocity, and / or acceleration of the imaging device, and reflect this to calculate the three-dimensional movement vector of the object.
[0183] For example, when the system described in this disclosure drives a gimbal motor (330) to rotate the axis of the gimbal, tracks a bicycle, which is an object moving to the right relative to the system, and calculates the three-dimensional motion vector of the bicycle, it is possible for both the motion vector from the motor (330) and the motion vector in the video (image) to be reflected in that three-dimensional motion vector.
[0184] To calculate the three-dimensional position and three-dimensional movement vector of the above object, the standard focal length of the imaging device (200) and / or the standard focal length of the overall image or the image of the object can be used. The relationship between the distance between the imaging device (200) and the object, the standard focal length, the three-dimensional height of the object (actual height), the height of the image (i.e., the vertical size of the image), the height of the object in the image, and the height of the imaging device (200) is as shown in Equation 1 below.
[0185]
number
[0186] Using equation 1, it is possible to calculate the distance from the imaging device (200) to the object, or the three-dimensional height of the object.
[0187] One of the references used to measure the distance from the imaging device (200) to an object is the distance reference line, which can be set using concentric spheres based on the system of this disclosure, for example, when two or more objects of the same length are detected. This line is composed of a reference plane circle and a measurement plane circle, as described later. The steps for setting it are as follows:
[0188] Figure 7c is a conceptual diagram illustrating a method for generating and using a reference plane circle and a measurement plane circle in a video (image) processing method based on one embodiment of this disclosure.
[0189] According to Figure 7c, first, a reference plane circle (750) is generated along a plane direction perpendicular to gravity. When determining whether it is perpendicular to gravity, at least one (or more) of the accelerometer, gyroscope, and / or the gravity horizontal object can be used.
[0190] Here, the reference plane circle (750) corresponds to the shape of the cross-section of the surface of the concentric sphere (740) with the imaging device (200) as the origin, that is, the concentric spherical surface, when cut in the xy plane.
[0191] Next, from among the objects of equal length, we detect objects of the same length that are identical to each other.
[0192] If we refer to the set of multiple points on the concentric spheres in Figure 7c, separated by an angle (θ) from the z-axis, as a measurement plane circle, we measure the difference in angular (θ) coordinate values between these objects by utilizing the fact that objects of the same length are interposed between these measurement plane circles at each angle (θ).
[0193] By doing so, it is possible to measure the distance from the imaging device (200) to each object by utilizing the differences in the above-mentioned angular coordinate values. In this case, at least one (or more) optical characteristic of the standard focal length and the lens and / or aperture size of the imaging device can be used as an auxiliary.
[0194] On the other hand, the models discussed in this disclosure, including the action discrimination model (M434), can be generated by supervised learning or reinforcement learning. It is well known that each action of an object, along with corresponding analysis values (text or classification index) and video information, can be used as training data for supervised learning. Furthermore, it is also possible to use a reinforcement learning method in which action discrimination model (M434) outputs action information to the user, and the user provides affirmative or negative feedback, which is then used to modify the model.
[0195] context The detailed classification model (M430) as a set can further include a context model (M435) that infers context from the overall video (image) described above. Context describes the state of an object and the purpose of an action. For example, if a video (image) of a person peeling an orange with a knife in a kitchen is input, the context model (M435) can output the sentence "A person in a kitchen is peeling an orange with a knife," or a corresponding signal, as context.
[0196] Behavior refers to an action of an object that does not reflect context. For example, both a person running on a basketball court and a person running on a treadmill are considered to be "running." Such behaviors can be combined with context, as will be discussed later, to determine an action that takes context into account. According to this, the former corresponds to the action of "playing basketball," and the latter corresponds to the action of "using a treadmill."
[0197] Specifically, a context can be an object context that includes at least one of the actions and states of individual objects shown in the overall image, as well as contextual interaction objects, which are other objects detected to interact with the individual objects through the aforementioned action.
[0198] Furthermore, the context can be a spatial context that includes, inferred from each of the spatial objects and individual objects other than the spatial objects that are visible in the overall video (image), a location which is the type of spatial object, at least one of the actions and states of each of the individual objects, an actor which is an individual object corresponding to the subject of the action, and a context interaction object which is an object detected to interact with the individual object through the action.
[0199] In one embodiment, the synthesis of actions and contexts to determine context-aware actions can be realized through supervised learning using an artificial neural network model. For example, an artificial neural network model can be trained using training data in which a sequence of images, i.e., a scene from a video, is used as input data, and the linguistic analysis (i.e., data expressed in language) including the actions of objects and spatial context is labeled with the correct output data.
[0200] This is as described in the non-patent literature paper Wu, Z., Yao, T., Fu, Y., & Jiang, Y.-G. (2017). Deep learning for video classification and captioning. In Frontiers of Multimedia Research (pp.3-29). ACM.
[0201] On the other hand, for a more specific method of determining the actions of each individual object, you can refer to Wu, C.-Y., Girshick, R., He, K., Feichtenhofer, C., & Krahenbuhl, P. (2020). A Multigrid Method for Efficiently Training Video Models. In 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE. https: / / doi.org / 10.1109 / cvpr42600.2020.00023.
[0202] On the other hand, there are cases where an action is heterogeneous in its spatial context, and in such cases, there is room to analyze the action in a way that does not fit the context. For example, an object that is a person running in a cafe can be analyzed not as the action of "running," but as the action of "being busy." Such inconsistencies occur because the action does not adequately reflect the context, or the context does not adequately reflect the action. In other words, actions can be inferred from context, and context can also be inferred from actions. Therefore, a cyclical process is sometimes required, in which context is inferred from actions, and actions are inferred again from context. This can be achieved, for example, by neural network models such as circulating neural networks (RNNs) or bidirectional LSTMs, as disclosed in the non-patent literature paper Ullah, A., Ahmad, J., Muhammad, K., Sajjad, M., & Baik, SW (2018). Action Recognition in Video Sequences using Deep Bi-Directional LSTM With CNN Features. IEEE Access, 6, 1155-1166. https: / / doi.org / 10.1109 / access.2017.2778011.
[0203] Priority The following details the priority of objects, which is closely related to the importance of objects mentioned above. Here, importance is a value used to assign priority, and it is usually possible to calculate the importance of each object and then assign a priority order based on that importance.
[0204] Object priorities can include privilege-based priorities, which are assigned to each of the aforementioned objects based on their respective privileges. This involves specifying priorities based on the privileges of each classified object; for example, in the system described in this disclosure, if a particular user has the highest privileges on an object, it is possible to assign a higher priority to that specific user.
[0205] The priority order based on the above permissions can be calculated by performing the following steps: first, determining whether the object is an authorized user who has been predetermined to be able to use the computing device, based on the characteristics of the object analyzed using the detailed classification model (M430); and second, if the object is an authorized user, setting the predetermined ranking for that user to the priority order based on the above permissions for the object.
[0206] For example, since the above-mentioned authority holder is expected to be a person, it is possible to first detect objects that are people and then set permission-based priorities only for those detected objects that are people.
[0207] Furthermore, the computing device described in this disclosure can identify individuals from among multiple people using the aforementioned advanced classification information. To identify these individuals, it is possible to utilize the characteristics of objects, such as the color of an object, the characteristics of a part of that object, or the object records of the aforementioned object.
[0208] On the other hand, object priorities may include classification-based priorities, which are assigned to each of multiple object sets, each set of objects that are distinguished by at least one characteristic, which includes at least the category of the object or the in-depth classification information of the object.
[0209] The classification-based priority can be a predetermined priority assigned to categories or advanced classification information such as people, animals, and other objects, but the system's authorized user can set it manually. As another example, the computing device described in this disclosure can also automatically and dynamically set the priority for each category or advanced classification information based on the system's authorized user's usage patterns.
[0210] Next, the priority of objects may include a size-based priority assigned to each of a set of objects, each set of objects distinguished by at least one characteristic, including the size of the object.
[0211] The priority order based on the above size can be calculated by the detailed classification module (4300) performing the following steps: obtaining the object segmentation of the above object; selecting objects from the above object group whose object segmentation size occupies a ratio equal to or greater than a predetermined ratio in the overall image; and setting the priority order based on the size of the object segmentation of the selected objects. This method assigns a higher priority to objects that are larger or that are closer to the system of this disclosure, as the proportion of the object segmentation of the above object in the overall image increases.
[0212] Furthermore, the priority of objects may include an action-based priority assigned to each of a set of objects, each of which is divided by at least one state, the state of which includes at least the action of the object described above.
[0213] Specifically, the detailed classification module (4300) can search for objects that perform a specific action and assign a priority to that action.
[0214] For example, if the action of a person falling is detected, it is possible to assign the highest priority to that fallen person. Another example is if there is a person-like object that gives commands through specific gestures, such as swiping in the air; in this case, it is possible to assign a high priority to that commanding person-like object.
[0215] As a final example, object priority may be a context-based priority assigned to each of a set of objects, each of which is distinguished by at least one state that includes the context of the object mentioned above.
[0216] Based on the prioritization criteria described above, the detailed classification module (4300) performs state analysis or spatial analysis on the overall video (image) to infer object context and spatial context from the overall video (image). Based on whether the object is an actor or a context-interacting object, which is detected as an actor or an object that interacts with the actor through the actions of the actor, based on at least one of the inferred object context and spatial context, it is possible to set a relatively high priority for the actor or the context-interacting object.
[0217] For example, if the spatial context is "people of child age are playing baseball on a field," then it is possible to assign a relatively high priority to people who are detected to be holding a baseball bat.
[0218] The steps and processes described in this disclosure do not mean that they must be performed in the order they are described, unless there is a logical inconsistency or the context otherwise defines them. It is obvious to the average person that each of the steps and processes can be performed simultaneously or not simultaneously.
[0219] Furthermore, while the aforementioned steps may be performed only once, preferably, as described above, in order to acquire temporally continuous video (images), the above steps can be performed in real time and / or iteratively.
[0220] In other words, the video (image) processing method of this disclosure may further include a step (S4000) to return to the video (image) acquisition step (S1000) in order to acquire a new overall video (image), and at this time, the video (image) acquisition step (S1000) can be executed again by controlling the available resources with the tracking controller (4400).
[0221] Here, the tracking controller (4400) is a module, similar to the video (image) input module (4100), the object analysis module (4200), and the detailed classification module (4300), so that the system of this disclosure performs the function of realizing the movement of the camera (200), i.e., tracking, so that the camera (200) faces a target object or target space, where available resources refer to hardware and / or software that enable such tracking.
[0222] In one embodiment, the tracking controller (4400) may include a target position determination module (4420) that determines a target position and magnification corresponding to the composition captured by the imaging device (200), and a tracking resource controller (4440) that controls the available resources, i.e., hardware and software resources, to acquire the entire image of the target position.
[0223] Target position determination module Figures 8a to 8c are flowcharts illustrating, as examples, several methods used to determine the target position in the image processing method described in this disclosure.
[0224] In one embodiment shown with reference to Figure 8a, the target position determination module (4420) can perform the following steps: obtaining the object priority of at least one candidate object shown in the overall image (S4110a); obtaining the object segmentation of the candidate object (S4120a); identifying a context interaction object, which is an object detected to interact with the candidate object in accordance with the context inferred in the overall image (S4130a); determining the composition to be captured by the imaging device (200) so as to include at least one target object, which is a candidate object with an object priority of a predetermined rank or higher, and the context interaction object of the target object (S4140a); and determining the target position and magnification by extending the composition along the direction of the target object (S4150a).
[0225] The identification of context interaction objects can be performed using a context model (M435). As disclosed in the non-patent literature paper Shin, D., & Kim, I. (2018). Deep Image Understanding Using Multilayered Contexts. Mathematical Problems in Engineering, 2018, 1-11., the aforementioned context model (M435) can be generated by supervised learning.
[0226] More specifically, the additional composition prediction step (S4150a) may include the step of acquiring the direction of the target object (S4152a) and the step of acquiring the velocity of the target object (S4154a). As mentioned above, the direction of the target object can be acquired based on the relative positions between multiple partial objects contained within the target object, and the velocity of the target object, i.e., the 3D motion vector, can be calculated based on the 2D motion vector acquired from the video (image) and the information from the sensor (340).
[0227] Next, the additional composition prediction step (S4150a) may further include the step of calculating the influence range of the target object, which is the range in which the target object can interact with the target object, based on the direction and velocity of the target object (S4156a), and the step of determining the target position and magnification by expanding the composition in accordance with the influence range (S4158a).
[0228] Here, the scope of influence refers to the range of space that the contextual interaction object may occupy, at least temporarily, within a predetermined time range starting from the present, where the target object may occupy or interact with; in other words, the range of space that the target object can physically contact, for example, the range within which a part of a person's body can physically contact, or the range within which a contextual interaction object can exist that the target object can send and receive signals through its eyes, ears, mouth, etc., for example, the field of view (FOV) within which the target object can send and receive signals.
[0229] For example, in stage (S4156a), the influence range of the eye can be 20m and the influence range of the hand can be 1m. This can be used when adjusting the composition in combination with the action information (action classification) of each object. For example, if the target object is "a child standing at a baseball plate with a bat in hand," in stage (S4158a), the composition can be adjusted to include the direction of the child's gaze and the object of the bat held in the child's hand.
[0230] On the other hand, in another embodiment shown with reference to Figure 8b, the target position determination module (4420) can determine the composition to be photographed based on the priority of objects. Specifically, the target position determination module (4420) can perform the steps of: obtaining the object priority of at least one candidate object shown in the overall image (S4120b); determining at least one target object which is a candidate object having an object priority of a predetermined rank or higher (S4140b); and determining the composition to be photographed by the shooting device (200) so as to include the target object (S4160b).
[0231] As shown in Figure 8c, in yet another embodiment that reflects the priority of objects, the target position determination module (4420) takes the following steps: (S4120c) obtains the object priority of at least one candidate object shown in the overall image; (S4140c) generates a virtual segmentation including the candidate object according to the object priority of the candidate object, such that the area of the virtual segmentation increases as the object priority increases; and (S4140c) increases the area of the virtual segmentation by a predetermined increment function. It is possible to perform the steps of determining the equilibrium state of the attractive force between the virtual segmentation, which increases monotonically accordingly, the first repulsive force between the candidate objects, which decreases monotonically according to the distance between the candidate objects by a predetermined first decreasing function, and the second repulsive force between the candidate object and the frame boundary of the overall image by a predetermined second decreasing function (S4160c), and determining the target position and magnification corresponding to the composition by calculating the target center, which is the central point of at least one candidate object, in the equilibrium state (S4180c).
[0232] Tracking Resource Controller The tracking resource controller (4440) may include an image (picture) conversion controller (4442) that controls the acquisition of new images (pictures) from images (pictures) acquired from the camera (200) without controlling the attitude of the gimbal (300) and the camera (200), and a frame conversion controller (4444) that controls the gimbal (300) to which the camera (200) is mounted and which has one or more rotation axes for controlling the attitude of the camera (200), and the computing device (100), thereby controlling the camera (200) to acquire the new images (pictures).
[0233] The frame conversion controller (4444) and the video (image) conversion controller (4442) can operate in a complementary manner, but the tracking resource controller (4440) can first acquire a new video (image) of the desired resolution from the frame conversion controller (4444), and if that acquisition fails, it can control the acquisition of a new video (image) through the reconstruction of the video (image) by the video (image) conversion controller (4442).
[0234] In one specific embodiment, the video (image) conversion controller (4442) can perform the following steps: determine whether or not there are idle resources available for video (image) reconstruction (S4310); if such idle resources exist, load the acquired entire video (image) or a portion of the entire video (image) into memory as the original video (image) (S4320); and depending on the target position and magnification, crop the original video (image) or, if the resolution of the original video (image) does not reach a predetermined threshold, perform super-resolution to acquire a new video (image) that matches the target position and magnification (S4330).
[0235] It is common knowledge among engineers that super-resolution can be achieved using neural networks such as autoencoders.
[0236] On the other hand, the frame conversion controller (4444) may include a gimbal controller (4444a) that controls the direction of the shooting device (200) by operating one or more rotation axes of the gimbal (300) to achieve the target position (i.e., to reach the target position), a zoom controller (4444b) that controls the zoom-in and zoom-out of the shooting device (200) to reach the magnification, and a front and rear shooting device controller (4444c) that controls the available resources to capture the environment while reducing the operation of the hardware, and may include at least the gimbal controller (4444a).
[0237] The imaging device (200) can be configured in two or more units, but for example, it can be an imaging device (200) mounted on the front and rear of a portable terminal (100). The front and rear imaging device controller (4444c) can use the front imaging device (200a; not shown) and the rear imaging device (200b; not shown) to scan the entire space surrounding the system, that is, the surrounding space, and the imaging device (200) to be used can be selected from the front and rear imaging devices, taking into consideration the user's convenience.
[0238] For example, the front and rear imaging device controller (4444c) can select the front imaging device (200a) as the imaging device used to acquire the user's video (image) when the user is looking at the display of a portable terminal (100). In another example, the front and rear imaging device controller (4444c) can, after recognizing an object as a front imaging device (200a), control the device to recognize the object as a rear imaging device (200b) with a higher resolution if a video (image) with a higher resolution than that of the front imaging device (200a) is required for that object.
[0239] Furthermore, the front and rear imaging device controller (4444c) can use the front imaging device (200a) and the rear imaging device (200b) simultaneously or sequentially to acquire images of the surrounding space while minimizing the rotation of the gimbal (300) axis.
[0240] On the other hand, the frame transition controller (4444) can control the orientation of the camera (200) to either the portrait or landscape direction, according to the aspect ratio (the ratio of the width to the height) of the specific object or partial object contained within that specific object, so that the composition includes at least one specific object or partial object. For example, this can be done by rotating the roll axis (R) of the gimbal (300). In the aspect ratio, the width can mean perpendicular to the gravity vector. In the aspect ratio, the height can mean horizontal to the gravity vector. Therefore, the aspect ratio can be calculated by taking a portion of a reference plane as the width. The reference plane can be a plane perpendicular to the gravity vector, and / or a plane in contact with the floor (ground). However, the reference plane is not limited to this, and various planes can be set as the reference plane. The aspect ratio (width to height) can also remain consistent (for example, greater than or equal to 1, or less than or equal to 1) depending on the object's orientation.
[0241] Specifically, the frame transition controller (4444) can control the orientation of the shooting device to either the vertical or horizontal direction so that object segmentation of at least one specific object or specific partial object, or an object box containing at least one specific object or specific partial object, does not come into contact with the frame boundary of the overall image.
[0242] Figures 9a and 9b are flowcharts illustrating, as examples, several methods used to control the orientation of the imaging device in the image processing method described in this disclosure.
[0243] Using Figure 9a, an embodiment of the frame transition controller (4444) for controlling the orientation of the above-mentioned shooting device will be described in more detail. The frame transition controller (4444) controls the gimbal to move in the opposite direction to the position of the first contact when the object segmentation or object box of at least one specific object or specific partial object makes first contact with the frame boundary of the overall image, within the range in which the center point of the object segmentation or object box is located within the composition. In the step of resolving the first contact (S4210a), and after the above movement, if the object segmentation or the object box makes a second contact with the frame boundary on the opposite side of the position of the first contact, it is possible to perform the step of resolving the second contact (S4220a) through at least one of the following: (i) switching the orientation of the imaging device from one of the vertical (portrait) and horizontal (landscape) directions to the other, and (ii) controlling the zoom-in and zoom-out of the imaging device.
[0244] Using Figure 9b, another embodiment of the frame conversion controller (4444) will be specifically described. In the step of acquiring the ratio characteristics of at least one specific object or specific partial object, the ratio characteristics are predetermined for each category of the specific object or specific partial object, or calculated for each category by measurement of the specific object or specific partial object, and include category-specific ratio characteristics, which are the horizontal-to-vertical ratio (a value obtained by dividing the horizontal length by the vertical length), and predetermined for each detailed classification of the specific object or specific partial object. Step (S4210b) includes a detailed classification ratio characteristic, which is the ratio of width to height (a value obtained by dividing the width by the height) calculated by measuring the specific object or specific partial object described above, and the ratio of width to height (a value obtained by dividing the width by the height) of the composition determined by the target position determination module described above. It is also possible to perform a step (S4220b) in which it is determined whether or not to perform a roll rotation, which is an operation that switches the orientation of the imaging device from either the vertical (portrait) direction or the horizontal (landscape) direction to the other, based on the ratio characteristic described above, and to adjust the orientation of the imaging device depending on whether or not the roll rotation is performed.
[0245] For example, the aspect ratio of a television (the ratio of the width to the height) is greater than 1, so the orientation of the camera (200) can be adjusted horizontally. The aspect ratio of a standing person (the ratio of the width to the height) is less than 1, so the orientation of the camera (200) can be adjusted vertically.
[0246] Following step (S4220b), the frame transition controller (4444) of this embodiment may, if the object segmentation or object box of at least one specific object or specific partial object comes into contact with the frame boundary of the overall image, perform an additional step (S4230b) by resolving the contact through at least one of the following: (i) switching the orientation of the shooting device from either portrait or landscape to the other, and (ii) controlling the zoom-in and zoom-out of the shooting device, thereby ensuring that at least one specific object or specific partial object is included in the composition. In this case, if the object segmentation or object box does not come into contact with the frame boundary, the process ends at step (S4220b) and step (S4230b) is not performed.
[0247] OCR (Optical Character Recognition) The following describes the detailed classification model (M430) and its application to OCR. The classification model (M430) may further include an OCR model (M436) that reads text displayed on the surface of the object. Therefore, the processor (120) of the computing device (100) can use the OCR model (M436) to obtain the characters contained in the object from the object and the corresponding analytical information.
[0248] Accordingly, the aforementioned step (S3400) may include a step (S3500) in which the detailed classification module (4300) performs OCR using the OCR model (M436) of the detailed classification model (M430) to calculate the text of the object as part of the object's characteristics.
[0249] To explain step (S3500) in more detail, the OCR model (M436) can first include a step (S3520) of determining whether or not characters are displayed on the surface of the object.
[0250] At stage (S3520), it is known that objects on which characters are displayed can be identified by detecting characters using deep learning or conventional OCR technology.
[0251] Following step (S3520), the OCR model (M436) can perform an OCR execution step (S3540) in which, if characters are displayed on the surface of the object, it performs OCR on the displayed characters (e.g., the entire character and / or character set), and inputs (saves) the resulting text, which is the result of the OCR, as a characteristic of the object.
[0252] According to one embodiment of the present disclosure, the OCR model (M436) can perform the steps of determining whether or not characters are displayed on the surface of an object, and, if it is determined that characters are displayed on the surface of the object, performing OCR on the entire set of characters.
[0253] According to another embodiment of the present disclosure, the OCR model (M436) can perform the steps of determining whether or not characters are displayed on the surface of an object, and, if it is determined that characters are displayed on the surface of the object, performing OCR for each character set.
[0254] Methods for classifying character sets At least one of the processor (120) or the OCR model (M436) is capable of classifying the character set to which a character belongs based on at least one of the character type, form, size, or arrangement. Furthermore, the processor (120) is capable of classifying character sets that have a high probability of being highly relevant in the analysis of a given context.
[0255] Possible methods for distinguishing between character sets to which a character belongs or character sets to which a context belongs include the following. For example, at least one of the processor (120) or OCR model (M436) can distinguish between character sets based on the characteristics of the character set.
[0256] The characteristics of a character set can include one or more of the following: the language type of the characters, line spacing, character spacing, character width (e.g., the ratio of width to height), size, boldness, color, typeface (font), style, characters associated with the beginning or end of a sentence, spacing (e.g., the space between at least one boundary of the following sentence, object, or partial object), or the position where characters appear in an object. Therefore, at least one of the processor (120) or OCR model (M436) can classify character sets using one or more of the following features: the language type of the characters, line spacing, character spacing, character width (e.g., the ratio of width to height), size, boldness, color, typeface (font), style, characters associated with the beginning or end of a sentence, spacing (e.g., the space between at least one boundary of the following sentence, object, or partial object), or the position where characters appear in an object.
[0257] For example, at least one of the processor (120) or OCR model (M436) can distinguish between complete sentences or contexts depending on the language to which the characters belong. When the characters to be recognized are multilingual, such as on the front of a product instruction manual, at least one of the processor (120) or OCR model (M436) can distinguish between complete sentences or contexts depending on the language to which each character belongs.
[0258] To give another example, at least one of the processor (120) or OCR model (M436) can distinguish paragraphs based on at least one of line spacing, character spacing, character width, or color, given that the text is of the same size and font.
[0259] To give yet another example, at least one of the processors (120) or OCR model (M436) can divide character sets based on the top of a paragraph or the position where the paragraph begins. The top of a paragraph or the position where the paragraph begins can also indicate the character or title of the paragraph using bold style or different sized characters.
[0260] Objectification of character sets The processor (120) and / or the OCR model (M436) can classify character sets based on the "method for classifying character sets" described above. The processor (120) and / or the OCR model (M436) can recognize character set objects based on the character sets. A character set object can be a character set set set as a single target. A character set object can contain at least one of the following: character images, text information, or characteristics of a character set. For example, a character set object can contain character images before OCR or natural language understanding (processing). Another example is that a character set object can contain text information after OCR or natural language understanding (processing). In this case, the character set object can also be separated into multiple character set objects after OCR or natural language understanding (processing).
[0261] The processor (120) and / or the OCR model (M436) are capable of distinguishing one or more character sets displayed on the surface of an object or partial object. The processor (120) and / or the OCR model (M436) are also capable of performing contextual analysis based on at least one piece of information relating to the characteristics or state of the object or partial object on which the character sets are displayed. The characteristics and / or state of the object and / or partial object may be included in the characteristics of the character sets.
[0262] The processor (120) and / or the OCR model (M436) can use the type and / or location of the object or partial object on which the character set is displayed as additional information for natural language understanding (processing) and / or contextual analysis.
[0263] For example, if the characters displayed on the shirt (object or partial object) and the characters displayed on the car (object or partial object) are the same, the processor (120) and / or the OCR model (M436) can use the characters displayed on either the shirt or the car (e.g., a brand name) as additional information. In this case, the processor (120) and / or the OCR model (M436) can determine (or determine that there is a relatively high probability that) that the characters displayed on the car reflect the characteristics of the object and / or the characteristics of the overall context, and can use the characters displayed on the car as additional information.
[0264] The processor (120) and / or the OCR model (M436) can acquire the characteristics of an object based on the positions where multiple characters are displayed, if the object contains multiple characters. Multiple characters may have different meanings depending on their position. For example, if characters are displayed on the front of a shirt (object or partial object) and on a label on the shirt (e.g., a care label), the processor (120) and / or the OCR model (M436) can determine that the characters on the shirt label reflect the characteristics of the object (or have a relatively high probability of reflecting them), and can use the characters on the shirt label as additional information.
[0265] The processor (120) and / or OCR model (M436) can perform OCR and / or natural language understanding (processing) on each character set object for the objectified character set. The processor (120) and / or OCR model (M436) can perform OCR and / or natural language understanding (processing) on one or more character set objects based on the characteristics of the character set. In this case, the processor (120) and / or OCR model (M436) can also perform OCR and / or natural language understanding (processing) on one or more character set objects sequentially or simultaneously based on the characteristics of the character set.
[0266] The unit of OCR and / or natural language understanding (processing) is not limited to the entirety of the characters on the surface of an object or partial object; the smallest unit can be a character or character set classified according to the characteristics of the character set described above. On the other hand, the OCR execution stage (S3540) can be performed using a variety of methods.
[0267] For example, step (S3540) of performing OCR on characters (e.g., the whole character and / or character set) by the processor (120) of the computing device (100) may include steps of: extracting at least one video (image) sample from an object on which characters are displayed; determining the boundary line of the area on which characters are displayed from at least one video (image) sample; generating a character display video (image) (e.g., a whole character display video (image) and / or a character set display video (image)) from at least one video (image) sample based on at least one boundary line or boundary point belonging to the boundary line; and performing OCR on the character display video (image).
[0268] The video (image) sample can be a video (image) pattern that exists at at least one position among the boundary lines or / and boundary points. The video (image) pattern can include at least one of some characters, the boundary parts of characters, a part of a character or / and the background. The background can mean a video (image) pattern that does not constitute a character.
[0269] The steps of generating a character display video (image) based on at least one of the boundary lines or boundary points (which are points belonging to the boundary line) among at least one video (image) sample, and performing OCR on the character display video (image) can include the step of obtaining, as a boundary marker, a video (image) pattern that is located at at least one of the boundary lines or boundary points among at least one video (image) sample, and the step of generating, as a partial video (image) included in the video (image) of the object on which the character is displayed, a character display video (image) that is a video (image) including the boundary marker. The boundary marker can include at least one of a start marker and / or an end marker. The start marker can be a video (image) pattern corresponding to the first character of the character. The end marker can be a video (image) pattern corresponding to the last character of the character.
[0270] The step of obtaining, as a boundary marker, a video (image) pattern that is located at at least one of the boundary lines or boundary points among at least one video (image) sample can include the step of using a tracking controller (4400) to obtain a video (image) of a character that includes a start marker and has a character recognition rate equal to or higher than a threshold value, and the step of determining whether an end marker is included in the video (image) of the character.
[0271] And when the end marker is included in the video (image) of the character, the processor (120) can determine the video (image) of the character as the character display video (image).
[0272] When the video (image) of the character does not contain an end marker, the processor (120) can use the tracking controller (4400) to obtain an additional video (image) of the character that includes the marker next to the last marker among the boundary markers included in the video (image) of the character and has a character recognition rate equal to or higher than a threshold value. The processor (120) can combine the additional video (image) of the character with the video (image) of the character to generate a combined video (image) of the character. When the combined video (image) of the character contains an end marker, the processor (120) can determine the combined video (image) of the character as the character display video (image).
[0273] On the other hand, as another example, the step (S3540) of performing OCR on a character by the processor (120) of the computing device (100) includes using the tracking controller to obtain a video (image) of an object that includes the start of the character and has a character recognition rate equal to or higher than a threshold value, performing OCR on the area of the first sentence in the video (image) of the object, and performing natural language understanding (NLU; natural language understanding) on the first text, which is the primary result of OCR, to calculate a first semantic value numerical value, which is a numerical value of the semantic value.
[0274] When the first semantic value numerical value is equal to or higher than the threshold value, the processor (120) can determine the first text as the result text, which is the result of OCR.
[0275] When the first semantic value numerical value is less than the threshold value, the processor (120) can perform OCR on the area of the next sentence after the area of the first sentence. The processor (120) can perform natural language understanding on the second text, which is the primary result of OCR, to calculate a second semantic value numerical value, which is a numerical value of the semantic value. When the second semantic value numerical value is equal to or higher than the threshold value, the processor (120) can determine the second text as the result text, which is the result of OCR.
[0276] The area for the first sentence can be defined as the area occupied by the sentence that is detected as being placed first in the character arrangement scheme based on the language used in the sentence.
[0277] The computing device (100) can be linked with the imaging device (200) and the gimbal (300). For example, the computing device (100) includes the imaging device (200) and can be linked with the gimbal (300) that controls the attitude of the imaging device (200).
[0278] The processor (120) of the computing device (100) can use the tracking controller (4400) to operate one or more rotation axes of the gimbal (300) and control the direction of the camera (200).
[0279] The processor (120) can acquire video (images) by zooming in and / or zooming out of the imaging device (200) using the tracking controller (4400).
[0280] Figures 10a to 10c are flowcharts illustrating multiple methods used to perform OCR in the video (image) processing method of this disclosure, and Figures 11a to 11d are diagrams shown as examples to illustrate multiple methods for performing OCR in the video (image) processing method of this disclosure.
[0281] First, referring to Figures 10a and 11a, the OCR execution step (S3540a) in the embodiment using boundary markers begins with a step (S3542a) of extracting a video (image) sample from an object on which the above-mentioned characters are displayed (for example, a paper document of reference numeral 1110), followed by a step (S3544a) of defining the boundary line (for example, a closed curve of reference numeral 1130) of the area (area) on which the characters are displayed from the video (image) sample.
[0282] Here, the image sample is an image pattern located at the boundary or boundary point, and this image pattern can be composed of some characters, the boundaries of characters, parts of characters, or a background. The boundary refers to the boundary of a group of characters shown on the surface of an object, and the boundary point refers to a set of points that can constitute a boundary. Here, the background refers to an image pattern that does not constitute characters on its own.
[0283] The OCR execution step (S3540a) of this embodiment further includes a boundary marker acquisition step (S3546a) following step (S3544a), in which an image pattern located at the boundary line or a point belonging to the boundary line is acquired from the image sample as a boundary marker. In this step, among the boundary markers, the image pattern corresponding to the first character of the entire set of characters can be called a leading marker, and among the boundary markers, the image pattern corresponding to the last character of the entire set of characters can be called a trailing marker.
[0284] Boundary markers function as markers to identify the coordinate region of the space occupied by each character belonging to the overall character set in the object image.
[0285] The boundary portion of a character can be used as a boundary marker, that is, one of the image patterns located at the boundary line or a boundary point belonging to the boundary line. The boundary portion of a character refers to a part of the image that constitutes a part of a character that extends from a sentence along the horizontal or vertical direction of that sentence. For example, In the sentence "JPEG0007851414000002.jpg630", "JPEG0007851414000003.jpg66" "JPEG0007851414000004.jpg45" and, "JPEG0007851414000005.jpg65" "JPEG0007851414000006.jpg63" can be described as a boundary marker, which is part of the character's boundary.
[0286] Furthermore, a portion (or the entirety) of a character can be used as a boundary marker. For example, in the sentence "ABCDEFG", A and G can be used as boundary markers, which are portions of characters located at the boundary point and / or boundary line.
[0287] The acquisition of a leading or trailing marker reflects the character arrangement scheme based on the language of the sentence that the entire set of characters constitutes. For example, in OCR for books written in Korean and English, as shown in Figure 11, the image sample (1132) in the upper left of the area where the characters are displayed can be the leading marker, and the image sample (1134) in the lower right of that area can be the trailing marker.
[0288] The boundary marker acquisition step (S3546a) may more specifically include a step (S3546a-1) in which the tracking controller (4400) controls available resources to acquire images of characters that include the leading marker and have a character recognition rate equal to or greater than a predetermined threshold.
[0289] Here, character recognition rate refers to the ratio in which a specific character is recognized as specific text in a video (image) of characters. For example, in a distant video (image) showing an object with characters on it, it is thought that the presence of some characters can be estimated using deep learning, etc., but it may be difficult to determine what those characters are due to reasons such as the size of the distant video (image) being small or the resolution being low. In other words, the character recognition rate by OCR may be low.
[0290] In other words, the video (image) of characters with the above-mentioned character recognition rate being equal to or higher than a predetermined threshold indicates that the resolution is sufficient for OCR. In the execution of this stage (S3546a-1), in order to obtain an enlarged video (image) of characters with a sufficiently high resolution, the system of this disclosure can control the zoom-in, zoom-out, and rotation of the imaging device (200) through the tracking controller (4400). For example, referring to FIG. 11a, in a composition including both the paper document (1110) and the person (1120), the rotation control (1160) and zoom-in control (1170, 1180) of the imaging device (200) are shown as examples to make the paper document (1110) come to the center of the video (image).
[0291] Next to stage (S3546a-1), this embodiment of the boundary marker acquisition stage (S3546a) includes an end marker determination stage (S3546a-2) for determining whether the end marker is included in the acquired video (image) of characters, and if the end marker is included in the acquired video (image) of characters, the video (image) of the characters is acquired as the entire character display video (image). If the end marker is not included in the acquired video (image) of characters, the tracking controller (4400) controls the available resources to search for the marker next to the last marker among the boundary markers shown in the video (image) of the characters, and acquire an additional video (image) of characters that includes the next marker and has a character recognition rate equal to or higher than a predetermined threshold. The additional video (image) of characters is combined with the video (image) of the characters to include a stage (S3546a-3) that re-executes from the end marker determination stage (S3546a-2).
[0292] The control of the available resources in stage (S3546a-3) includes control of the gimbal (300). For example, the control of the gimbal (300) for a book written in Korean and English can be set as control (for example, the rotation control of reference numeral 1190 in FIG. 11b) that assists the imaging device (200) to scan from left to right and from top to bottom in the area where the characters are displayed.
[0293] Furthermore, it is obvious to any ordinary engineer that various image stitching techniques can be used to combine the images of characters in step (S3546a-3). Referring to Figure 11b, it is possible to combine images using multiple overlapping image patterns (1136a, 1136b) from two or more images taken at temporal or spatial intervals.
[0294] Following the boundary marker acquisition step (S3546a) described above, the OCR execution step (S3540a) of this embodiment may include a step (S3468a) in which OCR is performed on the overall character display image, which is a partial image containing the image of the object on which the above characters are displayed, and which includes both the leading marker and the trailing marker.
[0295] On the other hand, according to Figures 10b and 11c, in the second embodiment of the OCR execution step (S3540), the OCR execution step (S3540b) begins with a step (S3542b) of extracting a video (image) sample from the object on which the characters are displayed, and following step (S3542b), includes a step (S3544b) of defining the boundary line (1130) of the area (area) on which the characters are displayed from the video (image) sample.
[0296] Next, the OCR execution stage (S3540b) includes the step (S3546b) of acquiring at least one video (image) pattern (1142a, 1142b) from the above video (image) samples that is not located at the above boundary line as a division marker, and performing OCR on the overall text display video (image), which is an image (image) generated by combining the video (image) of characters using the division marker.
[0297] In step (S3546b), a division marker refers to overlapping image patterns (1142a, 1142b) in two or more images taken at temporal or spatial intervals. By finding the same division marker in different images, it is possible to combine images using that division marker, that is, to stitch the images together.
[0298] Finally, referring to Figures 10c and 11d, in the third embodiment of the OCR execution step (S3540), the OCR execution step (S3540c) includes a step (S3542c) in which the tracking controller (4400) controls available resources so as to acquire the object image (video) that includes the beginning of the entire set of characters and whose character recognition rate is above a predetermined threshold, that is, so that the object image (video) has a predetermined resolution that allows OCR to be performed.
[0299] As a policy for the tracking controller (4400) that performs step (S3542c), it is possible to add not only the condition that the object video (image) has a predetermined resolution that allows OCR to be performed, but also the condition that the object video (image) should contain as many characters as possible at that resolution.
[0300] Next, the OCR execution stage (S3540c) further includes the step (S3544c) of performing OCR on the area (1152) of the first sentence in the acquired object image (picture).
[0301] Here, the area for the first sentence refers to the area occupied by the sentence that was detected as being placed first, considering the general arrangement of characters based on the language used in the sentence.
[0302] For example, when this step (S3544c) is performed for the first time, the area of the first sentence can be determined by scanning for delimiters that indicate the end of a sentence, such as ",", ".", and "?". Once such sentence symbols have been scanned, the area from the first character to its delimiter can be assumed to be the area of the first sentence (1152).
[0303] As another example, when this step (S3544c) is performed for the first time, the area of the first sentence (1152) can also be determined as the area occupied by a sentence composed of multiple characters following the first character, starting from the size of the first character and sequentially comparing the sizes of subsequent characters, up to the point where the size of the character changes to a predetermined threshold.
[0304] As yet another example, when this step (S3544c) is first performed, the area of the first sentence (1152) can also be determined as the area occupied by the sentence composed of multiple characters following the first character, starting from the size of the character interval between the first character and the next character, and then sequentially comparing the sizes of subsequent character intervals, until the size of that character interval changes to more than a predetermined threshold.
[0305] Next, the OCR execution stage (S3540c) further includes a semantic value calculation stage (S3546c) in which natural language understanding (NLU) is performed on the first text, which is the primary result of OCR in stage (S3544c), to calculate a semantic value number, which is a numerical value of the semantic value of the first text.
[0306] Next, the OCR execution stage (S3540c) further includes a stage (S3548c) in which, if the semantic value value in stage (S3546c) is equal to or greater than a predetermined threshold, the OCR is completed by acquiring the first text as the result text; if the semantic value value is less than the predetermined threshold, the OCR is performed on the area of the next sentence (1154) following the area of the first sentence, and the execution is restarted from the semantic value value calculation stage (S3546c).
[0307] While the calculation of semantic value can be done cumulatively, natural language understanding is not limited to applying only to a single sentence on which OCR has been performed. Rather, it can be applied again to a larger area of text by combining a new sentence with a sentence that has already undergone OCR.
[0308] For this purpose, step (S3548c) can sequentially and repeatedly perform OCR on new sentence areas, including the next sentence that exists after the delimiter in the area of the first sentence mentioned above.
[0309] Of course, even if OCR is performed on the last sentence, if the semantic value does not reach a predetermined threshold, it is possible to detect that there is no text consisting of all the characters mentioned above.
[0310] Thus, by performing the aforementioned steps described in this disclosure, a computing device that processes video (images) of an object can grasp a variety of information about the object detected from that video (image).
[0311] The multiple components shown in the drawings are, for the sake of clarity, examples of implementation in a single computing device, such as a portable terminal. However, it is obvious that the computing device (100) that performs the method of the present invention can also be configured in a way that multiple devices work together. Therefore, each step of the method of the present invention described above can be performed not only by a portable terminal but also by a gimbal (300) that has a built-in communication unit and processor. Furthermore, it is obvious that the process can be performed either directly by a single computing device or by the single computing device supporting the execution of other computing devices that work in conjunction with it.
[0312] As described above, the methods and apparatus of this disclosure, in all embodiments and variations thereof, are capable of recognizing and tracking one or more objects using video (images), actively acquiring information about objects and the environment, and in particular, being able to acquire information relating to the state of an object from a distance using video (images), identifying objects that an object interacts with using video (images), acquiring higher-resolution detailed video (images) of a part of an object from a distance, and being able to grasp characters printed on an object or output using other means such as a display from a distance, thus having the advantage of enabling remote input using video (images) in portable computing devices.
[0313] Based on the descriptions of various embodiments in this disclosure, a person ordinary in the art will clearly understand that the method and / or process of the present invention, and its various steps, can be implemented by hardware, software, or any combination of hardware and software adapted to a particular use. The hardware may include a general-purpose computer and / or a dedicated computing device, or a specific computing device or a special form or component of a specific computing device. The process can be implemented by one or more processors having internal and / or external memory, such as a microprocessor, a controller, such as a microcontroller, an embedded microcontroller, a microcomputer, an ALU (arithmetic logic unit), a digital signal processor, such as a programmable digital signal processor, or other programmable device. In addition, or as an alternative, the above process can be performed by application-specific integrated circuits (ASICs), programmable gate arrays, such as FPGAs (field programmable gate arrays), PLUs (programmable logic units), or Programmable Array Logic (PALs), or any other device capable of executing and responding to instructions, any other device that can be configured to process electronic signals, or a combination of multiple devices. The processing unit can execute an operating system (OS) and one or more software applications that run on the operating system. The processing unit can also access data, store data, manipulate data, process data, and generate data in response to software execution.For the sake of clarity, some parts of the explanation describe the use of a single processing unit, but a person with ordinary knowledge in the art will understand that a processing unit can include multiple processing elements and / or multiple types of processing elements. For example, a processing unit can include multiple processors or one processor and one controller. Other processing configurations, such as a parallel processor, are also possible.
[0314] Software can include computer programs, code, instructions, or a combination of one or more of these, and can configure a processing unit to perform a desired operation, or independently or collectively, it can instruct the processing unit. Software and / or data can be permanently or temporarily embodyed in certain types of machines, components, physical devices, virtual devices, computer storage media or devices, or transmitted signal waves for analysis by a processing unit or for providing instructions or data to a processing unit. Software can be distributed across computer systems connected via a network, and can be stored and executed in a distributed manner. Software and data can be stored on one or more machine-readable recording media.
[0315] Furthermore, the subject matter of the technical solution of the present invention or the portion contributing to the prior art can be embodied in the form of program instructions that can be executed by a variety of computer components and can be recorded on a machine-readable medium. The machine-readable medium can include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the machine-readable recording medium can be specifically designed and configured for the embodiment, or they can be used as known to a person of ordinary skill in the field of computer software. Examples of machine-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs, DVDs, and Blu-ray®; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include those that can be saved, compiled, or interpreted for execution on a machine capable of executing any other program instruction, not just one of the aforementioned devices, but a processor, a processor architecture, a heterogeneous combination of different hardware and software, or any other machine capable of executing such instructions. These instructions can be created using a structured programming language like C, an object-oriented programming language like C++, or a high-level or low-level programming language (assembly language, hardware technology language, and database programming language and technology). Therefore, this includes not only machine code and bytecode, but also high-level language code that can be executed by a computer using an interpreter, etc.
[0316] Accordingly, in one embodiment of the present invention, when the aforementioned methods and combinations thereof are executed by one or more computing devices, the methods and combinations thereof can be implemented as executable code that performs each step. In another embodiment, the above methods can be implemented as a system that performs the above steps, and the methods can be distributed in various ways across multiple devices, or all functions can be integrated into one dedicated, standalone device or other hardware. In yet another embodiment, the means for performing the above-described process and the associated steps can include any of the aforementioned hardware and / or software. All such sequential combinations and arrangements are intended to fall within the scope of this disclosure.
[0317] For example, the hardware device described above can be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa. The hardware device may include a processor such as an MPU, CPU, GPU, or TPU, which is coupled with memory such as ROM / RAM for storing program instructions and configured to execute the instructions stored in the memory, and may also include a communication unit that can exchange signals with external devices. The hardware device may also include a keyboard, mouse, or other external input device for receiving instructions created by the developer.
[0318] Although the present invention has been described above using specific components and other details, as well as limited embodiments and drawings, these are provided merely to facilitate a general understanding of the invention. The present invention is not limited to the above-described embodiments, and a person with ordinary skill in the art to which the present invention belongs can make various modifications and variations from this description.
[0319] Accordingly, the concept of the present invention is not limited to the embodiments described above, and the scope of the present invention includes not only the claims attached to this disclosure, but also all equivalent or equivalent modifications of these claims. For example, suitable results can be obtained even if the described techniques are performed in a different order than described, and / or if the components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or substituted or replaced by other components or equivalents. As stated above, equivalent or equivalent modifications include, for example, logically equivalent methods that yield the same results as when the method according to the present invention is performed. Therefore, the true meaning and scope of the present invention should not be limited by the above examples, but should be understood in the broadest sense permitted by law. [Modes for carrying out the invention]
[0320] As described above, the relevant information was stated based on the best mode for carrying out the invention. [Industrial applicability]
[0321] It can be used in devices, systems, etc., that process video (images) acquired from a camera.
Claims
1. A method for processing video (images) performed by a computing device including a processor, wherein the method is: The stage of acquiring video (images); A step of obtaining analysis information corresponding to the objects contained in the video (image) from the video (image) using an object analysis model; and A step of obtaining the characters contained in the object from the corresponding analysis information using an OCR model; Includes, The aforementioned OCR model is A step of determining whether or not characters are displayed on the surface of the object; The step of extracting at least one video (image) sample from the object on which the aforementioned characters are displayed; A step of determining the boundary of the area where text is displayed from the aforementioned at least one video (image) sample; From the at least one video (image) sample, an image (image) pattern located at at least one boundary line or boundary point belonging to the boundary line is acquired as a boundary marker; A step of generating a character display video (image) which is a partial video (image) included in the video (image) of the object on which the character is displayed, and which includes the boundary marker; and A step in which OCR is performed on the aforementioned text display video (image); Execute, The aforementioned boundary marker is A leading marker which is a video (image) pattern corresponding to the first character of the aforementioned characters; or A trailing marker is a video (image) pattern that corresponds to the last character of the aforementioned characters; at least one of the following: method.
2. In claim 1, The aforementioned video (image) sample is a video (image) pattern that exists at at least one position among the boundary line or the boundary point, and The aforementioned video (image) pattern includes at least one of the following: some characters, the boundaries of the characters, a part of the characters, or the background. method.
3. In Claim 1, The step of acquiring a video (image) pattern located at at least one of the boundary lines or boundary points from among the at least one video (image) sample as the boundary marker is: A step of using a tracking controller to acquire a video (image) of characters that includes the leading marker and whose character recognition rate is above a threshold; and A step of determining whether or not the aforementioned character image contains the aforementioned trailing marker; including, method.
4. In claim 3, If the image of the character includes the trailing marker, the step of determining the image of the character as the character display image; Further including, method.
5. In claim 3, If the image of the character does not include the trailing marker, the tracking controller is used to acquire an additional image of the character that includes the boundary marker following the last marker in the image of the character, and whose character recognition rate is equal to or greater than the threshold; The steps of combining the aforementioned additional character images with the character images to generate a combined character image; and If the combined character image includes the trailing marker, the step of determining the combined character image as the character display image; Further including, method.
6. A method for processing video (images) performed by a computing device including a processor, wherein the method is: The stage of acquiring video (images); A step of obtaining analysis information corresponding to the objects contained in the video (image) from the video (image) using an object analysis model; and A step of obtaining the characters contained in the object from the corresponding analysis information using an OCR model; Includes, The aforementioned OCR model is A step of determining whether or not characters are displayed on the surface of the object; The step of extracting at least one video (image) sample from the object on which the aforementioned characters are displayed; A step of determining the boundary of the area where text is displayed from the aforementioned at least one video (image) sample; From the at least one video (image) sample, an image (image) pattern located at at least one boundary line or boundary point belonging to the boundary line is acquired as a boundary marker; A step of generating a character display video (image) which is a partial video (image) included in the video (image) of the object on which the character is displayed, and which includes the boundary marker; and A step in which OCR is performed on the aforementioned text display video (image); Execute, From the stage of determining whether or not characters are displayed on the surface of the object, A step of using a tracking controller to acquire video (image) of the object that includes the beginning of the character and whose character recognition rate is above a threshold; The step of performing the above OCR on the area of the first sentence in the image of the object; and This step involves performing natural language understanding (NLU) on the first text, which is the primary result of the aforementioned OCR, and calculating a first semantic value number, which is a numerical value of semantic value; including, method.
7. In claim 6, If the first semantic value is equal to or greater than the threshold, the first text is determined to be the result text which is the result of the OCR; Further including, method.
8. In claim 6, If the first semantic value is less than the threshold, the OCR is performed on the area of the next sentence following the area of the first sentence; A step of performing natural language comprehension on the second text, which is the primary result of the aforementioned OCR, and calculating a second semantic value value, which is a numerical value of semantic value; and If the second semantic value is greater than or equal to the threshold, the second text is determined to be the result text which is the result of the OCR; Further including, method.
9. In claim 1, The computing device works in conjunction with the shooting device and the gimbal. The step of acquiring the aforementioned video (image) is: A step of using a tracking controller to operate one or more rotation axes of the gimbal and control the direction of the shooting device; and The step of acquiring the video (image) by zooming in or zooming out of the shooting device using the tracking controller; including, method.
10. A non-temporary computer-readable medium containing a computer program, wherein the computer program causes a computing device to execute a method for processing images, and the method is: The stage of acquiring video (images); A step of obtaining analysis information corresponding to the objects contained in the video (image) from the video (image) using an object analysis model; and A step of obtaining the characters contained in the object from the corresponding analysis information using an OCR model; Includes, The aforementioned OCR model is A step of determining whether or not characters are displayed on the surface of the object; The step of extracting at least one video (image) sample from the object on which the aforementioned characters are displayed; A step of determining the boundary of the area where text is displayed from the aforementioned at least one video (image) sample; From the at least one video (image) sample, an image (image) pattern located at at least one boundary line or boundary point belonging to the boundary line is acquired as a boundary marker; A step of generating a character display video (image) which is a partial video (image) included in the video (image) of the object on which the character is displayed, and which includes the boundary marker; and A step in which OCR is performed on the aforementioned text display video (image); Execute, The aforementioned boundary marker is A leading marker which is a video (image) pattern corresponding to the first character of the aforementioned characters; or A trailing marker is a video (image) pattern that corresponds to the last character of the aforementioned characters; at least one of the following: Non-temporary computer-readable media, including computer programs.
11. In computing devices, Processor; and Communications Department; Includes, The aforementioned processor, Acquire video (images), Using an object analysis model, analysis information corresponding to the objects contained in the video (image) is obtained from the video (image), and Using an OCR model, the characters contained in the object are obtained from the object and the corresponding analysis information. The aforementioned OCR model is Determine whether or not characters are displayed on the surface of the object. From the object on which the aforementioned characters are displayed, extract at least one video (image) sample. From the aforementioned at least one video (image) sample, determine the boundary of the area where the text is displayed. From the aforementioned at least one video (image) sample, a video (image) pattern located at at least one boundary line or boundary point belonging to the boundary line is acquired as a boundary marker. A character display video (image) is generated which is a partial video (image) included in the video (image) of the object on which the character is displayed, and which includes the boundary marker. Perform OCR on the aforementioned text display video (image), The aforementioned boundary marker is A leading marker which is a video (image) pattern corresponding to the first character of the aforementioned characters; or A trailing marker is a video (image) pattern that corresponds to the last character of the aforementioned characters; at least one of the following: Computing device.
Citation Information
Patent Citations
Document picture input device
JP1997161043A
Device, method and program for extracting document area from image
JP2010231686A