A method for stabilizing video using multiple data sources, a computer program, and a computing device.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WESTWORLD CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-08-03
AI Technical Summary
【0031】 本開示は、複数のデータを用いて映像を安定化させる方法に係り、より具体的には、精度は高いが途中でデータの欠損が発生し得る第1データ、及び、リアルタイムデータを提供して物体の動きを正確に把握するのに有用であるが測定値が時間の経過とともに実際の値から徐々に乖離する問題がある第2データについて、互いに欠損が存在する部分を相互に補完するように上記複数のデータを併用することにより、より正確且つ有効に映像を安定化できる。
Smart Images

Figure 2026125590000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for stabilizing video using a plurality of data. More specifically, the present disclosure relates to a method for stabilizing video by using a plurality of data in such a way that mutually complementary portions of missing data are complemented with each other, for first data that has high accuracy but may experience data loss during the process, and second data that is useful for accurately grasping the movement of an object by providing real-time data but has a problem that the measured value gradually deviates from the actual value over time.
Background Art
[0002] In conventional video and content production environments, the process of stabilizing camera video has played an extremely important role in digital content production. Specifically, various camera shooting videos used in video production environments have variations in individual distortion curvatures and position information, so there has been a problem that the quality of the captured video deteriorates due to inappropriate afterimages and shaking. In addition, data obtained from different sensors included in camera video to track the camera video have different time codes from each other, and there has been a limitation that it is difficult to use them together. For example, in the case of data obtained through an optical tracking system, 6DoF tracking information can be used for tracking by acquiring the rotation (roll, pitch, yaw) and position (x, y, z coordinates) of an object, but there has been a problem that it is impossible to acquire rapidly changing rotation and acceleration information.
[0003] Therefore, different from conventional camera tracking methods, there is an increasing need for a method that can more accurately and efficiently stabilize video by complementing missing main tracking data with sub-tracking data and complementing missing sub-tracking data with main tracking data.
[0004] On the other hand, while this disclosure is devised at least based on the aforementioned technical background, the technical issues or objectives of this disclosure are not limited to solving the aforementioned problems or issues. In other words, this disclosure can cover a variety of technical issues related to the content described below, in addition to the technical issues mentioned above. [Overview of the project] [Problems that the invention aims to solve]
[0005] This disclosure relates to a method for stabilizing video using multiple data sets, more specifically, to stabilizing video by using multiple data sets in combination so that each data set mutually complements the missing parts of the first data set, which has high accuracy but may experience data loss during processing, and the second data set, which is useful for accurately understanding the movement of objects by providing real-time data but has the problem that the measured values gradually deviate from the actual values over time.
[0006] Furthermore, according to this disclosure, while the loss of the first data may be due to environmental factors in addition to the reasons mentioned above, it is possible to compensate for this loss based on the second data which operates using a different technical method. Moreover, errors caused by the technology itself for acquiring the second data also occur in different parts than those in the first data. Thus, the parts of the data that are missing can be mutually compensated for.
[0007] However, the technical challenges that this disclosure aims to solve are not limited to those mentioned above, and as described below, a variety of technical challenges may be included within the scope that is obvious to an average engineer. [Means for solving the problem]
[0008] A method is disclosed that is performed by a computing device based on one embodiment of the present disclosure for achieving the aforementioned problems. The method may include the steps of: acquiring first data relating to position in an image; acquiring second data relating to the motion of the image; acquiring combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and acquiring a final image based on the combined data and the image.
[0009] Alternatively, the step of acquiring the first data relating to the position in the above video may include at least one of the following: the step of acquiring the first-1 data relating to the rotation of an object included in the above video; or the step of acquiring the first-2 data relating to the position of an object included in the above video.
[0010] Alternatively, the step of acquiring the second data relating to the motion of the above video may include at least one of the following: the step of acquiring the second-first data relating to the acceleration of the objects included in the above video; the step of acquiring the second-second data relating to the angular velocity of the objects included in the above video; or the step of acquiring the second-third data relating to the direction of the objects included in the above video.
[0011] Alternatively, the step of acquiring first data relating to position in the above video may further include a step of acquiring first time-series data relating to the above first data, and the step of acquiring second data relating to movement in the above video may further include a step of acquiring second time-series data relating to the above second data.
[0012] As an alternative, the step of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data may include the steps of: obtaining first corrected data based on the first data, the first time series data and the second time series data; obtaining second corrected data based on the second data, the first time series data and the second time series data; and obtaining combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data.
[0013] Alternatively, the step of obtaining first corrected data based on the first data, the first time series data, and the second time series data may include the step of obtaining combined time series data by matching the time series of the first time series data and the second time series data; and the step of obtaining second corrected data based on the first data and the combined time series data may include the step of obtaining second corrected data based on the second data, the first time series data, and the second time series data.
[0014] As an alternative, the step of obtaining combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data may include at least one of the following steps: obtaining combined data by supplementing errors accumulated over time in the second corrected data based on the first corrected data; or obtaining combined data by supplementing missing data in the first corrected data based on the second corrected data.
[0015] Alternatively, the step of obtaining combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data may further include a step of obtaining filtered combined data by filtering the combined data.
[0016] Alternatively, the step of obtaining filtered combined data by filtering the combined data may include at least one of the following: a step of obtaining filtered combined data by filtering the combined data based on a Kalman filter; or a step of obtaining filtered combined data by filtering the combined data based on a complementary filter.
[0017] Alternatively, the step of acquiring the final image based on the combined data and the video may include a step of correcting for shaking of objects included in the video based on the combined data and then acquiring the final image.
[0018] Alternatively, the step of acquiring the final image based on the combined data and the video may include a step of correcting the angle in a portion of the video based on the combined data to acquire the final image.
[0019] A computer program stored on a computer-readable storage medium is disclosed based on one embodiment of the present disclosure for achieving the aforementioned problems. When the computer program is executed on one or more processors, it causes the one or more processors to perform an operation to improve the accuracy of the image, and the operation may include: an operation to acquire first data relating to the position in the image; an operation to acquire second data relating to the movement of the image; an operation to acquire combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and an operation to acquire a final image based on the combined data and the image.
[0020] As an alternative, the operation to acquire the first data relating to the position in the above video may include at least one of the following: the operation to acquire the first-1 data relating to the rotation of an object included in the above video; or the operation to acquire the first-2 data relating to the position of an object included in the above video.
[0021] As an alternative, the operation to acquire the second data relating to the movement of the above video may include at least one of the following: an operation to acquire the second-first data relating to the acceleration of an object included in the above video; an operation to acquire the second-second data relating to the angular velocity of an object included in the above video; or an operation to acquire the second-third data relating to the direction of an object included in the above video.
[0022] Alternatively, the operation to acquire first data relating to position in the above video may further include the operation to acquire first time-series data relating to the above first data, and the operation to acquire second data relating to movement in the above video may further include the operation to acquire second time-series data relating to the above second data.
[0023] As an alternative, the operation of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data may include: obtaining first corrected data based on the first data, first time series data and second time series data; obtaining second corrected data based on the second data, first time series data and second time series data; and obtaining combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data.
[0024] Alternatively, the operation to obtain the first corrected data based on the first data, the first time series data, and the second time series data may include the operation to obtain combined time series data by matching the time series of the first time series data and the second time series data; and the operation to obtain the first corrected data based on the first data and the combined time series data may include the operation to obtain the second corrected data based on the second data, the first time series data, and the second time series data.
[0025] As an alternative, the operation of obtaining combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data may include at least one of the following: an operation of obtaining combined data by supplementing the errors accumulated over time in the second corrected data based on the first corrected data; or an operation of obtaining combined data by supplementing missing data in the first corrected data based on the second corrected data.
[0026] As an alternative, the operation of obtaining combined data by complementing the second corrected data based on the first corrected data and complementing the first corrected data based on the second corrected data may further include an operation of obtaining filtered combined data by filtering the combined data.
[0027] As an alternative, the operation of obtaining filtered combined data by filtering the combined data may include at least one of: an operation of obtaining filtered combined data by filtering the combined data based on a Kalman filter; or an operation of obtaining filtered combined data by filtering the combined data based on a complementary filter.
[0028] As an alternative, the operation of obtaining a final video based on the combined data and the video may include an operation of obtaining the final video by correcting the shake of an object included in the video based on the combined data.
[0029] As an alternative, the operation of obtaining a final video based on the combined data and the video may include an operation of obtaining the final video by correcting an angle in a partial area of the video based on the combined data.
[0030] A computing device based on an embodiment of the present disclosure for solving the above problems is disclosed. The device includes at least one processor; and a memory, and the at least one processor is configured to obtain first data related to a position in a video; obtain second data related to the movement of the video; obtain combined data by complementing the second data based on the first data and complementing the first data based on the second data; and obtain a final video based on the combined data and the video.
Advantages of the Invention
[0031] This disclosure relates to a method for stabilizing video using multiple data sets, more specifically, to a first data set that is highly accurate but may experience data loss during processing, and a second data set that is useful for accurately understanding the movement of objects by providing real-time data but has the problem that the measured values gradually deviate from the actual values over time. By using these multiple data sets in combination so that they mutually complement each other's missing parts, video can be stabilized more accurately and effectively.
[0032] On the other hand, the effects of this disclosure are not limited to those described above, and may include a variety of effects within the scope that is obvious to an ordinary engineer, as described below. [Brief explanation of the drawing]
[0033] [Figure 1] Figure 1 is a block diagram of a computing device for stabilizing video using multiple data, based on one embodiment of the present disclosure. [Figure 2] Figure 2 is a schematic diagram showing a network function based on one embodiment of the present disclosure. [Figure 3] Figure 3 is a flowchart illustrating a method for stabilizing video using multiple data points, based on one embodiment of the present disclosure. [Figure 4a] Figure 4a is a schematic diagram illustrating the process of acquiring first data relating to position in a video, based on one embodiment of the present disclosure. [Figure 4b] Figure 4b is a schematic diagram illustrating the process of acquiring second data related to the motion of a video, based on one embodiment of the present disclosure. [Figure 5] Figure 5 is a schematic diagram illustrating a process for obtaining combined data by supplementing first data with second data and supplementing second data with first data, based on one embodiment of the present disclosure. [Figure 6]Figure 6 is a schematic diagram illustrating the process of acquiring a final image based on combined data and video, according to one embodiment of the present disclosure. [Figure 7] Figure 7 is a schematic diagram illustrating the process of obtaining the final image by correcting the angle in a portion of the image based on combined data, according to one embodiment of the present disclosure. [Figure 8] Figure 8 is a simplified and general schematic diagram of an exemplary computing environment that can embody the embodiments of this disclosure. [Modes for carrying out the invention]
[0034] Various embodiments are described below with reference to the drawings, where similar drawing numbers are used to represent similar components. Various descriptions are provided herein to facilitate understanding of this disclosure. However, these embodiments can certainly be carried out without these specific descriptions.
[0035] In this specification, terms such as “component,” “module,” and “system” refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or software execution. For example, a component may be, but is not limited to, a processing procedure executed on a processor, a processor, an object, an execution thread, a program, and / or a computer. For example, both an application running on a computing device and the computing device itself can be components. One or more components may reside in a processor and / or an execution thread, and one component may be localized within one computer or distributed across two or more computers. Such components may also be executed from a variety of computer-readable media, each containing a variety of data structures. Components may communicate locally and / or remotely, for example, by signals containing one or more data packets (e.g., data transmitted over a network such as the Internet, through data and / or signals from one component interacting with other components in a local system or distributed system).
[0036] The term "or" is used with the intention of meaning an implicational "or," not an exclusive "or." That is, unless specifically specified and contextually clear, "X uses A or B" means one of the natural implicational substitutions. In other words, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" can be any of these. Furthermore, the terms "and / or" in this specification should be understood to refer to and include all possible combinations of one or more of the related items discussed.
[0037] Furthermore, the term “includes” as a predicate and / or as a modifier should be understood to mean that the feature and / or component in question exists. However, the term “includes” as a predicate and / or as a modifier should be understood not to exclude the existence or addition of one or more other further features, components and / or groups thereof. Also, where the number is not specifically identified or where it is not clear from the context to indicate a singular form, “singular” in this specification and claims should generally be interpreted to mean “one or more.”
[0038] Furthermore, the phrase "at least one of A or B" should be interpreted as meaning "including only A," "including only B," or "a combination of A and B."
[0039] Those skilled in the art should further recognize that the various exemplary logical blocks, configurations, modules, circuits, means, logic, and algorithmic stages described herein as relating to the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interoperability between hardware and software, various exemplary components, blocks, configurations, means, logic, modules, circuits, and stages have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and design constraints of the overall system. Skilled technicians can implement the described functionality in various ways for individual specific applications; however, decisions regarding such implementation should not be construed as departing from the scope of this disclosure.
[0040] The descriptions relating to the embodiments shown herein are provided so that a person with ordinary skill in the art of this disclosure may utilize or practice the invention. Various modifications to such embodiments are obvious to a person with ordinary skill in the art of this disclosure, and the general principles defined herein can be applied to other embodiments without departing from the scope of this disclosure. Therefore, this disclosure is not limited by the embodiments shown herein, but should be interpreted in the broadest sense consistent with the principles and novel features shown herein.
[0041] This disclosure allows for the interchangeable use of network functions, artificial neural networks, and neural networks.
[0042] Figure 1 is a block diagram of a computing device for stabilizing video using multiple data, based on one embodiment of the present disclosure.
[0043] The configuration of the computing device (100) shown in Figure 1 is merely a simplified example. In one embodiment of this disclosure, the computer device (100) may include other configurations for implementing the computing environment of the computer device (100), and it is also possible to configure the computer device (100) using only some of the disclosed configurations.
[0044] The computer device (100) may include a processor (110), memory (130), and a network unit (150).
[0045] In one embodiment of the present disclosure, the processor (100) may consist of one or more cores and may include processors for data analysis and deep learning, such as a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU). The processor (110) can read computer programs stored in memory (130) and execute data processing for machine learning in one embodiment of the present disclosure. Based on one embodiment of the present disclosure, the processor (110) can perform calculations for training a neural network. In deep learning (DL), the processor (110) can perform calculations for training a neural network, such as processing input data for training, extracting features from input data, calculating errors, and updating the weights of the neural network using backpropagation.
[0046] At least one of the CPU, GPGPU, and TPU of the processor (110) can process network function training. For example, both the CPU and GPGPU can perform network function training and data classification using network functions. In one embodiment of this disclosure, the processors of multiple computing devices can be used together to perform network function training and data classification using network functions. Furthermore, in one embodiment of this disclosure, the computer program executed in the computing device can be a program that can be executed on the CPU, GPGPU, or TPU.
[0047] According to one embodiment of the present disclosure, the memory (130) can store any form of information generated or determined by the processor (110) and any form of information received by the network unit (150).
[0048] In one embodiment of the present disclosure, the memory (130) may include at least one type of storage medium from among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, magnetic disk, and optical disk. The computing device (100) may also operate in conjunction with web storage that performs the storage function of the memory (130) over the internet. The above descriptions of memory are illustrative and the present disclosure is not limited thereto.
[0049] In one embodiment of the present disclosure, the network unit (150) can use a variety of wired communication systems such as public switched telephone networks (PSTN), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and local area networks (LANs). Furthermore, the network unit (150) in this specification can utilize a variety of wireless communication systems such as CDMA (Code Division Multi Access), TDMA (Time Division Multi Access), FDMA (Frequency Division Multi Access), OFDMA (Orthogonal Frequency Division Multi Access), SC-FDMA (Single Carrier-FDMA), and other systems. In this disclosure, the network unit (150) can be configured regardless of the communication mode, such as wired or wireless, and can be composed of various communication networks such as Personal Area Networks (PANs) and Wide Area Networks (WANs). The network may also be the well-known World Wide Web (WWW), and wireless transmission technologies used for short-range communication, such as Infrared Data Association (IrDA) and Bluetooth (registered trademark), can be utilized. The technologies described in this disclosure can also be used in the other networks mentioned above.
[0050] Figure 2 is a schematic diagram showing a network function according to one embodiment of the present disclosure.
[0051] Throughout this specification, the terms artificial intelligence model, artificial intelligence-based model, computational model, neural network, network function, and neural network may be used interchangeably.
[0052] A neural network can consist of a set of interconnected computational units, which can generally be called nodes. These nodes are sometimes also called neurons. A neural network consists of at least one node. The nodes (or neurons) that make up a neural network can be interconnected by one or more links.
[0053] Within a neural network, one or more nodes connected via links can form a relative input-output node relationship. The concepts of input and output nodes are relative; any node that is an output node to another node can be an input node to another node, and vice versa. As mentioned above, the input-output node relationship can be generated around links. One input node can be connected to one or more output nodes via links, and vice versa.
[0054] In a relationship between input and output nodes connected via a single link, the data of the output node can be determined based on the data input to the input node. Here, the link interconnecting the input and output nodes may have weights. The weights can be variable and can be changed by the user or algorithm to enable the neural network to perform desired functions. For example, if one or more input nodes are interconnected to a single output node by their respective links, the output node's value can be determined based on the values input to the input nodes connected to the output node and the weights set for the links corresponding to each input node.
[0055] As mentioned earlier, a neural network consists of one or more nodes interconnected via one or more links, forming input-output node relationships within the neural network. The characteristics of a neural network can be determined by the number of nodes and links within the neural network, the relationships between nodes and links, and the weight values assigned to each link. For example, if there are two neural networks with the same number of nodes and links but different link weight values, the two neural networks may be perceived as different from each other.
[0056] A neural network can consist of a set of one or more nodes. A subset of the nodes that make up a neural network can form a layer. Some of the nodes that make up a neural network can form a layer based on their distance from the first input node. For example, a set of nodes that are n in distance from the first input node can form an n-layer. The distance from the first input node can be defined by the minimum number of links that must be traversed to reach that node from the first input node. However, such a definition of a layer is arbitrary for illustrative purposes, and the difference in the number of layers within a neural network can be defined in a different way than described above. For example, a node layer can also be defined by its distance from the final output node.
[0057] In one embodiment of this disclosure, a collection of neurons or nodes can be defined as a layer.
[0058] The initial input node can refer to one or more nodes in a neural network that receive data directly without going through links in relation to other nodes. Alternatively, it can refer to a node in a neural network that does not have other input nodes connected by links in relation to other nodes based on links. Similarly, the final output node can refer to one or more nodes in a neural network that do not have other output nodes in relation to other nodes. Furthermore, hidden nodes can refer to nodes in a neural network that are neither the initial input node nor the final output node.
[0059] A neural network according to one embodiment of this disclosure may have the same number of nodes in the input layer as the number of nodes in the output layer, and the number of nodes may decrease as the network progresses from the input layer to the hidden layer, and then increase again. Another neural network according to another embodiment of this disclosure may have fewer nodes in the input layer than the number of nodes in the output layer, and the number of nodes may decrease as the network progresses from the input layer to the hidden layer. Yet another neural network according to yet another embodiment of this disclosure may have more nodes in the input layer than the number of nodes in the output layer, and the number of nodes may increase as the network progresses from the input layer to the hidden layer. Yet another neural network according to yet another embodiment of this disclosure may be a neural network that combines the aforementioned neural networks.
[0060] A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to the input and output layers. Using deep neural networks, it is possible to grasp the latent structures of data. That is, the latent structures of photographs, text, videos, audio, protein sequence structures, gene sequence structures, peptide sequence structures, and music (e.g., what objects are in a photograph, what is the content and emotion of the text, what is the content and emotion of the audio, etc.), and / or the degree of binding affinity between peptides and MHCs. Deep neural networks can include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siam networks, generative adversarial networks (GANs), and transformers. The aforementioned description of deep neural networks is merely illustrative, and this disclosure is not limited thereto.
[0061] The artificial intelligence-based model described herein can be represented by a network structure of any of the aforementioned structures, including an input layer, a hidden layer, and an output layer.
[0062] The neural networks that can be used in the artificial intelligence-based models described herein can be trained using at least one of the following methods: supervised learning, unsupervised learning, semi-supervised learning, transfer learning, active learning, or reinforcement learning. Learning a neural network can be the process by which the neural network applies knowledge to perform a specific action.
[0063] Neural networks can be trained to minimize the error in their output. Training a neural network involves repeatedly inputting training data, calculating the network's output and target error for the training data, and updating the weights of each node in the neural network by backpropagating the error from the output layer to the input layer to reduce the error. In supervised learning, training data with the correct answer labeled is used (i.e., labeled training data), while in unsupervised learning, the training data may not have the correct answer labeled. For example, in supervised learning for data classification, the training data may be data with each training data point labeled with a category. The error can be calculated by inputting the labeled training data into the neural network and comparing the neural network's output (category) with the labels on the training data.
[0064] As another example, in unsupervised learning for data classification, the error can be calculated by comparing the input training data with the output of the neural network. The calculated error is backpropagated in the neural network (i.e., from the output layer to the input layer), and the connection weights of each node in each layer of the neural network may be updated according to the backpropagation. The amount of change in the connection weights of each node to be updated can be determined by the learning rate. The calculation of the neural network on the input data and the backpropagation of the error can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of iterations of the neural network's learning cycle. For example, a high learning rate can be used in the early stages of learning to increase efficiency by allowing the neural network to quickly achieve a certain level of performance, while a low learning rate can be used in the later stages of learning to improve accuracy.
[0065] In neural network training, training data can generally be a subset of real-world data (i.e., data to be processed using the trained neural network). Therefore, there can be training cycles where errors on the training data decrease, but errors on the real-world data increase. Overfitting is a phenomenon where the training data is excessively trained, leading to increased errors on the real-world data. For example, a neural network trained to recognize a yellow cat as a cat may fail to recognize a cat of a different color as a cat; this is a type of overfitting. Overfitting can act as a cause of increased errors in machine learning algorithms. Various optimization methods can be used to prevent such overfitting. To prevent overfitting, methods such as increasing the amount of training data, regularization, dropout (deactivating some of the network nodes during the training process), and the use of a batch normalization layer can be applied.
[0066] Figure 3 is a flowchart showing a method for stabilizing video using multiple data, based on one embodiment of the present disclosure.
[0067] A computing device (100) according to one embodiment of the present disclosure can directly acquire or receive "information for stabilizing video using multiple data" from an external system. The external system can be a server, database, etc., that stores and manages information for stabilizing video using multiple data. The computing device (100) can use the information acquired directly or received from the external system as "input data for stabilizing video using multiple data".
[0068] According to one embodiment of the present disclosure, the computing device (100) can acquire first data relating to position in the video (S110). For example, the computing device (100) can acquire first-1 data relating to the rotation of an object included in the video, or first-2 data relating to the position of an object included in the video. For example, the computing device (100) can acquire the first-1 data consisting of the roll, pitch, and yaw angles rotating around each axis included in the video, and can acquire the first-2 data indicating movement (position) in the x, y, and z axis directions. On the other hand, the computing device (100) can acquire first time-series data relating to the first data. In this regard, the first data relating to position in the video may be collected at different frequencies from the second data relating to the motion of the video, which will be described later. In such cases, the time series differs between the data, and the first data cannot be utilized in the mutual complementation process. Therefore, the computing device (100) can acquire first time-series data relating to the first data in order to accurately match the first data and the second data described later on the time axis. On the other hand, the first data relating to the position in the video can be used in the process by which the computing device (100) complements the second data described later, and a detailed explanation therein will be given later with reference to Figure 5.
[0069] According to one embodiment of the present disclosure, the computing device (100) can acquire second data relating to the motion of the image in step S110 (S120). Specifically, the computing device (100) can acquire second-1 data relating to the acceleration of an object included in the image, second-2 data relating to the angular velocity of an object included in the image, and second-3 data relating to the direction of an object included in the image. On the other hand, the computing device (100) can acquire second time-series data relating to the second data. In this regard, the second data relating to the motion of the image may be collected at different frequencies from the first data relating to the position of the image. In such cases, the time series differs between the first and second data, and the second data cannot be utilized in the complementary process. Therefore, the computing device (100) can acquire second time-series data relating to the second data in order to accurately match the second data and the first data on the time axis. On the other hand, the second data relating to the movement of the above-mentioned video can be used in the process by which the computing device (100) complements the first data relating to the position in the above-mentioned video. A detailed explanation of this will be given later with reference to Figure 5.
[0070] According to one embodiment of the present disclosure, the computing device (100) can acquire combined data by supplementing the second data acquired through step S120 based on the first data acquired through step S110, and supplementing the first data based on the second data (S130). At this time, the computing device (100) can acquire first corrected data based on the first data, the first time series data, and the second time series data, and can acquire second corrected data based on the second data, the first time series data, and the second time series data. In this regard, if the time series of the first data and the second data are different, it is difficult for the computing device (100) to perform the steps of supplementing the second data based on the first data or supplementing the first data based on the second data. Therefore, the computing device (100) can acquire combined time series data by matching the time series of the first time series data and the second time series data, and acquire first corrected data based on the first data and the combined time series data. Similarly, the computing device (100) can acquire the second corrected data based on the second data and the combined data, and the first corrected data and the second corrected data can represent data with the same frequency.
[0071] Subsequently, the computing device (100) can obtain combined data by supplementing the second corrected data based on the first corrected data and supplementing the first corrected data based on the second corrected data. For example, the computing device (100) can obtain combined data by supplementing the errors accumulated over time in the second corrected data based on the first corrected data, and can obtain combined data by supplementing missing data in the first corrected data based on the second corrected data. On the other hand, in order to improve the accuracy of the combined data, the computing device (100) can obtain filtered combined data by filtering the combined data. For example, the computing device (100) can obtain filtered combined data by filtering the combined data based on a Kalman filter, or obtain filtered combined data by filtering the combined data based on a complementary filter. However, the embodiments described above in the filtering process of this disclosure are not limited to the filters described above, and various examples can be used. On the other hand, in the process of acquiring the final image based on the combined data and video images acquired above, a computing device (100) can be used, and a detailed explanation thereof will be given later with reference to Figure 6.
[0072] According to one embodiment of the present disclosure, the computing device (100) can acquire a final image based on the combined data acquired through step S130 and the image from step S110 (S140). For example, the computing device (100) can acquire the final image by correcting the shaking of objects included in the image based on the combined data. As another example, the computing device (100) can acquire the final image by correcting the angle in a part of the image based on the combined data. Specifically, the computing device (100) can acquire the final image by correcting the angle in the inner frustum region of the image captured by the camera without changing the entire image. In this regard, the computing device (100) can acquire information relating to the absolute position of the image through the first data and information relating to the relative movement of the image through the second data. Therefore, by acquiring the final image based on the combined data in which the first and second data complement each other, the computing device (100) can acquire a stable image by utilizing the advantages of both the first and second data, even in environments where real-time performance is required or where sensor data is limited. On the other hand, a detailed explanation of the process by which the computing device (100) acquires the final image based on the acquired combined data and the image will be given later with reference to Figures 6 to 7.
[0073] Figure 4a is a schematic diagram illustrating the process of acquiring first data relating to position in a video, based on one embodiment of the present disclosure, and Figure 4b is a schematic diagram illustrating the process of acquiring second data relating to motion in a video, based on one embodiment of the present disclosure.
[0074] First, referring to Figure 4a, the computing device (100) can acquire first data (10) relating to the position in the image. In this case, the image can mean the result captured by the camera, and the image can contain at least one object. Specifically, the computing device (100) can acquire first-1 data (10-1) relating to the rotation of an object included in the image, or first-2 data (10-2) relating to the position of an object included in the image. For example, the computing device (100) can acquire first-1 data (10-1) consisting of the roll, pitch, and yaw angles rotating around each axis included in the image, and can acquire first-2 data (10-2) indicating movement (position) in the x, y, and z axis directions. In this case, the roll can mean the amount of rotation around the x axis, the pitch can mean the amount of rotation around the y axis, and the yaw angle can mean the amount of rotation around the z axis. Furthermore, the computing device (100) can acquire first time-series data relating to the first data (10). In this regard, the first data (10) relating to the position in the video may be collected at different frequencies from the second data relating to the motion of the video, which will be described later. In such cases, the time series differs between the data, and the first data (10) cannot be used in the mutual complementation process. Therefore, the computing device (100) can acquire first time-series data relating to the first data (10) in order to accurately match the first data (10) and the second data described later on the time axis. On the other hand, the first data (10) relating to the position in the video can be used in the process by which the computing device (100) complements the second data described later, and a detailed explanation therein will be given later with reference to Figure 5.
[0075] Next, referring to Figure 4b, the computing device (100) can acquire second data (20) related to the motion of the image. Specifically, the computing device (100) can acquire second-first data (20-1) related to the acceleration of the object included in the image, second-second data (20-2) related to the angular velocity of the object included in the image, and second-third data (20-3) related to the direction of the object included in the image. For example, the computing device (100) can acquire second-first data (20-1) related to the acceleration of the object included in the image by estimating movement (change in position) based on gravity and acceleration through an accelerometer, and can acquire second-second data (20-2) related to the angular velocity of the object included in the image by measuring the angular velocity (changes in roll, pitch, and yaw) through a gyroscope. Alternatively, the computing device (100) can acquire second-third data (20-3) by measuring the motion (acceleration, rotation) of the object included in the image using a geomagnetic sensor. However, the second data (20) is not limited to the example shown in Figure 4b above, and may include a variety of examples related to the movement of the image. In this regard, since the second data (20) includes information related to the relative movement of the image, the computing device (100) can obtain a more stable image by using the second data (20) in the process of correcting image distortion caused by camera movement or shaking. On the other hand, the computing device (100) can acquire second time-series data related to the second data (20). In this regard, the second data (20) related to the movement of the image may be collected at different frequencies from the first data (10) related to the position of the image. In such cases, the time series differs between the first data (10) and the second data (20), and the second data (20) cannot be used in the mutual complementation process.Therefore, the computing device (100) can acquire second time-series data relating to the second data (20) in order to precisely match the second data (20) and the first data (10) on the time axis. On the other hand, the second data (20) relating to the movement of the video can be used in the process by which the computing device (100) complements the first data (10) relating to the position in the video, and a detailed explanation of this will be given later with reference to Figure 5.
[0076] Figure 5 is a schematic diagram illustrating a process for obtaining combined data by supplementing first data with second data and supplementing second data with first data, based on one embodiment of the present disclosure.
[0077] Referring to Figure 5, the computing device (100) can acquire combined data (30) by supplementing the acquired second data (20) based on the acquired first data (10), and supplementing the first data (10) based on the second data (20). At this time, the computing device (100) can acquire first corrected data (10') based on the first data (10), the first time series data, and the second time series data. Furthermore, the computing device (100) can acquire second corrected data (20') based on the second data (20), the first time series data, and the second time series data. In this regard, if the time series of the first data (10) and the second data (20) are different, it is difficult for the computing device (100) to perform the steps of supplementing the second data (20) based on the first data (10) or supplementing the first data (10) based on the second data (20). Therefore, the computing device (100) can synchronize the time series of the first time series data and the second time series data to obtain combined time series data, and obtain the first corrected data (10') based on the first data (10) and the combined time series data. Similarly, the computing device (100) can obtain the second corrected data (20') based on the second data (20) and the combined time series data, and the first corrected data (10') and the second corrected data (20') can mean data with the same frequency.
[0078] Subsequently, the computing device (100) can obtain the combined data (30) through a mutual complementation process in which it complements the second corrected data (20') based on the first corrected data (10') and complements the first corrected data (10') based on the second corrected data (20'). For example, the computing device (100) can obtain the combined data (30) by complementing the errors accumulated over time in the second corrected data (20') based on the first corrected data (10'), and can obtain the combined data (30) by complementing the missing data contained in the first corrected data (10') based on the second corrected data (20'). In this regard, the first corrected data (10') is accurate in determining the position of objects and the position of images relative to the second corrected data (20'), but does not contain information about relative movement, so data loss may occur along the way, and it may have the disadvantage of being sensitive to changes in lighting. Furthermore, the second corrected data (20') can be updated in real time at a higher frequency than the first corrected data (10'), allowing it to provide data at a much faster speed than a camera. It also has the advantage of correcting for shaking that occurs when the object being photographed moves, allowing for clearer tracking of the object. However, a drift problem may occur where the values gradually deviate from the actual values over time. Therefore, the computing device (100) can compensate for any data loss or instability that may occur in the first corrected data (10') with the second corrected data (20'), which can be updated in real time, and can compensate for any drift problems in the second corrected data (20') by using the first corrected data (10'). As a result, the combined data (30) obtained in this way is a result in which the shortcomings of the two sets of data (10' and 20') are mutually compensated for, and the computing device (100) can obtain more accurate results in the video stabilization process or camera tracking process by using the combined data (30).
[0079] On the other hand, the computing device (100) can obtain filtered combined data by filtering the combined data (30) in order to improve the accuracy of the combined data (30). For example, the computing device (100) can obtain filtered combined data by filtering the combined data (30) based on a Kalman filter, or by filtering the combined data (30) based on a complementary filter. In this regard, if the computing device (100) utilizes a Kalman filter in the filtering process for the combined data (30), it can calculate a more accurate estimate based on the reliability of the two data (10 and 20) sources. Alternatively, if the computing device (100) utilizes a complementary filter in the filtering process for the combined data (30), the gyroscope has high accuracy over short periods, and the accelerometer has high accuracy over long periods, so using the data from the gyroscope and accelerometer complementaryly has the advantage of being easy to implement and requiring less computation. However, the embodiments described above in the filtering process of this disclosure are not limited to the filters described above, and a variety of other examples may be used. On the other hand, a computing device (100) can be used in the process of acquiring the final image based on the acquired combined data (30) and the image, and a detailed explanation thereof will be given later with reference to Figure 6.
[0080] Figure 6 is a schematic diagram illustrating the process of acquiring a final image based on combined data and video, according to one embodiment of the present disclosure.
[0081] Referring to Figure 6, the computing device (100) can acquire the final image (11') based on the acquired combined data (30) and the acquired image (11). At this time, the image (11) can include images acquired through a camera during the content shooting process, or images previously stored in a database. For example, the computing device (100) can acquire the final image (11') by correcting the shaking of objects included in the image (11) based on the combined data (30). In this regard, the first data (10) used in the process of acquiring the combined data (30) can include information relating to the absolute position of the image (11), and the second data (20) can include information relating to the relative movement of the image (11). Therefore, the computing device (100) uses the combined data (30), which is obtained by mutually complementing the first data (10) and the second data (20), to stabilize the video (11), and through this, acquires the final video (11'). This allows the device to acquire a stabilized video by utilizing the advantages of both the first data (10) and the second data (20), which provide different types of information, even in environments where real-time performance is required or where sensor data is limited, thereby significantly improving the accuracy and stability of camera tracking.
[0082] Figure 7 is a schematic diagram illustrating the process of obtaining the final image by correcting the viewpoint of a portion of the image based on combined data, according to one embodiment of the present disclosure.
[0083] Referring to Figure 7, the computing device (100) can acquire the final image by correcting (32-2) the angle in a portion of the image (11) (32-1) based on the acquired combined data (30). Specifically, in a virtual production environment, the computing device (100) can acquire the final image by correcting the angle in the inner frustum region (32-1 and 32-2) of the image captured by the camera, without rendering the entire image output to the background, and by performing rendering only on the inner frustum correction region (32-2) of the entire image (11) through the image rendering machine (31). In this regard, the first data (10) used in the process of acquiring the combined data (30) can include information relating to the absolute position of the video (11), and the second data (20) can include information relating to the relative movement of the video (11). Therefore, the computing device (100) uses the combined data (30), which is obtained by mutually complementing the first data (10) and the second data (20), to stabilize the internal frustum region (32-1 to 32-2), which is the camera's shooting area, within the video (11), and performs rendering only on that region to acquire the final video (11'). This allows for the rapid acquisition of stabilized captured video by utilizing the advantages of both the first data (10) and the second data (20), which provide different types of information, even in environments where real-time performance is required or where sensor data is limited, thereby significantly improving the accuracy and stability of camera tracking.
[0084] Disclosed is a computer-readable medium storing a data structure according to one embodiment of the present disclosure. A data structure may mean an organization, management, and storage of data that enables efficient access to and modification of the data. A data structure may mean an organization of data for solving a particular problem (e.g., data analysis, data retrieval, data storage, data modification). A data structure can also be defined as physical or logical relationships between data elements designed to support a particular data processing function. Logical relationships between data elements may include linking relationships between user-defined data elements. Physical relationships between data elements may include actual relationships between data elements that are physically stored in a computer-readable medium (e.g., permanent storage). Specifically, a data structure may include a collection of data, relationships between data, and functions or instructions that can be applied to data. A well-designed data structure allows a computing device to perform operations with minimal use of the computing device's resources. Specifically, a computing device can improve the efficiency of operations, reads, inserts, deletes, compares, exchanges, and retrieves through a well-designed data structure.
[0085] Data structures can be classified into linear and nonlinear data structures depending on their form. A linear data structure may be one in which only one data item is linked after another. Linear data structures can include lists, stacks, queues, and deques. A list can represent a set of data items that have an internal order. A list can include linked lists. A linked list may be a data structure in which data is linked in a linear fashion, with each item having a pointer. Pointers in a linked list may contain linking information to preceding and succeeding data. A linked list can be represented as a single linked list, a double linked list, or a circular linked list, depending on its form. A stack may be a data list structure in which data can be accessed in a restricted manner. A stack may be a linear data structure in which data can only be processed (e.g., inserted or deleted) from one end of the data structure. Data stored in a stack may be a last-in, first-out (LIFO) data structure. A queue is a data structure that allows restricted access to data, and unlike a stack, it can be a first-in, first-out (FIFO) data structure. A deck can be a data structure that allows data to be processed from both ends of the data structure.
[0086] A nonlinear data structure can be a structure in which multiple data are concatenated after a single data. A nonlinear data structure can include a graph data structure. A graph data structure can be defined as a vertex and an edge, where an edge can contain a line connecting two distinct vertices, and can include a graph data structure tree. A tree data structure can be a data structure in which there is only one path connecting two distinct vertices among the multiple vertices contained in the tree. In other words, it can be a data structure that does not form a loop in a graph data structure.
[0087] Throughout this specification, the terms artificial intelligence-based model, computational model, neural network, network function, and neural network may be used interchangeably. Hereafter, the term neural network will be used consistently. A data structure may include a neural network, and such a data structure may be stored in a computer-readable medium. The data structure may also include pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. The data structure may include any of the components in the disclosed configuration. That is, the data structure may include all or any combination of the following: pre-processed data for processing by the neural network, data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for learning the neural network. Beyond the configurations described above, data structures including neural networks can contain any other information that determines the properties of the neural network. Furthermore, data structures can contain, and are not limited to, all forms of data used or generated during the computational processes of the neural network. Computer-readable media can include computer-readable recording media and / or computer-readable transmission media. A neural network can be structured as a collection of interconnected computational units, which may generally be referred to as nodes. These nodes may also be referred to as neurons.A neural network consists of at least one node.
[0088] A data structure may include data to be input to a neural network. A data structure containing data to be input to a neural network may be stored on a computer-readable medium. Data to be input to a neural network may include training data input during the training process of the neural network and / or input data to be input to a neural network after training is complete. Data to be input to a neural network may include pre-processed data and / or data to be pre-processed. Pre-processing may include data processing processes to prepare data for input to a neural network. Therefore, a data structure may include data to be pre-processed and data generated by pre-processing. The data structures described above are merely examples, and this disclosure is not limited thereto.
[0089] The data structure may include the weights of a neural network (in this specification, weights and parameters may be used interchangeably). The data structure containing the weights of a neural network may be stored in a computer-readable medium. A neural network may have multiple weights. The weights may be variable and can be varied by the user or algorithm to enable the neural network to perform desired functions. For example, if one or more input nodes are interconnected to an output node by their respective links, the output node can determine the data value output from the output node based on the values input to the input nodes connected to the output node and the weights set on the links corresponding to each input node. The data structures described above are merely examples, and this disclosure is not limited thereto.
[0090] As an unrestricted example, weights may include weights that are variable during the learning process of a neural network and / or weights after the neural network has finished learning. Weights that are variable during the learning process of a neural network may include weights at the start of a learning cycle and / or weights that are variable during a learning cycle. Weights that have finished learning a neural network may include weights after a learning cycle has finished. Thus, a data structure containing the weights of a neural network may include a data structure containing weights that are variable during the learning process of a neural network and / or weights after the neural network has finished learning. Therefore, the aforementioned weights and / or each combination of weights shall be included in the data structure containing the weights of a neural network. The aforementioned data structures are illustrative only, and this disclosure is not limited thereto.
[0091] A data structure containing neural network weights can be stored in a computer-readable medium (e.g., memory, hard disk) after undergoing a serialization process. Serialization may be a process of storing the data structure in the same or other computing device and later reconstructing it into a usable form. The computing device can serialize the data structure and send and receive the data over a network. The serialized data structure containing neural network weights can be reconstructed in the same or other computing device by deserialization. The data structure containing neural network weights is not limited to serialization. Furthermore, the data structure containing neural network weights may include data structures that enhance computational efficiency while minimizing the use of computing device resources (e.g., B-Tree, R-Tree, Trie, m-way search tree, AVL tree, Red-Black Tree in nonlinear data structures). The foregoing are merely examples, and this disclosure is not limited thereto.
[0092] The data structure may include the hyperparameters of the neural network. Furthermore, the data structure containing the neural network's hyperparameters may be stored in a computer-readable medium. The hyperparameters may be variable variables that can be changed by the user. Examples of hyperparameters may include the learning rate, cost function, number of training cycles, weight initialization (e.g., setting the range of weight values to be initialized), and the number of Hidden Units (e.g., the number of hidden layers, the number of nodes in the hidden layers). The data structures described above are merely examples, and this disclosure is not limited thereto.
[0093] Figure 8 is a simplified and general schematic diagram of an exemplary computing environment in which embodiments of the present disclosure can be realized.
[0094] While it has been stated that this disclosure can generally be embodied by computing devices, those skilled in the art will understand that this disclosure can also be embodied in combination with computer executable instructions and / or other program modules that can be run on one or more computers, and / or as a combination of hardware and software. Generally, modules as defined herein include routines, programs, components, data structures, and so on, that perform a specific task or implement a specific abstract data type. Furthermore, those skilled in the art will understand that the methods disclosed herein can be implemented in configurations of other computer systems, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, as well as personal computers, handheld computing devices, microprocessor boards, or programmable consumer electronics, and so on (each of which can operate in conjunction with one or more associated devices).
[0095] The embodiments described herein can further be implemented in a distributed computing environment in which a task is performed by remote processing units connected via a communication network. In a distributed computing environment, program modules can reside in both local and remote memory storage devices.
[0096] Computers include a variety of computer-readable media. Any media accessible by a computer can be computer-readable, but such computer-readable media include volatile and non-volatile media, transient and non-transitory media, and portable and non-portable media. By example, rather than by limitation, computer-readable media may include computer-readable storage media and computer-readable transmission media. Computer-readable storage media include volatile and non-volatile media, transient and non-transitory media, portable and non-portable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD (digital video disk) or other optical disk storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other media that can be accessed by a computer and used to store information.
[0097] Computer-readable transmission media typically include all information transmission media that implement computer-readable instructions, data structures, program modules, or other data on a modulated data signal, such as a carrier wave or other transport mechanism. The term modulated data signal means a signal in which one or more of its characteristics have been set or modified to encode information within the signal. By example, rather than by limitation, computer-readable transmission media include wired media such as wired networks or direct-wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of any of the aforementioned media is also included in the scope of computer-readable transmission media.
[0098] An exemplary environment (1100) is shown that realizes various aspects of this disclosure, including a computer (1102), the computer (1102) including a processor (1104), system memory (1106), and a system bus (1108). The system bus (1108) connects system components, including (but not limited to) system memory (1106), to the processor (1104). The processor (1104) can be any processor from a variety of commercial processors. Dual processors and other multiprocessor architectures can also be used as the processor (1104).
[0099] The system bus (1108) can be one of several types of bus structures that can be further interconnected to a local bus using any of the following: a memory bus, a peripheral bus, and various commercial bus architectures. System memory (1106) includes read-only memory (ROM) (1110) and random access memory (RAM) (1112). The basic input / output system (BIOS) is stored in non-volatile memory (1110) such as ROM, EPROM, or EEPROM, and this BIOS includes basic routines that support the exchange of information between multiple components within the computer (1102) during startup, etc. RAM (1112) may also include high-speed RAM such as static RAM for caching data.
[0100] The computer (1102) also includes an internal hard disk drive (HDD) (1114) (e.g., EIDE, SATA)—this internal hard disk drive (1114) can also be configured for external use in a suitable chassis (not shown), a magnetic floppy disk drive (FDD) (1116) (e.g., for reading from and writing to a portable diskette (1118)), and an optical disk drive (1120) (e.g., for reading CD-ROM disks (1122) or for reading from and writing to other high-capacity optical media such as DVDs). The hard disk drive (1114), magnetic disk drive (1116), and optical disk drive (1120) can be connected to the system bus (1108) by a hard disk drive interface (1124), a magnetic disk drive interface (1126), and an optical drive interface (1128), respectively. The interface (1124) for implementing external drives includes, for example, at least one or both of the following: USB (Universal Serial Bus) or IEEE 1394 interface technology.
[0101] These drives and computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, and so on. In the case of a computer (1102), the drives and media correspond to storing any data in a suitable digital format. While the above description of computer-readable storage media refers to HDDs, portable magnetic disks, and portable optical media such as CDs or DVDs, those skilled in the art will understand that other types of computer-readable storage media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, and so on, can also be used in exemplary operating environments, and that any of these media may contain computer-executable instructions for performing the methods of the present disclosure.
[0102] Numerous program modules, including an operating system (1130), one or more application programs (1132), other program modules (1134), and program data (1136), can be stored in the drive and RAM (1112). All or part of the operating system, applications, modules, and / or data can also be cached in RAM (1112). It will be understood that this disclosure can be implemented by various commercially available operating systems or combinations of operating systems.
[0103] The user can input commands and information to the computer (1102) through one or more wired or wireless input devices, such as a keyboard (1138) and a pointing device such as a mouse (1140). Other input devices (not shown in the diagram) may include a microphone, IR remote control, joystick, gamepad, stylus pen, touchscreen, and so on. These and other input devices are often connected to the processing unit (1104) through an input device interface (1142) connected to the system bus (1108), but they can also be connected through other interfaces such as parallel ports, IEEE1394 serial ports, game ports, USB ports, IR interfaces, and so on.
[0104] A monitor (1144) or other type of display device is also connected to the system bus (1108) through an interface such as a video adapter (1146). In addition to the monitor (1144), the computer generally includes other peripheral output devices such as speakers, printers, and so on (not shown in the illustration).
[0105] A computer (1102) can operate in a networked environment by utilizing logical connections to one or more remote computers (1148), such as multiple remote computers (1148), via wired and / or wireless communication. The multiple remote computers (1148) can be workstations, server computers, routers, personal computers, portable computers, microprocessor-based entertainment devices, peer devices, or other typical network nodes, and generally include many or all of the components described for a computer (1102), although for simplification only a memory storage device (1150) is illustrated. The illustrated logical connections include wired and wireless connections in a short-range network (LAN) (1152) and / or a larger network, such as a long-range network (WAN) (1154). Such LAN and WAN networking environments are common in offices and companies, facilitating enterprise-wide computer networks such as intranets, all of which can connect to global computer networks, such as the Internet.
[0106] When used in a LAN networking environment, the computer (1102) connects to the local network (1152) via a wired and / or wireless network interface, or via an adapter (1156). The adapter (1156) facilitates wired or wireless communication to the LAN (1152), which also includes a wireless access point installed therein to communicate with the wireless adapter (1156). When used in a WAN networking environment, the computer (1102) may include a modem (1158), connect to a communication server on the WAN (1154), or have other means of establishing communication through the WAN (1154), such as via the Internet. The modem (1158), which can be internal or external, and wired or wireless, connects to the system bus (1108) via a serial port interface (1142). In a networked environment, a program module or part thereof described for a computer (1102) can be stored in a remote memory / storage device (1150). While the illustrated network connection is illustrative, it is readily apparent that other means of establishing communication links between multiple computers may be used.
[0107] The computer (1102) operates to communicate with any wireless device or unit that is arranged and operates wirelessly, such as a printer, scanner, desktop and / or portable computer, PDA (portable data assistant), communication satellite, any equipment or location relating to a wirelessly discoverable tag, and a telephone. This includes at least Wi-Fi and Bluetooth wireless technologies. Thus, the communication may be a predefined structure like a conventional network, or simply ad hoc communication between at least two devices.
[0108] Wi-Fi (Wireless Fidelity) enables internet access and other connectivity without a wired connection. Wi-Fi is a wireless technology, similar to cell phones, that allows devices like computers to send and receive data indoors and outdoors, i.e., anywhere within the range of a base station. Wi-Fi networks use IEEE 802.11 (a, b, g, etc.) wireless technology to provide secure, reliable, and high-speed wireless connectivity. Wi-Fi can be used to connect computers to each other, to the internet, and to wired networks (using IEEE 802.3 or Ethernet). Wi-Fi networks can operate in unlicensed 2.4 and 5 GHz wireless bands at data rates such as 11 Mbps (802.11a) or 54 Mbps (802.11b), or in products that include both bands (dual band).
[0109] A person with ordinary skill in the art of this disclosure will understand that information and signals can be represented using any variety of different techniques and methods. For example, the data, instructions, commands, information, signals, bits, symbols and chips referenced in the foregoing description can be represented by voltage, current, electromagnetic waves, magnetic fields, etc. or particles, optical fields, etc. or particles, or any combination thereof.
[0110] A person with ordinary skill in the art of this disclosure will understand that the various exemplary logic blocks, modules, processors, means, circuits, and algorithmic stages described herein can be implemented by electronic hardware, various forms of programs or design code (referred to herein for convenience as “software”), or a combination of all of these. To illustrate this interoperability of hardware and software, various exemplary components, blocks, modules, circuits, and stages have been generally described above with respect to their functions. Whether such functions are implemented in hardware or software depends on the design constraints imposed on a particular application and the overall system. A person with ordinary skill in the art of this disclosure can implement the functions described in various ways for individual specific applications, but such decisions should not be construed as departing from the scope of this disclosure.
[0111] The various embodiments described herein can be realized by methods, apparatus, or manufactured articles using standard programming and / or engineering techniques. The term “manufactured article” includes computer programs, carriers, or media accessible from any computer-readable device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). The various storage media described herein also include one or more devices and / or other machine-readable media for storing information.
[0112] It should be understood that the specific order or hierarchical structure of the multiple stages in the presented process is an example of an exemplary approach. It should be understood that, based on design priorities, the specific order or hierarchical structure of the stages in the process can be rearranged within the scope of this disclosure. The appended method claims provide a variety of stage elements in sample order, but are not limited to the specific order or hierarchical structure shown.
[0113] The descriptions relating to the embodiments provided herein are provided so that any person with ordinary skill in the art of the present disclosure may utilize or implement the present disclosure. Various variations of such embodiments are readily apparent to a person with ordinary skill in the art of the present disclosure, and the general principles defined herein can be applied to other embodiments without departing from the scope of the present disclosure. Accordingly, the present disclosure is not limited by the embodiments provided herein and should be interpreted in the broadest sense consistent with the principles and novel features provided herein.
Claims
1. A method for causing one or more processors in a computing device to perform an operation to improve the accuracy of video, The stage of acquiring first data related to position in the video; Steps to acquire the first time-series data related to the first data; The step of acquiring second data relating to the movement of the aforementioned video; Steps to acquire the second time-series data related to the second data; A step of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and A step of acquiring the final image based on the combined data and the video; including, method.
2. In the method according to claim 1, The step of acquiring first data relating to the position in the aforementioned video is: The step of acquiring the first-1 data relating to the rotation of an object included in the aforementioned video; or Steps to acquire first and second data relating to the positions of objects included in the aforementioned video; This includes at least one of the following: method.
3. In the method according to claim 1, The step of acquiring the second data relating to the movement of the aforementioned video is as follows: Step 2-1: Acquiring data relating to the acceleration of an object included in the aforementioned video; The step of acquiring the second-second data relating to the angular velocity of an object included in the aforementioned video; or Steps to acquire second- and third data relating to the orientation of objects included in the aforementioned video; This includes at least one of the following: method.
4. In the method according to claim 1, The step of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data is: A step of obtaining first corrected data based on the first data, the first time series data, and the second time series data; A step of obtaining a second corrected data based on the second data, the first time series data, and the second time series data; and The process includes the step of obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data. method.
5. In the method according to claim 4, The step of obtaining first corrected data based on the first data, the first time series data, and the second time series data is: A step of obtaining combined time series data by matching the time series of the first time series data and the second time series data; and This includes the step of obtaining first corrected data based on the first data and the combined time series data, The step of obtaining second corrected data based on the second data, the first time series data, and the second time series data is: This includes the step of obtaining a second corrected data based on the second data and the combined time series data, method.
6. In the method according to claim 4, The step of obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data, is: A step of obtaining the combined data by supplementing the errors accumulated over time in the second corrected data based on the first corrected data; or A step of obtaining the combined data by supplementing the missing data in the first corrected data based on the second corrected data; This includes at least one of the following: method.
7. In the method according to claim 4, The step of obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data, is: The process further includes the step of obtaining filtered combined data by filtering the combined data. method.
8. In the method according to claim 7, The step of obtaining filtered combined data by filtering the combined data is as follows: The step of obtaining filtered combined data by filtering the combined data based on a Kalman filter; or The step of obtaining filtered combined data by filtering the combined data based on a complementary filter; This includes at least one of the following: method.
9. In the method according to claim 1, The step of acquiring the final video based on the combined data and the video is as follows: The process includes the step of correcting the shaking of objects included in the video based on the combined data and obtaining the final video, method.
10. A method for improving the accuracy of video, which is performed by one or more processors of a computing device, The stage of acquiring first data related to position in the video; The step of acquiring second data relating to the movement of the aforementioned video; A step of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and This includes the step of correcting the angle in a portion of the video based on the combined data to obtain the final video. method.
11. A computer program stored on a computer-readable storage medium, wherein, when the computer program is executed by one or more processors, it causes the one or more processors to perform an operation to improve the accuracy of the image, and the operation is: An operation to acquire first data related to position in the video; An operation to acquire the first time series data related to the first data; An operation to acquire second data related to the movement of the aforementioned video; An operation to acquire second time-series data related to the second data; An operation to obtain combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and An operation to acquire the final image based on the combined data and the image; including, A computer program stored on a computer-readable storage medium.
12. In the computer program described in claim 11, The operation to acquire the first data relating to the position in the aforementioned video is as follows: An operation to acquire the 1-1 data relating to the rotation of an object included in the aforementioned video; or An operation to acquire first and second data relating to the position of objects included in the aforementioned video; This includes at least one of the following: A computer program stored on a computer-readable storage medium.
13. In the computer program described in claim 11, The operation to acquire the second data relating to the movement of the aforementioned video is as follows: An operation to acquire 2-1 data relating to the acceleration of an object included in the aforementioned video; An operation to acquire the second-second data relating to the angular velocity of an object included in the aforementioned video; or An operation to acquire second- and third data relating to the orientation of objects included in the aforementioned video; This includes at least one of the following: A computer program stored on a computer-readable storage medium.
14. In the computer program described in claim 11, The operation of obtaining combined data by supplementing the second data based on the first data and supplementing the first data based on the second data is: An operation to obtain first corrected data based on the first data, the first time series data, and the second time series data; An operation to obtain a second corrected data based on the second data, the first time series data, and the second time series data; and The operation includes obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data. A computer program stored on a computer-readable storage medium.
15. In the computer program described in claim 14, The operation of obtaining the first corrected data based on the first data, the first time series data, and the second time series data is as follows: An operation to obtain combined time series data by matching the time series of the first time series data and the second time series data; and This includes the operation of acquiring first corrected data based on the first data and the combined time series data, The operation of obtaining second corrected data based on the second data, the first time series data, and the second time series data is as follows: This includes an operation to acquire a second corrected data based on the second data and the combined time series data, A computer program stored on a computer-readable storage medium.
16. In the computer program described in claim 14, The operation of obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data, is: An operation to obtain the combined data by supplementing the errors accumulated over time in the second corrected data based on the first corrected data; or An operation to obtain the combined data by supplementing any missing data in the first corrected data based on the second corrected data; This includes at least one of the following: A computer program stored on a computer-readable storage medium.
17. In the computer program described in claim 14, The operation of obtaining combined data by supplementing the second corrected data based on the first corrected data, and supplementing the first corrected data based on the second corrected data, is: The operation further includes applying a filter to the aforementioned combined data to obtain filtered combined data. A computer program stored on a computer-readable storage medium.
18. In the computer program described in claim 17, The operation of obtaining filtered combined data by filtering the combined data is as follows: An operation to obtain filtered combined data by filtering the combined data based on a Kalman filter; or The operation of obtaining filtered combined data by filtering the combined data based on a complementary filter; This includes at least one of the following: A computer program stored on a computer-readable storage medium.
19. In the computer program described in claim 11, The operation to acquire the final image based on the combined data and the image is as follows: This includes the operation of correcting the shaking of objects included in the video based on the combined data and acquiring the final video, A computer program stored on a computer-readable storage medium.
20. A computer program stored on a computer-readable storage medium, wherein, when the computer program is executed by one or more processors, it causes the one or more processors to perform an operation to improve the accuracy of the image, and the operation is: An operation to acquire first data related to position in the video; An operation to acquire second data related to the movement of the aforementioned video; An operation to obtain combined data by supplementing the second data based on the first data and supplementing the first data based on the second data; and This includes an operation to correct the angle in a portion of the video based on the combined data and to acquire the final video. A computer program stored on a computer-readable storage medium.
21. A computing device, at least one processor; and memory Includes, The at least one processor is First data relating to position in the video is acquired, The first time series data relating to the first data is acquired, Second data relating to the movement of the aforementioned video is acquired; Acquire the second time-series data related to the second data mentioned above, By supplementing the second data with the first data and supplementing the first data with the second data, combined data is obtained; and The system is configured to acquire a final image based on the combined data and the video. Computing device.