Multi-view 2d pose matching method for real-time 3D pose estimation using mobile phone

WO2026177346A1PCT designated stage Publication Date: 2026-08-27ACTNOVA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/095196
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2025-04-10
Publication Date
2026-08-27

Smart Images

  • Figure KR2025095196_27082026_PF_FP_ABST
    Figure KR2025095196_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present specification relates to a multi-view 2D pose matching method for real-time 3D pose estimation by a server, the method comprising the steps of: acquiring images from two or more terminals; acquiring extrinsic parameters of the terminals on the basis of the images; acquiring inertial measurement unit (IMU) data of the terminals on the basis of the images; and reflecting the IMU data in the extrinsic parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-view 2D Pose Registration Method for Real-time 3D Pose Estimation Using a Mobile Phone

[0001] The present specification relates to a multi-view 2D pose matching method for real-time 3D pose estimation using two or more mobile phones.

[0002]

[0003] The process of estimating poses (key-points, skeletons) is crucial in animal behavior analysis. Pose data is key information for understanding animal movement and behavior, enabling more precise analysis of behavioral patterns. In particular, utilizing pose data can significantly reduce data complexity compared to image data (e.g., height x width x 3) and is advantageous for high-speed computational tasks such as real-time processing. Furthermore, analysis based on pose data can demonstrate robust performance even under optical conditions such as changes in background images or lighting, allowing for stable animal behavior analysis in various environments.

[0004] Existing animal pose estimation technologies have mostly been developed based on 2D pose estimation. However, 2D pose data has limitations in fully explaining an animal's posture or movement. Furthermore, since it is difficult to obtain sufficient 3D pose data of animals, there are many constraints in training 3D pose estimation models. On the other hand, utilizing 3D poses has the advantage of significantly improving the accuracy and reliability of analysis by enabling the derivation of features independent of the animal's body scale or viewpoints. For this reason, acquiring 3D pose data is considered a critical task in research and experimental fields that analyze animal behavior.

[0005] However, there are several issues with acquiring 3D pose data. First, it is difficult to obtain consistent data from multiple viewpoints, and there is a lack of synchronization to maintain alignment between cameras. Additionally, camera parameters required for 3D pose estimation are not calibrated in advance, which can affect accuracy. Finally, there is a possibility that data quality may degrade due to environmental constraints such as lighting, background, and distance. These issues have a significant impact on the accuracy and reliability of 3D pose estimation.

[0006]

[0007] The purpose of the present specification is to implement a method of photographing an animal from multiple viewpoints using a mobile phone owned by a general user and reconstructing a 3D pose based on the above through an advanced software algorithm.

[0008] The technical problems that this specification aims to solve are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this specification belongs from the detailed description of the specification below.

[0009] One aspect of the present specification may include, in a multi-view 2D pose matching method for real-time 3D pose estimation, a server acquiring images from two or more terminals; acquiring extrinsic parameters of the terminals based on the images; acquiring Inertial Measurement Unit (IMU) data of the terminals based on the images; and reflecting the IMU data in the extrinsic parameters.

[0010] Additionally, the step of obtaining extrinsic parameters of the terminal may include: a step of extracting 2D key-points of a target from the image; a step of estimating a Fundamental Matrix based on the image and the 2D key-points; a step of obtaining internal parameters of the terminal; and a step of obtaining extrinsic parameters based on the Fundamental Matrix and the internal parameters.

[0011] In addition, the above 2D key-points may include joint information extracted through a pose estimation algorithm.

[0012] In addition, the step of estimating the Fundamental Matrix may be based on the RANSAC (Random Sample Consensus) algorithm and the 8-point algorithm.

[0013] In addition, the step of reflecting the IMU data to the external parameter may include a step of correcting the conversion relationship between the IMU data acquired for each frame of the image and the external parameter of the frame.

[0014] In addition, the two or more terminals can capture the target at various viewpoints to generate an image.

[0015] Additionally, it may further include the step of reconstructing the 3D pose of the object based on the 2D key-points, the internal parameters, and the external parameters.

[0016] Additionally, if the number of frames of the processed image exceeds a preset number, the method may further include the step of updating the external parameters.

[0017] Another aspect of the present specification is a server for performing multi-view 2D pose matching for real-time 3D pose estimation, comprising: a communication module; a memory; and a processor for functionally controlling the communication module and the memory, wherein the processor acquires images from two or more terminals through the communication module, acquires extrinsic parameters of the terminals based on the images, acquires Inertial Measurement Unit (IMU) data of the terminals based on the images, and can reflect the IMU data in the extrinsic parameters.

[0018]

[0019] According to an embodiment of the present specification, a method can be implemented to photograph an animal from multiple viewpoints using a mobile phone owned by a general user and to reconstruct a 3D pose based on this using an advanced software algorithm.

[0020] The effects obtainable in this specification are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which this specification belongs from the description below.

[0021]

[0022] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification.

[0023] FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification.

[0024] FIG. 3 illustrates a multi-view 2D pose matching method to which the present specification can be applied.

[0025] FIG. 4 illustrates a method for obtaining external parameters to which the present specification can be applied.

[0026] FIG. 5 illustrates a 3D pose extraction method to which the present specification can be applied.

[0027] The accompanying drawings, included as part of the detailed description to aid in understanding the present specification, provide embodiments of the present specification and explain the technical features of the present specification together with the detailed description.

[0028]

[0029] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols will be assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this specification.

[0030] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0031] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0032] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0033] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0034]

[0035] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification.

[0036] The above electronic device (100) may include a wireless communication unit (110), an input unit (120), a sensing unit (140), an output unit (150), an interface unit (160), a memory (170), a control unit (180), and a power supply unit (190), etc. Since the components illustrated in FIG. 1 are not essential for implementing the electronic device, the electronic device described herein may have more or fewer components than those listed above.

[0037] More specifically, among the above components, the wireless communication unit (110) may include one or more modules that enable wireless communication between the electronic device (100) and a wireless communication system, between the electronic device (100) and another electronic device (100), or between the electronic device (100) and an external server. Additionally, the wireless communication unit (110) may include one or more modules that connect the electronic device (100) to one or more networks.

[0038] This wireless communication unit (110) may include at least one of a broadcast receiving module (111), a mobile communication module (112), a wireless internet module (113), a short-range communication module (114), and a location information module (115).

[0039] The input unit (120) may include a camera (121) or video input unit for inputting a video signal, a microphone (122) or audio input unit for inputting an audio signal, and a user input unit (123, e.g., a touch key, a mechanical key, etc.) for receiving information from a user. Voice data or image data collected from the input unit (120) may be analyzed and processed into a control command by the user.

[0040] The sensing unit (140) may include one or more sensors for sensing at least one of information within the electronic device, information about the surrounding environment surrounding the electronic device, and user information. For example, the sensing unit (140) may include at least one of a proximity sensor (141), an illumination sensor (142), a touch sensor, an acceleration sensor, a magnetic sensor, a gravity sensor (G-sensor), a gyroscope sensor, a motion sensor, an RGB sensor, an infrared sensor (IR sensor: infrared sensor), a fingerprint sensor (finger scan sensor), an ultrasonic sensor, an optical sensor (e.g., see camera (121)), a microphone (see 122), a battery gauge, an environmental sensor (e.g., a barometer, a hygrometer, a thermometer, a radiation detection sensor, a heat detection sensor, a gas detection sensor, etc.), and a chemical sensor (e.g., an electronic nose, a healthcare sensor, a biometric sensor, etc.). Meanwhile, the electronic device disclosed in this specification can utilize information sensed by at least two of these sensors in combination.

[0041] The output unit (150) is intended to generate output related to sight, hearing, or touch, and may include at least one of a display unit (151), an audio output unit (152), a haptic module (153), and an optical output unit (154). The display unit (151) may form a layered structure with a touch sensor or be formed integrally to implement a touch screen. Such a touch screen functions as a user input unit (123) that provides an input interface between the electronic device (100) and the user, and at the same time can provide an output interface between the electronic device (100) and the user.

[0042] The interface section (160) serves as a passage for various types of external devices connected to the electronic device (100). This interface section (160) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. In response to an external device being connected to the interface section (160), the electronic device (100) can perform appropriate control related to the connected external device.

[0043] Additionally, the memory (170) stores data that supports various functions of the electronic device (100). The memory (170) can store a number of application programs (or applications) running on the electronic device (100), data for the operation of the electronic device (100), and commands. At least some of these application programs may be downloaded from an external server via wireless communication. Also, at least some of these application programs may exist on the electronic device (100) from the time of shipment for the basic functions of the electronic device (100) (e.g., phone incoming and outgoing functions, message receiving and outgoing functions). Meanwhile, the application programs may be stored in the memory (170), installed on the electronic device (100), and driven by the control unit (180) to perform the operation (or function) of the electronic device.

[0044] In addition to operations related to the application program, the control unit (180) typically controls the overall operation of the electronic device (100). The control unit (180) can provide or process appropriate information or functions to the user by processing signals, data, information, etc. that are input or output through the components described above, or by running an application program stored in memory (170).

[0045] Additionally, the control unit (180) can control at least some of the components examined together with FIG. 1 in order to run an application program stored in memory (170). Furthermore, the control unit (180) can operate at least two or more of the components included in the electronic device (100) in combination with each other to run the application program.

[0046] The power supply unit (190) receives external power and internal power under the control of the control unit (180) and supplies power to each component included in the electronic device (100). This power supply unit (190) includes a battery, and the battery may be a built-in battery or a replaceable battery.

[0047] At least some of the above components may operate in cooperation with each other to implement the operation, control, or control method of an electronic device according to various embodiments described below. Additionally, the operation, control, or control method of the electronic device may be implemented on the electronic device by running at least one application program stored in the memory (170).

[0048] In this specification, the electronic device (100) may be collectively referred to as a server, and the server may include a cloud server. Additionally, the terminal may include all or part of the configuration of the electronic device (100), and may include a tablet PC.

[0049]

[0050] FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification.

[0051] The AI ​​device (20) may include an electronic device including an AI module capable of performing AI processing, or a terminal including the AI ​​module. Additionally, the AI ​​device (20) may be configured to be included as at least a part of the configuration of the electronic device (100) shown in FIG. 1 to perform at least a part of the AI ​​processing together.

[0052] The above AI device (20) may include an AI processor (21), memory (25) and / or a communication unit (27).

[0053] The above AI device (20) is a computing device capable of learning a neural network and can be implemented as various electronic devices such as a terminal, desktop PC, laptop PC, tablet PC, etc.

[0054] The AI ​​processor (21) can train a neural network using a program stored in memory (25). In particular, the AI ​​processor (21) may include a large-scale pre-trained pose estimation model. For example, the pose estimation model can predict key points of an object in video frames acquired from a terminal in real time.

[0055] Meanwhile, the AI ​​processor (21) that performs the functions described above may be a general-purpose processor (e.g., CPU), but may be an AI-dedicated processor for artificial intelligence learning (e.g., GPU, graphics processing unit).

[0056] The memory (25) can store various programs and data required for the operation of the AI ​​device (20). The memory (25) can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SDD). The memory (25) is accessed by the AI ​​processor (21), and the AI ​​processor (21) can perform reading / writing / modification / deletion / updating of data. Additionally, the memory (25) can store a neural network model (e.g., a deep learning model) generated through a learning algorithm for data classification / recognition according to one embodiment of the present specification.

[0057] Meanwhile, the AI ​​processor (21) may include a data learning unit that learns a neural network for data classification / recognition. For example, the data learning unit may learn a deep learning model by acquiring training data to be used for learning and applying the acquired training data to a deep learning model.

[0058] The communication unit (27) can transmit the AI ​​processing results by the AI ​​processor (21) to an external electronic device.

[0059] Here, external electronic devices may include other terminals.

[0060] Meanwhile, although the AI ​​device (20) illustrated in FIG. 2 is described by functionally separating it into an AI processor (21), memory (25), and communication unit (27), the aforementioned components may be integrated into a single module and referred to as an AI module or an artificial intelligence (AI) model.

[0061] FIG. 3 illustrates a multi-view 2D pose matching method to which the present specification can be applied.

[0062] Referring to FIG. 3, the server can receive video data in real time from two or more terminals. The server may include a pose estimation AI model and can estimate the 3D pose of an object using the video data received from the terminals.

[0063] Here, the terminal may refer to a device such as a mobile phone and can acquire image data to estimate the 3D pose of a target. The terminal may include a separate application for this purpose and generally includes a built-in camera and an IMU (Inertial Measurement Unit) sensor. Through this, the terminal can capture images of the target from various viewpoints and collect motion data. For example, the accurate 3D pose of the target can be estimated through multi-angle information obtained by utilizing multiple terminals.

[0064] However, when using multiple terminals for 3D pose estimation, quality degradation may occur due to insufficient alignment between cameras, ambiguity of internal and external parameters, and difficulty in synchronizing between IMU data and image data. Below, a multi-view 2D pose alignment method to solve this problem is exemplified.

[0065] The server acquires images of a target from two or more terminals (S3010). For example, the server can acquire images of the same target from multiple terminals (mobile phones). The terminals capture the target from various views, and the images thus acquired may contain information from multiple angles. The reason for collecting images from multiple views is that data from at least two different views is required for 3D pose estimation. Two or more terminals can transmit the acquired images of the same target to the server through cameras.

[0066] The server obtains external parameters of the terminal based on the acquired image (S3020). For example, the external parameters are values ​​representing the position and direction of the terminal's camera, which can represent the relative relationship between the target and the camera.

[0067] FIG. 4 illustrates a method for obtaining external parameters to which the present specification can be applied.

[0068] Referring to FIG. 4, the server can repeat the external parameter acquisition task described below as many times as the total number of cameras of the terminal to acquire external parameters of the terminal.

[0069] The server extracts 2D key-points of the target from the video (S4010). The server can extract 2D key-points (e.g., 11 coordinates), which are the main feature points of the target, from video frames captured by the camera of each terminal. For example, key-points are coordinate data that accurately represents the target, mainly indicating joints or specific locations of an object, and can be used as basic data for calculating 3D poses. The key-points are

[0070] Joint information automatically extracted through computer vision algorithms may be included. Joints are important data for estimating the 3D pose of an object, and positional information between joints can simplify the process of matching the same object at each camera viewpoint. For example, if the positions of specific joints, such as the right wrist or left shoulder, are specified through a pose estimation algorithm, the accuracy and complexity of the matching process can be increased.

[0071] However, the number of joint points may be limited, and there is a possibility that they may not be sufficiently suitable for the RANSAC (Random Sample Consensus) and 8-point algorithms described later. To compensate for this, key points other than joints can be utilized if necessary. For example, if the subjects consist of multiple individuals (multiple mice), a key point matching process between individuals can be performed. Through this, key points extracted from multiple viewpoints can be aligned. This process can be performed iteratively, and key-points obtained from different viewpoints based on multi-view data can be comprehensively utilized.

[0072] The server estimates the Fundamental Matrix based on images and 2D key-points (S4020). For example, the server can estimate the Fundamental Matrix based on the Random Sample Consensus (RANSAC) algorithm and the 8-point algorithm.

[0073] 1. RANSAC Algorithm

[0074] RANSAC is an iterative sampling algorithm designed to remove outliers from given data and derive an optimal model. Since some key-points in image data may be mismatched or contain errors, failing to remove them can compromise the reliability of the Fundamental Matrix calculation. Therefore, the server can utilize the RANSAC algorithm to first randomly select a certain number (e.g., 8) of key-points from the image data and calculate the Fundamental Matrix based on these selected key-points. Subsequently, it calculates the error of the generated model for the remaining image data and classifies data where the error falls within an acceptable threshold as inliers. By repeating this process, the server can select the model containing the most inliers as the optimal Fundamental Matrix.

[0075] 2.8-point algorithm

[0076] The server can compute the Fundamental Matrix using the 8-point algorithm, utilizing key-points selected in each iteration of RANSAC. This algorithm estimates the Fundamental Matrix using eight corresponding points between two cameras. For example, the server can normalize each corresponding point to reduce the influence of scale and position, and construct a system of linear equations based on the coordinates of the corresponding points. Subsequently, the Fundamental Matrix can be computed using Singular Value Decomposition (SVD). The server can condition the computed Fundamental Matrix to adjust its rank to 2.

[0077] Thus, the estimated Fundamental Matrix defines the geometric relationship between the two cameras and can express the epipolar constraint between the two images. This implies that when the two cameras observe the same object, a point in one image must lie on an epipolar line in the other image. The Fundamental Matrix can be used to calculate the Projection Matrix and Extrinsic Parameters in subsequent steps.

[0078] The server obtains the intrinsic parameters of the terminal (S4030). The server can obtain the unique intrinsic parameters of each terminal (camera). The intrinsic parameters may include physical characteristics unique to the camera, such as the camera's focal length, principal point, and lens distortion coefficient. The server can obtain the intrinsic parameters through predefined camera model information or calibration data.

[0079] The server obtains extrinsic parameters of the terminal based on the Fundamental Matrix and intrinsic parameters (S4040). The server can calculate extrinsic parameters of the camera by combining the Fundamental Matrix and intrinsic parameters. Extrinsic parameters represent the position and orientation of the camera and can mathematically express the relative relationship between the camera and the target in 3D space. For example, a Projection Matrix is ​​generated through intrinsic parameters and extrinsic parameters, which can convert the camera's 2D image data into actual 3D spatial coordinates.

[0080] Referring again to FIG. 3, the server acquires IMU (Inertial Measurement Unit) data of the terminal based on the video (S3030). An IMU is an inertial measurement device used to measure the movement and direction of the terminal, and generally includes an accelerometer, a gyroscope, and a magnetometer. IMU data can provide important information that allows real-time identification of the terminal's state, such as whether it is stationary, moving, or rotating. For example, IMU data can be collected in real-time through the built-in sensors of each terminal. More specifically, while the camera is capturing video, the IMU can track the camera's movement by providing synchronized data. This data is collected on a frame-by-frame basis and can be linked to the video by measuring changes that occur when the camera rotates or moves. For example, by identifying the camera's rotation angle or direction of movement in a specific frame through IMU data, the server can calculate how the video data of that frame reflects changes in the actual environment.

[0081] The server incorporates IMU data into external data (S3040). For example, IMU data measures the movement and orientation of the camera and, by combining it with external parameters, can more precisely define the geometric relationship between the camera and the target.

[0082] The server can utilize IMU data to more precisely adjust the camera's external parameters. To this end, the server can perform IMU-Camera Calibration through a process that is repeated a preset number of times (e.g., 20 to 30 times) over multiple frames. This process compares and analyzes the camera's external parameters collected at each moment with the amount of change in the IMU sensor to correct for and match the difference between the two values.

[0083] During the calibration process, the server compares the IMU data measured in each frame (e.g., acceleration, angular velocity, etc.) with the extrinsic parameters of that frame (e.g., camera position and orientation), calculates the transformation relationship between the two values, and corrects it. This process is intended to minimize the difference in coordinate systems between the IMU and the camera and to maintain accurate extrinsic parameters during the real-time 3D pose estimation process.

[0084] For example, when the camera rotates or moves within each frame, the extrinsic parameters of that frame can be updated based on changes in acceleration and angular velocity measured by the IMU sensor. This process continuously adjusts the correlation between IMU data and camera extrinsic parameters to maintain accurate geometric relationships with respect to camera movement. This iterative calibration process can contribute to reducing cumulative errors between frames and minimizing distortion of extrinsic parameters caused by environmental changes or movement.

[0085] FIG. 5 illustrates a 3D pose extraction method to which the present specification can be applied.

[0086] Referring to Fig. 5, the server can accurately maintain the target coordinates in 3D space even when the camera is moving through external parameters updated in S3040.

[0087] The server performs 3D pose reconstruction based on the target's 2D key-points and the terminal's internal and external parameters (S5010). For example, the server can reconstruct the target's 3D pose using 2D key-point data extracted from each terminal. To do this, the server can align the same target from different viewpoints using multi-view data. Multi-view data alignment is a process of matching data based on the same target captured from different viewpoints. Since the 2D key-points captured by each camera are the result of observing the same target from different viewpoints, they appear differently depending on the position and direction of each camera. During the multi-view alignment process, the server can define the geometric relationship between each camera and match corresponding key-points by utilizing the Fundamental Matrix and epipolar constraints. In this process, the server combines the terminal's internal and external parameters to generate a Projection Matrix, through which the 2D data collected from each camera's viewpoint can be converted into coordinates in 3D space.

[0088] Such a Projection Matrix reflects the internal and external characteristics of the camera to convert 2D image coordinates into coordinates in actual 3D space or to project 3D data onto a 2D plane. For example, the server can use triangulation techniques to integrate data obtained from multiple camera viewpoints and accurately calculate the positions occupied by the target's key points in 3D space.

[0089] When the number of frames of the video being processed by the server exceeds a preset number of frames, the server updates the external parameters of the terminal (S5020). The server can update the external parameters of each terminal after a certain frame interval (e.g., 30 frames) has elapsed. This can be performed through the procedure of FIG. 3 described above. Through this, the server can prevent the possibility that the relationship between the camera and the target may be distorted due to errors accumulated over time. During the update process, the external parameters are recalculated based on new frame data, thereby maintaining accuracy. During the update process, the server can utilize IMU data to reflect changes caused by the movement and rotation of the camera. At the same time, the reliability of the external parameters can be enhanced by re-aligning multi-view data (2D key-points observed from multiple viewpoints). Such iterative updates ensure the stability of 3D pose data over a long period of time and minimize errors caused by environmental changes or camera movement.

[0090] The foregoing specification may be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include Hard Disk Drives (HDDs), Solid State Disks (SSDs), Silicon Disk Drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, optical data storage devices, etc., and also include implementations in the form of carrier waves (e.g., transmission over the Internet). Accordingly, the above detailed description should not be interpreted restrictively in all respects and should be considered exemplary. The scope of this specification should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of this specification are included within the scope of this specification.

[0091] Furthermore, although the above description has focused on the services and embodiments, this is merely illustrative and does not limit the scope of this specification. Those skilled in the art will understand that various modifications and applications not exemplified above are possible without departing from the essential characteristics of the services and embodiments. For example, each component specifically shown in the embodiments may be modified and implemented. Differences related to such modifications and applications should be interpreted as being included within the scope of this specification as defined in the appended claims.

Claims

1. In a multi-view 2D pose matching method for real-time 3D pose estimation, a server A step of acquiring images from two or more terminals; A step of obtaining extrinsic parameters of the terminal based on the above image; A step of acquiring IMU (Inertial Measurement Unit) data of the terminal based on the above image; and A step of reflecting the above IMU data to the above external parameter; A matching method including 2. In Paragraph 1, The step of obtaining extrinsic parameters of the above terminal A step of extracting 2D key-points of a target from the above image; A step of estimating a Fundamental Matrix based on the above image and the above 2D key-points; A step of obtaining internal parameters of the above terminal; and A step of obtaining the external parameters based on the above Fundamental Matrix and the above internal parameters; A matching method including 3. In Paragraph 2, The above 2D key-points are A matching method comprising joint information extracted through a pose estimation algorithm.

4. In Paragraph 2, The step of estimating the above Fundamental Matrix is A matching method based on the RANSAC (Random Sample Consensus) algorithm and the 8-point algorithm.

5. In Paragraph 2, The step of reflecting the above IMU data to the above external parameter A step of correcting the transformation relationship between the IMU data acquired for each frame of the above image and the external parameters of the frame; A matching method including 6. In Paragraph 5, The above two or more terminals are A registration method for generating images by photographing the above-mentioned object from various viewpoints.

7. In Paragraph 6, A step of reconstructing a 3D pose of the object based on the 2D key-points, the intrinsic parameters, and the extrinsic parameters; A matching method that further includes 8. In Paragraph 7, A step of updating the external parameters when the number of frames of the video being processed exceeds a preset number; A matching method that further includes 9. A server that performs multi-view 2D pose matching for real-time 3D pose estimation, Communication module; Memory; The above communication module, and a processor for functionally controlling the above memory; Includes, The above processor Through the above communication module, images are acquired from two or more terminals, and A server that acquires extrinsic parameters of the terminal based on the above image, acquires IMU (Inertial Measurement Unit) data of the terminal based on the above image, and reflects the IMU data in the extrinsic parameters.