System and method for synchronizing camera feed of two or more cameras
The AI-driven synchronization method for HMD cameras addresses synchronization issues by generating virtual frames and transforming image frames, improving accuracy and immersive experience in HMD devices.
Patent Information
- Application Number
- PCT/KR2025/009206
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Synchronization issues between multiple cameras in head-mounted display (HMD) devices lead to inaccurate head-tracking and degraded immersive experiences due to factors like technical malfunctions, environmental interference, and configuration errors, causing discrepancies in virtual reality rendering.
A system and method using a pre-trained artificial intelligence (AI) model to synchronize camera feeds by generating a virtual image frame, determining significant feature points, estimating transformation parameters, and transforming image frames to align them accurately, assisted by inertial measurement units (IMUs) for precise synchronization.
Improves synchronization accuracy, enhancing the immersive experience by correcting synchronization errors and ensuring accurate virtual reality rendering in HMD devices.
Smart Images

Figure KR2025009206_08012026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR SYNCHRONIZING CAMERA FEED OF TWO OR MORE CAMERAS
[0001] The present disclosure relates to the field of simultaneous localization and mapping (SLAM) in head-mounted display (HMD) devices. For example, a system and a method for synchronizing the camera feed of two or more cameras of an HMD device to improve performance of SLAM techniques of the HMD.
[0002] Simultaneous Localization and Mapping (SLAM) is a technology that enables an HMD device to create a real-time map of the surrounding environment while simultaneously tracking their position and orientation within that map. Said technology leverages a combination of multiple cameras and sensors such as inertial measurement unit (IMU) sensors to generate an accurate and dynamic representation of the surrounding environment.
[0003] For example, based on acceleration data and angular velocity data obtained through an inertial measurement unit and image data acquired via the multiple cameras, the HMD device can estimate the user's movement and gaze, and generate a real-time three-dimensional map of the surrounding environment.
[0004] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.
[0005] According to an embodiment of the present disclosure, disclosed herein is a method for synchronizing camera feed of two or more cameras of a head-mounted device (HMD) device. the method includes receiving at a current timestamp, a first image frame captured using a first camera and a second image frame captured using a second camera followed by generating, using a pre-trained artificial intelligence (AI) model, a virtual second image frame corresponding to the second camera at the current timestamp. Further, the method includes determining one or more significant feature points, among a plurality of feature points, with a confidence score equal to or greater than a predefined threshold associated with the accuracy of each pixel in the virtual second image frame, and corresponding significant feature points in the second image frame. Thereafter, the method includes estimating transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame. Furthermore, the method includes transforming each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame.
[0006] According to an embodiment of the present disclosure, disclosed herein is a system for synchronizing camera feed of two or more cameras of an HMD device. The system includes a memory storing at least one instruction, and at least one processor comprising a processing circuitry. The at least one processor is configured to individually or collectively execute the at least one instruction stored in the memory to receive at a current timestamp, a first image frame captured using a first camera and a second image frame captured using a second camera followed by generating, using a pre-trained artificial intelligence (AI) model, a virtual second image frame from the second camera at the current timestamp. Further, the processor is configured to determine one or more significant feature points in the virtual second image frame and corresponding significant feature points in the second image frame, and estimate transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame. Furthermore, the processor is configured to transform each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame.
[0007] According to an embodiment of the disclosure, a computer-readable recording medium having recorded thereon a computer program, which, when executed by a computer, performs at least one of the above-disclosed embodiments of the operation method may be provided.
[0008] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.
[0009] The foregoing and other features of embodiments will become more apparent from the following detailed description of embodiments when read in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements.
[0010] FIG. 1 is a pictorial diagram depicting a plurality of feature points detected within an exemplary image captured using one of the multiple cameras of the HMD device, according to an existing art;
[0011] FIG. 2 is a pictorial diagram depicting capturing images of the scene using two cameras from different points of view, according to an existing art;
[0012] FIG. 3A is a schematic diagram depicting an ideal scenario of synchronized cameras, according to an existing art;
[0013] FIG. 3B is a schematic diagram depicting an exemplary scenario of rotation and transformation error, according to an existing art;
[0014] FIG. 3C is a schematic diagram depicting an exemplary scenario of synchronization error, according to an existing art;
[0015] FIG. 4 is a pictorial diagram illustrating discrepancies resulting from the synchronization errors between two cameras, according to an existing art;
[0016] FIG. 5 is a block diagram depicting an HMD device implementing a system for synchronizing the camera feed of two or more cameras of the HMD device, according to embodiments of the present disclosure;
[0017] FIG. 6A is a block diagram depicting an operational flow of the plurality of modules for synchronizing the camera feed of the two or more cameras of the HMD device, according to one or more embodiments of the present disclosure;
[0018] FIG. 6B illustrates an operational flow associated with training pipeline of the AI model for synchronizing the camera feed of the two or more cameras of the HMD device, according to one or more embodiments of the present disclosure;
[0019] FIG. 7A is a pictorial diagram illustrating the operational flow of the plurality of modules using exemplary image frames, according to one or more embodiments of the present disclosure;
[0020] FIG. 7B is a pictorial diagram illustrating the operational flow of the plurality of modules using exemplary image frames, according to one or more embodiments of the disclosure;
[0021] FIG. 8 is a schematic diagram depicting an exemplary implementation of the synchronization module, according to one or more embodiments of the present disclosure; and
[0022] FIG. 9 is a block diagram depicting a method for synchronizing the camera feed of two or more cameras of the HMD device, according to one or more embodiments of the present disclosure.
[0023] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0024] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.
[0025] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.
[0026] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element". Furthermore, the use of the terms "one or more" or "at least one" feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, "there needs to be one or more..." or "one or more elements is required".
[0027] Reference is made herein to some "embodiments". It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.
[0028] Use of the phrases and / or terms including, but not limited to, "first embodiment", "a further embodiment", "an alternate embodiment", "one embodiment", "an embodiment", "multiple embodiments", "some embodiments", "other embodiments", "further embodiment", "furthermore embodiment", "additional embodiment" or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, or in the context of more than one embodiment, or in the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.
[0029] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.
[0030] The terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by "comprises... a" does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.
[0031] A detailed methodology is explained in the following paragraphs of the disclosure.
[0032] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.
[0033] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit "1' are shown at least in Fig. 1. Similarly, reference numerals starting with digit "2" are shown at least in Fig. 2.
[0034] FIG. 1 is a pictorial diagram depicting a plurality of feature points detected within an exemplary image captured using one of the multiple cameras of the HMD device, according to an existing art.
[0035] In the implementation of the SLAM technique, the real-time map of the surrounding environment is generated from visual data obtained from the HMD device by detecting and matching feature points within multiple images captured using the multiple cameras. A feature point in the present context refers to a distinct and identifiable point in an image that is expressive in texture. For instance, a pixel may be identified a corner (i.e., a feature point) when a distinct intensity change occurs across a contiguous section of the surrounding pixels. For example, in an image 101 of a surrounding as depicted in FIG. 1 and captured using one of the multiple cameras of the HMD device, a plurality of feature points 103a through 103d may be detected using one or more predetermined feature point detection techniques, such as Features from Accelerated Segment Test (FAST). However, the present disclosure is not limited thereto, and the plurality of feature points may also be detected using different feature point detection algorithm applied to the image 101.
[0036] In an embodiment of the disclosure, a first feature point 103a and a second feature point 103b may be identified as corner points of a frame included in the image 101, respectively. A third feature point 103c and a fourth feature point 103d may be identified as points clearly distinguishable from the background or other objects as the boundary of the armrests of a chair included in the image 101, respectively. However, the disclosure is not limited thereto, and additional feature points may also be detected in the image 101, in addition to the plurality of feature points in FIG. 1.
[0037] FIG. 2 is a pictorial diagram depicting capturing images of the scene using two cameras from different points of view, according to an existing art.
[0038] The detection and matching of feature points is further used for relative pose calculation. The relative pose may be understood as a transformation between two camera positions. That is, the relative pose may be defined by rotation and translation between two camera positions. For instance, as depicted in FIG. 2, two cameras, camera 0 and camera 1 are used to capture images of the same scene from different points of view resulting in images depicting different views of the same scene.
[0039] A relative pose, such as the rotation and translation between two camera positions, may be calculated by detecting a plurality of feature points from two images representing different points of view of the same scene and matching the corresponding feature points.
[0040] Further, as an analogy, consider traveling from location A to location B. When finding the way to location B, different landmarks on the way are noted down. Thereafter, when returning to location A or if the same is traveled again, the same landmarks are found to determine if the correct route is being followed. Thus, landmarks help in navigation.
[0041] In a similar manner, the SLAM technique operates by performing localization and mapping processes through the identification and storage of multiple visual landmarks along with corresponding representations, i.e., visual description of the landmark. The corresponding visual description of each landmark is used for recognizing and differentiating different landmarks from one another. This enables keeping track of the different landmarks which helps in navigating through the map of the surroundings. Such implementation of the SLAM technique in utilized in various real-life scenarios.
[0042] For example, consider a user wearing an HMD device while working at a desk and watching a video. Further, consider when the user goes away from the desk, the video is stopped and is resumed when the user returns to the desk. In said scenario, the understanding of the return of the user to the desk is improved through the localization and mapping of specific landmarks, such as the desk, using the SLAM technique.
[0043] However, multiple cameras are prone to synchronization issues due at least one of various factors such as technical malfunction, natural environmental factors, connectivity issues, power interruptions, hardware wear and tear, interference, and configuration errors. The synchronization issues result in inaccurate head-tracking causing discrepancies in virtual reality rending and interactions therewith. Such discrepancies ultimately degrade the immersive experience provided by the HMD device. Said synchronization issues and discrepancies associated therewith are described below in conjunction with FIG. 3A- FIG. 3C
[0044] FIG. 3A is a schematic diagram depicting an ideal scenario of synchronized cameras. FIG. 3B is a schematic diagram depicting a scenario of rotation and transformation error. FIG. 3C is a schematic diagram depicting a scenario of synchronization error. FIG. 3A, FIG. 3B, and FIG. 3C are described in conjunction with each other for the sake of brevity and ease of reference.
[0045] Referring to FIG. 3A, a feature point p on the image plane 301 associated with the camera 0, corresponds to a landmark X in the three-dimensional (3D) world. Further, the landmark X can be found in the image plane 303 associated with camera 1 as feature point q'. The feature point q' is obtained as a reprojection of the feature point p which is a 3D projection of the landmark X. Said reprojection to obtain feature point q' is performed based on a known rotation (R) and transformation (T) between the two cameras camera 0 and camera 1. Additionally, a feature point q corresponding to the feature point p can also be found on the image plane 301 associated with the camera 1 using a predetermined feature point detection technique. Any distance between the point q and the point q' is known as reprojection error (RPE).
[0046] Initially, the feature point p on the image plane 301 is found using the predetermined feature point detection technique. Further, the corresponding landmark X to is found using appearance matching of the feature point p and the feature point q with an available list of 3D landmarks within the surrounding environment that have already been identified in previous iterations of SLAM. Thereafter, the corresponding landmark X is projected to image plane 303 based on the known R and T between the camera 0 and the camera 1 to obtain the projected feature point q'. Further, the feature point q on the image plane 303 corresponding to the feature point p is found using the predetermined feature point detection technique.
[0047] As depicted in FIG. 3A, in an ideal scenario, when the known R and T between the camera 0 and the camera 1 is accurate with no synchronization issues, the feature point q and the projected feature point q' are the same. In an ideal scenario, synchronization issues do not occur.
[0048] Referring to FIG. 3B, in a scenario, when the R and T are inaccurate, the image plane of the camera 1 is calculated to be in a new and inaccurate location, represented as an image plane 303' which is incorrect. Further, the projection of the landmark X to obtain the feature point q' is performed using the inaccurate R and T. Consequently, the feature point q' is formed at an incorrect location with respect to the correct position of the image plane 301. Thus, since the location of the feature point q' is erroneous, the feature point p and the feature point q' are not formed at the same location in the image of the surrounding environment. This results from the RPE which, in the present scenario, is equivalent to a distance between the ideal feature point q and the erroneously projected feature point q'. Even if there is no synchronization issue, the ideal feature q and the incorrectly projected feature point q' may not match if the R and T between the camera 0 and the camera 1 are inaccurate. In such a case, this problem can be resolved by correcting the inaccurate R and T between the camera 0 and the camera 1.
[0049] Referring to FIG. 3C, depicted is another scenario, when the R and T are accurate, but the synchronization error is present between the camera 0 and the camera 1. In the current implementation of the SLAM technique, the multiple cameras of the HMD device are synchronized using timestamps. In the figure, the image plane 301 and image plane 303 represent the positions of the camera 0 and the camera 1 at time instance T0. Further, the image plane 301' and the image plane 303' represent the position of the camera 0 and the camera 1 at time instance T1. In an ideal case, both the cameras click the image at the same time instance i.e., T0. However, when there is synchronization error, the camera 0 may click at the time instance T0 while the camera 1 may click at the time instance T1 when the location of the camera 1 is changed as represented by image plane 303'. Thus, by the time camera 1 clicks the image, the device would have moved in some direction resulting in an incorrect pair of image plane 301 and image plane 303'.
[0050] Further, since the camera 1 has moved, the projected feature point q' is also formed at an incorrect location. Said location would have been correct for the camera 1 when both the camera 0 and the camera 1 would have clicked the image at the same time instance T0 to obtain a correct pair of image plane 301 and image plane 303. Further, in the present scenario, the location of the feature point q is found in the image plane 303' obtained at time instance T1. Consequently, the projected feature point q' and the feature point q are not formed at the same location resulting in the RPE equivalent to the distance between the ideal feature point q and the erroneously projected feature point q' in the present scenario.
[0051] FIG. 4 is a pictorial diagram illustrating discrepancies resulting from the synchronization errors between two cameras, according to an existing art.
[0052] The above-mentioned discrepancies resulting from the synchronization issues is further illustrated in FIG. 4 with respect to the image 101 of the surrounding depicted in FIG. 1. As shown in FIG. 4, the image frame 101 is captured using the camera 0 at time instance T0 and the feature points 103a through 103d are found using the predetermined feature point detection technique. Further, the image frame 401 is captured using camera 1 at time instance T1. Therefore, the feature points 103a through 103d are translated to erroneous feature points 403a through 403d due to the synchronization error between the camera 0 and the camera 1.
[0053] Therefore, in view of the above-mentioned problems, it is advantageous to provide an improved system and method that can overcome the above-mentioned problems and limitations associated with the synchronization of multiple cameras and virtual reality (VR) rendering in the HMD devices. The present disclosure provides a method and system to overcome synchronization errors between such cameras.
[0054] FIG. 5 is a block diagram depicting a head mounted display (HMD) device 500 implementing a system 501 for synchronizing camera feed of two or more cameras 503 of the HMD device 500, according to embodiments of the present disclosure. The HMD device 500 may correspond to a device which enables immersive experiences by providing users with an interactive view of virtual and physical environments. In an example, the HMD device 500 may include various technologies, such as, but not limited to, virtual reality (VR), augmented reality (AR), mixed reality (MR), and extended reality (XR).
[0055] According to one or more embodiments of the present disclosure, the system 501 may include the two or more cameras 503, a processor 505, a memory 507, an artificial intelligence (AI) model 509, an inertial measurement unit (IMU) 511, and a plurality of modules 513.
[0056] According to one or more embodiments of the present disclosure, the two or more cameras 503 may be outward-facing cameras configured to capture images of a surrounding environment of a user wearing the HMD 500. In an embodiment, the two or more cameras 503 may be stereo camera pairs mounted in fixed positions relative to each other on the front end of the HMD 500. Further, the stereo camera pairs may be configured to operate in synchronization enabling simultaneous capturing of images by each camera of the two or more cameras 503.
[0057] Furthermore, the two or more cameras 503 may find the depth of a point in three dimension (3D) using stereo matching. Further, corresponding images captured by the stereo pairs are aligned using stereo rectification. The stereo rectification is the process of transforming a point in a corresponding left image captured by a left camera of the stereo pairs of the two or more cameras 503. Said transformation projects the point to a line on a corresponding right image captured by a right camera of the stereo pairs. The stereo rectification process is performed to ensure that corresponding points in the image pairs corresponding to the stereo pairs of the two or more cameras 503 are aligned and lie on the same corresponding epipolar lines.
[0058] According to one or more embodiments of the present disclosure, the two or more cameras 503 may include a first camera 503a and a second camera 503b. In an embodiment, the first camera 503a may be considered as a reference camera and an image frame captured by the first camera 503a at a current timestamp may be referred to as a reference image frames for the current timestamp. In an embodiment, the second camera 503b may refer to a group of one or more cameras which may be considered as secondary cameras to the reference camera. Further, the image frames captured by the second camera 503b may be referred to as secondary image frames. In an embodiment, the current timestamp may refer to a present frame in which the image is captured through the camera.
[0059] According to one or more embodiments of the present disclosure, when a synchronization error is present among the two or more cameras 503 of the HMD device 500, the second camera 503b may lag with respect to the first camera 503a. Further, the secondary image frames captured by the second camera 503b at the current timestamp may not align with the reference image frame captured by the first camera 503a at the current timestamp. The synchronization error may lead to erroneous virtual reality (VR) rendering using the HMD device 500. Therefore, an objective of the present disclosure is to provide a method, implemented by the system 501, for synchronizing the camera feed of the two or more cameras 503 of the HMD device 500.
[0060] In an example, the processor 505 may be a single processing unit or a number of units, all of which could include multiple computing units. The processor 505 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logical processors, virtual processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 505 is configured to fetch and execute computer-readable instructions and data stored in the memory 507.
[0061] The memory 507 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory or Random Access Memory (RAM), such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.
[0062] In an embodiment of the disclosure, instructions, a data structure, and a program code readable by the at least one processor 505 may be stored in the memory 507. According to an embodiment of the disclosure, the memory 507 may be at least one memory. According to disclosed embodiment, operations performed by the at least one processor 505 may be implemented by executing the instructions or codes of a program stored in the memory 507. According to an embodiment of the disclosure, the memory 507 may not exist separately but may be included in the at least one processor 505. The memory 507 may store instructions or program codes for performing functions or operations of the HMD device 500. The instructions, algorithm, data structure, program code, and application program stored in the memory 507 may be implemented in, for example, programming or scripting languages such as C, C++, Java, python, assembler, and the like.
[0063] At least one of a plurality of operations of the system 501 may be implemented through the AI model 509. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor 505. In an embodiment, the AI model 509 may be included in the memory 507.
[0064] The processor 505 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).
[0065] According to an embodiment of the disclosure, the at least one processor 505 may include circuitry such as a system on chip (SoC) or an integrated circuit (IC). According to an embodiment of the disclosure, the at least one processor 505 may execute various types of modules stored in the memory 507. The at least one processor 505 may execute at least one instruction that constitutes the various types of modules stored in the memory 507. By executing a program or at least one instruction stored in the memory 507, the at least one processor 505 may process data according to predefined operation rules or an AI model.
[0066] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or AI model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0067] Here, being provided through learning means that, by applying a learning technique to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in the user device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.
[0068] According to one or more embodiments of the present disclosure, the pre-trained AI model 509 (hereinafter interchangeably referred as 'AI model') refers to an AI model pre-trained to generate a virtual image frame to be generated by a camera of the two or more cameras 503. In an example, the pre-trained AI model 509 may be a generative AI model. According to one or more embodiments of the present disclosure, the operations of the pre-trained AI model 509 are triggered using a prediction module 513a of the plurality of modules 513 as described below in greater details in the forthcoming paragraphs.
[0069] The AI model 509 may consist of a plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through the calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
[0070] The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to decide or predict. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0071] As mentioned above, the AI model 509 may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic AI model with multiple pieces of training dataset by a training technique. A function associated with the AI model 509 may be performed through the non-volatile memory, the volatile memory, and the processor.
[0072] The IMU 511 may be used to measure the acceleration, orientation, and rotation of the user wearing the HMD device 500. The IMU 511 may comprise a combination of accelerometers, gyroscopes, and magnetometers coupled with each other to provide information associated with the position and orientation of the user head in 3D space.
[0073] As an example, the plurality of modules 513 may include a program, a subroutine, a portion of a program, a software component, or a hardware component capable of performing a stated task or function. As used herein, the plurality of modules 513 may be implemented on a hardware component such as a server independently of other modules, or a module can exist with other modules on the same server, or within the same program. The plurality of modules 513 may be implemented on a hardware component, such as processor one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. The plurality of modules 513 when executed by the processor(s) 505 may be configured to perform any of the functionalities discussed herein.
[0074] In an embodiment, the plurality of modules 513 may be implemented using the AI model 509 which may include a plurality of neural network layers. At least one of the plurality of modules 513 may be implemented to thereby achieve execution of the present subject matter's mechanism through an AI model.
[0075] The plurality of modules 513 may include a set of instructions that may be executed according to the embodiments of the present disclosure to synchronize camera feed of two or more cameras 503 of the HMD device 500. The plurality of modules may include the prediction module 513a, a determination module 513b, an estimation module 513c, a transformation module 513d, and a synchronization module 513e. In an embodiment, the plurality of modules 513 may be included in the memory 507. The 'module' included in the memory 507 may be implemented as software, such as instructions, an algorithm, a data structure, or program code.
[0076] According to one or more embodiments, the prediction module 513a may be configured to generate a virtual second image frame from the second camera upon receiving a first image frame captured using a first camera and a second image frame captured using a second camera. According to one or more embodiments, the received first image frame and the second image frame may be captured at the current timestamp. Further, the determination module 513b may be configured to determine one or more significant feature points in the virtual second image frame and corresponding significant feature points in the second image frame.
[0077] Further, the estimation module 513c may be configured to estimate transformation parameters associated with the one or more significant feature points. Furthermore, the transformation module 513d may be configured to transform each feature in the second image frame by applying the estimated transformation parameters. Moreover, the synchronization module 513e may be configured to synchronize the first camera 503a and the second camera 503b according to one or more embodiments of the present disclosure. The plurality of modules 513 are described in detail in the forthcoming paragraphs in conjunction with FIG. 6A and 6B, and FIGS. 7A and 7B.
[0078] FIG. 6A is a block diagram depicting an operational flow of the plurality of modules 513 for synchronizing the camera feed of the two or more cameras 503 of the HMD device 500, according to one or more embodiments of the present disclosure. FIG. 6B illustrates an operational flow associated with training pipeline of the AI model 509 for synchronizing the camera feed of the two or more cameras 503 of the HMD device 500, according to one or more embodiments of the present disclosure. FIG. 7A is a pictorial diagram illustrating the operational flow the plurality of modules 513 using image frames, according to one or more embodiments of the present disclosure. FIG. 7B is a pictorial diagram illustrating the operational flow the plurality of modules 513 using image frames, according to one or more embodiments of the disclosure. FIGS. 6A and 6B are described in conjunction with the FIGS. 7A and 7B for the sake of brevity and ease of reference.
[0079] Referring to FIG. 6A, at operation 601, the processor 505 receives, at the current timestamp, camera feed from the plurality of cameras 503 of the HMD device 500. According to one or more embodiments of the present disclosure, the camera feed may comprise a first image frame received at the current timestamp and a second image frame received at the current timestamp. Further, the first image frame may be captured using a first camera 503a while the second image frame may be captured using the second camera 503b. According to one or more embodiments, the first image frame may be referred to as an actual first image frame. Also, the second image frame may be referred to as an actual second image frame. According to one or more embodiments, the camera feed may also comprise one or more previously received image frames captured using the plurality of cameras 503.
[0080] Referring to FIG. 7A, in an example, a first image frame 701 may correspond to the actual first image frame and a second image frame 703 may correspond to the actual second image frame. Furthermore, a group of one or more image frames 705 may correspond to the one or more previously received image frames captured using the plurality of cameras 503. The group of one or more image frames 705 may be a set of images captured by the first camera 503a and the second camera 503b, synchronized in one or more previous frames. In an embodiment, the previous frames may refer to a frame prior to the current frame corresponding to the current timestamp.
[0081] In an example, the second camera 503b may be lagging with respect to the first camera 503a due to the synchronization error among the two or more cameras 503. According to one or more embodiments of the present disclosure, the first image frame 701 may be considered as the reference image frame at the current timestamp and the second image frame 703 may be considered as the secondary image frame at the current timestamp. Further, the secondary image frame may not align with the reference image frame due to the synchronization error.
[0082] Referring to FIG. 6A, at operation 602, the prediction module 513a, predicts a virtual second image frame from the second camera at the current timestamp. According to one or more embodiments, the virtual second image frame may correspond to an image frame that would have been captured by the second camera 503b at the current timestamp in an ideal scenario where the first camera 503a and the second camera 503b are synchronized.
[0083] According to one or more embodiments, the prediction module 513a may be configured to generate the virtual second image frame using the pre-trained AI model 509 based on the first image frame, the second image frame, and the one or more previously received image frames. Referring to FIG. 7A, the image frame 707 may correspond to the virtual second image frame predicted using the pre-trained AI model 509.
[0084] Further, the virtual second image frame may be generated, using the pre-trained AI model 509, with corresponding confidence score for each pixel. The confidence of each pixel is influenced by a contrast of each pixel against the surrounding environment, visibility of each pixel in the image frame, and a consistency associated with observation of each pixel. Table 1 below is a representation of difference in characteristics of high confidence pixels and low confidence pixels.
[0085] According to one or more embodiments, the AI model 509 may be pre-trained AI model configured to predict a confidence score for each of a plurality of pixels included in the virtual second image frame. The AI model may be trained to predict the confidence score for each pixel either by using image frames repeatedly generated through multi-sampling, or based on differences from a ground truth during the training process.
[0086]
[0087] Table 1
[0088] Referring to FIG. 7A, the high confidence pixels are represented by blocks 709 in the virtual image frame 707 generated using the AI model 509. In an embodiment, the blocks 709 may refer to a region that includes a set of pixels with high confidence.
[0089] Referring to FIG. 6A, at operation 603, the determination module 513b determines one or more significant feature points in the virtual second image frame and corresponding significant feature points in the second image frame. According to one or more embodiments, a plurality of feature points may be determined using a predetermined feature point detection technique. Examples of the predetermined feature point detection technique may include but are not limited to, features from accelerated segment test (FAST), Harris corner detection, Shi-Tomasi corner detection, scale-invariant feature transform (SIFT), and speeded-up robust features (SURF).
[0090] Further, the plurality of feature points include the one or more significant feature points, and one or more non-significant feature points. According to one or more embodiments, the determination module 513b may be configured to determine the one or more significant feature points, among the plurality of feature points, corresponding to one or more pixels with a confidence score equal to or greater than a predefined threshold associated with the accuracy of each pixel in a generated image frame. Furthermore, the determination module 513b may determine the one or more non-significant feature points, among the plurality of feature points, corresponding to one or more pixels with a confidence score less than the predefined threshold associated with accuracy of each pixel in a generated image frame.
[0091] However, the disclosure is not limited thereto, and the determination module 513b may be configured to perform feature point detection within the block 709 to detect one or more significant feature points. The plurality of significant feature points may be detected from one or more high-confidence pixels included in the block 709. It should be noted, however, that not all of the one or more high-confidence pixels are detected as significant feature points.
[0092] In FIG. 7A, one or more significant feature points 711 determined in the virtual second image frame 707 by the determination module 513b are illustrated. Additionally, in FIG. 7A, one or more significant feature points 713 determined in the second image frame 703 corresponding to the second image frame by the determination module 713b.
[0093] Further, the estimation module 513c may be configured to estimate transformation parameters associated with the one or more significant feature points. In an embodiment, the transformation parameters may refer to a transformation matrix. Referring to FIG. 6, at operation 604, the estimation module 513c matched the one or more significant feature points 711 in the virtual second image frame 707 and the corresponding one or more significant feature points 713 in the second image frame 703. In an example, as shown in FIG. 7A, the one or more significant feature points 711 in the virtual second image frame 707 is matched with the corresponding one or more significant feature points 713 in the second image frame 703.
[0094] Referring to FIG. 6A, in response to matching the one or more significant feature points 711 in the virtual second image frame 707 and the corresponding one or more significant feature points 713 in the second image frame 703, the estimation module 513c, at operation 605, estimates transformation parameters associated with the one or more significant feature points. In FIG. 7A, the feature matching between the one or more significant feature points 711 in the virtual second image frame 707 and the corresponding one or more significant feature points in the second image frame 703 result in estimation of the transformation parameters 605. The estimation module 513c may calculate a transformation parameters representing a geometric relationship between one or more significant feature points 711 in the virtual second image frame 707 and one or more significant feature points 713 in the second image frame 703.
[0095] In an embodiment, the transformation parameters 605 may be homographic transformation parameters. In an embodiment, the estimation module 513c may obtain homographic transformation parameters using an algorithm such as Direct Linear Transformation (DLT) or RANdom Sample Consensus (RANSAC). In an embodiment, the homographic transformation parameters may be a transformation matrix that may be used to estimate the coordinate information of one or more significant feature points 711 in the virtual second image frame 707 by performing matrix multiplication with the coordinate information of one or more significant feature points 713 included in the second image frame 703.
[0096] Further, the transformation module 513d may be configured to transform the second image frame such that the transformed second image frame is synchronized with the first image frame. As shown in FIG. 6A, at operation 606, the transformation module 513d applies the estimated transformation parameters to the second image frame. Further, at operation 607, the transformation module 513d obtains the transformed second image frame upon application of the estimated transformation parameters to the second image frame. In an embodiment, the transformed second image frame is obtained by applying the homographic transformation parameters to the second image frame.
[0097] Further, as shown in FIG. 7B, the transformed second image frame 717 is obtained by transforming each a plurality of feature points in the second image frame 703. In an embodiment, the transformation module 513d may apply the transformation parameters to transform both the significant feature points and the non-significant feature points in the second image frame 703 to generate the transformed second image frame 717. In an example, the plurality of feature points in the second image frame 703 are represented by the block 715, and a plurality of feature points in the transformed second image frame 717 are represented by the block 715'. According to one or more embodiments, each of the plurality of feature points in the second image frame 703 is transformed to corresponding the plurality of feature points to obtain the transformed second image frame 717.
[0098] According to one or more embodiments, since the transformation parameters are obtained based on the significant feature points determined from high-confidence score pixels, it is possible to acquire highly accurate transformation parameters even when using the virtual second image frame 707 generated by the AI model. In addition, by synchronizing the first camera and the second camera using the transformed second image frame 717, obtained by transforming the second image frame 703 based on the highly accurate transformation parameters, the synchronization accuracy can also be improved.
[0099] According to one or more embodiments, the synchronization module 513e obtains inertial data associated with an inertial measurement unit (IMU) sensor 511 in the HMD device 500. According to one or more embodiments, the obtained inertial data may include, but is not limited to, acceleration data, rotation data, and orientation data associated with the IMU sensor 511. Referring to FIG. 6A, at operation 608, the synchronization module 513e may determine time delta between the second image frame and the first image frame based on the inertial data and the transformation parameters.
[0100] Further, the synchronization module 513e synchronizes the two or more cameras 503 of the HMD device 500 by adjusting one or more camera parameters of the second camera 503b. According to one or more embodiments, the one or more camera parameters of the second camera 503b are adjusted when the time delta is greater than or equals to a predefined synchronization error threshold. In an example, the one or more camera parameters include, but are not limited to, camera exposure time, frame rate, rolling shutter speed, auto-focus time, and post-processing time. In an example, adjusting the one or more camera parameters may comprise adjusting the exposure time of the second camera 503b as depicted in FIG. 8.
[0101] Referring to FIG. 6B, an operational flow associated with training of a combination of two conditional diffusion models 616a and 616b for generating the virtual secondary camera frame 618 is disclosed. The diffusion models 616a and 616b may collectively correspond to a generative AI model 617, which may refer to the AI model 509 as explained above. To prepare a first diffusion model input 610, a Gaussian noise to a secondary camera frame 612 may be added. The first diffusion model input 610 may then be provided to the first conditional diffusion AI model 616a. Further, the first conditional diffusion AI model 616a may also receive a primary / reference camera frame 613 along with calibration parameters 614 associated with the camera, such as camera extrinsic parameters (rotation, translation), camera intrinsic parameters, etc. as conditioning input.
[0102] In response to receiving the diffusion input 610, the reference camera frame 613, and the camera calibration parameters 614, the first conditional diffusion AI model 616a outputs an intermediate embedding vector which includes current spatial information of stereo camera setup and current view information. The output from the first conditional diffusion AI model 616a is provided as a primary input to the second conditional diffusion AI model 616b. Further, one or more recently sampled frames 619 from the prior secondary camera frames 615 and prior calculated poses 620 (based on one or more saved poses 621) are provided as a conditioning inputs to the second conditional diffusion AI model 616b for scene information and spatial information of the device in the environment associated with the camera. In response, the second conditional diffusion AI model 616b generates the virtual second image frame 618.
[0103] Further, the final synced camera frames are sent to a differentiable SLAM model which estimates the pose using the synced camera data. This estimated pose is compared with the ground truth poses for loss computation. The loss is computed as the log loss of relative pose error between estimated pose and the ground truth pose. This loss is propagated back for the joint optimization of the diffusion models 616a and 616b.
[0104] However, the present disclosure is not limited thereto, and the AI model may include various types of generative artificial intelligence models capable of generating the virtual secondary frame 618.
[0105] FIG. 8 is a schematic diagram depicting an implementation of synchronization module 513e, according to one or more embodiments of the present disclosure. As shown in the figure, a second camera 503b is lagging with respect to the first camera 503a due to synchronization error between the first camera 503a and the second camera 503b. Consequently, when the first camera 503a and the second camera 503b each capture an image at the current timestamp, due to the synchronization error, the camera exposure timestamps of the first camera 503a and the second camera 503b may differ. That is, when the first camera 503a captures the first image frame during the camera exposure timestamps T3-T0, the second camera 503b captures the second image frame during the camera exposure timestamps T4-T1 due to the synchronization error as depicted at block 801.
[0106] In response to the delay in capturing of the corresponding image frames by the first camera 503a and the second camera 503b, while the first image frame is obtained at the timestamp T5, the second image frame is obtained at timestamp T6. According to embodiments of the present disclosure, the time delta between the second image frame and the first image frame is determined to be timestamp duration T6-T5.
[0107] Further, when the time delta is greater than or equals to the predefined synchronization error threshold, the synchronization module 513e may fix the synchronization error by adjusting the one or more camera parameters of the second camera 503b. In an example, the camera exposure time of the second camera 503 may be adjusted such that the second image frame is captured during timestamp duration T3-T1 such that the second image frame is obtained after processing at the same timestamp, i.e., T5, when the first image frame is obtained as depicted at block 803. This results in synchronization of the first camera 503a and the second camera 503b.
[0108] FIG. 9 is a block diagram depicting a method 900 for synchronizing camera feed of two or more cameras of the HMD device 500, according to one or more embodiments of the present disclosure. The method 900 includes a series of operations 901 through 908 executed by one or more components of the HMD device 500, in particular the processor 505.
[0109] At step 901, the processor 505 receives at a current timestamp, the first image frame captured using the first camera 503a and the second image frame captured using the second camera 503b.
[0110] At step 902, the processor 505 generates, using the pre-trained AI model 509, the virtual second image frame from the second camera 503b at the current timestamp. According to one or more embodiments, the processor 505 generates the virtual second image frame based on the first image frame, the second image frame, and the one or more previously received image frames. According to one or more embodiments, the processor 505 predicts the virtual second image frame with corresponding confidence score for each pixel.
[0111] At step 903, the processor 505 determines the one or more significant feature points in the virtual second image frame and corresponding significant feature points in the second image frame. According to one or more embodiments, the processor 505 determines the one or more significant feature points, among the plurality of feature points, corresponding to the one or more pixels with confidence score greater than the predefined threshold.
[0112] At step 904, the processor 505 estimates the transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame.
[0113] At step 905, the processor 505 transforms each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame. At step 905, in an embodiment, the processor 505 transforms the one or more feature points in the second image frame by applying the estimated transformation parameters.
[0114] At step 906, the processor 505 obtains inertial data associated with an inertial measurement unit (IMU) sensor in the HMD device 500.
[0115] At step 907, the processor 505 determines the time delta between the second image frame and the first image frame based on the inertial data and the transformation parameters.
[0116] At step 908, the processor 505 adjusts the one or more camera parameters of the second camera when the time delta is greater than or equals to the predefined synchronization error threshold.
[0117] At least by virtue of the aforesaid, the present subject matter at least provides the following advantages:
[0118] The systems and methods described herein enable improved stereo matching by using the first image frame and the transformed second image frame. The improved stereo matching leads to improved depth calculation which results improved identification of 3D landmarks. Finally, the improved 3D landmarks identification results in improved implementation of SLAM technique.
[0119] Further, the systems and methods described herein prevent various issues (e.g., reduced SLAM technique accuracy, degraded user pose estimation, inaccurate display position of virtual objects) that may arise due to synchronization problems between multiple cameras included in the HMD device, by ensuring that the multiple cameras included in the HMD device are synchronized with each other.
[0120] Further, the systems and methods described herein synchronize the two or more cameras by analyzing the camera data instead of synchronizing using the timestamps.
[0121] Further, the systems and methods described herein improve the synchronization of current visual data and prevent future occurrences of synchronization errors.
[0122] Further, the systems and methods described herein result in improved accuracy in estimation of user poses by ensuring the input to SLAM technique are of improved quality.
[0123] To address the above-described technical problems, an embodiment of the disclosure provides a method for synchronizing camera feed of two or more cameras of a head-mounted display (HMD) device. The method includes receiving at a current timestamp, a first image frame captured using a first camera, and a second image frame captured using a second camera. The method includes generating a virtual second image frame from the second camera at the current timestamp using a pre-trained artificial intelligence (AI) model. The method includes determining one or more significant feature points in the virtual second image frame and corresponding significant feature points in the second image frame. The method includes estimating transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame. The method includes transforming each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame.
[0124] According to an embodiment of the disclosure, the generating a virtual second image frame includes generating, using the pre-trained AI model, the virtual second image frame based on the first image frame, the second image frame, and one or more received image frames, which were synchronized in one or more previous frames.
[0125] According to an embodiment of the disclosure, the generating a virtual second image frames includes predicting, using the pre-trained AI model (509), the virtual second image frame with corresponding confidence score for each pixel included in the virtual second image.
[0126] According to an embodiment of the disclosure, the determining one or more significant feature points includes detecting a plurality of feature points in the virtual second image frame and a plurality of feature points in the second image frame. The determining one or more significant feature points includes determining the one or more significant feature points, among the plurality of feature points detected in the virtual second image frame, corresponding to the one or more pixels with confidence score greater than the predefined threshold. The determining one or more significant feature points includes determining the corresponding significant feature points, among the plurality of feature points detected in the second image frame.
[0127] According to an embodiment of the disclosure, the transforming each feature in the second image frame includes transforming the plurality of feature points detected in the second image frame by applying the estimated transformation parameters.
[0128] According to an embodiment of the disclosure, the method further includes obtaining inertial data associated with an inertial measurement unit (IMU) sensor in the HMD device. The method further includes determining a time delta between the second image frame and the first image frame based on the inertial data and the transformation parameters. The method further includes adjusting one or more camera parameters of the second camera when the time delta is greater than or equals to a predefined synchronization error threshold.
[0129] In this application, unless specifically stated otherwise, the use of the singular includes the plural, and the use of "or" means "and / or". Furthermore, the use of the terms "including" or "having" is not limiting. Any range described herein will be understood to include the endpoints and all values between the endpoints. Features of the disclosed embodiments may be combined, rearranged, omitted, etc., within the scope of the invention to produce additional embodiments. Furthermore, certain features may sometimes be used to advantage without a corresponding use of other features.
[0130] While at least one embodiment has been presented in the foregoing detailed description, it should be appreciated that a vast number of variations exist
Claims
1.A method (900) for synchronizing camera feed of two or more cameras of a head-mounted display (HMD) device (500), the method (900) comprising:receiving at a current timestamp, a first image frame captured using a first camera (503a), and a second image frame captured using a second camera (503b);generating, using a pre-trained artificial intelligence (AI) model (509), a virtual second image frame corresponding to the second camera at the current timestamp;determining one or more significant feature points, among a plurality of feature points, with a confidence score equal to or greater than a predefined threshold associated with the accuracy of each pixel in the virtual second image frame, and corresponding significant feature points in the second image frame;estimating transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame; andtransforming each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame.2.The method (900) as claimed in claim 1, wherein generating the virtual second image frame comprises:generating, using the pre-trained AI model (509), the virtual second image frame based on the first image frame, the second image frame, and one or more received image frames, which were synchronized in one or more previous frames.3.The method (900) as claimed in any one of claims 1 or 2, wherein the generating comprises predicting, using the pre-trained AI model (509), the virtual second image frame with corresponding confidence score for each pixel included in the virtual second image.4.The method (900) as claimed in claim 3, wherein the determining comprisesdetecting a plurality of feature points in the virtual second image frame and a plurality of feature points in the second image frame;determining the one or more significant feature points, among the plurality of feature points detected in the virtual second image frame, corresponding to the one or more pixels with confidence score greater than the predefined threshold; anddetermining the corresponding significant feature points, among the plurality of feature points detected in the second image frame.5.The method (900) as claimed in claim 4, wherein the transforming comprises transforming the plurality of feature points detected in the second image frame by applying the estimated transformation parameters.6.The method (900) as claimed in any one of claims 1 to 5, further comprising:obtaining inertial data associated with an inertial measurement unit (IMU) sensor in the HMD device (500);determining a time delta between the second image frame and the first image frame based on the inertial data and the transformation parameters; andadjusting one or more camera parameters of the second camera when the time delta is greater than or equals to a predefined synchronization error threshold.7.A system (501) for synchronizing camera feed of two or more cameras of a head-mounted device (HMD) device (500), the system (501) comprising:a memory (507) storing at least one instruction; andat least one processor (505), comprising a processing circuitry;wherein the at least one processor (505) is configured to individually or collectively execute the at least one instruction stored in the memory to:receive at a current timestamp, an first image frame captured using a first camera (503a), and an second image frame captured using a second camera (503b);generate, using a pre-trained artificial intelligence (AI) model (509) (509), an virtual second image frame from the second camera at the current timestamp;determine one or more significant feature points, among a plurality of feature points, with a confidence score equal to or greater than a predefined threshold associated with the accuracy of each pixel in the virtual second image frame and corresponding significant feature points in the second image frame;estimate transformation parameters associated with the one or more significant feature points by matching the one or more significant feature points in the virtual second image frame and the corresponding one or more significant feature points in the second image frame; andtransform each feature in the second image frame by applying the estimated transformation parameters such that the transformed second image frame is synchronized with the first image frame.8.The system (501) as claimed in claim 7, wherein for generating the virtual second image frame, the processor (505) is configured to:generate, using the pre-trained AI model (509), the virtual second image frame based on the first image frame, the second image frame, and one or more received image frames synchronized in previous frame.9.The system (501) as claimed in any one of claims 7 or 8, wherein for generating, the processor (505) is configured to predict, using the pre-trained AI model (509), the virtual second image frame with corresponding confidence score for each pixel included in the virtual second image.10.The system (501) as claimed in claim 9, wherein for determining, wherein the processor (505) is configured to:detect a a plurality of feature points in the virtual second image frame and a plurality of feature points in the second image frame;determine the one or more significant feature points, among the plurality of feature points detected in the virtual second image frame, corresponding to the one or more pixels with confidence score greater than the predefined threshold; anddetermine the corresponding significant feature points, among the plurality of feature points detected in the second image frame.11.The system (501) as claimed in claim 10, wherein for transforming, wherein the processor (505) is configured to the transforming comprises transforming the plurality of feature points detected in the second image frame by applying the estimated transformation parameters.12.The system (501) as claimed in any one of claims 7 to 11, wherein the processor (505) is further configured to:obtain inertial data associated with an inertial measurement unit (IMU) sensor in the HMD device (500);determine a time delta between the second image frame and the first image frame based on the inertial data and the transformation parameters; andadjust one or more camera parameters of the second camera when the time delta is greater than or equals to a predefined synchronization error threshold.13.A computer-readable recording medium having recorded thereon a computer program, which, when executed by a computer, performs the operation method of one or claims 1 through 7.
Citation Information
Patent Citations
A SLAM fusion method and system based on multiple camera types
CN110517216B
Video perspective method based on AR technology
CN114742977A
Self callbrating camera system
US10347009B1
Time synchronized cameras for multi-camera event videos
US20200396392A1
Customer-based video feed
US20210374428A1