User position recognition apparatus and method
The user location recognition device and method effectively address the limitations of existing technologies by using camera installation information to accurately determine a user's 3D space coordinates and direction, offering improved accuracy and activity zones through the combination of general and depth camera modes.
Patent Information
- Application Number
- PCT/KR2023/020772
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
Existing methods for recognizing a user's location in 3D space, such as eye tracking technology, are limited as they only provide visible views and lack depth information.
A device and method that utilize camera installation information to recognize a user's location in 3D space, employing a combination of first and second cameras with and without depth information to obtain accurate 3D space coordinates and direction of the user.
The solution accurately recognizes the user's position in 3D space, providing a wider activity zone when using a general camera mode and more precise head 6DoF results when using a depth camera mode.
Smart Images

Figure KR2023020772_19062025_PF_FP_ABST
Abstract
Description
User location recognition device and method
[0001] The present disclosure relates to a device and method for recognizing a user's location, and more particularly, to a device and method for recognizing a user's location in a three-dimensional space based on camera installation information.
[0002] To provide various services based on a user's location, it's necessary to know the user's location. For example, a service could be provided that places virtual objects in virtual reality or augmented reality based on the user's location and orientation.
[0003] Traditionally, eye tracking technology has been used to identify a user's location. However, this method has the disadvantage of tracking the user's viewpoint and only providing the view visible from that location.
[0004] The present disclosure is proposed to solve the above-mentioned problems and various problems related thereto, and the purpose of the present disclosure is to provide a user location recognition device and method that obtains location information of a user in 3D space based on camera installation information and provides various services based on the obtained user location information.
[0005] Another object of the present disclosure is to provide a user location recognition device and method that recognizes the 3D spatial coordinates and direction of a user's location based on camera installation information and provides various services based on the recognized 3D spatial coordinates and direction of the user.
[0006] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0007] In order to achieve the above-described purpose and other advantages, a user location recognition device including a display unit for displaying an image may be installed in the user location recognition device, and may include a camera unit including at least one of a first camera without depth information and a second camera including depth information, a storage unit for storing camera installation information of at least one of the first camera and the second camera, a user location recognition unit for obtaining location information of the user included in the image based on an image acquired by the camera unit, landmark information of the user included in the image, and the camera installation information, and a display control unit for controlling an image displayed on the display unit based on the location information of the user recognized by the user location recognition unit.
[0008] According to embodiments, the user's location information may include at least one of user head location information indicating 3D spatial coordinates of the user's head and user head direction information indicating the direction of the user's head.
[0009] According to embodiments, the camera installation information may include at least camera height information indicating a height from the floor to the camera, camera angle information indicating an angle of inclination of the camera, or field of view information of the camera.
[0010] According to embodiments, the user position recognition unit operates in a first mode for obtaining three-dimensional space coordinates of the user's head based on the user's landmark information in the image acquired by the first camera and the camera installation information, and in the first mode, an absolute distance from the user position recognition device to the user's feet is obtained, depth information is obtained as a relative position from the user's feet to the head based on the absolute distance, and two-dimensional coordinates of the user's head are obtained based on the depth information, and then the three-dimensional space coordinates of the user's head are obtained by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
[0011] According to embodiments, the user location recognition unit operates in a second mode for obtaining three-dimensional space coordinates of the user's head based on landmark information of the user in an image obtained by the first camera, depth information of an image obtained by the second camera, and the camera installation information, and in the second mode, two-dimensional coordinates of the user's head are obtained based on the depth information, and the three-dimensional space coordinates of the user's head are obtained by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
[0012] According to embodiments, the user location recognition unit can obtain the three-dimensional spatial coordinates of the user's head by aligning the image acquired by the second camera to the image acquired by the first camera in the second mode.
[0013] According to embodiments, the user position recognition unit may operate in the second mode to obtain three-dimensional spatial coordinates of the user's head when the user is within the field of view of the second camera, and may operate in the first mode to obtain three-dimensional spatial coordinates of the user's head when the user is outside the field of view of the second camera but within the field of view of the first camera.
[0014] According to embodiments, the field of view of the first camera may be wider than the field of view of the second camera.
[0015] According to embodiments, the user location recognition unit further includes an image control unit that provides an image acquired by the first camera to a landmark detection unit to obtain landmark information of the user, and the image control unit sets a first area and a second area in an image input from the first camera, the second area including the first area, and when there is no user in the first area of the image or the user leaves the second area, a portion of the remaining area excluding the first area of the image can be masked and provided.
[0016] According to embodiments, the image control unit can provide an original image without masking when a user is present in a first area of an input image.
[0017] According to embodiments, a method for recognizing a user location of a user location recognition device including a display unit for displaying an image, wherein at least one of a first camera without depth information and a second camera including depth information is installed, may include a step of acquiring location information of a user included in an image based on camera installation information of at least one of the first camera and the second camera, an image acquired by at least one of the first camera and the second camera, and landmark information of the user included in the image, and a step of controlling an image displayed on the display unit based on the location information of the user recognized by the user location recognition unit.
[0018] According to embodiments, the user's location information may include at least one of user head location information indicating 3D spatial coordinates of the user's head and user head direction information indicating the direction of the user's head.
[0019] According to embodiments, the camera installation information may include at least camera height information indicating a height from the floor to the camera, camera angle information indicating an angle of inclination of the camera, or field of view information of the camera.
[0020] According to embodiments, the user location recognition step may include a step of operating in a first mode to obtain three-dimensional spatial coordinates of the user's head based on the user's landmark information in the image acquired by the first camera and the camera installation information.
[0021] According to embodiments, the step of operating in the first mode may include the step of obtaining an absolute distance from a user position recognition device to the user's feet, the step of obtaining depth information as a relative position from the user's feet to the head based on the absolute distance, the step of obtaining two-dimensional coordinates of the user's head based on the depth information, and the step of obtaining three-dimensional space coordinates of the user's head by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
[0022] According to embodiments, the user location recognition step may include a step of operating in a second mode to obtain three-dimensional spatial coordinates of the user's head based on landmark information of the user in the image acquired by the first camera, depth information of the image acquired by the second camera, and camera installation information.
[0023] According to embodiments, the step of operating in the second mode may include the step of obtaining two-dimensional coordinates of the user's head based on the depth information, and the step of obtaining three-dimensional space coordinates of the user's head by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
[0024] According to embodiments, the step of operating in the second mode may include aligning an image acquired by the second camera to an image acquired by the first camera and then obtaining three-dimensional spatial coordinates of the user's head.
[0025] According to embodiments, the user position recognition step may operate in the second mode to obtain three-dimensional space coordinates of the user's head when the user is within the field of view of the second camera, and may operate in the first mode to obtain three-dimensional space coordinates of the user's head when the user is outside the field of view of the second camera but within the field of view of the first camera.
[0026] According to embodiments, the field of view of the first camera may be wider than the field of view of the second camera.
[0027] According to embodiments, the user location recognition step may further include an image control step of controlling an image acquired by the first camera to provide the image to a landmark detection unit to obtain user landmark information.
[0028] According to embodiments, the image control step sets a first area and a second area in an image input from the first camera, the second area includes the first area, and when there is no user in the first area of the image or the user leaves the second area, a portion of the remaining area excluding the first area of the image can be masked and provided.
[0029] According to embodiments, the image control step can provide an original image without masking when a user is present in a first area of an input image.
[0030] The device and method according to the embodiments can accurately recognize the user's position in a 3D space coordinate system by recognizing the distance between the camera and the user using camera-based image data and camera installation information, and obtaining the user's 3D space coordinates and / or direction based on the information.
[0031] The device and method according to the embodiments have the effect of operating in a general camera mode or a depth camera mode depending on the user's location when recognizing the user's location using information from a general camera without depth information and information from a depth camera with depth information, thereby obtaining a wider activity zone than in the depth camera mode when operating in the general camera mode, and obtaining a Head 6DoF result more precise than in the general camera mode when operating in the depth camera mode.
[0032] The device and method according to the embodiments can selectively recognize and control the user by utilizing the user's 3D space coordinates and / or direction information for user recognition.
[0033] In addition to the technical effects explicitly mentioned above, effects that are obvious to those skilled in the art can also be inferred from the entire description of this specification.
[0034] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments.
[0035] FIG. 1 is a drawing showing an example of a user location recognition device according to embodiments.
[0036] FIG. 2(a) and FIG. 2(b) are drawings showing examples of camera installation information and foot landmarks according to embodiments.
[0037] FIG. 3 is a drawing showing an example of a camera projection model according to embodiments.
[0038] FIG. 4 is a flowchart showing an example of obtaining 3D spatial coordinates of a user's head according to embodiments.
[0039] FIG. 5(a) and FIG. 5(b) are drawings showing an example of obtaining the absolute distance from the TV plane to the user's feet according to embodiments.
[0040] FIG. 6(a) to FIG. 6(c) are drawings showing an example of obtaining Head Z as a relative position from a user's feet to the head according to embodiments.
[0041] FIG. 7 is a drawing showing an example of converting a coordinate system of a relative position from a user's feet to his / her head according to embodiments.
[0042] Fig. 8 is a drawing showing an example of obtaining Head X, Y, and Z of the world coordinate system according to embodiments.
[0043] FIGS. 9(a) to 9(c) are drawings showing examples of a process for recognizing the position of a user's head in a second mode according to embodiments.
[0044] Figures 10(a) and 10(b) are drawings showing examples of operation in the third mode according to embodiments.
[0045] Fig. 11 is a flowchart showing an example of obtaining the direction of a user's head in a first mode according to embodiments.
[0046] Figures 12(a) to 12(d) are drawings showing specific examples of obtaining the direction of the user's head in the first mode.
[0047] FIGS. 13(a) to 13(e) are drawings showing examples of image control methods according to embodiments.
[0048] FIGS. 14(a) to 14(f) are drawings showing other examples of image control methods according to embodiments.
[0049] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be assigned the same reference numbers, and redundant descriptions thereof will be omitted. It should be noted that the following embodiments are intended only to concretize the present disclosure and do not limit or restrict the scope of the present disclosure. Anything that a specialist in the technical field to which the present disclosure pertains can easily infer from the detailed description and embodiments of the present disclosure is interpreted as falling within the scope of the present document.
[0050] The detailed description in this document should not be construed in any way as restrictive, but rather as illustrative. The scope of this document should be determined by a reasonable interpretation of the appended claims, and all changes within the scope of equivalents herein are intended to be included within the scope of this document.
[0051] Meanwhile, the blocks of the attached block diagram or the steps of the flowchart may be implemented directly in hardware, implemented as software modules executed by hardware, or implemented by a combination of these. In addition, the blocks of the attached block diagram or the steps of the flowchart may be interpreted as computer program instructions that are loaded into the processor or memory of a data processing device such as a general-purpose computer, a special-purpose computer, a portable notebook computer, or a network computer to perform designated functions. Since these computer program instructions may be stored in a memory provided in a computer device or a computer-readable memory, the functions described in the blocks of the block diagram or the steps of the flowchart may be produced as a product that includes a command means for performing the same. In addition, each block or each step may represent a module, segment, or part of code that includes one or more executable instructions for performing a specific logical function(s). Furthermore, in some alternative embodiments, the functions mentioned in the blocks or steps may be executed out of the specified order. For example, two blocks or steps shown in succession may be performed substantially simultaneously, or may be performed in reverse order, or in some cases, some blocks or steps may be performed with some blocks or steps omitted.
[0052] The present disclosure is intended to recognize a user's location in 3D space based on camera installation information and to provide various services based on the recognized user's location.
[0053] In the present disclosure, the user location can be recognized based on position information of the user's head and / or orientation information of the user's head.
[0054] In the present disclosure, position information of the user's head can be expressed as 3D space coordinates such as X, Y, and Z, and direction information of the user's head can be expressed as rotation information such as Roll, Pitch, and Yaw. For example, 3D space coordinates such as X (i.e., horizontal), Y (i.e., vertical), and Z (i.e., height) can indicate where the user's head is located in 3D space. In addition, rotation information such as Roll, Pitch, and Yaw can indicate in which direction the user's head is rotating. That is, Roll represents the angle at which the user's head rotates left and right, Pitch represents the angle at which the user's head moves up and down, and Yaw represents the angle at which the user's head turns left and right. For example, if the X, Y, and Z values are (2, 3, 1), this means that the user's head moves 2 units in the X axis, 3 units in the Y axis, and 1 unit in the Z axis. Also, if the Roll, Pitch, and Yaw values are (30 degrees, 20 degrees, and 45 degrees), this means that the user's head has rotated 30 degrees forward and backward, 20 degrees up and down, and 45 degrees left and right.
[0055] In this disclosure, the translational degrees of freedom for the three translational directions (X, Y, Z) are referred to as positional (X, Y, Z) 3DoF, and the rotational degrees of freedom for the three rotational directions (Roll, Pitch, Yaw) are referred to as orientational (Roll, Pitch, Yaw) 3DoF.
[0056] The present disclosure can obtain 6DOF (Degrees Of Freedom) based on the position information of the user's head and the direction information of the user's head. That is, 6DOF is the sum of the degrees of freedom for movement in three movement directions (X, Y, Z) and the degrees of freedom for rotation in three rotation directions (Roll, Pitch, Yaw), and represents the degrees of freedom in 3D space based on three axes. For example, the degrees of freedom for movement represent the degrees to which the user's head can move forward and backward, left and right, and up and down, and the degrees of freedom for rotation represent the degrees to which the user's head can rotate around the horizontal and vertical axes. In the present disclosure, the degrees of freedom for movement can be used to accurately track the user's position in 3D space, and the degrees of freedom for rotation can be used to track the user's gaze direction or head rotation, etc.
[0057] According to embodiments, the present disclosure can recognize a user's location using information from a general camera. In the present disclosure, a general camera is a camera that cannot measure the depth of an object, and in the present disclosure, an RGB camera is used as an example. This is an example to help those skilled in the art understand, and any camera that cannot measure depth can be used. That is, the present disclosure can recognize the user's location in 3D space coordinates by recognizing the absolute distance (Z) using only an RGB camera based on the camera's installation information. In the present disclosure, the mode that recognizes the user's location using information from a general camera is referred to as the RGB mode. In addition, the RGB mode may be referred to as a general mode, a general camera mode, or a first mode.
[0058] According to embodiments, the present disclosure can recognize the location of a user using information from a depth camera. In the present disclosure, the depth camera is a camera capable of measuring the depth of an object (e.g., a user), and in the present disclosure, as an embodiment, the depth camera (or 3D camera) is a ToF (Time of Flight) camera. This is an example to help those skilled in the art understand, and any camera capable of measuring depth can be used. In addition, in the present disclosure, as an embodiment, the FoV (Field of View) of a general camera is wider than that of a depth camera. In the present disclosure, a mode for recognizing the location of a user using information from a depth camera is referred to as a ToF mode. In addition, the ToF mode may be referred to as a depth mode, a depth camera mode, or a second mode.
[0059] According to embodiments, the present disclosure can recognize the location of a user by using both information from a general camera and information from a depth camera. That is, although the FoV of a depth camera is generally narrower than that of an RGB camera, it can measure depth, and thus can capture an object (e.g., a user) with higher precision than an RGB camera. Accordingly, the present disclosure can provide high-precision Head 6DoF within the FoV of the depth camera, and furthermore, since the RGB camera can recognize the Head 6DoF of the user even when the user is out of the FoV of the depth camera, the activity zone can be expanded. The present disclosure will refer to a mode that recognizes the location of a user by using both information from a general camera and information from a depth camera as a blend mode. In addition, the blend mode may be referred to as a blending mode or a third mode.
[0060] The present disclosure can provide a Window Perspective View service as a UX (User Experience) that allows for a more immersive viewing of the metaverse world on a 2D large screen based on the recognized position by recognizing the position of the user's head in 3D space by applying one of the first to third modes.
[0061] The present disclosure can provide a service that recognizes the posture of a user's head in 3D space by applying one of the first to third modes, and uses the recognized posture as an expression of the user's intention to use a target service from now on.
[0062] In addition, vision interaction requires a solution for recognizing when and who a user is, as there is no physical input device such as a remote control. The present disclosure can provide a service that can selectively recognize and control users by using position and orientation information of head 6DoF.
[0063] FIG. 1 is a drawing showing an example of a user location recognition device according to embodiments.
[0064] FIG. 1 may include a first camera unit (110) including at least one general camera, a second camera unit (120) including at least one depth camera, a storage unit (140) for storing camera installation information, a user location recognition unit (150) for recognizing a user's location based on information of the first camera unit (110) and / or information of the second camera unit (120), and a display control unit (160) for controlling a service provided to a display unit based on the recognized user's location.
[0065] The above first and second camera units (110, 120) and the user location recognition unit (150) can be connected via a wired or wireless network (130).
[0066] According to embodiments, the wired network may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-232 (recommended standard232), power line communication, or plain old telephone service (POTS).
[0067] According to embodiments, the wireless network may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless broadband module.
[0068] Additionally, the wireless network may include a short-range communication module, and the short-range communication module may support short-range communication using at least one of Bluetooth, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0069] The user location recognition device according to the embodiments may be any image display device capable of displaying images and providing services or controlling displayed services based on the recognized user location. The present disclosure describes, in one embodiment, a television as the user location recognition device. This is merely an example to aid understanding by those skilled in the art, and the user location recognition device may also be a monitor capable of displaying images, a refrigerator, a washing machine, or the like.
[0070] In the present disclosure, at least one general camera of the first camera unit (110) and / or at least one depth camera of the second camera unit (120) is installed at one or more locations on the television, as an example. For example, at least one general camera and / or at least one depth camera may be installed at at least one location among the top, bottom, left, or right of the television. For example, in the case of the top of the television, it may be installed at the center of the top or at a location other than the center.
[0071] In the present disclosure, the storage unit (140) stores camera installation information, as an example. According to embodiments, the camera installation information may include camera installation height, camera tilt angle, camera FoV, etc.
[0072] In addition, the storage unit (140) may store programs for signal processing and control within the user location recognition unit (150) and the display control unit (160), and may store signal-processed images, voices, or data signals. According to embodiments, the storage unit (140) may include at least one type of storage medium among, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (for example, an SD or XD memory, etc.), RAM, and ROM (for example, an EEPROM).
[0073] The following describes a method for recognizing a user's location in a user location recognition unit (150).
[0074] The present disclosure first describes a process for obtaining position information of a user's head (i.e., position (X, Y, Z) 3DoF) according to RGB mode, ToF mode, and blend mode for Head 6DOF, and then describes a process for obtaining orientation information of the user's head (i.e., orientation (Roll, Pitch, Yaw) 3DoF) according to RGB mode, ToF mode, and blend mode.
[0075] Position (X, Y, Z) 3DoF
[0076] 1) RGB mode
[0077] In the present disclosure, when only the first camera unit (110) is equipped on the television, the user position recognition unit (140) recognizes the user's head position (X, Y, Z) (i.e., 3D spatial position) within a 2D image captured using at least one RGB camera of the first camera unit (110) even without depth information, and estimates how far the user is from the actual television. In the present disclosure, the user position recognition unit (140) may recognize the user's 3D spatial position using only the first camera unit (110) even when both the first camera unit (110) and the second camera unit (120) are equipped.
[0078] At this time, the storage unit (140) stores camera installation information such as the installation height, tilt angle, and FoV of at least one RGB camera. In one embodiment, this camera installation information is predetermined according to the camera installation location at the time of product shipment and stored in the storage unit (140). In another embodiment, some of the camera installation information may be input by the user.
[0079] FIG. 2(a) and FIG. 2(b) are drawings showing examples of camera installation information and foot landmarks according to embodiments.
[0080] FIG. 3 is a drawing showing an example of a camera projection model according to embodiments.
[0081] Figure 4 is a flowchart showing an example of obtaining 3D spatial coordinates of a user's head according to embodiments.
[0082] The following describes the process of obtaining the 3D spatial coordinates of the user's head (Head (X, Y, Z)) using a 2D image captured by a general camera and camera installation information in the user location recognition unit (150) with reference to FIGS. 2 to 4.
[0083] In Fig. 2(a), the camera height (212) represents the height from the floor to the camera, and the camera tilt angle (213) represents the tilt angle of the camera. In the present disclosure, the camera height is used in the same meaning as the camera installation height, and the camera tilt angle is also referred to as the camera pitch angle. That is, the camera tilt angle represents the rotation in the vertical plane, and the camera pitch angle represents the rotation in the horizontal plane. In addition, the FoV is divided into the horizontal FoV and the vertical FoV, and represents the field of view angle for the image captured by the camera. For example, the horizontal FoV represents the field of view angle from left to right, and represents how much of an area can be seen in the horizontal direction. The vertical FoV represents the field of view angle from top to bottom, and represents how much of an area can be seen in the vertical direction. Fig. 2(a) shows the vertical field of view angle (vertical FoV angle) (214).
[0084] In one embodiment of the present disclosure, a user in an image captured by a camera includes one or more landmarks for user location recognition. These landmarks are provided by a landmark detection unit and represent specific points on each part of the body, typically corresponding to joints or bones.
[0085] In this disclosure, the landmark detection unit utilizes, as an example, the landmark detection algorithm used in the MediaPipe Pose library. The MediaPipe Pose library is an open-source framework developed by Google and is a tool for estimating body poses based on machine learning and deep learning technologies. This is an example to facilitate understanding by those skilled in the art, and any algorithm capable of extracting landmarks can be applied.
[0086] For example, when an image containing a user is processed by the MediaPipe Pose library of the landmark detection unit, the user's landmarks contained in the image may be displayed together on the user image, or only the landmarks may be displayed, as shown in FIG. 2(b). For example, the MediaPipe Pose model may estimate 33 landmarks for major body parts of a person, and the foot landmark in FIG. 2(b) may be one of the 33 landmarks. That is, in the present disclosure, the foot landmark represents a specific point or landmark on the user's foot, as shown in FIG. 2(b).
[0087] If the user in the image captured by the first camera unit (110) is as shown in Fig. 2(a), the absolute distance (211) from the TV plane to the user's foot is calculated (S411). In the present disclosure, the absolute distance (211) from the TV plane to the user's foot is referred to as Foot Z. In one embodiment, the user in the image includes one or more landmarks.
[0088] Then, Head Z is obtained from the World Coordinate System (S412) based on the relative position from the feet to the head, and then converted to the Camera Coordinate System (S413). For example, Head Z in the Camera Coordinate System can be obtained by rotating Head Z along the Z-axis by the camera tilt angle.
[0089] And, Head X,Y of the camera coordinate system (i.e., 2D coordinates of the user's head) are obtained from Head Z of the camera coordinate system (S414). For example, in a camera projection model such as Fig. 3, Head X,Y of the camera coordinate system can be obtained using the Head Z value.
[0090] In Fig. 3, 312 is Head Z in the camera coordinate system (meters), 313 is the focal length, and 314 is the pixel position in the camera view of the head. 312, 313, and 314 in Fig. 3 are known values from step S414. Here, the focal length represents the distance from the origin (0,0,0) to the image plane.
[0091] And, in the camera projection model of Fig. 3, (x, y) is the camera coordinate system (pixel) in pixel units, and (X, Y) is the camera coordinate system (meter) in meter units. The camera coordinate system (pixel) is a coordinate system expressed in pixel units of the image, and each pixel represents the coordinates for the horizontal and vertical directions of the image. Since this camera coordinate system (pixel) is expressed in pixel units according to the size of the image, it represents a relative size. Meanwhile, the camera coordinate system (meter) is a coordinate system that expresses coordinates in actual two dimensions in meters.
[0092] According to embodiments, Head X,Y (311) of the camera coordinate system in the camera projection model can be obtained as in the following mathematical expression 1. Here, the camera coordinate system is in meters as an embodiment.
[0093] [Mathematical Formula 1]
[0094]
[0095] That is, if we know the Head Z (312) of the camera coordinate system in meters, the focal length (313), and the pixel position (314) in the camera view of the head, we can calculate the Head X,Y (311) of the camera coordinate system in meters using a proportional equation such as Equation 1.
[0096] The present disclosure can obtain Head X,Y,Z of the world coordinate system (i.e., 3D space coordinates of the user's head) by rotating Head X,Y of the camera coordinate system obtained in this way along the X axis by the camera tilt angle (S415).
[0097] In the present disclosure, the camera coordinate system (i.e., @CameraCoordinateSystem(X,Y,Z)) is a local coordinate system centered on the camera itself. That is, the center of the camera is the origin, and coordinates are defined according to the direction, position, and rotation of the camera. In addition, the world coordinate system (i.e., @WorldCoordinateSystem(X,Y,Z)) is a coordinate system representing the entire 3D space. All objects have a position and orientation in the world coordinate system. That is, the camera coordinate system is generally used to define relative coordinates from the viewpoint of the camera, and the world coordinate system represents the entire 3D space, and is used to define the overall position and interaction of various objects or the camera itself. That is, by recognizing Foot Z with the camera installation conditions (e.g., 212, 213, 214) and the y value of the foot landmark, the 3D space coordinates of the user's head (i.e., Head position (X, Y, Z)) can be recognized even in a situation where there is no depth information. For example, if a television is a metaverse display device, the present disclosure can recognize the 3D spatial position of the user's head by calculating the absolute distance between the meta window and the user's feet. In this meta window environment, the TV screen can serve as the meta window.
[0098] The following is a detailed description of each step (S411-S415) of Fig. 4.
[0099] FIG. 5(a) and FIG. 5(b) are diagrams showing an example of calculating the absolute distance from the TV plane to the user's foot according to embodiments. That is, FIG. 5(a) and FIG. 5(b) are detailed embodiments of step S411 of calculating Foot Z, which is the absolute distance (211) from the TV plane to the foot.
[0100] That is, assuming that the user's feet are touching the floor, the distance (211) from the TV where the camera is installed to the user's feet and the position of the head from the user's feet are calculated.
[0101] At this time, the camera vertical FoV angle (214) is one of the camera installation information, i.e., the camera setting value. In addition, the camera pitch angle (217) may be an IMU (Inertial Measurement Unit) measurement value or may be directly input by the user. For example, the camera pitch angle (or the camera's up-down angle or camera tilt angle) (217) may be measured by combining the IMU's gyroscope (i.e., angular velocity sensor) and accelerometer (i.e., acceleration sensor). The camera pitch angle (217) is generally defined based on the camera center line. The camera pitch angle (217) indicates how much the camera is tilted in the vertical direction. In addition, the camera center line indicates the center of the light rays exiting through the camera lens, and this camera center line is generally aligned to face the center of the image sensor.
[0102] In addition, the camera height (212) may be stored in advance in the storage unit (140) through a single pre-calibration as one of the camera installation information, or may be directly input by the user or additionally generated (Ideation). Here, the calibration may be performed by the user at a predetermined distance from the meta window (i.e., the television plane). That is, in step S411, the camera height (212), the camera vertical FoV angle (214), and the camera pitch angle (217) are known values.
[0103] Next, we obtain the y values of the user's feet (e.g., both index toes) from these known values and the user's landmarks in the image (foot y or y foot), the angle from the camera center line to the user's foot (i.e., foot angle from camera center line) (215) can be calculated from the acquired foot y value.
[0104] Referring to Fig. 2(b), the angle (215) from the camera center line to the user's foot can be obtained from the y value of the foot as shown in the following mathematical expression 2.
[0105] [Equation 2]
[0106] Foot angle from camera center line =tan -1 ((y foot -0.5) / h,
[0107] h = 1 / (2 x tan (camera vertical FoV angle / 2))
[0108] Here, y foot represents the y value of the user's foot, and h is a value corresponding to the same unit as pose y.
[0109] Then, using the camera pitch angle (217) and the foot angle (215) from the camera center line, the foot angle from the camera vertical line (216) is obtained as shown in the following mathematical expression 3.
[0110] [Equation 3]
[0111] Foot angle from camera vertical line (216) = 90 degrees (90 ˚ ) - Camera pitch angle (217) - Foot angle from camera center line (215)
[0112] That is, since the camera pitch angle (217) + foot angle from the camera center line (215) + foot angle from the camera vertical line (216) = 90 degrees, the foot angle (216) from the camera vertical line can be obtained by subtracting the camera pitch angle (217) and the foot angle (215) from the camera center line from 90 degrees.
[0113] And, the absolute distance (211) from the TV plane to the user's foot, i.e., the distance from the TV front to the user's foot (foot distance from TV front), Foot Z, can be obtained as in the following mathematical expression 4 based on the camera height (212) and the foot angle (216) from the camera vertical line.
[0114] [Equation 4]
[0115] Foot Z = Camera Height * tan (foot angle from camera vertical (216))
[0116] As above, when Foot Z, which is the absolute distance (211) from the TV plane to the user's foot, is obtained, Head Z, which is the relative position of the user's head from Foot Z, is obtained in the world coordinate system.
[0117] FIGS. 6(a) to 6(c) are diagrams showing an example of obtaining a relative position from a user's feet to their head according to embodiments. That is, FIGS. 6(a) to 6(c) are detailed embodiments of step S412 of obtaining Head Z in the world coordinate system as a relative position from a user's feet to their head based on Foot Z and camera installation information.
[0118] The information required in step S412 is the camera height (212), the camera pitch angle (217), and Foot Z (211). Foot Z (211) is the absolute distance from the TV plane to the user's foot, which is obtained based on step S411 and FIGS. 5(a) and 5(b). That is, the camera height (212), the camera pitch angle (217), and Foot Z (211) are values known in step S412.
[0119] Next, we can obtain Head Z (511) in the world coordinate system using these known values.
[0120] First, the coordinates of the foot (512) can be obtained using the Foot Z (211) and the camera height (212) as in the following mathematical expression 5.
[0121] [Equation 5]
[0122] Foot coordinates (512) = (X, camera height (212), Foot Z (211)) @ world coordinate system
[0123] Here, X has no effect, so don't care.
[0124] Then, a vector from the user's feet to the eyes is obtained based on landmarks representing the user's body parts (e.g., Fig. 6(b)).
[0125] For example, as in Fig. 6(b), the camera is installed with a pitch angle (217), so that the skeleton of a standing person is measured as if it is tilted forward in the camera coordinate system.
[0126] Therefore, the vector from the foot to the eye is rotated and compensated for by the camera pitch angle as shown in Fig. 6(c), thereby converting the vector in the camera coordinate system into a vector (513) in the world coordinate system. For example, if the camera pitch angle (θ) is -20 degrees, θ becomes 20 in Fig. 6(c). That is, in Fig. 6(c), θ represents the camera pitch angle, x', y', z' represent world coordinates, and x, y, z represent camera coordinates.
[0127] Then, the coordinates of the feet (512) and the vector (513) of the world coordinate system from the feet to the eyes are added to obtain the relative position from the feet to the head, i.e., Head Z in the world coordinate system.
[0128] FIG. 7 is a diagram showing an example of converting Head Z from the world coordinate system to the camera coordinate system, based on the relative position from the user's feet to the head, according to embodiments. That is, FIG. 7 is a detailed embodiment of step S413 of converting Head Z from the world coordinate system to the camera coordinate system, based on the relative position from the feet to the head.
[0129] The information required in step S413 is the camera pitch angle (217) and the Head Z (511) of the world coordinate system obtained in step S412, which are values known in step S413.
[0130] The present disclosure can convert Head Z (511) of the world coordinate system into Head Z (516) of the camera coordinate system using these known values.
[0131] That is, Head Z (516) of the camera coordinate system can be obtained by rotating and compensating Head Z (511) of the world coordinate system by the camera pitch angle. In other words, Head Z in the camera coordinate system can be found out by rotating by the camera tilt angle. For example, Head X, Y, Z (here, X is don't care) of the world coordinate system can be converted to Head X, Y, Z of the camera coordinate system by rotating and compensating by the camera pitch angle, and Head Z of the camera coordinate system can be obtained from this. Here, the reason for converting the world coordinate system to the camera coordinate system is to obtain X and Y based on the Z value in the camera coordinate system.
[0132] Then, in step S414, Head X,Y of the camera coordinate system are calculated based on Head Z of the camera coordinate system, focal length, and pixel position of the head in the camera view as described based on FIG. 3.
[0133] Fig. 8 is a drawing showing an example of obtaining Head X, Y, Z of a world coordinate system according to embodiments. That is, Fig. 8 is a detailed embodiment of step S415 of converting Head X, Y of a camera coordinate system into Head X, Y, Z of a world coordinate system.
[0134] That is, the Head X,Y of the camera coordinate system obtained in step S414 is rotated by the camera tilt angle (or camera pitch angle) around the X-axis to obtain the Head X,Y,Z (611) of the world coordinate system. Fig. 6 (c) is an example of a determinant for obtaining the Head X,Y,Z of the world coordinate system by rotating the Head X,Y of the camera coordinate system by the camera tilt angle (or camera pitch angle) around the X-axis.
[0135] By going through these processes, we can obtain the 3D space coordinates (Head X, Y, Z) of the user's head (i.e., the position of the user's head in 3D space) in the world coordinate system from a 2D image captured by an RGB camera without depth information. In other words, we can accurately estimate whether the user is actually 1m or 2m away from the TV plane. In other words, we can know how far the user is from the TV.
[0136] According to embodiments, the display control unit (160) can control a service displayed on the display unit or a service to be displayed based on the user location recognized by the user location recognition unit (150).
[0137] For example, if a television equipped with a user location recognition device is a metaverse display device, the display control unit (160) can zoom in or out on an image being displayed on the display unit (e.g., a screen or a screen) based on the user's location or movement recognized by the user location recognition unit (150). For example, the image being displayed can be zoomed in as the user moves toward the front of the TV, and zoomed out as the user moves away, or vice versa.
[0138] 2) ToF mode
[0139] The following describes the process of obtaining the 3D spatial coordinates of the user's head using information from a depth camera including depth information in the user location recognition unit (150).
[0140] According to embodiments, camera installation information required in ToF mode includes camera tilt angle (or camera pitch angle), camera FoV, camera installation height, etc. In addition, one embodiment detects user landmarks required in ToF mode through the MediaPipe Pose library of the landmark detection unit.
[0141] In the present disclosure, since a depth camera, for example, a ToF camera, has depth information, the process of obtaining Head Z (i.e., steps S411-S413 in FIG. 4) can be omitted.
[0142] That is, in ToF mode, the Head Z of the camera coordinate system can be directly obtained from the depth information of the image captured by the depth camera.
[0143] And, Head X,Y of the camera coordinate system are obtained from Head Z of the camera coordinate system. For example, in a camera projection model such as Fig. 3, Head X,Y of the camera coordinate system can be obtained using the Head Z value. And, Head X,Y of the camera coordinate system obtained in this way can be rotated by the camera tilt angle around the X-axis to obtain Head X,Y,Z of the world coordinate system (i.e., 3D space coordinates of the user's head).
[0144] At this time, RGB images are required in ToF mode to detect the user's landmarks using the MediaPipe Pose library of the landmark detection unit. For this purpose, both an RGB camera and a ToF camera may be required in ToF mode.
[0145] Accordingly, when both a ToF camera and an RGB camera are installed in a TV, since the two cameras acquire images from different positions and directions, the user position recognition unit (150) can find a head representative point (e.g., between the eyebrows) in the landmark area of each eye in the RGB image in the ToF image (X, Y, Z), align it based on the RGB, and then convert it to a world coordinate system, as follows.
[0146] First, an RGB image containing the user's pose is input into the MediaPipe Pose library to detect 33 landmarks according to the user's pose. Fig. 9(a) is a diagram showing examples of the 33 pose landmarks provided by the MediaPipe Pose library.
[0147] When the user's landmark included in the RGB image is detected (or recognized), the center position of the binocular landmarks (2, 5) among the 33 landmarks is set as the head representative point, as shown in Fig. 9(b). At this time, the Boundary(X, Y) distance is set considering the low-resolution ToF camera. Fig. 9(b) is a diagram showing an example of a head representative point based on 33 pose landmarks provided by the MediaPipe Pose library. For example, in the present disclosure, the Boundary(X, Y) distance can be determined as a value obtained by dividing the binocular distance by 2 (i.e., Boundary(X, Y) distance = 1 / 2 of the binocular distance). At this time, the binocular distance (Eye_Distance) can be determined as a value obtained by subtracting the 2nd landmark (e.g., left eye) from the 5th landmark (i.e., right eye) as shown in Fig. 9(b).
[0148] Then, a ToF image matching the glabellar region of the RGB image (i.e., a representative point of the head) within the boundary condition (e.g., Boundary(X, Y) distance) is searched. In this case, X and Y are obtained as the center of gravity of the candidate group, and the depth is obtained as an average, as an example.
[0149] In a camera projection model such as Fig. 9(c), the X, Y of the head representative point searched in the ToF image based on the depth are converted into the RGB camera coordinate system (meters).
[0150] Fig. 9(c) is a diagram showing an example of a camera projection model. In Fig. 9(c), Z represents depth, and the focal length represents the distance from the origin (0,0,0) to the image plane. In addition, (x,y) are camera coordinates in pixels (pixels), and (X,Y) are camera coordinates in meters (meters).
[0151] And, alignment with the RGB camera is performed. For example, alignment can be performed by matching the origin through rotation and translation along the X, Y, and Z axes.
[0152] Additionally, the camera coordinate system can be transformed into the world coordinate system by compensating for the rotation by the camera tilt angle or camera pitch angle.
[0153] The following describes the RGB-ToF conversion model.
[0154] The first embodiment of the RGB-ToF conversion model can be performed when the alignment of the ToF camera with respect to the RGB camera is not correct after the physical camera placement. For example, this can be performed when there is no rotation about the camera Z-axis. Furthermore, an actively aligned camera module does not perform this process.
[0155] The second embodiment of the RGB-ToF conversion model is a method that returns a ToF pixel and depth when a target RGB pixel (X, Y) is input using Extrinsic Calibration. For example, Extrinsic Calibration is a method that matches a pixel acquired from an RGB camera to a corresponding pixel of a ToF camera when an RGB camera and a ToF camera are installed in different locations.
[0156] The following describes how to calculate the camera installation height in ToF mode. For example, the camera installation height (Camera_Height) can be calculated as follows based on the absolute distance (Foot_Z) from the TV plane to the user's foot, the vector from the user's foot to the eye (Foot2Eye_Vector), and the Head XYZ@world coordinate system obtained in steps S411 to S415 of FIG. 4.
[0157] In the present disclosure, when Head_XYZ@world coordinate system is (Head_X, Head_Y, Head_Z), Foot2Eye_Vector@world coordinate system can be obtained by subtracting Foot_Vector @world coordinate system from Eye_Vector @world coordinate system (Foot2Eye_Vector@world coordinate system = Eye_Vector @world coordinate system - Foot_Vector @world coordinate system), and Foot_Z @world coordinate system can be obtained by subtracting Foot2Eye_Vector_Z from Head_Z (i.e., Foot_Z @world coordinate system = Head_Z - Foot2Eye_Vector_Z).
[0158] And, the camera installation height (Camera_Height) can be obtained by subtracting Foot2Eye_Vector_Y from Head_Y as in mathematical expression 6.
[0159] [Equation 6]
[0160] Camera_Height = Head_Y - Foot2Eye_Vector_Y
[0161] Additionally, in the present disclosure, the camera installation height (Camera_Height) may be directly input by the user, may be estimated through calibration, or may be stored in advance in the storage unit (140) when the TV is shipped.
[0162] 3) Blend mode
[0163] The following describes how the user location recognition unit (150) recognizes the user location in blend mode.
[0164] In the present disclosure, the blend mode is a combination of the RGB mode and the ToF mode. Typically, the field of view of the RGB camera is wider than that of the ToF camera. In other words, the field of view of the ToF camera is a narrow angle. Therefore, in the present disclosure, the user position recognition unit (150) operates in the ToF mode when the user is within the field of view of the ToF camera, and operates in the RGB mode when the user is outside the field of view of the ToF camera but within the field of view of the RGB camera, thereby recognizing the 3D spatial position of the user's head, which is referred to as the blend mode in the present disclosure. In contrast, the ToF mode can also utilize information from both the RGB camera and the ToF camera, but the ToF mode recognizes the user's position only when the user is within the field of view of the ToF camera.
[0165] In order for the user location recognition unit (150) to operate in blend mode, a camera installation height is required. In the present disclosure, one embodiment utilizes the camera installation height obtained in ToF mode. For example, camera installation height information can be calculated as shown in FIG. 6 and stored in the storage unit (140) when a user is present within the FoV of the ToF camera.
[0166] Additionally, there is a possibility that deviations may occur in the calculated camera installation height depending on the performance of the landmark detected through the MediaPipe Pose library of the landmark detection unit. To this end, the present disclosure adds a camera installation height tracking function (i.e., Camera_Height Tracking) to the user location recognition unit (150), as an example.
[0167] In addition, the user position recognition unit (150) operates in RGB mode even when the user moves away from the camera based on distance within the FoV of the ToF camera, or when the user is not recognized or performance deteriorates at the FoV boundary.
[0168] When the user position recognition unit (150) of the present disclosure operates in blend mode, a more precise Head 6DoF result than the RGB mode can be obtained by using information from a ToF camera with a narrow FoV area, and a wider activity zone than the ToF mode can be obtained by using information from an RGB camera with a wide FoV.
[0169] The following describes how to obtain the 3D spatial coordinates of the user's position (Position (X, Y, Z)) when operating in RGB mode and ToF mode in blend mode.
[0170] For example, if the camera installation height (Camera_Height) information is input in advance and stored in the storage unit (140), the user position recognition unit (150) can operate in RGB mode regardless of the ToF mode. If the camera installation height (Camera_Height) information obtained in the ToF mode is utilized, the user position recognition unit (150) can operate in RGB mode after initially operating in the ToF mode.
[0171] And, when the user location recognition unit (150) operates in ToF mode, the camera installation height (Camera_Height) is calculated and stored in the storage unit (140). At this time, tracking of the camera installation height is also included.
[0172] Additionally, the user position recognition unit (150) can determine a report mode. For example, a mode in which the Position (X, Y, Z) is within a normal range (e.g., the Z value is 0 or higher, etc.) and does not change abruptly compared to the Position of N-1 frames can be selected and reported.
[0173] The following describes the calculation and tracking of camera installation height in blend mode. That is, as an embodiment, the camera installation height calculation formula (see Equation 6) in ToF mode is utilized even in blend mode. In addition, as an embodiment, after calculating the camera installation height within the active zone, the camera installation height is processed in a smoothing manner. For example, among the landmarks detected through the Mediapipe Pose library, the Z and Y values based on the center of the hip (e.g., Left / Right Hip, Landmark_23, Landmark_24) require smoothing since there is a difference between the actual user posture and height.
[0174] In this way, since the FoV of the camera limits the activity area, the user location recognition unit (150) can operate in ToF mode to provide a high-precision user location when the user is in the central area, and can operate in RGB mode with a wide FoV to provide a wider activity area than the ToF mode when the user is in the outer area.
[0175] FIG. 10(a) and FIG. 10(b) are diagrams showing an example of blend mode operation according to embodiments.
[0176] For example, Fig. 10(a) and Fig. 10(b) are cases where the distance between the TV and the target object (user) is 2 m, the center point of the TV screen height is the human eye height (1.65 m), and the camera tilt angle is assumed to be 31.4 degrees.
[0177] Also, the horizontal FoV of the ToF camera is assumed to be 69 / 90 degrees, the vertical FoV is assumed to be 45 / 68 degrees, and the horizontal FoV of the RGB camera is assumed to be 103.8, and the vertical FoV is assumed to be 86.6.
[0178] Fig. 10(a) is a side view diagram illustrating an example where a user is positioned 2 m from the TV plane under the above conditions. Furthermore, Fig. 10(b) illustrates an example where the system operates in ToF mode or RGB mode depending on the user's position under these conditions. That is, if the user is within the horizontal FoV of the TOF camera (e.g., 60 degrees or 90 degrees), the system operates in ToF mode, and if the user is outside the 90-degree range, the system operates in RGB mode.
[0179] In the present disclosure, when the user position recognition unit (150) operates in blend mode, in RGB mode, all steps S411 to S415 described in FIG. 4 are performed to obtain the 3D space coordinates (position (X, Y, Z)) of the user's head, and in ToF mode, since there is depth information, steps S411 to S413 are omitted and steps S414 to S415 are performed to obtain the 3D space coordinates (position (X, Y, Z)) of the user's head.
[0180] So far, we have explained how to recognize the user's location when the camera is fixed.
[0181] If the camera rotates left and right or up and down, the user's position can be recognized by applying one of the aforementioned RGB mode, ToF mode, or blend mode while changing the camera tilt angle and / or camera pitch angle according to the movement of the camera.
[0182] In addition, so far, the process of obtaining the user's head position information (i.e., position (X, Y, Z) 3DoF) for Head 6DOF according to RGB mode, ToF mode, and blend mode has been described.
[0183] The following describes the process of obtaining the user's head orientation information (i.e., orientation (Roll, Pitch, Yaw) 3DoF) for Head 6DOF according to RGB mode, ToF mode, and blend mode. That is, position (X, Y, Z) 3DoF is information that can determine the location of the user's head, and orientation (X, Y, Z) 3DoF is information that can determine the direction in which the user's head identified by position (X, Y, Z) 3DoF is looking. Therefore, Head 6DOF, which can be expressed as position (X, Y, Z) 3DoF and orientation (X, Y, Z) 3DoF, allows us to determine the direction in which the user is looking at from which location.
[0184]
[0185] Orientation (X, Y, Z) 3DoF
[0186] 1) RGB mode
[0187] Fig. 11 is a flowchart showing an example of obtaining the direction of a user's head in RGB mode according to embodiments.
[0188] That is, the user position recognition unit (150) includes a step of obtaining a face frame (S711) when in RGB mode, a step of obtaining a rotation matrix (S712), and a step of converting the rotation matrix into a quaternion (S713), as an example.
[0189] FIGS. 12(a) to 12(d) are detailed embodiments of steps S711 to S713 of FIG. 11.
[0190] Figure 12(a) shows the definition of axes and rotation information (Roll, Pitch, Yaw) used to indicate the direction information of the user's head. For example, Roll means tilting the head to both sides with respect to the facing axis, and corresponds to tilting. Pitch means moving the head up and down with respect to the eye2eye axis, and corresponds to nodding. Yaw means turning the head to the right and left with respect to the head top axis, and corresponds to turning. In the present disclosure, the facing axis is defined as the axis along which the user looks straight ahead, the eye2eye axis is defined as the axis from eye to eye, and the head top axis is defined as the axis from the chin to the crown of the head.
[0191] In this disclosure, Roll represents the angle at which the user's head rotates left and right, Pitch represents the angle at which the user's head moves up and down, and Yaw represents the angle at which the user's head rotates left and right. In this disclosure, the rotational degrees of freedom for the three rotation directions (Roll, Pitch, Yaw) are referred to as orientation (Roll, Pitch, Yaw) 3DoF. That is, orientation (X, Y, Z) 3DoF can indicate which direction the user is looking or what posture the user is in.
[0192] According to embodiments, the user position recognition unit (150) can create two vectors with three points corresponding to the left eye, right eye, and mouth (mouse) among the landmarks detected through the Mediapipe Pose library, and then configure three axes (e.g., facing axis, eye2eye axis, head top axis) to obtain the orientation of the user's head. The present disclosure assumes that the landmarks of the eyes and mouth exist on the same plane in order to create a face frame in step S711.
[0193] Figure 12(b) shows an example of creating an Eye2Mouse vector using the left eye and mouth points within a landmark, and an Eye2Eye vector using the left eye and right eye points.
[0194] Then, the Eye2Eye axis vector is created using the Eye2Eye vector (or Eye2Eye unit vector), the Facing axis vector is created using the cross product of the Eye2Eye vector and the Eye2Mouth vector as a unit vector, and the Head Top axis vector is created using the cross product of the Facing axis and the Eye2Eye vector as a unit vector.
[0195] Then, in step S712, a rotation vector (rotation matrix) is obtained as in Fig. 12(c) using the vectors generated as landmarks of the two eyes and mouth. This is to know how much the user's head has rotated around the corresponding axis. In Fig. 12(c), Fx_current, Fy_current, and Fz_current are known values corresponding to each axis (eye2eye axis, head top axis, facing axis), and Fx_init, Fy_init, and Fz_init are also known values as initial values, so the rotation values (r11-r33) can be obtained by applying the two known values.
[0196] In step S713, the rotation values (r11-r33) obtained from the rotation matrix are applied as in Fig. 12(d) to convert the rotation matrix into a quaternion. The quaternion consists of four numbers (qw (r32-r23) / 4qw (r13-r31) / 4qw (r21-r12) / 4qw), where qw is a scalar part and (r32-r23) / 4qw (r13-r31) / 4qw (r21-r12) / 4qw is a vector part. The present disclosure can estimate the orientation (orientation (X, Y, Z) 3DOF) of the user's head in RGB mode based on this quaternion.
[0197] 2) ToF mode
[0198] In one embodiment, the user position recognition unit (150) finds points matching the eye and mouth landmarks in the RGB image in the ToF image (X, Y, Z) in the ToF image, aligns them based on RGB, and then converts the aligned ToF image (X, Y, Z) to the world coordinate system to obtain the orientation (Roll, Pitch, Yaw) 3DOF.
[0199] That is, the user position recognition unit (150) inputs an RGB image into the Mediapipe Pose library to detect 33 landmarks. Then, among the 33 detected landmarks, the ToF pixels that match the positions of the binocular and mouth landmarks are converted into an RGB camera model and converted into a world coordinate system.
[0200] Then, as described in RGB mode, the face frame is obtained, and the rotation matrix is obtained based on this, and then the rotation matrix is converted to a quaternion. By doing this, the orientation (orientation (X, Y, Z) 3DOF) of the user's head can be estimated in ToF mode.
[0201] The following describes the RGB-ToF conversion model.
[0202] The first embodiment of the RGB-ToF conversion model can be performed when the alignment of the ToF camera with respect to the RGB camera is not correct after the physical camera placement. For example, this can be performed when there is no rotation about the camera Z-axis. Furthermore, an actively aligned camera module does not perform this process.
[0203] The second embodiment of the RGB-ToF conversion model is a method that returns a ToF pixel and depth when a target RGB pixel (X, Y) is input using Extrinsic Calibration. For example, Extrinsic Calibration is a method that matches a pixel acquired from an RGB camera to a corresponding pixel of a ToF camera when an RGB camera and a ToF camera are installed in different locations.
[0204] 3) Blend mode
[0205] As described above, the blend mode is a combination of the RGB mode and the ToF mode. That is, the user position recognition unit (150) operates in the ToF mode utilizing depth information when the user is within the field of view of the ToF camera, and operates in the RGB mode when the user is outside the field of view of the ToF camera but within the field of view of the RGB camera. In addition, the user position recognition unit (150) operates in the RGB mode even when the user's eyes and mouth are not visible depending on the user's head posture.
[0206] Also, the method of obtaining orientation (Roll, Pitch, Yaw) 3DOF when operating in RGB mode and the method of obtaining orientation (Roll, Pitch, Yaw) 3DOF when operating in ToF mode have been explained above, so they will be omitted here to avoid redundant explanation.
[0207] In the present disclosure, Head 6DoF can be configured by combining position information of the user's head (position (X, Y, Z) 3DOF) indicating the 3D space coordinates of the user's head and orientation information of the user's head (orientation (Roll, Pitch, Yaw) 3DOF) indicating the direction of the user's head. In other words, the position and direction of the user's head in 3D space can be known by Head 6DOF.
[0208] In the present disclosure, the user position or user position information may be composed of position information of the user's head (position (X,Y,Z) 3DOF) indicating 3D space coordinates of the user's head, or may be composed of position information of the user's head (position (X,Y,Z) 3DOF) indicating 3D space coordinates of the user's head and orientation information of the user's head (orientation (Roll, Pitch, Yaw) 3DOF) indicating the direction of the user's head (Head 6DoF).
[0209] Additionally, in the present disclosure, the user location recognition unit (150) may operate in one of RGB mode, ToF mode, and blend mode when recognizing the user location.
[0210] In the present disclosure, when the user position recognition unit (150) operates in blend mode, more precise user position information (e.g., Head 6DoF result) than in RGB mode can be obtained by information from a ToF camera with a narrow FoV area, and a wider activity zone than in ToF mode can be obtained by information from an RGB camera with a wide FoV.
[0211] So far, we've described a method for recognizing a user's location within an image. To do this, we use one or more landmarks within the image, extracted from a landmark detection library like the MediaPipe Pose library. For example, MediaPipe Pose can detect 33 landmarks representing major human body parts.
[0212] The present disclosure may further include an image control unit that manipulates a portion of an image captured by an RGB camera and / or a ToF camera and provides it to the landmark detection unit to enable the landmark detection unit to more accurately recognize a user within the image. In the present disclosure, the image control unit may be included in the user location recognition unit (150) or may be configured as a separate module or component.
[0213] That is, the present disclosure is to control the input image of a landmark detection unit (e.g., a machine learning tool) that recognizes users and landmarks, so that the landmark detection unit recognizes only users with intention.
[0214] If the landmark detection unit is Mediapipe Pose, Mediapipe Pose recognizes the person with the highest confidence throughout the entire scene in 'detection mode' and detects the landmark of the recognized person. The confidence of a person's recognition in Mediapipe Pose is determined by various factors such as lighting, clothing type, clothing color, location, size, overlapping with other objects, and body part obscuration.
[0215] And, in the 'tracking mode' after recognizing a person, that is, a region of interest (ROI) is set and the person is followed and recognized as long as the reliability of the person initially recognized is maintained.
[0216] However, when there are multiple people (two or more) in the camera or in the video (image) captured by the camera, there is no interface for specifying who to recognize.
[0217] The present disclosure provides, as an embodiment, a method of setting at least one area within an image captured by a first camera unit (110) and / or a second camera unit (120), and then providing a mask to a portion of an image (i.e., an image) provided to a landmark detection unit based on information within the set at least one area.
[0218] The present disclosure sets a first area (811) and a second area (813) within an image acquired from a first camera unit (110) and / or a second camera unit (120) as shown in FIG. 13 or FIG. 14. Here, the image is set as an example for providing to a landmark detection unit.
[0219] In the present disclosure, the first region (811) is a core region of interest for initially recognizing a user in a user-unrecognized state, and the second region (813) is an region for continuously determining whether the user is within the region of interest in a user-recognized state, as an example.
[0220] FIGS. 13(a) to 13(e) are drawings showing examples of image control methods according to embodiments.
[0221] That is, when an image is input for user location recognition, the image control unit detects that a user (i.e., an interaction target) is within the first area (811), and the image is provided to the landmark detection unit as an original image without a mask.
[0222] Meanwhile, when an image is input for user location recognition, if there is no user (i.e., interaction target) within the first area (811), the image control unit places a mask on a portion of the input image as shown in Fig. 13(a).
[0223] For convenience of explanation, the present disclosure will refer to a masked area within an image as a masked area, and an unmasked area as a non-masked area. In Fig. 13(a), the black portion is the masked area. In the present disclosure, the non-masked area includes a first area (811), and is larger than the first area (811), as an example. In addition, the second area (813) includes the first area (811), and is larger than the first area (811), as an example.
[0224] Since there is no user to be recognized in the masked image in the present disclosure, the masked image may or may not be provided to the landmark detection unit.
[0225] According to embodiments, the image control unit masks a portion of the input image even when the user is outside the first area (811), as shown in FIG. 13(b).
[0226] At this time, if the user enters the first area (811) as shown in Fig. 13(c), the image is not masked and is provided to the landmark detection unit as the original image. That is, in this case, the user is recognized.
[0227] Additionally, as shown in Fig. 13(d), the user within the image can leave the second area (813).
[0228] If the user leaves the second area (813), as shown in Fig. 13(e), a mask is applied to a portion of the image. The user is then kept unrecognized until the user re-enters the first area (811). In this case, as an example, even if another user (e.g., a bystander) enters the first area (811), the user is not recognized. In other words, by utilizing the mask, the user can be prevented from being mistakenly identified as a bystander until the interaction target (i.e., the user) enters the first area (811).
[0229] In this way, when a user appears within the first area (811), the image control unit provides the original image (i.e., image) to the landmark detection unit without a mask, and when the user leaves the second area (813), the image is masked again to prevent misrecognition by people around the user.
[0230] FIGS. 14(a) to 13(f) are drawings showing other examples of image control methods according to embodiments.
[0231] That is, when an image is input for user location recognition, the image control unit detects that a user (i.e., an interaction target) is within the first area (811), and the image is provided to the landmark detection unit as an original image without a mask.
[0232] Meanwhile, when an image is input for user location recognition, the image control unit masks a portion of the input image as shown in Fig. 14(a) if there is no user (i.e., an interaction target) within the first region (811). In the present disclosure, since there is no user to be recognized in the masked image, the masked image may or may not be provided to the landmark detection unit.
[0233] According to embodiments, the image control unit masks a portion of the input image even when the user is outside the first area (811), as shown in FIG. 14(b).
[0234] At this time, if the user enters the first area (811) as shown in Fig. 14(c), the image is not masked and is provided to the landmark detection unit as the original image. That is, in this case, the user is recognized.
[0235] Additionally, the image control unit does not perform periodic detection of the first area (811) when the user is within the first area (811), as shown in Fig. 14(d).
[0236] However, as shown in Fig. 14(e), if the user is within the second area (813) but not within the first area (811), that is, if the user is within the second area (813) but outside the first area (811), the image control unit can periodically detect the first area (811) to determine whether a new user appears in the first area (811). At this time, a new third area may be designated instead of the first area (811) to recognize the new user, and it may be determined whether the user enters the third area.
[0237] If a new user appears in the first area (811) or the third area, the interaction target is replaced with the new user, and if there is no new user, the existing user is continuously recognized.
[0238] Also, the replaced user can leave the second area (813).
[0239] If the replaced user leaves the second area (813), as shown in Fig. 14(f), a mask is applied to a portion of the image. The user remains unrecognized until the user re-enters the first area (811). In this case, as an example, even if another user (e.g., a bystander) enters the first area (811), the user is not recognized. In other words, by utilizing the mask, the bystander is prevented from being mistaken for the user until the interaction target (i.e., the user) enters the first area (811).
[0240] In this way, even if a user is replaced, if the replaced user appears within the first area (811), the image control unit provides the original image (i.e., image) to the landmark detection unit without a mask, and if the replaced user leaves the second area (813), the image is masked again to prevent misrecognition by surrounding people.
[0241] According to embodiments, the landmark detection unit recognizes a user from an image provided by the image control unit, detects one or more landmarks of the recognized user, and provides the detected landmarks to the user location recognition unit (150). At this time, the landmark detection unit recognizes only the user with intent through image control of the image control unit, and can detect one or more landmarks of the recognized user.
[0242] In the present disclosure, when the landmark detection part is a Mediapipe pose, the Mediapipe can be changed by adding a 'calculator that can modify input_video' to the input terminal of the Mediapipe solution. That is, by modifying input_video, the left / right / top / bottom can be received as input packets of the calculator, and the aforementioned image mask can be applied. In addition, the Head 6DoF algorithm of the present disclosure can adjust the image mask size and operate a user recognition scenario by passing the input packet value to the Mediapipe Calculator.
[0243] As described above, the present disclosure utilizes camera-based vision data to recognize a user's location in a 3D spatial coordinate system. This data recognizes the distance between the camera and the user, and calculates the user's 3D spatial coordinates and / or orientation based on this information. At this time, the distance between the user and the camera is calculated using camera installation information. Furthermore, for selective user recognition using vision data, the input image is controlled to selectively recognize users located at specific locations.
[0244] Furthermore, the present disclosure can recognize the user's precise spatial coordinates and orientation (Head 6DoF) using only an RGB camera. That is, the present disclosure recognizes the user's position in spatial coordinates by recognizing the absolute distance (Z) using only an RGB camera based on the camera's installation information.
[0245] In addition, the present disclosure provides a blend mode using information from both a depth camera (i.e., a 3D camera) and an RGB camera, thereby providing high-precision Head 6DoF within the FoV of the depth camera and recognizing Head 6DoF even when the user is out of the FoV of the depth camera, thereby expanding the activity area.
[0246] In addition, the present disclosure can selectively recognize and control users by utilizing the position and direction information of Head 6DoF obtained by the aforementioned method. That is, by utilizing the information of Head 6DoF, it is possible to additionally check whether the user has entered a specific area or, even if the user has entered, whether the user is looking at the display with Orientation 3DoF (considering only the direction of the head, not the gaze), and to apply a technology for recognizing the user when the user has entered a specific area while looking at the display. In addition, the present disclosure can determine whether the user has been recognized by adding a user-specific gesture to the user recognition method and user direction (Head Orientation 3DoF).
[0247] Accordingly, the present disclosure can provide a window perspective view with a more immersive UX of the metaverse world viewed on a 2D large screen by recognizing the user's head position (X, Y, Z) in space, and can utilize the user's head posture (X, Y, Z) in space as an expression of the user's intention to use the target device from now on. In addition, based on the user position obtained by applying the above-described method, it is possible to determine whether the user exists in a specific area viewed by the camera, and if so, to actively expand, and if the user disappears from the expanded active area, to return to the initial state (so that only a specific area can be viewed).
[0248] As described above, when the user location recognition unit (150) recognizes the user location by applying the methods described above, the display control unit (160) can provide various services based on the recognized user location.
[0249] For example, if a television equipped with a user location recognition device is a metaverse display device, the display control unit (160) can zoom in or out an image being displayed on the display unit (e.g., a screen or a screen) according to the user's movement recognized by the user location recognition unit (150). For example, the image being displayed may be zoomed in as the user moves toward the front of the TV, and may be zoomed out as the user moves away, or vice versa. In addition, depending on the direction in which the user is looking, a previously invisible scene (or object) may become visible, or a previously visible scene (or object) may disappear. In other words, just as a portion of the content visible to the user changes depending on the user's position or direction when the user looks outside or inside through an actual window in reality, a portion of the image displayed on the TV may change depending on the position and direction of the user's head. This is also referred to as a window perspective view in the present disclosure. In this way, the user can experience a VR effect through the TV without having to wear a VR device.
[0250] Each of the parts, modules, or units described above may be software, processors, or hardware parts that execute sequential execution processes stored in memory (or storage units). Each of the steps described in the embodiments described above may be performed by processors, software, or hardware parts. Each of the modules / blocks / units described in the embodiments described above may operate as a processor, software, or hardware. In addition, the methods presented in the embodiments may be implemented as code. This code may be written on a processor-readable storage medium and thus may be read by a processor provided by an apparatus.
[0251] For convenience of explanation, this specification has been described separately in each drawing. However, it is also possible to design new embodiments by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by those skilled in the art, is also within the scope of the present disclosure.
[0252] The device and method according to the present disclosure are not limited to the configuration and method of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.
[0253] In this disclosure, terms such as "first" and "second" may be used to describe various components of the disclosure. However, the various components according to the disclosure should not be interpreted in a limited manner by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not necessarily mean the same user input signals unless the context clearly indicates otherwise.
[0254] The terminology used to describe the present disclosure is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. In addition, the term "and / or" is used to mean all possible combinations of terms. The term "comprises" or "includes" describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as "if" or "when" used to describe the present disclosure are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0255] The best mode for carrying out the invention has been specifically described.
[0256] It will be apparent to those skilled in the art that various modifications and variations can be made to the present embodiments without departing from the spirit or scope of the present embodiments. Accordingly, the present embodiments are intended to include modifications and variations of the present embodiments provided they come within the scope of the appended claims and their equivalents.
Claims
1. In a user position recognition device including a display unit that displays an image, A camera unit installed in the user location recognition device and including at least one of a first camera without depth information and a second camera including depth information; A storage unit that stores camera installation information of at least one of the first camera and the second camera; A user location recognition unit that obtains location information of the user included in the image based on the image acquired by the camera unit, the user's landmark information included in the image, and the camera installation information; and A user location recognition device including a display control unit that controls an image displayed on the display unit based on the user location information recognized by the user location recognition unit.
2. In paragraph 1, A user location recognition device, wherein the user's location information includes at least one of the user's head location information indicating the 3D space coordinates of the user's head and the user's head direction information indicating the direction of the user's head.
3. In the second paragraph, the camera installation information A user position recognition device including at least camera height information indicating a height from the floor to the camera, camera angle information indicating an inclination angle of the camera, or field of view information of the camera.
4. In the third paragraph, the user location recognition unit It operates in a first mode for acquiring three-dimensional spatial coordinates of the user's head based on the user's landmark information in the image acquired by the first camera and the camera installation information. A user position recognition device that, in the first mode, obtains an absolute distance from a user position recognition device to the user's feet, obtains depth information as a relative position from the user's feet to the head based on the absolute distance, obtains two-dimensional coordinates of the user's head based on the depth information, and then rotates the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information to obtain three-dimensional space coordinates of the user's head.
5. In paragraph 4, the user location recognition unit It operates in a second mode for acquiring three-dimensional space coordinates of the user's head based on the user's landmark information in the image acquired by the first camera, the depth information of the image acquired by the second camera, and the camera installation information. A user position recognition device that obtains two-dimensional coordinates of the user's head based on the depth information in the second mode, and obtains three-dimensional spatial coordinates of the user's head by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
6. In paragraph 5, the user location recognition unit A user position recognition device that, in the second mode, aligns an image acquired by the second camera with an image acquired by the first camera and then acquires three-dimensional spatial coordinates of the user's head.
7. In paragraph 6, the user location recognition unit A user position recognition device that operates in the second mode to obtain three-dimensional space coordinates of the user's head when the user is within the field of view of the second camera, and operates in the first mode to obtain three-dimensional space coordinates of the user's head when the user is outside the field of view of the second camera but within the field of view of the first camera.
8. In paragraph 7, A user location recognition device wherein the field of view of the first camera is wider than the field of view of the second camera.
9. In paragraph 6, the user location recognition unit It further includes an image control unit that provides an image acquired by the first camera to a landmark detection unit to obtain the user's landmark information. A user location recognition device in which the image control unit sets a first area and a second area in an image input from the first camera, the second area including the first area, and when there is no user in the first area of the image or the user leaves the second area, a part of the remaining area excluding the first area of the image is masked and provided.
10. In the 9th paragraph, the image control unit A user location recognition device that provides an original image without masking when a user is present in the first area of an input image.
11. A method for recognizing a user's location in a user's location recognition device, wherein at least one of a first camera without depth information and a second camera including depth information is installed and a display unit for displaying an image is included. A step of obtaining location information of a user included in the image based on camera installation information of at least one of the first camera and the second camera, an image obtained by at least one of the first camera and the second camera, and landmark information of the user included in the image; and A user location recognition method including a step of controlling an image displayed on the display unit based on the user location information recognized by the user location recognition unit.
12. In paragraph 11, A method for recognizing a user location, wherein the user location information includes at least one of the user head location information indicating the 3D space coordinates of the user's head and the user head direction information indicating the direction of the user's head.
13. In paragraph 12, the camera installation information is A user position recognition method comprising at least camera height information indicating a height from the floor to the camera, camera angle information indicating an inclination angle of the camera, or field of view information of the camera.
14. In the 13th paragraph, the user location recognition step A step of operating in a first mode for obtaining three-dimensional space coordinates of a user's head based on the user's landmark information in the image obtained by the first camera and the camera installation information, The step of operating in the above first mode is: A step of obtaining the absolute distance from the user's location recognition device to the user's feet, A step of obtaining depth information based on the relative position from the user's feet to the head based on the above absolute distance. A step of obtaining two-dimensional coordinates of the user's head based on the above depth information, and A user position recognition method comprising a step of obtaining three-dimensional space coordinates of the user's head by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
15. In the 14th paragraph, the user location recognition step A step of operating in a second mode for acquiring three-dimensional space coordinates of the user's head based on the user's landmark information in the image acquired by the first camera, depth information of the image acquired by the second camera, and the camera installation information, The step of operating in the above second mode is: A step of obtaining two-dimensional coordinates of the user's head based on the above depth information, and A user position recognition method comprising a step of obtaining three-dimensional space coordinates of the user's head by rotating the two-dimensional coordinates of the user's head along the X-axis by the camera angle information in the camera installation information.
16. In paragraph 15, A method for recognizing a user position, wherein the step of operating in the second mode comprises aligning an image acquired by the second camera to an image acquired by the first camera and then acquiring three-dimensional spatial coordinates of the user's head.
17. In paragraph 16, the user location recognition step A user position recognition method, wherein when the user is within the field of view of the second camera, the method operates in the second mode to obtain three-dimensional space coordinates of the user's head, and when the user is outside the field of view of the second camera but within the field of view of the first camera, the method operates in the first mode to obtain three-dimensional space coordinates of the user's head.
18. In paragraph 17, A method for recognizing a user location, wherein the field of view of the first camera is wider than the field of view of the second camera.
19. In paragraph 16, the user location recognition step It further includes an image control step for controlling an image acquired by the first camera to obtain the user's landmark information and providing it to the landmark detection unit. A user location recognition method in which the image control step sets a first area and a second area in an image input from the first camera, the second area including the first area, and when there is no user in the first area of the image or the user leaves the second area, a part of the remaining area excluding the first area of the image is masked and provided.
20. In the 19th paragraph, the image control step A user location recognition method that provides an original image without masking when a user is present in the first area of an input image.
Citation Information
Patent Citations
Apparatus and method for presenting display of 3D image using head tracking
KR101313797B1
Ergonomic human computer interface
KR1020080055622A
Method and apparatus of recognizing location of user
KR1020110123532A
Apparatus and method for tracking position using webcam
KR1020120071287A
Coordinate Calculation Acquisition Device using Stereo Image and Method Thereof
KR1020160002510A