Camera Environment Drawing
By receiving camera video data frames, identifying the character's bone structure and calculating the spindle, and generating the environment's 3D floor plane, the problem that existing cameras cannot draw the environment is solved, and efficient environmental drawing of ordinary cameras is achieved.
Patent Information
- Application Number
- CN201980088346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-10
- Filing Date
- 2019-12-31
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2039-12-31
AI Technical Summary
Existing cameras cannot generate an environment's 3D floor plane without a predetermined knowledge of camera position, positioning, or angle and rely on expensive hardware sensors or specific arrangements.
By receiving video data frames captured by the camera, identifying the character's skeleton structure, computing the main axis includes the foot endpoint, and drawing a 3D floor plane based on these positions. Using the character's skeleton tracking and assumptions such as the floor plane being a plane, the character's height is constant, etc., to generate the floor plane of the environment.
Without relying on hardware depth sensors, the environment can be generated, which improves the efficiency and applicability of the camera, reduces costs, and realizes environmental drawing on ordinary cameras.
Smart Images

Figure CN113272818B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to environment rendering. Background Art
[0002] Cameras are becoming ubiquitous in modern society. Businesses of all sizes use cameras, whether for security, inventory tracking, traffic cameras, or other purposes. However, these cameras are limited to image acquisition and do not generate any environmental mapping unless they are pre-programmed, part of a proprietary network, or expensive to set up. Current technologies for mapping the environment rely on specific hardware sensors (e.g., depth sensors) or a specific arrangement or knowledge of the camera's position or positioning in order to map the environment. Summary of the Invention
[0003] One aspect of the present invention relates to a device for mapping an environment that can be observed via a camera, the device comprising a processor; and a memory communicatively coupled to the processor, the memory comprising instructions that, when executed, cause the processor to perform the following operations: receive a sequence of frames of video data captured by the camera; identify a person within the sequence of frames; calculate the principal axes of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axes including foot endpoints; and map a 3D floor plane of the area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames.
[0004] Another aspect of the present invention relates to a method for mapping an environment that can be observed via a camera, the method comprising: receiving a sequence of frames of video data captured by a camera at a processor; identifying a person within the sequence of frames using the processor; calculating the principal axes of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axes including foot endpoints; generating a map of a 3D floor plane of the area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames using the processor; and outputting the map.
[0005] Another aspect of the present invention relates to an apparatus for drawing an environment that can be observed via a camera, the apparatus comprising: a unit for receiving a sequence of frames of video data captured by a camera; a unit for identifying a person within the sequence of frames; a unit for calculating the principal axes of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axes including foot endpoints; a unit for drawing a 3D floor plane of the area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames; and a unit for outputting the 3D floor plane for display. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] In the accompanying drawings, which are not necessarily drawn to scale, the same reference numerals may describe the same components in different views. The same reference numerals with different letter suffixes may represent different instances of the same components. The accompanying drawings generally illustrate various embodiments discussed in this document by way of example and not limitation.
[0007] Figure 1 A diagram illustrating an environment including a camera and showing the camera visible in an image captured by the camera, in accordance with some embodiments.
[0008] Figure 2 Shown are two frames of an environment observable via a camera, according to some embodiments.
[0009] Figure 3 A floor map generated according to some embodiments is shown.
[0010] Figure 4 A flow chart illustrating a technique for mapping a camera environment in accordance with some embodiments is shown.
[0011] Figure 5 An example of a block diagram of a machine on which any one or more of the techniques discussed herein may be performed is generally shown, in accordance with some embodiments. DETAILED DESCRIPTION
[0012] Described herein are systems and methods for mapping an environment observable via a camera. The systems and methods described herein can map an environment without any prior knowledge of non-intrinsic characteristics of the camera, such as position or orientation (e.g., angle). The camera can be a camera without a hardware-based depth sensor. The camera can be a black and white camera, a color camera (e.g., a red-green-blue (RGB) camera), a laser, an infrared camera, a sensory capture device, a sensor-based detector, and the like.
[0013] The systems and methods described herein use information from the movement of objects within images captured by a camera to map an environment observable via the camera. The movement of an object (e.g., a person) can be used to determine the floor plane of the environment observable via the camera. The floor plane can be a 2D plane within a 3D environment. The floor plane can be determined by observing the object in two or more different images (e.g., frames) captured by the camera. The orientation of the object can be automatically determined based on the images, including a first portion of the object that is in contact with or adjacent to the floor of the environment, and a second portion of the object that is opposite the first portion. For example, when the object is a person, one or both feet of the person can be identified as the first portion in the first frame and the second frame, for example, using skeletal tracking of the person, while the head or upper torso of the person is the second portion. The orientation of the person can be used to help track the person's movement throughout the environment. In an example, the position of one or both feet of the person (or, for example, the average position between the person's feet) can be used to map the floor. The two or more frames do not need to be consecutive or in chronological order. The systems and techniques described herein use at least two frames that include the object in different positions (e.g., shifted).
[0014] In an example, the object may be an animal (e.g., a dog), a robotic device (e.g., an autonomous vacuum cleaner), etc. Any object may be used that has a determinable height (e.g., the height may be estimated or assumed based on an image of the object, such as the estimated height of a person), that generally remains fixed, and that is in contact with the floor (e.g., like a foot, the object does not necessarily always contact the floor, but does contact the floor when walking).
[0015] After mapping the floor plane, other aspects of the environment visible in the image captured by the camera can be determined. For example, a moving heat map can be generated, the location of furniture (e.g., furniture on which a person sits or lies, or furniture not used to support a person, such as a bookshelf or table) or entry or exit points (e.g., doors) can be generated. The floor plane with or without additional aspects of the environment can be displayed on a user interface, such as a display device.
[0016] The floor plane of an environment observable via a camera can be generated without any predetermined information, for example by using skeletal tracking to estimate the position of one or both feet of a person (or the position between or adjacent to one or both feet of a person). For example, knowledge of information within the environment or information related to the camera position, placement, or angle may be unknown, but the floor plane can still be determined using the systems and methods described herein. Prior techniques for determining environmental information relied on predetermined knowledge of some aspect of the environment or the camera position, placement, or angle. The present system and method improves the efficiency of cameras used for security, tracking, inventory, and the like by generating information that cannot be determined otherwise (e.g., the floor plane). The present system and method also improves camera efficiency by not requiring any predetermined knowledge or setup, thereby allowing the use of cheaper cameras, after-sales solutions, or the generation of the floor plane without requiring technical expertise in generating the floor plane.
[0017] The following sections describe in detail systems and methods that can be used to infer the relationship between a camera (e.g., a single, fixed-position RGB camera) and the floor plane observed by the camera using only intrinsic characteristics of the camera and observations of image-space markers of objects (e.g., human skeletons captured by the camera in the scene). In an example, skeleton tracking can be performed.
[0018] Figure 1 A diagram is shown that includes a camera 102 and illustrates an environment 106 visible in images captured by the camera, according to some embodiments. A person 104 is visible within the environment 106. The person 104 is tracked as it moves throughout the environment 106, and based on this data, a floor plane of the environment 106 is generated. In another example, objects other than the person can be tracked as they move to generate the floor plane.
[0019] In an example, camera 102 is a camera without a hardware-based depth sensor. For example, camera 102 can be a typical RGB camera. For example, camera 102 can include one or more charge-coupled device (CCD) sensors, complementary metal oxide semiconductor (CMOS) sensors, N-type metal oxide semiconductor (NMOS) or other light sensors. In other examples, camera 102 can include two or more cameras. In some examples, camera 102 can be a black and white camera or a grayscale camera, or an infrared camera.
[0020] Camera 102 can detect moving objects in environment 106. The moving object can be a person 104. As person 104 moves through environment 106, person 104 can be tracked. This can be accomplished using skeletal tracking of person 104. Environment 106 can be identified as a plurality of pixels in an image captured by camera 102. The pixels can be classified as being part of person 104 or as being background. Based on the pixels identified as being part of person 104, joints and bones can be identified for person 104. The skeleton can include line segments connecting the identified joints.
[0021] As person 104 moves throughout environment 106, different pixels may be identified as corresponding to person 104. Based on the joint or bone pixels, pixels corresponding to the floor plane may be identified based on the movement of person 104 over time. For example, a time sequence or similarity of image data may be used to identify floor plane pixels. The time sequence may include a time sequence of images, where floor plane pixels are determined based on the sequence, or the time sequence may use a specific confidence level of image data captured by a camera to identify floor plane pixels. Similarity of image data may include using multiple cameras and time aligning the outputs (e.g., pixels of person 104 at a specific time) and identifying similarities to determine floor plane pixels.
[0022] The skeleton and joints can be used to extract the lower extremities from image data of a moving object captured over time by camera 102. Based on the skeleton and the lower extremities, pixels or points that may be next to the object can be identified (e.g., a foot pixel or point that may be next to or on a floor pixel can be identified).
[0023] As person 104 moves, different floor pixels or points can be identified. These pixels or points can be identified using an assumed height and projected geometry of person 104. Assuming the height is constant as person 104 moves, different pixel distances from the distal ends of lower limb pixels to pixels of the head, shoulders, upper body, and so on can be used to determine the distance moved by person 104. In another example, instead of or in addition to the assumed height, orthogonality of person 104 to the floor can be assumed.
[0024] A 2D or 3D floor plan of the space being viewed (within environment 106) can be based on a floor inferred from skeletal information extracted from image data of a moving object (e.g., person 104) captured over time by camera 102. In an example, a heat map of the movement of person 104 or other objects can be created after the floor plan is generated.
[0025] Figure 2Two frames 200A and 200B of an environment 206 of a camera 202 are shown, according to some embodiments. To determine the floor plane (and the depth of the floor at an arbitrary scale) from the camera 102 (which can be a single camera in a fixed position), skeletal observations are used along with one or more assumptions about the person and the floor. For example, one or more of the following can be assumed: 1. The floor is planar; 2. People do not change size over a short time span; 3. Gravity exists and is constant within the observation area; 4. People present in the environment are oriented parallel to the gravity vector (e.g., the person is standing and walking so that their head is substantially directly above their feet); 5. The feet are in contact with, near, or very close to the floor.
[0026] Camera 202 can be used to capture images at different times, such as frames 200A and 200B. A set of skeletal observations of a single individual (e.g., person 204) over a constrained time period can be generated based on the images. In an example, using the assumptions listed above, a floor plane or a 3D person trajectory can be generated based on the set of skeletal observations.
[0027] For illustrative purposes, frames 200A and 200B include details that may be generated after being captured by camera 202. For example, person 204 is shown with identified primary axis 208 and foot endpoints 210A (which may be points between the feet of person 204, or another location such as the intersection of a line segment connecting the person's feet with the primary axis). Primary axis 208 and foot endpoints 210A may be generated from frames 200A and 200B, but primary axis 208 and foot endpoints 210A do not appear in images originally captured from camera 202.
[0028] Primary axis 208 is the axis of character 204 when the character is standing, walking, running, or otherwise moving (e.g., on the floor from frame 200A to 200B). In an example, primary axis 208 is a 3D line segment extending from a point in the upper body of character 204 (e.g., the character's head, shoulders, chest, etc.) to the floor directly below (or to a point at or between the character's feet). Other examples of primary axis 208 may be from the character's 204 neck to the floor, from the character's 204 center of mass to the floor, from the head to the midpoint between the character's hips or knees, etc. In an example, primary axis 208 may be defined as being substantially perpendicular to the floor (e.g., not at an angle, such as from the neck to the right foot, as the orientation of the line segment from the foot to the neck may change during different phases of a step). In an example, primary axis 208 may be any axis where the size and orientation of the axis of character 204 remain relatively constant during typical movements of standing and walking. In an example, when the character 204 is not standing or walking, the primary axis 208 may not exist or those frames or images may be ignored.
[0029] The assumptions listed above can be restated in terms of the principal axes 208. For example, Assumption 2 can be restated as the length of the principal axes 208 (e.g., in 3D) of a given individual remains approximately constant across all temporally proximate observations of that individual (e.g., from frame to frame). Assumption 3 can be restated as all apparent principal axes are subject to the same gravity. Assumption 4 can be restated as the principal axes 208 are approximately parallel to the gravity vector, with the foot endpoint 210A at the "lower" end of the principal axes 208 (e.g., closest to the floor). Assumption 5 can be restated as the foot endpoint 210A always lies close to the floor plane in space.
[0030] These assumptions about the floor plane, the principal axes 208, and the relationship between them provide sufficient information to compute the floor plane from only 2D skeletal observations of the person 204 using a single camera 202. A sequence of images from the camera 202 can be used to generate a set of skeletal observations O.
[0031] For each 2D labeled bone in O, the image space localization of the projection of the foot endpoints 210A and 210B of the principal axis can be calculated. In examples, the localization can be generated using geometric estimation, machine learning, the "midpoint of the foot" approximation described above, etc. The set O can be used to generate a set of matched axis endpoints A, where each element a in A contains a pair of 2D localizations (h2:=head localization, f2:=foot endpoint, e.g., 210A or 210B) for the corresponding element o in O.
[0032] For all principal axes a in A, a.h2 and a.f2 are 2D points corresponding to the image plane projections of the unknown 3D points a.h3 and a.f3, respectively. Using projective geometry, the ray on which the 3D point is located can be determined by the 2D projection and the observation camera 202. The principal axis a can be, for example, a 3D vector defined by a point and a ray (e.g., direction and length). Other definitions of the principal axis can include a line segment connected by two specified endpoints (e.g., 3D anchor points).
[0033] Using the assumption that the floor is planar and that foot endpoints 210A and 210B tend to lie on or near the floor, the floor plane can be determined as the 3D plane that maximally contains (or lies below, which is equivalent to within some scale difference) foot positioning a.f3 for all a's in A.
[0034] In this example, a.h3 and a.f3 are separated by a constant distance for all a in A, for example, based on Assumption 2. In this example, the scale of the floor plane determination can be relative to other objects in the environment. For example, an absolute scale may not be generated. Therefore, it may not be necessary to know or determine the numerical distance between the 3D endpoints. The principal axis can use this feature of the system without loss of generality, where for all a in A, ||a.h3 - a.f3|| = 1.
[0035] Now let's turn to the following assumption: given a unit vector g representing the direction of gravity (e.g., pointing toward the center of the Earth), the dot product of g and (a.f3 - a.h3) is equal to 1 (in this example, the distance between the head and the feet can be defined as 1). These two unit vectors are identical. Using the assumption that g is constant for all a in A, the third constraint on the system is: for all a and b in A, (a.f3 - a.h3)·(b.f3 - b.h3) = 1.
[0036] Given an observation camera 202, in the environment, for all a's in A, a.h3 projects to a.h2 and a.f3 projects to a.f2. For all a's in A, ||a.h3 - a.f3|| = 1. For all a's and b's in A, (a.f3 - a.h3)·(b.f3 - b.h3) = 1.
[0037] In an example, the floor plane is estimated based on a close approximation of the solution to these constraints. In an example, an optimizer is generated to iteratively apply (try to fit) ||a.h3 - a.f3|| = 1 and (a.f3 - a.h3) (b.f3 - b.h3) = 1, while ensuring that a.h3 projects to a.h2 and a.f3 projects to a.f2. In another example, a general nonlinear optimization tool can be used to estimate the floor plane based on these formulas.
[0038] After determining an optimized or estimated solution to the system described above, the floor plane can be recovered from the 3D positioning of the feet (e.g., endpoints 210A and 210B). The floor plane can be defined by a point p and a normal vector n. In an example, for all a in A, the floor plane can be the result of setting p to the average of a.f3 and n to the average of (a.h3 - a.f3). In another example, p and n can be determined based on the value of a single principal axis (e.g., 208 in frame 200A), and then a random sample consensus (RANSAC) or similar algorithm can be used to find which axis produces the best-fitting floor plane.
[0039] The techniques described herein can be used to generate a floor plane using very few elements of A and the corresponding O. For example, when A contains at least two elements a and b, a floor plane can be generated such that ||a.h2 — a.f2|| ≠ ||b.h2 — b.f2||. Since measurements can be noisy and assumptions are approximate, it can be useful to make multiple observations to prevent erroneous inputs from corrupting the output.
[0040] In another example, the technique can relax or not use some of these constraints or assumptions, or substitute others, while still achieving high-quality results when generating a floor plane. For example, by relaxing the requirement for a single individual, multiple people can be tracked. In another example, a simplification of Assumption 4 above to the following definition can be used: for all a in A, (a.h3 - a.f3) is orthogonal to the view direction of camera 202. In this example, adjusting the parameters of the nonlinear optimizer to allow it to "bend" the constraints to resolve the resulting inconsistencies can incorporate these differences in the assumptions to still generate a floor plane.
[0041] Figure 3 A floor map 300 generated according to some embodiments is shown. Floor map 300 can be an output of the systems and techniques described herein. Floor map 300 shows a 3D or 2D representation of a floor plane within an environment. Points or pixels 302A, 302B, ..., 302N represent foot endpoints generated using the systems and techniques described herein. Floor map 300 can be a best fit or approximation of points or pixels 302A-302N. In an example, the floor plane is defined by a point (e.g., 302A) and a vector (e.g., a line connecting points 302B to 302N). Based on the floor plane, floor map 300 can be generated, for example, based on the location of a character moving throughout the environment. Other outputs can include floor plane coordinates (e.g., points and vectors), camera orientation (e.g., position or angle) relative to the floor plane (e.g., six-axis orientation of x, y, z, and pitch, yaw, and roll relative to the "ground"). In an example, a heat map can be generated based on floor plane 300 and data on the character's movement throughout the environment, showing, for example, where the character is sitting, standing, stopped, or moving.
[0042] Figure 4 A flow chart illustrating a technique 400 for mapping a camera environment according to some embodiments is shown. The technique 400 includes an operation 402 of receiving a sequence of frames or images of video data, such as captured by a camera. In an example, the camera lacks any hardware-based depth sensing capabilities. The technique 400 includes an operation 404 of identifying a person or other object (e.g., a moving object or a movable object) within the sequence of frames.
[0043] Technique 400 includes an operation 406 for calculating the principal axis of the character (e.g., including the foot endpoints) using the character's skeletal structure. Other operations may replace operation 406, such as using different object recognition techniques to identify the object (e.g., including the object's orientation). In this example, the principal axis is a line segment extending from the character's head to the midpoint between the character's feet.
[0044] Technique 400 includes an operation 408 of generating a map or rendering of a 3D space. The map may include a 3D floor plane of the area captured in the sequence of frames based on the major axes in each frame in the sequence of frames (e.g., based on the positions of the foot endpoints identified in each frame in the sequence of frames). The camera may be a single camera, and the 3D floor plane is rendered based only on data received from the single camera. The 3D floor plane may be rendered based on the assumption that the major axes of a character vary based on the distance from the camera while the height of the character remains constant. For example, the 3D floor plane may be rendered based on the assumption that the major axes are parallel to the gravity vector. Technique 400 may include determining an angle of orientation of the camera relative to the 3D floor plane of the area.
[0045] Figure 5 An example of a block diagram of a machine 500 is generally shown, according to some embodiments, on which any one or more of the techniques (e.g., methodologies) discussed herein may be performed. In alternative embodiments, the machine 500 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine 500 may operate as a server machine, a client machine, or both in a server-client network environment. In an example, the machine 500 may function as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. The machine 500 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, a switch, or a bridge, or any machine capable of executing (sequentially or otherwise) instructions specifying actions to be taken by the machine. Furthermore, while a single machine is shown, the term "machine" should also be understood to include any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein (e.g., cloud computing, software as a service (SaaS), other computer cluster configurations).
[0046] As described herein, examples may include or operate on a logic unit, a plurality of components, modules, or mechanisms. A module is a tangible entity (e.g., hardware) that is capable of performing a specified operation when in operation. A module includes hardware. In an example, the hardware may be specifically configured to perform a specific operation (e.g., hard-wired). In an example, the hardware may include a configurable execution unit (e.g., a transistor, a circuit, etc.) and a computer-readable medium containing instructions, wherein the instructions configure the execution unit to perform a specific operation when in operation. Configuration may occur under the guidance of the execution unit or a loading mechanism. Therefore, when the device is operating, the execution unit is communicatively coupled to the computer-readable medium. In this example, the execution unit may be a member of more than one module. For example, during operation, the execution unit may be configured by a first instruction set to implement a first module at a point in time, and reconfigured by a second instruction set to implement a second module.
[0047] The machine (e.g., a computer system) 500 may include a hardware processor 502 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 504, and a static memory 506, some or all of which may communicate with each other via an interconnection link (e.g., a bus) 508. The machine 500 may also include a display unit 510, an alphanumeric input device 512 (e.g., a keyboard), and a user interface (UI) navigation device 514 (e.g., a mouse). In an example, the display unit 510, the alphanumeric input device 512, and the UI navigation device 514 may be a touch screen display. The machine 500 may additionally include a storage device (e.g., a drive unit) 516, a signal generating device 518 (e.g., a speaker), a network interface device 520, and one or more sensors 521 (e.g., a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors). The machine 500 may include an output controller 528, such as a serial (e.g., Universal Serial Bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).
[0048] The storage device 516 may include a non-transitory machine-readable medium 522 on which is stored one or more data structures or instructions 524 (e.g., software) embodying any one or more of the techniques or functionality described herein or utilized by these techniques or functionality. The instructions 524 may also reside, completely or at least partially, within the main memory 504, within the static memory 506, or within the hardware processor 502 during execution of the instructions by the machine 500. In an example, one or any combination of the hardware processor 502, the main memory 504, the static memory 506, or the storage device 516 may constitute a machine-readable medium.
[0049] Although machine-readable medium 522 is illustrated as a single medium, the term “machine-readable medium” may include a single medium or multiple media (eg, a centralized or distributed database, or associated caches and servers) configured to store one or more instructions 524 .
[0050] The term "machine-readable medium" may include any medium that is capable of storing, encoding, or carrying instructions for execution by the machine 500, and the instructions cause the machine 500 to perform any one or more of the techniques of the present disclosure; or that is capable of storing, encoding, or carrying data structures used by or associated with the instructions. Non-limiting examples of machine-readable media may include solid-state memory and optical and magnetic media. Specific examples of machine-readable media may include non-volatile memory (e.g., semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM))) and flash memory devices; magnetic disks (e.g., internal hard disks and removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0051] Instructions 524 may also be sent or received using a transmission medium over a communications network 526 via a network interface device 520 utilizing any of a variety of transmission protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Example communications networks may include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), a mobile telephone network (e.g., a cellular network), a plain old telephone (POTS) network, and a wireless data network (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards (known as ), IEEE 802.16 series of standards (known as )), IEEE 802.15.4 family of standards, peer-to-peer (P2P) networks, and the like. In an example, the network interface device 520 may include one or more physical jacks (e.g., Ethernet, coaxial cable, or telephone jacks) or one or more antennas for connecting to the communication network 526. In an example, the network interface device 520 may include multiple antennas to communicate wirelessly using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) technology. The term "transmission medium" should be understood to include any intangible medium capable of storing, encoding, or carrying instructions for execution by the machine 500 and including digital or analog communication signals, or other intangible media for facilitating the communication of such software.
[0052] Each of these non-limiting examples can stand on its own or can be combined with one or more of the other examples in various permutations or combinations.
[0053] Example 1 is a device for mapping an environment observable via a camera, the device comprising: a processor; and a memory communicatively coupled to the processor, the memory comprising instructions that, when executed, cause the processor to: receive a sequence of frames of video data captured by a camera; identify a person within the sequence of frames; calculate a principal axis of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axis including foot endpoints; and map a 3D floor plane of an area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames.
[0054] In Example 2, the inventive subject matter according to Example 1 includes wherein the camera lacks any hardware-based depth sensing capabilities.
[0055] In Example 3, the inventive subject matter of Examples 1-2 includes: wherein the primary axis is a line segment extending from the character's head to a point midway between the character's feet.
[0056] In Example 4, the inventive subject matter of Examples 1-3 includes wherein the camera is a single camera, and the 3D floor plane is rendered based solely on data received from the single camera.
[0057] In Example 5, the inventive subject matter according to Examples 1-4 includes wherein the 3D floor plane is rendered based on the assumption that the major axes of the character vary based on the distance from the camera while the height of the character remains constant.
[0058] In Example 6, the inventive subject matter according to Examples 1-5 includes wherein the 3D floor plane is drawn based on the assumption that the major axes are parallel to the gravity vector.
[0059] In Example 7, the inventive subject matter of Examples 1-6 includes wherein the instructions further cause the processor to determine an angle of orientation of the camera relative to a 3D floor plane of the area.
[0060] Example 8 is a method for mapping an environment observable via a camera, the method comprising: receiving, at a processor, a sequence of frames of video data captured by a camera; identifying, using the processor, a person within the sequence of frames; calculating the principal axes of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axes including foot endpoints; generating, using the processor, a map of a 3D floor plane of the area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames; and outputting the map.
[0061] In Example 9, the inventive subject matter according to Example 8 includes wherein the camera lacks any hardware-based depth sensing capabilities.
[0062] In Example 10, the inventive subject matter of Examples 8-9 includes wherein the primary axis is a line segment extending from the character's head to a point midway between the character's feet.
[0063] In Example 11, the inventive subject matter of Examples 8-10 includes wherein the camera is a single camera, and the 3D floor plane is rendered based solely on data received from the single camera.
[0064] In Example 12, the inventive subject matter according to Examples 8-11 includes wherein the 3D floor plane is rendered based on the assumption that the major axes of the character vary based on the distance from the camera while the height of the character remains constant.
[0065] In Example 13, the inventive subject matter according to Examples 8-12 includes wherein the 3D floor plane is drawn based on the assumption that the major axes are parallel to the gravity vector.
[0066] In Example 14, the inventive subject matter of Examples 8-13 includes determining an angle of an orientation of a camera relative to a 3D floor plane of the area.
[0067] Example 15 is a device for drawing an environment observable via a camera, the device comprising: a unit for receiving a sequence of frames of video data captured by a camera; a unit for identifying a person within the sequence of frames; a unit for calculating the principal axes of the person based on a skeletal structure determined for the person in each frame in the sequence of frames, the principal axes including foot endpoints; a unit for drawing a 3D floor plane of the area captured in the sequence of frames based on the positions of the foot endpoints identified in each frame in the sequence of frames; and a unit for outputting the 3D floor plane for display.
[0068] In Example 16, the inventive subject matter according to Example 15 includes wherein the camera is a single camera, and the 3D floor plane is rendered based solely on data received from the single camera.
[0069] In Example 17, the inventive subject matter according to Examples 15-16 includes wherein the camera lacks any hardware-based depth sensing capabilities.
[0070] In Example 18, the inventive subject matter of Examples 15-17 includes: wherein the primary axis is a line segment extending from the character's head to a point midway between the character's feet.
[0071] In Example 19, the subject matter of Examples 15-18 includes wherein the 3D floor plane is rendered based on the assumption that the major axes of the character vary based on the distance from the camera while the height of the character remains constant.
[0072] In Example 20, the inventive subject matter according to Examples 15-19 includes determining an angle of an orientation of a camera relative to a 3D floor plane of the area.
[0073] Example 21 is at least one machine-readable medium comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform operations to implement any of Examples 1-20.
[0074] Example 22 is an apparatus comprising means for implementing any one of Examples 1-20.
[0075] Example 23 is a system for implementing any one of Examples 1-20.
[0076] Example 24 is a method for implementing any one of Examples 1-20.
[0077] The method examples described herein may be at least partially machine or computer-implemented. Some examples may include a computer-readable medium or machine-readable medium encoded with instructions, which are operable to configure an electronic device to perform the methods described in the above examples. The implementation of these methods may include code, for example, microcode, assembly language code, higher-level language code, etc. Such code may include computer-readable instructions for performing various methods. The code may form part of a computer program product. In addition, in an example, the code may be tangibly stored on one or more volatile, non-transitory or non-volatile tangible computer-readable media, for example during execution or at other times. Examples of these tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., compact disks and digital video disks), magnetic tapes, memory cards or memory sticks, random access memories (RAMs), read-only memories (ROMs), etc.
Claims
1. A device for mapping an environment observable via a camera, the device comprising: processor; as well as a memory communicatively coupled to the processor, the memory comprising instructions that, when executed, cause the processor to: receiving a sequence of frames of video data captured by the camera; identifying a person within the sequence of frames; calculating a principal axis of the character based on the skeletal structure determined for the character in each frame in the sequence of frames, the principal axis including a foot endpoint for each frame; as well as Based on the positions of the foot endpoints identified in at least two frames in the sequence of frames, a 3D floor plane of the area captured in the sequence of frames is drawn, wherein the at least two frames include foot endpoints in different positions that can be observed by the camera, thereby drawing the 3D floor plane based on the different positions of the foot endpoints.
2. The device according to claim 1, wherein The camera lacks any hardware-based depth sensing capabilities.
3. The device according to claim 1, wherein The main axis is a line segment extending from the head of the character to the midpoint between the feet of the character.
4. The device according to claim 1, wherein The camera is a single camera, and the 3D floor plane is rendered based only on data received from the single camera.
5. The apparatus according to claim 1, wherein The 3D floor plane is drawn based on the assumption that the principal axis of the character varies based on the distance from the camera, while the height of the character remains constant.
6. The apparatus according to claim 1, wherein The 3D floor plane is drawn based on the assumption that the major axes are parallel to the gravity vector.
7. The apparatus according to claim 1, wherein The instructions further cause the processor to determine an angle of orientation of the camera relative to the 3D floor plane of the area.
8. A method for mapping an environment observable via a camera, the method comprising: receiving, at a processor, a sequence of frames of video data captured by a camera; identifying, using the processor, a person within the sequence of frames; calculating a principal axis of the character based on the skeletal structure determined for the character in each frame in the sequence of frames, the principal axis including a foot endpoint for each frame; generating, using the processor, a map of a 3D floor plane of an area captured in the sequence of frames based on positions of the foot endpoints identified in at least two frames in the sequence of frames, wherein the at least two frames include foot endpoints in different positions observable by the camera, thereby rendering the 3D floor plane based on the different positions of the foot endpoints; as well as The map is output.
9. The method according to claim 8, wherein The camera lacks any hardware-based depth sensing capabilities.
10. The method according to claim 8, wherein The main axis is a line segment extending from the head of the character to the midpoint between the feet of the character.
11. The method according to claim 8, wherein The camera is a single camera, and the 3D floor plane is rendered based only on data received from the single camera.
12. The method according to claim 8, wherein The 3D floor plane is drawn based on the assumption that the principal axis of the character varies based on the distance from the camera, while the height of the character remains constant.
13. The method according to claim 8, wherein The 3D floor plane is drawn based on the assumption that the major axes are parallel to the gravity vector.
14. The method of claim 8, further comprising determining an angle of orientation of the camera relative to the 3D floor plane of the area.
15. A device for mapping an environment observable via a camera, the device comprising: means for receiving a sequence of frames of video data captured by a camera; means for identifying a person within said sequence of frames; means for calculating a principal axis of the character based on a skeletal structure determined for the character in each frame in the sequence of frames, the principal axis including a foot endpoint for each frame; means for rendering a 3D floor plane of an area captured in the sequence of frames based on positions of the foot endpoints identified in at least two frames in the sequence of frames, wherein the at least two frames include foot endpoints in different positions observable by the camera, thereby rendering the 3D floor plane based on the different positions of the foot endpoints; as well as Means for outputting the 3D floor plane for display.