Pose detection using hot data
By using thermopile sensor module and neural network technology, combined with low-resolution image processing, non-invasive posture detection in a private environment is achieved, solving the problem of privacy violations by traditional methods and improving the accuracy and privacy protection of detection.
Patent Information
- Application Number
- CN202380043635.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-30
- Filing Date
- 2023-02-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-02-27
AI Technical Summary
It is difficult for prior art to achieve non-invasive posture detection in private environments, especially in situations where monitoring falls and other activity analysis is required, traditional methods such as high-resolution cameras may collect personnel identity information, which violates privacy protection requirements.
The thermal data of humans is obtained by using the thermopile sensor module and combined with neural network technology, data is extracted from bounding boxes in the image to determine the human pose. The system can work effectively in low-resolution data environments, avoiding the collection of person identity information.
It realizes non-invasive posture detection in a private environment, can effectively monitor falls and other activities without invading privacy, and improves the accuracy and privacy protection of detection.
Smart Images

Figure CN119301428B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of U.S. Serial No. 17 / 708,493, filed on March 30, 2022, and titled "POSE DETECTION USING THERMAL DATA". U.S. Serial No. 17 / 708,493 is a partial continuation of U.S. Serial No. 17 / 516,954, filed on November 2, 2021, and titled "USER INTERFACE FOR DETERMINING LOCATION, TRAJECTORY AND BEHAVIOR", and claims the benefit and priority thereof. U.S. Serial No. 17 / 516,954 is a partial continuation of U.S. Serial No. 17 / 232,551, filed on April 16, 2021, and titled "THERMAL DATA ANALYSIS FOR DETERMINING LOCATION, TRAJECTORY AND BEHAVIOR". U.S. Serial No. 17 / 232,551 is a continuation of U.S. Serial No. 17 / 178,784 (also known as U.S. Patent No. 11,022,495, published on June 1, 2021), filed on February 18, 2021, and titled "MONITORING HUMAN LOCATION, TRAJECTORY AND BEHAVIOR USING THERMAL DATA". U.S. Serial No. 17 / 178,784 claims the priority of U.S. Provisional Serial No. 62 / 986,442, filed on March 6, 2020, and titled "MULTI - WIRELESS - SENSOR SYSTEM, DEVICE, AND METHOD FOR MONITORING HUMAN LOCATION AND BEHAVIOR". All of the above - mentioned disclosures are hereby incorporated by reference in their entirety for all purposes. Technical Field
[0003] The present disclosure generally relates to pose detection using thermal data, and more particularly to using pose information for fall detection and other activity analysis. Background Art
[0004] While some merchants may attempt to use basic machines to count the number of people entering and leaving a particular door leading to a store, such information is very limited for analyzing the movements of these people within the store. Merchants may be very interested in better understanding the movement, trajectories, and activities of customers within their stores. For example, a merchant may be interested in knowing whether a particular display in a specific aisle of the store attracts more customers to that aisle. Additionally, a merchant may be interested in knowing how many customers who walk along aisle #4 also walk along aisle #5 and how many customers who walk along aisle #4 skip aisle #5 and instead continue along aisle #6. This data can help merchants optimize their operations and maximize profits.
[0005] Merchants may also be interested in more comprehensively understanding the time-related general traffic patterns in their stores. To help merchants better allocate their own resources and optimize their business relationships with cooperating third parties, merchants may want to know the traffic patterns during the busiest times of the day, during the busiest days of the week, during specific months, and / or during specific years. Additionally, to help identify abnormal behavior and / or detect incidents in real time, merchants may want to obtain more information about the spatial and / or temporal patterns of traffic and occupancy levels.
[0006] Furthermore, to analyze the well-being of residents and determine whether residents are eligible to live independently, assisted living providers typically want to obtain the spatial and temporal movement data of tenants. For example, the provider may want to analyze the movement speed of a tenant based on the tenant's indoor location over time, calculate the total calories consumed based on the tenant's movement, and / or monitor the tenant's body temperature.
[0007] Detecting human poses may also be helpful for fall detection and other activity analysis. However, in private residences or other private environments involving privacy, pose detection is typically required. In this regard, users typically prefer that any data collection or scanners do not acquire or store personal identity information of people. Thus, a non-invasive technology is needed to implement such fall detection analysis in residences. In this regard, high-resolution cameras or other technologies may not be advisable because such cameras may use key points to acquire and / or store facial data or other personal identity information. Key points may include a subset of points on the human skeleton, which requires obtaining detailed information about the human. Identifying key points in low-resolution images is very difficult and typically impossible, so high-resolution technologies are typically required to identify key points. In contrast, residential environments tend to use low-resolution data. An example of low-resolution data may be thermal data. Thus, there is a need to use thermal data in combination with algorithms that can process the thermal data to provide pose detection.
[0008] In addition to cameras, other solutions have also been used for fall detection, such as watches (e.g., using accelerometers and gyroscopes), radar, or lidar. However, watch technology requires humans to actively wear the technology. In addition, the use of radar may result in false alarms (e.g., the radar may be triggered by a pet). In addition, radar technology generally cannot effectively distinguish between stationary people as part of the detection process. In addition, existing systems may require the analysis of densely aggregated data points, which typically results in lower accuracy. The analysis of such dense points may be more expensive because additional computing power may be required to distinguish points in a dense point cloud. Summary of the Invention
[0009] In various embodiments, the system may implement a method that includes: receiving, by a processor, an image of a human from a sensor; receiving, by the processor, the position (placement) of a bounding box on the image, where the bounding box contains pixel data of the human in the image; obtaining, by the processor, bounding box data from within the bounding box; and determining, by the processor, the pose of the human based on the bounding box data.
[0010] In various embodiments, the method may further include training, by the processor, a neural network to predict the position of the bounding box on the image. The method may further include training, by the processor, the neural network using pixel data of the human, thermal data of the human, and environmental data. The method may further include adjusting, by the processor, the algorithm of the neural network based on the environmental data, where the environmental data includes at least one of the following: environmental temperature, indoor temperature, floor plan, non-human thermal objects, gender of the human, age of the human, height of the sensor, clothing of the human, or body weight of the human.
[0011] In various embodiments, the sensor may obtain thermal data about the human. The user may indicate the position of the bounding box on the image. Determining the pose may be performed for a frame of the image captured by the sensor. Determining the pose may further include determining an aggregated pose across multiple frames over a period of time. The image may be part of a video clip of the human. Obtaining the bounding box data may include obtaining the bounding box data in at least one of the following cases: over time or during an initial calibration session. The pose may include at least one of the following: sitting, standing, lying down, exercising, dancing, running, or eating.
[0012] In various embodiments, the method may further include: determining a fall by a processor based on the aggregated posture changing from at least one of a standing posture or a sitting posture to a lying posture and the lying posture persisting for a certain amount of time. The method may include extracting discriminative features from multiple frames of an image by a processor using pattern recognition. The method may include restricting the resolution of an image by a processor based on at least one of the following: privacy issues, power consumption of a sensor, cost of pixel data, bandwidth of pixel data, computational cost, or computational bandwidth. The method may include labeling the posture of a human in an image by a processor.
[0013] In various embodiments, the method may further include: determining the temperature of a human in a space by a processor based on IR energy data regarding IR energy from the human; determining the position coordinates of the human in the space by a processor; comparing, by a sensor system, the position coordinates of the human with the position coordinates of a fixture; and determining, by the sensor system, that the human is a human being in response to the temperature of the object being within a range and in response to the position coordinates of the human being different from the position coordinates of the fixture. The method may include analyzing, by a processor, discriminative features of a pattern of a thermal feature of the top of the head of the human to determine the posture of the human. The method may include determining, by a processor, the trajectory of the human based on a change in temperature in pixel data, wherein the temperature is projected onto a pixel grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The subject matter of the present disclosure is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, a more complete understanding of the present disclosure may be best obtained by referring to the detailed description and the claims when considered in conjunction with the accompanying drawings.
[0015] Figure 1A is an exemplary schematic diagram of the main components of a sensor node as part of an overall system according to various embodiments.
[0016] Figure 1B is an exemplary schematic diagram of a gateway and a microprocessor as part of an overall system according to various embodiments.
[0017] Figure 2 is an exemplary data flow diagram according to various embodiments.
[0018] Figure 3 is an exemplary system architecture according to various embodiments.
[0019] Figure 4A and Figure 4B is an exemplary user interface according to various embodiments.
[0020] Figure 5 is an exemplary building layout according to various embodiments.
[0021] Figure 6 is an exemplary building layout according to various embodiments, which shows certain sensor nodes and the coverage area for each sensor node.
[0022] Figure 7 is an exemplary user interface according to various embodiments, which shows the response of the application to detecting sensor nodes in the physical space after the user successfully logs in to the sensor node.
[0023] Figure 8A and Figure 8B is an exemplary user interface according to various embodiments, prompting the user to upload any specific file that can be supplementary visual information representing their physical space to the application.
[0024] Figure 9 is an exemplary user interface according to various embodiments, which shows certain sensor nodes located in a space that can be labeled with names so that the user can understand the background information of each sensor node in their equivalent space.
[0025] Figure 10 is an exemplary user interface according to various embodiments, which shows the ability to set up and calibrate the sensor node.
[0026] Figures 11A to 11C shows an exemplary thermal signature pattern according to various embodiments, which includes a standing posture, a sitting posture, and a lying posture.
[0027] Figure 12 shows an exemplary posture inference process according to various embodiments.
[0028] Figure 13 shows an exemplary fall detection process according to various embodiments. DETAILED DESCRIPTION
[0029] In various embodiments, the system is configured to locate, track, and / or analyze the activities of a living being in an environment. The system does not require input of personnel biometric data. Although the present disclosure may discuss human activities, the present disclosure contemplates tracking any item that can provide infrared (IR) energy, such as an animal or any object, for example. Although the present disclosure may discuss indoor environments, the system can also track in outdoor environments (e.g., outdoor concert venues, outdoor amusement parks, etc.) or a mixed environment of outdoor and indoor environments.
[0030] As Figure 1A 、 Figure 1B and Figure 3As described in more detail below, in various embodiments, the system may include a plurality of sensor nodes 102, a gateway 135, a microprocessor 140, a computing module 350 (e.g., a cloud computing module), a database 360, and / or a user interface 370 (e.g., Figure 4A , Figure 4B , Figure 7 , FIG. 8, Figure 9 and Figure 10 ). Each sensor node 102 may include an enclosure 105, an antenna 110, a sensor module 115, a switch 120, a light emitting diode (LED) 125, and a power source 130.
[0031] In various embodiments, the sensor module 115 may be any type of sensor, such as a thermopile sensor module. The thermopile sensor module 115 may include, for example, a Heimann GmbH sensor module or a Panasonic AMG8833. Each sensor module 115 may be housed in the enclosure 105. The sensor module 115 is configured to measure temperature from a distance by detecting IR energy from an object (e.g., a living being). If the living being has a higher temperature, the living being will emit more IR energy. The thermopile sensing element in the thermopile sensor module 115 may include thermocouples on a silicon chip. The thermocouples absorb the IR energy and generate an output signal indicative of the amount of IR energy. Thus, a higher temperature causes more IR energy to be absorbed by the thermocouples, resulting in a higher signal output.
[0032] In various embodiments, the sensor node 102 interface may be wireless to help reduce the labor and material costs associated with installation. In various embodiments, each sensor node in the sensor nodes 102 may obtain power from any power source 130. The power source 130 may power one or more sensor nodes 102. Each sensor node in the sensor nodes 102 may be powered individually by a battery. The battery may be of low enough power to operate for over 2 years using a single battery (e.g., a 19wh battery). The battery 130 may include batteries from any manufacturer and / or PKCELL batteries, D-cell batteries, or any other battery type. The battery 130 may be housed within a battery holder (e.g., a Bulgin battery holder). The system may also measure the battery voltage of the battery 130 (e.g., a D-cell battery). An analog-to-digital converter located on-board with the antenna 110 (e.g., a Midatronics Dusty PCB antenna) may be used to measure the battery voltage. When a battery voltage measurement is obtained, the system may also add a timestamp to the battery voltage data.
[0033] By adding more sensor nodes 102 to the array of sensor nodes 102, the system can be extended to larger occupancy spaces. In various embodiments, sensor nodes 102 can be added dynamically, where an exemplary user interface for adding sensors is set forth in Figure 7 Specifically, a user can add or remove sensor nodes 102 from an established network of sensor nodes 102 at any time or any location. Sensor nodes 102 can be located anywhere as long as the sensor node 102 location is within the bandwidth of the gateway 135. As such, a large number of sensor nodes 102 can form a mesh network to communicate with the gateway 135. Sensor nodes 102 can communicate with the gateway 135 approximately simultaneously. Each new sensor node 102 connected to the gateway 135 can further extend the boundary of the mesh network, improving system stability and system performance. As such, a large number of sensor nodes 102 can form a mesh network to communicate with the gateway 135. Sensor nodes 102 can communicate with the gateway 135 approximately simultaneously. Each new sensor node 102 connected to the gateway 135 extends the boundary of the mesh network further, improving system stability and system performance. Sensor nodes 102 can be positioned or mounted on any part of a building or on any object. For example, sensor nodes 102 can be mounted into the ceiling, sidewall, or floor of a desired space using any fasteners known in the art. For a 2.5-meter-high ceiling, the distance between sensor nodes 102 can be spaced apart, for example, 4 meters.
[0034] In various embodiments, each sensor node 102 installed in any given space can have a unique number (e.g., MAC address) assigned to the sensor node 102. The system uses the unique number to create a structured network with the numbered sensor nodes 102. As shown in Figure 6 each sensor node can cover a different area and carry the unique number that can be accessed by a user both in a digital environment as well as in a physical environment. Specifically, the unique number is clearly printed on the enclosure 105 of the sensor node and is announced on the user's screen when the user is installing the sensor node. In this way, a user can identify whether the location of their physical sensor node matches the digital representation in the electronic space they have created in the installation application.
[0035] As shown in Figure 7 and Figure 8BAs described in , after adding and setting the sensor node 102 in the system, the sensor node 102 creates a profile. The user is prompted to either scan the QR code located on the sensor to automatically register the digital representation of the physical sensor or manually enter the MAC address of the sensor in question. The sensor node 102 profile may include the height of the sensor node 102, the MAC address, the relative position in space, surrounding objects, and / or background information. The system determines the extent of the coverage area of each sensor node based on the height of the sensor node 102 entered by the user as information in its profile. Some examples of background information may include the name of the room where the sensor is located, the name of the sensor itself (if the user wishes to assign a name), and the number assigned to the sensor. The user can upload files about the surrounding environment, such as Figure 8A as shown in . Such files may include, for example, PDF, JPG, PNG, 3DM, OBJ, FBX, STL, and / or SKP. The surrounding information may include furniture within the field of the sensor node 102, the building layout around the field of the sensor node 102, etc. The user can decide for themselves to add surrounding objects in the space. The sensor node may not register the surrounding objects; however, the surrounding objects can provide a richer visual environment for the user's own personal use. The sensor node 102 profile can be stored in the user's profile in the database 360.
[0036] The thermopile sensor module 115 can project the temperature of an object onto a grid. The grid can be an 8 pixel × 8 pixel grid, a 16 pixel × 16 pixel grid, or a 32 pixel × 32 pixel grid (64 pixels, 256 pixels, or 1024 pixels respectively). The thermopile sensor module 115 can be tuned to detect a specific thermal spectrum, allowing for the detection of an object with a standard temperature (such as a human body). The average normal body temperature is generally recognized as 98.6°F (37°C). However, normal body temperature can have a wide range, from 97°F (36.1°C) to 99°F (37.2°C). A higher temperature usually indicates an infection or disease. The system can detect such temperature differences because the sensor module 115 can have an accuracy of 0.5°C. In the case where multiple human bodies are in the same area, the thermopile sensor module 115 captures and processes each body as a different heat source. In particular, the system avoids overlapping body temperature readings from different bodies by including a calibration process built into the 3D front end (an exemplary user interface is shown in Figure 10 ).
[0037] As a calibration process (an exemplary user interface is shown in Figure 10a part of (shown in), through the interface of the application, the user is required to exit the coverage of all sensor modules 115, so that the system can automatically adjust the sensitivity of the sensor modules 115. Starting from the maximum sensitivity, the system gradually reduces its sensitivity until high-frequency noise is no longer detected. The absolute elimination of noise allows the body to be detected as a different heat source, and subsequently, two heat sources such as the human body to be detected as different and independent entities. For the spatial overlap readings between the "fields of view" of the two sensor modules 115, during the calibration process, the system identifies the overlapping area between the two sensor modules 115 and averages the common detections between the two sensor modules 115. In the case where an overlap is detected between more than two sensor modules 115, the system will average the overlap of each pair in sequence. For example, for the overlap between sensor modules 115A, B, and C, the system first averages A and B, and then continues to average the result of AB with C.
[0038] If the automatic calibration fails during the calibration process (the exemplary user interface is shown in Figure 10 ), the system will automatically generate a digital path between the installed sensor nodes, thus prompting the user to physically stand below each sensor node. When doing so, the sensor node 102 detects the movement of the user and the detection becomes visible within the application. If successful, the user is prompted to follow the path and thus complete the calibration of each sensor node 102 of the network. If any sensor does not respond as mentioned above, the user is prompted to digitally manipulate the sensitivity of the sensor module 115 by sliding a digital bar and accordingly adjusting the sensitivity level of the sensor node under discussion. More specifically, during the troubleshooting of the calibration process, the user is prompted to stand below the physical sensor, at each of the four corners of its field of view - one corner at a time. At each corner, which is default named A, B, C, and D in the application, the user is called to stay as long as the system detects their presence and successfully prints them on the digital avatar of the sensor. Once the position is detected, the application will ask the user which one of the four possible corners they are trying to mark.
[0039] In various embodiments, the thermopile sensor module 115 may detect an array of temperature readings. In addition to detecting the temperature of a living being based on pixel values, the thermopile sensor module 115 may also obtain the temperature of the environment in which the sensor module 115 is located. The system uses local information and global information - both for each pixel individually and for all the pixels of the entire network-forming sensor modules 115 - to determine what the background temperature field is. The thermopile sensor module 115 may obtain an independent temperature measurement of the sensor node 102 itself. The temperature of the sensor module 115 may be obtained using on-board thermocouples. The system may use the temperature of the sensor node 102 and / or the temperature of the sensor module 115 to give an assessment of the temperature profile of the monitored space. The on-board thermocouple measurement itself measures the temperature of the space at the location of the sensor. The system uses bilinear interpolation to estimate the temperature between sensor nodes 102 and / or sensor modules 115 in the space to approximate the temperature distribution. Additionally, in various embodiments, the system may measure and capture the temperature of the environment multiple times throughout the day in order to reduce the adverse effects of maintaining a fixed background temperature field for threshold calculation, thereby increasing the overall detection accuracy in real-world scenarios where the ambient temperature is dynamic.
[0040] In various embodiments, multiple sensor nodes 102 may provide information in real time to assist in generating real-time location, trajectory, and / or behavior analysis of human activities. By employing multiple sensor nodes 102 and based on the density of the network, the system may infer the trajectory of any moving object detected by the sensor nodes 102. As mentioned above, the thermopile sensor module 115 inside the sensor node is designed to measure temperature from a distance by detecting the infrared (IR) energy of an object. The higher the temperature, the more IR energy is emitted. The thermopile sensor module 115, which consists of small thermocouples on a silicon chip, absorbs the energy and generates an output signal. The output signal is a small voltage that is proportional to the surface temperature of the IR-emitting object in front of the sensor. Each thermopile sensor module 115 has 64 thermopiles, and each thermopile in the thermopile is sensitive to the IR energy emitted by the object. To determine the trajectory, in various embodiments, each sensor module 115 divides the area captured by the sensor module 115 into multiple pixels that are organized in a rectangular grid in a direction aligned with the 64 thermopiles, and each of the 64 thermopiles is associated with an 8×8 portion of the above grid. The system monitors the sequential change in the temperature of consecutive pixels. The system determines that such sequential change indicates the movement of a living being. The system records such movement as the formation of a trajectory in space. The more nodes there are in the network, the more accurate the inference about the trajectory, because the detected trajectory is not interrupted by "blind spots".
[0041] The computing engine analyzes human behavior and trajectories. For example, regarding occupancy control, the system can calculate the total number of people in a space to compare with established occupancy requirements. The system identifies all heat sources in the space monitored by the sensor module 115 and sums the number of all heat sources generated by people.
[0042] Regarding occupant temperature screening, the system can detect the presence of a person by capturing their body heat. Once such a detection is made, the temperature screening can include an automatic adjustment of the sensitivity of the sensor module 115. It should be noted that body temperature screening may be different from body position detection. Body temperature screening means detecting the elevated body temperature of the person being detected, such that the sensitivity requirements are higher than for just body position detection.
[0043] Regarding monitoring the body temperature of an occupant, by using a more detailed 32×32 grid in the sensor module 115 to read the temperature of the area near the eye socket, the system 100 can be capable of obtaining the user's body temperature in the near field at a distance of one (1) meter from the sensor node 102. To enable the system to locate the eye socket, the user may be required to directly gaze at the sensor node 102, allowing the sensor module 115 to detect the pixels with the highest temperature.
[0044] Regarding analyzing the movement speed of an occupant, the system can record the movement of a person under the network of sensor nodes 102. A series of "waypoints" are generated over time. The system uses the distance traveled based on the waypoint information and the time taken to travel that distance to calculate the movement speed of the user in question.
[0045] Regarding calculating the total calories burned based on the movement of an occupant, the user inputs information such as the occupant's weight, gender, and age into the system through the interface 370. The system can use the captured movement speed and distance (as mentioned above) to calculate the approximate calories burned during the time period of the captured movement.
[0046] Behavior analysis stems from the fact that by overlapping a structured network with the physical space, the captured data becomes contextualized. For example, the system can understand the shopping behavior of a moving body by cross - referencing the actual trajectories and dwell times captured by the sensor nodes 102 with a floor plan carrying information about the location of specific products and aisles, as Figure 5 and Figure 6 elaborated. In particular, in various embodiments, the user interface 370 allows the user to create a three - dimensional representation of the space in question, such as Figure 5 、 Figure 6 and Figure 9As shown. For example, the owner of a grocery store can record information about the location of an ice cream freezer or a product by naming or "tagging" each sensor node 102 with a specific product located within the field of the sensor node 102. Thus, if sensor node #1 detects IR energy for 30 seconds, then the system determines that a person has stayed within the field of sensor node #1 for 30 seconds. If sensor node #1 is tagged as being in front of the ice cream freezer, the system provides data that a person has stayed in front of the ice cream freezer for 30 seconds. Examples of system outputs are as Figure 4A and Figure 4B shown.
[0047] In various embodiments, and as Figure 1A and Figure 1B shown, each of the multiple sensor nodes 102 can interface with a module. The module can include, for example, an HTPA32D module or an HTPA16D module. The module can be a radio module. The module may be wireless. The module can be a hardware module.
[0048] The sensor node 102 can include a switch 120 (e.g., ALPS) that controls the power supply to the sensor node 102. The switch 120 can allow the manufacturer of the system to turn off the power supply to the sensor node 102 to save the battery 130 of the module during the entire transfer or transportation process from the manufacturer to the customer. After the sensor node 102 is delivered to the customer, the system can be installed, where the switch 120 for the sensor node 102 is turned on and kept on. If the customer closes the store or the system for a period of time, the customer can use the switch 120 to turn off the sensor node 102 to save battery life. The LED 125 on the sensor node 102 indicates the system status, such as, for example, the on mode and the off mode.
[0049] Figure 2The general data flow according to various embodiments is described. The sensor node 102 may receive raw thermal data from the environment (step 205). The raw thermal data is compressed to create compressed thermal data (step 210). The gateway 135 receives the compressed thermal data from the sensor node 102. The gateway 135 decompresses the compressed thermal data to create decompressed thermal data (step 215). The cloud computing module 350 on the server receives the decompressed thermal data and creates detection data (step 220). The post-processing calculation module on the server receives the detection data from the cloud computing module 350. The post-processing calculation module processes the detection data to create post-processed detection data (step 225). The post-processing calculation module sends the post-processed detection data to the database 360. The database 360 uses the post-processed detection data to create time series detection result data (step 230). The system applies a background analysis algorithm to the time series detection result data to create an analysis result (step 235). The system applies the API service 380 and the client APP to the analysis result to obtain 3D / 2D visualization of the data (step 240).
[0050] According to various embodiments, in Figure 3A general system architecture including more details about data streams is described. In hardware, sensor node 102 acquires raw data and then performs edge compression (step 305) and / or edge computing (step 310) to create Message Queuing Telemetry Transport (MQTT) raw data. The cloud computing module 350 on the server receives the MQTT raw data topic 1 from the sensor module 115 via the gateway 135. The cloud computing module 350 applies data stitching and decompression (step 315) to the MQTT raw data topic 1 to create the MQTT raw data topic 2. The cloud computing module 350 applies the core algorithm (step 320) to the MQTT raw data topic 2 to create the MQTT result topic 1. The cloud computing module 350 applies world coordinate remapping (step 325) to the MQTT result topic 1 to create the MQTT result topic 2. The cloud computing module 350 sends the MQTT raw data topic 1, the MQTT raw data topic 2, the MQTT result topic 1, and the MQTT result topic 2 to the database 360. influxDB receives the data. The database 360 applies background analysis (step 330) to the data via a background analysis algorithm to create background results (analysis results). The background results are stored in DynamoDB. DynamoDB also stores the sensor node 102 profile. The database 360 can apply additional background analysis in response to updates or additional settings to the sensor module 115 profile. The API 380 obtains the background results from the database 360. The API applies real-time raw data, real-time detections, historical raw data, historical detections, historical occupancy, historical traffic, and / or historical duration to the data (step 335). The API sends the results to the user interface 370 (e.g., on a client device). The user interface 370 provides visualization (step 340) (e.g., Figure 4A and Figure 4B ). The user interface 370 also provides a settings interface (step 345). The settings interface can provide updates to the sensor module 115 profile. The user interface 370 also provides a login function (step 350). The login function can include AuthO / Firebase.
[0051] More specifically, the sensor module 115 can collect sensor module 115 data, preprocess the data, and / or send the collected sensor module 115 data to the gateway 135. The module can include an on-board microprocessor 140. The raw data from the sensor module 115 can be saved in the RAM of the microprocessor 140. The RAM serves as the system's temporary memory. The microprocessor 140 is configured to preprocess the raw data by removing outliers from the raw data.
[0052] Specifically, the microprocessor 140 applies a defined statistical program to the raw data to obtain processed data. In various embodiments, this module uses firmware software for preprocessing. The firmware software determines outliers in the temperature readings in a statistical manner. For example, outliers can be defined by normalizing the data by subtracting each pixel value from the mean of the frame. The result is divided by the standard deviation of the frame. Using bilinear interpolation techniques, i.e., the product of the interpolations of adjacent values, pixel values that are more than three times the standard deviation above or less than one-third of the standard deviation below are removed and replaced. The pixel values are replaced rather than removed so that the input detection is similar before and after the program. This technique helps to fix minor data problems that may be caused by potential defects in the data quality of the sensor module 115. The combination of firmware software, circuit design, and drivers enables the system to run algorithms to determine the "regions of interest" representing human activities in the field of view of the sensor module 115 for each data frame in the data frames. The regions of interest are not pixels with a specific temperature, but pixels that have a different temperature (higher or sometimes lower) relative to their surrounding pixels. The regions of interest are then used to compress the processed data and prepare the compressed data for wireless transmission.
[0053] The system can include a rolling cache of ten data frames for preprocessing. More specifically, the firmware of the microprocessor 140 can use the latest ten data frames of the captured data for the preprocessing program and the postprocessing program. Due to the limited amount of RAM memory on board (e.g., the application requires 8 kb), the system may only process a subset of the data.
[0054] The data passing through the gateway 135 can be uploaded to the cloud computing module 350 on the server. The gateway 135 can be powered by any power source. In various embodiments, the gateway 135 is powered by a 110V outlet. The gateway 135 includes modules that are connected to a network (e.g., the Internet) via Ethernet, wifi, and / or cellular connections. Thus, the gateway 135 can upload data to any database 360, server, and / or cloud. The gateway 135 sends the preprocessed data and the compressed data to the computing engine in the cloud computing module 350, which in turn outputs the results to the databases 360s. The gateway 135 pulls operation commands from the server to perform administrative functions such as software updates, commanding the sensor module 115 to turn on and off, changing the sampling frequency, etc.
[0055] In various embodiments, the gateway 135 captures the compressed raw data in transit and sends the compressed raw data to an algorithm running on a processor (e.g., Raspberry Pi 4, model BCM2711), and then forwards the information to a server (e.g., cloud computing) for further processing. The processing of the data on the server includes decoding the compressed raw data, normalizing the temperature data of the sensor module 115 according to the firmware and environmental settings of each sensor module 115, detecting objects, classifying objects, performing a spatial transformation for world coordinate system positioning and fusion, multi-sensor module 115 data fusion, object tracking and trajectory generation, cleaning of outlier pixel-level readings, and other post-processing.
[0056] In various embodiments, the processing steps work with the decompressed raw data. Decoding the compressed raw data optimizes data transmission and the consumption level of the battery 130. In addition, the normalization of the temperature of the sensor node 102 to an appropriate temperature range enables the processing steps to adapt to various quality differences and environmental differences (which are expected from sensor nodes 102 located at different spots in space).
[0057] One of the core processing steps of the computing engine of the cloud computing module 350 is object detection and classification. This processing step detects the position of the object of interest in the frame and classifies the object as a person or different classes of objects (e.g., laptop, coffee cup, etc.). The spatial transformation from the local coordinate system to the world coordinate system enables the analysis to be "background-aware". With the spatial transformation, the system can compare and cross-reference the coverage of the spatial sensor module 115 with the actual floor plan and 3D model of the space under discussion. Multi-sensor module 115 data fusion unifies the data in the case of missing information or overlapping coverage between multiple sensor modules 115. As mentioned above, using various algorithms, object tracking and trajectory generation can distinguish multiple people from each other over time. Object tracking and trajectory generation provide a set of trajectories originating from the detected objects and people. The system uses such trajectories to determine behavioral analysis (e.g., the location and duration of stays), movement speed, and direction. The post-processing step addresses any minor inconsistencies in the detection and tracking algorithms. For example, when there are missing detections or gaps in the trajectory, the post-processing step helps piece the information together and repair any damaged trajectories.
[0058] The system can use a thermal application protocol interface (API), which can be located in the API layer in the system architecture, as Figure 3As described in. The API hosts real-time and historical occupancy data for spaces enhanced with the Sensor Module 115 solution. Built on REST, the API returns JSON responses and supports Cross-Origin Resource Sharing. The solution uses standard HTTP verbs to perform CRUD operations, and for error indication purposes, the API returns standard HTTP response codes. Additionally, namespaces are used to implement API versions, while token authentication is used to authenticate each API request. The API token on the dashboard is used to authenticate all API endpoints.
[0059] The token may be included in the authorization HTTP header, prefixed with the string literal "Token" and separated by a single space between the two strings. If the correct authorization header is not involved in the API call, a 403 error message is generated. HTTP Authorization header Authorization: Token YOUR_API_TOKEN. The endpoints use standard HTTP error codes. The response includes any additional information about the error.
[0060] The API lists low-level "Sensor Module 115 events" for the Sensor Module 115 and a period of time. Each Sensor Module 115 event includes a timestamp and a trajectory relative to the Sensor Module 115 in question. The trajectory does not have to be equal to any direction relative to the space (e.g., entrance or exit). This call should only be used to test the performance of the Sensor Module 115.
[0061] The API provides information about the total number of people entering a specific space each day over a period of a week. The analysis object with the data for the interval and the total number of entrances is nested in the interval object of each result. This call can be used to find out how many people visit the space on different days of the week.
[0062] The API records, tallies, and lists all exits of individuals from the space of interest for an entire day (or any 24-hour period). Each result is accompanied by a timestamp and a direction (e.g., -1). This call is used to find out when people leave the space.
[0063] The API provides information about the current and historical wait times at the entrance to a specific space at any given time during the day. The analysis object with the data for the interval and the estimated total wait duration is nested in the interval object of each result. This call is used to find out how many people are queuing to enter the space for different time spans.
[0064] Webhook subscriptions allow receiving callbacks on a specified endpoint on a server. For each space in which an event occurs, a webhook can be triggered after each event is received from one of the sensor modules 115. The system can create webhooks, retrieve webhooks, update webhooks, or delete webhooks. When a webhook is received, the JSON data will be similar to the space and sensor module 115 events in the previous section. It will have additional information: the current count of the associated space and the ID of the space itself. The direction field will be 1 for entry and -1 for exit. If any additional headers are configured for the webhook, the additional headers will be included in the POST request. An example of received Webhook data can be a single event occurring at a path connecting two spaces.
[0065] In various embodiments, the system can include one or more tools to help facilitate installation, help with software and hardware setup, provide more accurate detection, create virtual representations, visualize human movement, test devices, and / or troubleshoot devices. The system can include any type of software and / or hardware, such as, for example, one or more applications, GUIs, dashboards, APIs, platforms, tools, web-based tools, and / or algorithms. The system can be in the form of downloadable software. For example, the software can be in the form of an application downloaded from a website that can be used on a desktop or laptop computer. The software can also include a web application that can be accessed via a browser. Such a web application can be device-independent and adaptive such that the web application can be accessible on a desktop computer, laptop computer, or mobile device. The system can be obtained through licensing or subscription. One or more login credentials can be used to access the system partially or fully.
[0066] As used herein, a space can include the overall layout of an area that can be composed of one or more rooms. System functionality may affect different spaces separately. The system may associate multiple rooms within a space. A room can include any walled portion of a space (e.g., a meeting room) or an open area within a space (e.g., a hot-desking area or a corridor). Headcount can include the number of people entering and leaving a particular space or room within a given time frame. Occupancy can include the number of people inside a room or space at a particular time. Fixtures can include furniture (e.g., chairs, tables, etc.) or devices (e.g., washing machines, stoves, etc.).
[0067] In general, in various embodiments, the system can plan an installation by, for example, visualizing sensor positions, visualizing coverage in space using a 3D drag-and-drop interface, and / or knowing the number of sensors and / or hives that may be appropriate for a room or space. The system can achieve more accurate detection by, for example, analyzing the spatial context to distinguish human and inanimate objects (and other confounding factors). The system can obtain the spatial context by receiving input about the layout of the space such as 3D furniture, rooms, and markers / labels. Analysis of the spatial context may involve artificial intelligence and / or machine learning. AI is used to utilize the spatial context to understand the space so that artificial intelligence can accurately identify the presence, behavior, posture, and other specific activities of humans. The system can create a virtual representation of the real location (e.g., across any dashboard, set up applications that use algorithms and APIs to visualize spatial data, and any other applications). The virtual representation can be based on the tool receiving markers, labels, and / or names for sensors, rooms, and spaces. The virtual representation can be shown as a unique identifier in the dashboard and / or API.
[0068] In various embodiments, the system can visualize human movement by, for example, displaying the current frame and the previous frame (e.g., in the form of an image such as dots), displaying and listing the position coordinates within the context of the sensors, virtual space layout, and virtual fixtures, displaying the trajectory of a person, and / or displaying the posture of a person (e.g., standing, sitting, lying down). The system can test and troubleshoot the device by, for example, showing the user what the sensor is detecting. The user can then confirm that the type and location of the actual object correlate with the visual representation. The system can display what the sensor is detecting in a real-time and / or frame-by-frame representation of human presence and movement within the space. The system can also show when the sensor is online, offline, connected, and / or disconnected.
[0069] In various embodiments, the system can include functionality for creating a spatial layout. Creating a spatial layout may involve adding spaces and adding space names. The system then stores the space along with its name. The system provides the user with the ability to create and manage multiple spaces. Having several separate spaces can be useful when monitoring multiple floors in a building (e.g., the first floor, the second floor), a facility with multiple individual rooms (e.g., a senior living apartment), or multiple facilities located at different physical locations (e.g., a Boston laboratory, a San Francisco laboratory). In various embodiments, as part of the setup, the system can provide the ability to do one or more of the following: rename a space, add auto-align, add visual smoothing, add display local detection, add toolbars (e.g., a main toolbar, a side toolbar, etc.), add fixtures from a library (e.g., individual pieces of furniture from a furniture library), or return to the project library. The main toolbar can include functionality related to rooms, sensors, boxes, languages, and / or saving. Exemplary side toolbar functionality can include 2D or 3D, show or hide sensors, show or hide rooms, show or hide fixtures, etc. The toolbars and functionality can be described as being on the main toolbar or the side toolbar, but any functionality can be associated with any toolbar.
[0070] In various embodiments, the system can include functionality for adding rooms to match or be similar to a floor layout. The system may display a control panel (e.g., in response to selecting a room). The control panel can allow for changing dimensions, marking certain locations or features, and / or selecting a border color for each room. In response to selecting a room icon, the system can add one or more rooms to the space. In response to auto-align being activated, a moved object (e.g., a fixture, a sensor, or a room) can automatically snap to the edge of a nearby object of the same type. For example, a chair can automatically move next to another object to allow the user to more easily and quickly arrange objects in a neat and orderly manner to match a floor plan. The system can determine that a moved object is an object of the same type as an existing object based on a similar identifier or tag associated with each of the objects. In response to auto-align being deactivated, the user can freely and incrementally move an object that may not be aligned with similar objects.
[0071] In various embodiments, the system can include functionality for adding fixtures to a virtual space in a GUI (which can be related to fixtures present in a physical space). Fixtures can include, for example, furniture or equipment that the user can add to the space. Fixtures can allow the user to distinguish between rooms and allow the user to put the movement seen on the screen into context. The system can allow the user to add a fixture by selecting any of the furniture or equipment icons and then dragging the fixture to a specific location in a different room. The system also provides the user with functionality for virtually selecting a fixture and then deleting or rotating the fixture using panel controls. The system also provides the user with functionality for virtually adjusting the size or position of the fixture. In various embodiments, the user can input the specific coordinates of the furniture so that the system knows the size and location of the furniture. Additionally, when the user moves a fixture, the system can display the distance between the center point of the fixture and each of the four walls of the room where the fixture is placed. The coordinates of the furniture relative to the overall space can be stored in the system via an API. In this way, the user can extract this information from the backend as needed. The user can add any number of fixtures in the space or room as needed. The user can stack fixtures or furniture on top of other fixtures or furniture. Additionally, if the physical table is particularly large, the user can use multiple virtual tables to match its size. The system encodes, identifies, stores, and takes into account the presence and coordinates of the fixtures. These data are used by the system when determining whether a detection is human and whether the detection should count towards the occupancy or headcount of a room or space. The system includes functionality (e.g., using an API and algorithms) for encoding each fixture (e.g., table, door, etc.) with a fixture type. The user can select the fixture and the fixture type from the icons. The API can save the fixture and its coordinates on the system so that the algorithms can use this information to identify detections and behaviors. In response to the location selected by the user, the system records the x-y coordinates of the fixture center point, the fixture type, and the rotation (in degrees (e.g., 0, 90, 180, 270) degrees) from the center point. Rotation can include the action of rotating an object (room, sensor, fixture) 0, 90, 180, 270 degrees from its center point. Rotation can be achieved by selecting a circular arrow to rotate the object. The fixtures in the system can be set to be in a default pointing direction, which can be different from the direction of real furniture in the room. The user can rotate the virtual furniture model to 4 different directions (i.e., 0 degrees, 90 degrees, 180 degrees, 270 degrees from its default direction) along the center point of the model.
[0072] The system can show when people are on a piece of furniture, around a piece of furniture, or passing by a piece of furniture. The system stores names or icons associated with different fixtures. These names or icons include factors or rules that the system considers when analyzing the fixtures. For example, a bed icon can include the rule that a human can be on the bed, while a table icon can include the rule that a human will not be on the table. Examples of other fixtures (with the rule that a human will not be on the fixture) include tables, counters, stoves, refrigerators, dishwashers, sinks, radiators, counters, and / or washing machines. Examples of other fixtures (with the rule that a human can be on or in the fixture) include beds, sofas, chairs, toilets, and / or showers. The system can infer a person's activity based on the person being next to a piece of furniture for a specific amount of time. For example, if the system detects that a person is near the TV for an hour, the system can infer that the person likes the program being shown on the TV. This location and inferred information also provides valuable context for the algorithm to allow the system to infer daily activities and enable more accurate human detection. For example, for daily activities, if the detected object (represented, for example, by a purple sphere) is on or within the outline of a bed fixture, the system can infer that the human is sleeping. Then, the system can infer the person's sleep duration by determining how long the person is on the bed fixture. As another example, for more accurate human detection, the system can detect both human heat sources and non-human heat sources (such as stoves and laptops). If the detected heat source is found to be in the middle of a fixture such as a table, the system can identify that this is not a human. Therefore, the system can not count this detection as part of the occupancy data (e.g., the user can receive this occupancy data via an API or dashboard). The system can include a function for activating or deactivating the "Show partial detections" option to view or hide detections (such as purple spheres).
[0073] In various embodiments, the system may include a "visual smoothing" function. Disabling "visual smoothing" causes the detected spheres to be displayed exactly frame by frame at the specific coordinates as detected. Enabling "visual smoothing" causes the frame-by-frame movement of the purple spheres to be displayed as a smooth, continuous animation. More particularly, when the sensor detects an object in physical space, the system may create a "proxy" (e.g., a purple sphere) that appears at the corresponding location in virtual space. The system may also assign a "lifetime" to the proxy. The "lifetime" may be the length of time that the purple sphere appears in virtual space to indicate the detection. In various embodiments, this lifetime may be set to last 300 milliseconds, which matches (or is similar to) the interval at which the sensor sends new detections. The system uses new detections of the new position in the physical world to update the position of the purple sphere in virtual space. As mentioned, the system may display the purple sphere frame by frame or via visual smoothing. The purple sphere may appear to be blinking, but the sphere may actually be displaying real-time detections captured by the sensor every 300 milliseconds. The blinking effect is caused by the fact that the purple sphere changes from opaque to transparent within its 300-millisecond lifetime. If the sphere is more opaque, the detection is likely to be newer. If a person stands still under the actual sensor, the purple sphere will appear to blink in place. If a person moves under the actual sensor and visual smoothing is not enabled, the system may display a string of purple spheres. For example, each of the purple spheres may continuously change from opaque to transparent at each detection coordinate (or pixel) every 300 milliseconds. With visual smoothing enabled, only one purple sphere appears on the screen and appears to move linearly. During the 300-millisecond lifetime, the system searches for the next detection within a 1-foot radius and "refills" its lifetime to another 300 milliseconds. This means that the sphere continuously appears and moves from one point to another according to the detection coordinates (pixels).
[0074] In various embodiments, the system may include functionality for adding one or more enclosures and / or thermal sensors. Data from outside the room may still be stored via the API, but data from outside the room may not be displayed on the dashboard. The dashboard displays room-specific activities and occupancy within the space. Thus, the system instructs the user to place sensors and objects within the room. The enclosure may provide a gateway function to connect the thermal sensors and transmit data from the thermal sensors to a storage location (e.g., the cloud). For example, an enclosure may be connected to a certain number of sensors (e.g., up to 12 sensors), and then additional enclosures may be required for additional sensors. For ease of installation, each enclosure may be pre-configured with a set of sensors (e.g., loading the sensor IDs as sensor data into the enclosure's database). The pre-configuration process may include two parts. The sensors may be set to have the same NetID as the enclosure. This enables the sensors to connect to each other. The enclosure may obtain the sensor MAC addresses and sensor modes programmed into the configuration file. This enables the enclosure to correctly manage the sensor frame rate. The system may allow the user to add one or more enclosures to the space. The system may include functionality for adding an enclosure to the space by scanning a code (e.g., a QR code) or by entering the enclosure ID (e.g., the enclosure ID found on the sticker under the enclosure). During the process of adding enclosure data to the system, the system also receives the pre-configured sensor data (since the enclosure data has been pre-configured to include sensor data). The system uses the sensor data to display the "Sensor" icon for each enclosure. The system may display sensors from different enclosures, as the sensors may be color-coded to indicate which sensor belongs to a particular group and enclosure. Specific implementations may vary, but can generally be visually identified by the system. In response to receiving a selection of the sensor icon for an enclosure, all sensors associated with that enclosure will be displayed so that each sensor can be added to the space. The user can drag each virtual sensor to a location in the room to record room occupancy, or drag it onto a doorway to record the number of people entering or leaving through that doorway. Each sensor may be unique and may be identified with a unique address (e.g., a MAC address). As such, the user should place the virtual sensor in the same location, with the same orientation, and in the same room as the corresponding physical sensor.
[0075] In various embodiments, the system may include functionality for setting and / or calibrating virtual thermal sensors. Each virtual thermal sensor may appear on a display (e.g., as a square). In response to the virtual sensor being turned off, the virtual sensor may appear in a certain way (e.g., a black square) to indicate that the actual sensor is not detecting anything. In response to the virtual sensor being turned on, the virtual sensor is displayed as a grid (e.g., an 8×8 grid consisting of 64 squares). Each of the 64 squares may represent each pixel, and the color of each pixel may represent the temperature detected by the actual sensor at that point in the grid. The color of each pixel may include a range of chromaticities to indicate the temperature level. For example, the range of chromaticities may be from yellow (lower temperature) to red (higher temperature).
[0076] In response to receiving a selection of a virtual sensor, the system displays a control panel (e.g., located on the left side edge of the display). Using the control panel, the system may allow the user to set the height of the virtual sensor. For example, it may be important that an aisle in a grocery store is within the field of view, but a janitor's closet may be outside the field of view. The optimal (and preferably maximum) height of the actual sensor is about 3.2 meters. Such a height may provide the maximum coverage and best resolution required for human detection. The best resolution may be the ability to detect human presence from a thermal map image from the sensor based on an algorithm. The best resolution refers to the resolution at which the system can accurately and reliably detect human presence. The height of the virtual sensor corresponds to the height of the ceiling or the height on the wall to which the actual sensor will be attached. The higher the sensor is above the floor, the wider the coverage of the sensor over the floor. The closer the sensor is to the floor, the narrower the coverage of the sensor over the floor. Based on the height, the system can determine how much floor space the sensor is monitoring (covering).
[0077] The system may use the formula (90% * 2 * tan(30°) * height)^2. For example, a ceiling height of 3 meters may result in a coverage area of 3.03m × 3.03m on the floor. The standard supermarket ceiling height is 5.78m, the standard office ceiling height is 3.12m, and the standard door height is 2.43m. The system also allows the user to test different heights to confirm that certain areas are within (or outside) the field of view of the sensor. A sensor height of 110 inches (2.8m) may provide an effective coverage width of 106 inches (2.7m), a sensor height of 102 inches (2.6m) may provide an effective coverage width of 78 inches (2.0m), a sensor height of 95 inches (2.4m) may provide an effective coverage width of 63 inches (1.6m), and a sensor height of 87 inches (2.2m) may provide an effective coverage width of 56 inches (1.4m).
[0078] In various embodiments, the system can also allow the user to set the virtual sensor orientation to match the way the physical sensor is physically installed, to ensure an accurate representation between the virtual world (e.g., in a setup application) and the physical world. The sensor orientation in both the real world and the virtual world must be similar so that the visible detections also match. For example, a person standing in the northeast corner of a room must appear in the northeast corner of the corresponding virtual sensor pixel in the setup application. To assist in matching the sensor orientation, the physical sensor can include an arrow (e.g., on its mounting plate). When the user adds the physical sensor to the setup application as a virtual sensor, the user matches the direction of the virtual sensor arrow with the direction of the physical sensor arrow by rotating the virtual sensor.
[0079] In various embodiments, the system can include a function to view detections. "Detection" mainly refers to the detection of the presence of a person. The actual sensor can capture a heat map of a certain area, for example, at 3 to 5 frames per second. The system can detect the presence of a human by identifying areas of the heat map related to the body temperature (or "thermal signature") of a human. The system can represent the detection of a person (e.g., a purple sphere) on a display. The average normal body temperature of a person is generally accepted as 98.6°F (37°C). However, normal body temperature can have a wide range, from 97°F (36.1°C) to 99°F (37.2°C). Temperatures above 100.4°F (38°C) can mean that the person may have a fever due to an infection or illness.
[0080] In various embodiments, the detection process can include sensitivity adjustment. The system can receive data about the temperature in the space or room based on a thermometer located in the physical space or room. Sensitivity adjustment can involve improving the system's ability to detect the presence of a human in an environment where its temperature is close to the body temperature of a human. The detection is improved by changing a parameter related to the temperature difference between the detected (human temperature) and the surrounding environment. Increasing the sensitivity can involve minimizing this temperature difference so that the system can more easily detect the presence of a human in, for example, a very warm climate. In other words, in a cooler climate of 65 degrees, the system can more easily determine that any object that is 30 degrees or more above the normal room temperature is a human. However, in a warmer climate of 96 degrees, the system would need to detect an object with a temperature 2 to 3 degrees above the room temperature to consider the object as a human.
[0081] This detection information can be further processed by the system algorithm to re-check and ensure that the data finally sent to the API and dashboard is accurate. This further processing can include additional criteria to filter out detections that do not behave like humans. The system can filter out any objects with a temperature below the human body temperature range (e.g., from 97°F (36.1°C) to 99°F (37.2°C)). The filtering can include detections that do not move at all (e.g., a device like a stove) or stationary detections that seem to be on fixtures (e.g., in the middle of a table) where humans do not expect to be located. For example, the system can determine that a heat map displayed on a table is more likely to indicate a laptop. Additionally, the system can store the coordinates around each fixture such that detections of objects with coordinates overlapping the coordinates of any fixture are not counted. In other words, the coordinates around the fixture are blacklisted so that any detections within these coordinates are not counted.
[0082] In various embodiments, the detection can include the number of people. A sensor for determining the number of people entering and leaving a room (a people counting sensor) can be installed above the entrance door of the room (e.g., on the wall facing the interior of the room). The people counting sensor can use data associated with a virtual threshold (or door jamb line), which can be a certain distance from the door. The people counting sensor can only count detections of people who cross the door jamb line. For example, "In" is when a person crosses the door jamb line from left to right, and "Out" is when a person crosses the door jamb line from right to left. The door jamb line reduces false readings for people who might stick their heads inside the door to see what's inside the room but never fully enter the room.
[0083] In various embodiments, the system can include a function for 2D or 3D view settings. In response to receiving a selection of the 2D / 3D button, the system can display all (or any part) of the entire space in 2D or 3D form. The 3D view can provide a more "real" spatial context for users who are not familiar with floor plans. This view can also provide a more intuitive way to understand the space, detections, and data. In various embodiments, the system can include a function for displaying or hiding various features (such as, for example, sensors, rooms, or fixtures). The system can include functions for different viewing modes for editing the space, managing the space, and visually communicating data.
[0084] In various embodiments, the system can visually convey data regarding the people flow and / or dwell time over a certain period of time. The people flow can be conveyed based on the number of people entering and exiting a given doorway (or across the door jamb line) and entering a room or space. The people flow can also be conveyed by the paths or trajectory lines along which people move within the space. The system can include a head count view, in which the entry and exit numbers can be displayed on a virtual layout. The system can also include a trajectory view that displays multiple detections over a period of time, thereby forming a detected flow in the display. The detected flow can be used to determine the path of the people flow. The system can also display a linear movement trajectory, where the system creates a line through multiple detections to create a line representing the path of the people flow. The dwell time can be the amount of time a person spends at a set of coordinates in a room or space. The dwell time can be determined by measuring the amount of time a person is detected within a certain area. The system can infer or identify unique detections (the same person) based on location and trajectory, even if the system cannot determine whether it is the "same" person detected previously. The dwell time can be displayed as a heat map, where the darker the color, the more time spent in a specific area, and the lighter the color, the less time spent in a certain area.
[0085] As discussed above, traditional computer vision uses high-resolution images to develop human pose detection algorithms. High-resolution images can include, for example, recorded high-resolution video clips. Many of the existing algorithms detect key points (key points can be points on the human skeleton) to detect human poses from images, and the detection of key points requires a high-resolution camera system. Thus, it may not be feasible to detect key points using a system with reduced resolution. However, a system that can use reduced-resolution images may be important for promoting more privacy.
[0086] In various embodiments, the system can provide the function for pose detection by using a pose detection algorithm on reduced-resolution images. For example, the system can display whether a person is standing, sitting, lying down, or has fallen. Generally, the algorithm can include a sub-module for extracting a bounding box containing human pixels. The algorithm can also include a sub-module for detecting the pose of the human inside the bounding box. These algorithms can be algorithms based on deep neural network learning, and the algorithms based on deep neural network learning can be data-driven. The neural network can be a CNN (Convolutional Neural Network), which is an artificial neural network specifically designed for image recognition and processing to process pixel data.
[0087] Generally speaking, bounding the boundaries of a bounding box is one of the behaviors learned by a neural network. The neural network can be trained on images with bounding boxes that can contain all the pixels (or a subset of pixels) corresponding to a human in the image. The neural network can use this information to predict where the (multiple) bounding boxes should be on the image. More specifically, in various embodiments, the system can receive an image of a human. The input image / frame can be of low resolution, such as, for example, an 8×8, 32×32, or 64×64 pixel image. For each frame, a human annotator can first observe the pose of the human (e.g., the test subject) in the image, label the pose, and draw a bounding box around the person in the image. The human annotator can confirm the presence of the human in the image by viewing the recorded high-resolution video / image. However, this high-resolution image may not be used to train the algorithm.
[0088] In various embodiments, the user can label the pose with an integer. The user can enter an integer for a specific pose on the screen. The GUI can include a text box with fields that accept integers for one or more poses. The system can associate the integer with the pose in a database. For example, label 0 can represent a sitting pose, label 1 can represent a standing pose, label 2 can represent a lying pose, etc. In various embodiments, the user can draw a bounding box around the human in the image on the screen (e.g., using any type of device that accepts input on the GUI) such that the bounding box is stored as (x, y) coordinates that the system can recognize. The system can use such labels and manually annotated data (e.g., bounding boxes) to train a learning algorithm. The learning algorithm can be trained using, for example, training based on gradient descent. Based on the training, the learning algorithm can learn how to automatically create bounding boxes, collect data from within the bounding boxes, and determine the human pose from the collected data. The collected data used by the algorithm can be collected over time or during an initial calibration session while the system is running. The system can obtain environmental data from a specific environment where sensors are deployed, so the system can use the environmental data to adjust its algorithm based on the specific environment. The data from the environment can include parameters and / or variables from the environment, such as, for example, environmental temperature, indoor temperature, floor plan, non-human heat objects, gender of the human, age of the human, height of the installed sensors, clothing of the human, body weight of the human, etc. Pixel data, thermal data, and / or environmental data collected from the environment can be used to train the algorithm through machine learning or artificial intelligence.
[0089] Due to the low resolution, it may be difficult to determine the differences in image intensity between different pixels, so the algorithm may not attempt to determine features based on image intensity. Instead, the algorithm can focus on the discriminative features of the overhead thermal signature pattern from human presence. The CNN may implicitly define what constitutes a "discriminative feature". Each layer of the network defines a "feature map" learned through stochastic gradient descent. The user may not be aware of the features that the network considers important. The system can detect edges, curves, sharp contrasts in adjacent pixel values, etc. However, due to the nature of the neural network, the system may not have any (or very little) information about which features are important for differentiating poses.
[0090] As discussed above, in various embodiments, the system can extract regions that include the discriminative features of a human (e.g., the region is represented as a rectangle representing a bounding box). The bounding box can limit the amount of data analyzed by the system, and the box forces the algorithm to focus only on the human silhouette. At this step of the algorithm, the system can focus on the extracted bounding box to attempt to classify the human pose. The system can extract the human pose information for each frame. A frame can be a single image captured by one of the thermal cameras in the thermal camera array.
[0091] In various embodiments, the system can calculate an "aggregated pose" to help smooth the pose changes across multiple frames over a period of time. For example, the aggregated pose can be determined based on the pattern of a set of poses collected over a given time period (e.g., the pose that appears most frequently in the set). All the poses in the set may not be the same, but the pose aggregation method creates a consensus. This consensus can be referred to as the "aggregated pose". For example, the system can obtain a pose consensus every 5 seconds. The aggregated pose can include, for example, sitting, standing, or lying down. The system can determine an event as a fall based on certain changes in the pose or the pose persisting for a certain amount of time. For example, if the event consists of: (i) the aggregated pose changing from standing / sitting to lying down, and (ii) the lying down pose persisting for a certain amount of time (e.g., at least 30 seconds), then the event is determined to be a fall.
[0092] Using pattern recognition, fuzzy logic, artificial intelligence, and / or machine learning, the system can determine human poses based on similar thermal signatures. In various embodiments, with regard to pattern recognition, the system can extract discriminative features and / or patterns from the frame to help identify the pose contained in the frame. In various embodiments, with regard to artificial intelligence, and based on the collected human annotation data and / or frames, the system can automatically learn these patterns by using a CNN. As illustrated in FIG. 11, the patterns from human thermal signatures can be different, and different patterns can enable the system to classify human poses differently. For example, Figure 11Ashows thermal features indicating a standing position, Figure 11B shows thermal features indicating a sitting position, and Figure 11C shows thermal features indicating a lying position.
[0093] As Figure 12 described, in various embodiments, the system can perform a pose inference process. The system can obtain an input image (e.g., an 8×8, 32×32, or 64×64 pixel image) from an infrared sensor (step 1205). The upper and / or lower limits of the resolution can be based on privacy issues, sensor power consumption, data cost, data bandwidth, computational cost, and / or computational bandwidth. The system can apply a CNN to the input image (step 1210). The CNN can learn different filters at each subsequent layer of the network through stochastic gradient descent, thereby implicitly extracting important information from the image.
[0094] The system can create a vector of image features (step 1215). This vector can be an intermediate result consisting of the learned image features output by the CNN. A transformer encoder / decoder can be used (step 1220) to convert the vector of image features into box predictions (step 1225). DETR (DEtectionTRansformer) can use a traditional CNN backbone to learn a 2D representation of the input image. Before passing the image to the Transformer encoder, the model may flatten the image and supplement it with positional encoding. Then, the Transformer decoder can take a small fixed number of learned positional embeddings (e.g., object queries) as input and additionally attend to the encoder output. The system can pass each output embedding of the decoder to a shared feed-forward network (FFN), which can predict detections (e.g., classes and bounding boxes) or a "no object" class.
[0095] In various embodiments, the system may include box prediction, which can predict the size and location of a bounding box that includes distinguishing features of a human. The system may obtain the interior of the bounding box (step 1230). The system obtains the interior of the bounding box because anything outside the bounding box may not be of interest to the system, as the system is interested in recognizing the human pose of a human. In various embodiments, the system applies a CNN to the interior of the bounding box (step 1235). More specifically, the CNN applied to the interior of the bounding box may implicitly learn the features of the image given as training. These features may be invisible to a human user and may never be explicitly expressed by the neural network. Based on the CNN, the system creates a pose prediction for the image (step 1240). More specifically, during the training phase of the model, the system may extract features to distinguish human poses. This is prior knowledge used during the inference / test phase when using the trained model to distinguish various poses. The pose prediction may include, for example, sitting, standing, lying down, or any other pose or configuration of a human. The system may also detect other poses using a higher resolution (32×32 pixels or higher). For example, the system may use a higher resolution to obtain better performance such that the system can distinguish between a standing pose and a sitting pose. The system may also distinguish activities such as exercising, dancing, running, eating, etc.
[0096] As Figure 13 described, in various embodiments, the system may perform a fall detection process. The system may acquire an image (step 1305). For example, all images may be aggregated within 5 seconds. The system may benchmark the amount of time used to aggregate the images. If the amount of time used to aggregate the images is too short, the system may lose image details. If the amount of time used to aggregate the images is too long, the system may acquire noisier data in the image. The system may not make any assumptions about the position of the person in the frame. The person may be stationary or moving. The system acquires data about the pose of the person in each frame. The above-described pose inference process is performed on the image, as Figure 12 described (step 1310). The result of performing the pose inference algorithm on each frame is the pose prediction for that particular frame. After pose inference, the system may obtain many predicted "poses". Figure 12 Shown as "poses collected within 5 seconds". After passing through the "pose aggregation" module, the system may determine one "aggregated pose", whether it is standing / sitting or lying down. Then the system may use this information to further determine whether a fall has occurred.
[0097] In various embodiments, the system may perform pose aggregation (step 1315). More specifically, the system may determine the aggregated pose based on the pattern of the set of poses collected over a given period of time. For example, the pose that appears most frequently in the set. Although all the poses in the set may not be the same, the pose aggregation method creates a consensus, i.e., the aggregated pose. The system may determine the aggregated pose (step 1320). The aggregated pose may be determined to be standing or sitting (step 1325). The system may determine that the aggregated pose is lying down (step 1330). If the aggregated pose is lying down, the system looks at the database of previously aggregated pose data and determines whether the aggregated pose at the previous timestamp was standing or sitting. The system may compare the pose with the pose identified in the previous frame. The frames may be sorted by time. The system may determine that the aggregated pose has changed from standing or sitting to lying down and then the pose remains in the lying down position for at least 30 seconds (step 1335).
[0098] The system may store timestamps associated with each of the various actions. The system may check the difference between the start timestamp and the current timestamp to determine whether the time difference exceeds a threshold. The threshold may be pre-specified, pre-determined, dynamically adjusted, algorithm-based, etc. In various embodiments, if a fall exceeds the threshold amount of time, the system issues a fall alert. The fall alert may be sent to any other device via any communication means. For example, the system may send a signal over the Internet to an application on a relative's smartphone to notify the relative that a fall may have occurred. In addition to the alert, the system may also provide data about the person, location, facility, health information, demographic profile, fall history, etc.
[0099] Due to the nature of the CNN, the CNN may assign a certain type of pose. However, if the system is uncertain about a pose or a frame, the system may ignore some frames. For example, the system may use a confidence score of 0.7. A confidence score of 0.7 may indicate that if the system predicts with 70% confidence, then the pose is a fall to some extent. The system may also predict a fall with a confidence above or below 0.7.
[0100] The system can determine a potential fall (step 1340). For example, if a person is in bed or on a couch, the system can determine this movement as a "potential fall". However, being in bed or on a couch can be a normal movement, and this movement may not be a real fall. Thus, the system can perform post - processing on the potential fall (step 1345). In various embodiments, the post - processing can include filters, such as, for example, obscuring some spatial locations where a fall cannot reasonably occur (e.g., a bed). Another filter can include a confidence threshold for the detection of potential falls using an algorithm.
[0101] Based on the post - processing, the system can confirm the potential fall as a confirmed fall (step 1350). Specifically, after using one or more filters, the system can confirm the fall and provide a notification of the confirmed fall. For example, in a user interface / software product, the user can create and place virtual furniture (e.g., beds, chairs, closets, etc.) or occlusion areas in a settings application. Even if a "potential fall" is detected according to a machine - learning algorithm, the system can automatically blacklist these areas to prevent triggering a fall alert. For example, as discussed above, if a person lies or sleeps in bed, the machine - learning algorithm can detect a "possible fall". Since lying or sleeping in bed is a normal movement, this movement will not trigger a user - facing fall alert. However, if a person actually falls on the floor, the machine - learning algorithm can detect a "potential fall". Then, the system determines whether the potential fall is within or outside the blacklist area. If the potential fall is in the blacklist area, the system will not issue an alert. If the potential fall is outside the blacklist area, the system can determine that the fall is a "confirmed fall" and trigger (e.g., send) a user - facing fall alert.
[0102] The posture detection system can include many commercial applications and commercial benefits. For example, for senior living communities, the system can provide predictive insights and prescriptions. In this regard, the system can provide markers or notifications to caregivers for early intervention. Specifically, the system can measure and track frailty based on, for example, an analysis of changes in baseline movement patterns and frailty. The system can also detect and / or mark abnormal activities (e.g., staying in bed or in the bathroom for too long). The system can also detect and / or mark trips to the toilet at night.
[0103] Previous systems typically required active steps and the use of wearable devices or the completion of surveys. However, the current system can provide a private and non - intrusive value proposition. The system can also passively sense actual behavior regardless of changes in behavior.
[0104] Currently, frailty is measured in a clinic or doctor's office. A doctor can perform tests on a patient, which include a series of activities that may take from 5 to 10 minutes to complete. However, such a quick test does not provide a comprehensive understanding of the patient, and the test is conducted in an artificial setting where the patient may be more focused, more effortful, etc. In this regard, the frailty score may vary over time due to different scenarios and the efforts made by the patient. The current system improves the prior art frailty test by using "longitudinal" tracking (including tracking the movement of a person over time). Thus, the current system can provide a more comprehensive understanding of frailty over time.
[0105] Using the pose detection function, the system can provide notifications or reports about events (such as falls, abnormal behavior, etc.) to, for example, caregivers, building management systems, alarm systems, notification systems, and / or emergency response systems. The report can include data associated with the event (such as, for example, location, time of day, movements before and / or after the fall, nearby objects (such as furniture), objects held by the person (such as groceries, walkers, canes, other people), etc.). Compared with radar or other sensor devices, the system can be cheaper, installed faster, and installed more easily. The system can also analyze images to facilitate auditing and / or compliance. For example, the system can use images over a period of time to monitor or audit the care provided (such as, for example, the bed check was completed at 11 pm the previous night). The system can also use images over a period of time to measure the time spent providing care (such as, for example, the average number of minutes per day in the bathroom with a patient). The system can also be integrated with other systems such as scheduling systems and reporting systems. For example, the system can obtain data about when an employee starts a shift, when an employee ends a shift, the employee's name and / or identifier, the time claimed by the employee to assist a certain patient, etc. The system can compare the submitted data with the actual data obtained from images of resident activities and / or resident care to determine the accuracy of the submitted data.
[0106] The detailed description of each embodiment herein refers to the accompanying drawings and pictures, which illustrate the various embodiments by way of illustration. Although the various embodiments have been described in sufficient detail to enable those skilled in the art to practice the present disclosure, it should be understood that other embodiments may be implemented and logical and mechanical changes may be made without departing from the spirit and scope of the present disclosure. Accordingly, the detailed description presented herein is for illustrative purposes only and not for purposes of limitation. For example, the steps recited in any method or process description may be executed in any order and are not limited to the order presented. Additionally, any function or step may be outsourced to or performed by one or more third parties. Modifications, additions, or omissions may be made to the systems, devices, and methods described herein without departing from the scope of the present disclosure. For example, the components of the systems and devices may be integrated or separated. Additionally, the operations of the systems and devices disclosed herein may be performed by more, fewer, or other components, and the methods described may include more, fewer, or other steps. Further, these steps may be executed in any suitable order. As used in this document, "each" refers to each member of a set or each member of a subset of a set. Additionally, any reference to the singular includes multiple embodiments, and any reference to more than one component may include a singular embodiment. Although specific advantages have been recited herein, the various embodiments may include some, none, or all of the recited advantages. Systems and methods are provided.
[0107] In the detailed description herein, references to "various embodiments", "one embodiment", "an embodiment", "example embodiments", etc., mean that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. After reading the description, those skilled in the relevant art will be clear on how to implement the present disclosure in alternative embodiments.
[0108] This document has described benefits, other advantages, and solutions to problems with respect to specific embodiments. However, the benefits, advantages, solutions to problems, and any element that may cause any benefit, advantage, or solution to occur or become more apparent should not be construed as a critical, essential, or fundamental feature or element of the present invention. Accordingly, the scope of the present invention is defined solely by the appended claims, wherein the recitation of an element in the singular is not intended to mean "one and only one" unless explicitly stated to that effect, but rather "one or more." Also, when a phrase such as "at least one of A, B, or C" is used in the claims, it is intended to be interpreted to mean that A can exist alone in an embodiment, B can exist alone in an embodiment, C can exist alone in an embodiment, or any combination of elements A, B, and C can exist in a single embodiment; for example, A and B, A and C, B and C, or A and B and C. Additionally, any element, component, or method step in this disclosure is not intended to be dedicated to the public regardless of whether it is explicitly recited in the claims. No element of any claim of this document shall be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means for." As used herein, the term "comprising," "including," or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0109] A computer program (also referred to as computer control logic) is stored in the main memory and / or auxiliary memory. The computer program can also be received via a communication interface. Such a computer program, when executed, causes the computer system to perform the features discussed herein. In particular, the computer program, when executed, causes the processor to perform the features of the various embodiments. Thus, such a computer program represents the controller of the computer system.
[0110] These computer program instructions can be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed on the computer or other programmable data processing apparatus create means for implementing the functions specified in one or more flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means for implementing the functions specified in one or more flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more flowchart blocks.
[0111] In various embodiments, the software may be stored in a computer program product and loaded into a computer system using a removable storage drive, hard drive, or communication interface. When executed by a processor, the control logic (software) causes the processor to perform the functions of the various embodiments described herein. In various embodiments, the hardware components may take the form of an application specific integrated circuit (ASIC). Implementations of the hardware to perform the functions described herein will be apparent to those skilled in the relevant art.
[0112] As will be understood by one of ordinary skill in the art, the system may be embodied as a customization of an existing system, an add-on product, a processing device executing upgraded software, a stand-alone system, a distributed system, a method, a data processing system, an apparatus for data processing, and / or a computer program product. Thus, any part of the system or module may take the form of a processing device executing code, an Internet-based implementation, a fully hardware implementation, or an implementation combining aspects of the Internet, software, and hardware. Additionally, the system may take the form of a computer program product on a computer-readable storage medium having computer-readable program code means embodied in the storage medium. Any suitable computer-readable storage medium may be utilized, including a hard disk, CD-ROM, BLU-RAY optical storage device, magnetic storage device, etc.
[0113] In various embodiments, the components, modules, and / or engines of the system may be implemented as a micro-application or micro-app. Micro-applications are typically deployed in the context of a mobile operating system, including, for example a mobile operating system, an operating system, Operating systems, the operating systems of the company, etc. The micro application can be configured to utilize the resources of a larger operating system and associated hardware via a set of predefined rules that govern the operation of various operating systems and hardware resources. For example, in the case where the micro application desires to communicate with a device or network other than a mobile device or mobile operating system, the micro application can utilize the communication protocols of the operating system and associated device hardware under the predefined rules of the mobile operating system. Additionally, in the case where the micro application desires input from a user, the micro application can be configured to request a response from the operating system that monitors various hardware components and then transmits the detected input from the hardware to the micro application.
[0114] The system and method can be described herein in terms of functional block components, screen shots, optional selections, and various processing steps. It should be understood that such functional blocks can be implemented by any number of hardware and / or software components configured to perform the specified functions. For example, the system can employ various integrated circuit components such as memory elements, processing elements, logic elements, lookup tables, etc., which can perform various functions under the control of one or more microprocessors or other control devices. Similarly, the software elements of the system can be implemented in any programming or scripting language such as: C, C++, C#, JavaScript Object Notation (JSON), VBScript, Macromedia COLD FUSION, COBOL, the dynamic server pages of the company, assembly, PHP, awk, Visual Basic, SQL stored procedures, PL / SQL, any shell script, and Extensible Markup Language (XML), where various algorithms are implemented using any combination of data structures, objects, processes, routines, or other programming elements. Additionally, it should be noted that the system can employ any number of conventional techniques for data transfer, signal transmission, data processing, network control, etc. Further still, the system can be used to detect or prevent security issues with client-side scripting languages such as VBScript, etc.) or prevent client-side scripting languages such as VBScript, etc.).
[0115] The system and method are described herein with reference to screen shots, block diagrams, and flowchart illustrations of methods, apparatus, and computer program products according to various embodiments. It should be understood that each functional block in the block diagrams and flowchart illustrations, and combinations of functional blocks in the block diagrams and flowchart illustrations, can be implemented separately by computer program instructions.
[0116] Accordingly, the functional blocks of the block diagrams and flowcharts support combinations of devices for performing specified functions, combinations of steps for performing specified functions, and program instruction devices for performing specified functions. It should also be understood that each functional block of the block diagrams and flowcharts, as well as combinations of functional blocks in the block diagrams and flowcharts, can be implemented by a dedicated computer system based on hardware for performing the specified functions or steps, or by a suitable combination of dedicated hardware and computer instructions. In addition, the illustration of the process flow and its description can refer to user application programs, web pages, websites, web forms, prompts, etc. Practitioners will understand that the illustrated steps herein can be included in any number of configurations: including application programs, web pages, web forms, pop-ups the use of application programs, prompts, etc. It should also be understood that the multiple steps illustrated and described can be combined into a single web page and / or application program, but have been expanded for simplicity. In other cases, steps illustrated and described as a single processing step can be divided into multiple web pages and / or application programs, but have been combined for simplicity.
[0117] Middleware can include any hardware and / or software that is suitably configured to facilitate communication and / or processing of transactions between disparate computing systems. Middleware components are commercially available and are known in the art. Middleware can be implemented by commercially available hardware and / or software, by custom hardware and / or software components, or by a combination thereof. Middleware can reside in various configurations and can exist as an independent system or can be a software component residing on an Internet server. Middleware can be configured to process transactions between various components of an application server and any number of internal or external systems for any purpose disclosed herein. of Inc. (Armonk, NY) MQTM (formerly MQSeries) is an example of a commercially available middleware product. Enterprise service bus (“ESB”) applications are another example of middleware.
[0118] The computers discussed herein can provide a suitable website or other Internet-based graphical user interface accessible to users. In one embodiment, the company's Internet Information Services (IIS), Transaction Server (MTS) services, and database with the operating system, web server software, SQL database, and in combination with a commercial server. In addition, components such as software, database, software, software, software, software, software, etc. can be used to provide a database management system compliant with ActiveX Data Objects (ADO). In one embodiment, a web server is combined with an operating system, a database, and PHP, Ruby, and / or a programming language.
[0119] For the sake of brevity, traditional aspects of data networking, application development, and other functional aspects of the system (as well as components of the various operating components of the system) may not be described in detail herein. In addition, the connecting lines shown in the various figures included herein are intended to represent exemplary functional relationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may exist in an actual system.
[0120] In various embodiments, the methods described herein are implemented using the various specific machines described herein. As will be immediately understood by those skilled in the art, the methods described herein can be implemented using the following specific machines and those developed hereinafter in any suitable combination. In addition, as is clear from this disclosure, the methods described herein can result in various transformations of certain articles of manufacture.
[0121] In various embodiments, the system and various components can be integrated with one or more intelligent digital assistant technologies. For example, exemplary intelligent digital assistant technologies can include systems developed by a company, the GOOGLE system developed by Alphabet, Inc., the system of a company, the GOOGLE system, and the systems can each provide cloud-based voice-activated services that can assist with tasks, entertainment, general information, and even more. All AMAZON ECHO AMAZON and AMAZON TVs can all be accessed system. systems, GOOGLE systems and systems can receive voice commands via their voice activation technology, activate other functions, control smart devices, and / or collect information. For example, intelligent digital assistant technology can be used to interact with the following: music, email, text, phone, Q&A, residential improvement information, smart home communication / activation, games, shopping, creating to-do lists, setting alarms, streaming interactive podcasts, playing audiobooks, and providing weather, traffic, and other real-time information such as news. GOOGLE and systems can also allow users to access information about eligible transaction accounts associated with online accounts on all devices that support digital assistants.
[0122] The various system components discussed herein can include one or more of the following: a host server or other computing system including a processor for processing digital data; a memory coupled to the processor for storing digital data; an input digitizer coupled to the processor for inputting digital data; an application program stored in the memory and accessible by the processor for guiding the processor's processing of digital data; a display device coupled to the processor and the memory for displaying information sourced from the digital data processed by the processor; and multiple databases. The various databases used herein can include: client data; merchant data; financial institution data; and / or data useful for the operation of the system. As will be understood by those skilled in the art, a user computer can include an operating system (e.g., etc.) and various conventional support software and drivers typically associated with a computer.
[0123] The present system or any of its part(s) or function(s) can be implemented using hardware, software, or a combination thereof, and can be implemented in one or more computer systems or other processing systems. However, the operations performed by the embodiments can be referred to in terms such as matching or selecting, which are typically associated with intellectual operations performed by a human operator. In most cases, in any of the operations described herein, such capabilities of a human operator are not required or desirable. Instead, the operations can be machine operations, or any operation can be performed or enhanced by artificial intelligence (AI) or machine learning. AI generally can refer to the study of agents (e.g., machines, computer-based systems, etc.) that perceive the world around them, form plans, and make decisions to achieve their goals. The basis of AI includes mathematics, logic, philosophy, probability, linguistics, neuroscience, and decision theory. Many fields fall within the scope of AI, such as computer vision, robotics, machine learning, and natural language processing. Useful machines for performing the various embodiments include general-purpose digital computers or similar devices.
[0124] In various embodiments, the embodiments are directed to one or more computer systems capable of performing the functions described herein. The computer system includes one or more processors. The processors are connected to a communication infrastructure (e.g., communication bus, crossover bar, network, etc.). Various software embodiments are described in accordance with this exemplary computer system. After reading this description, those skilled in the relevant art will be clear on how to use other computer systems and / or architectures to implement the various embodiments. The computer system can include a display interface that forwards graphics, text, and other data from the communication infrastructure (or from a frame buffer not shown) for display on a display unit.
[0125] The computer system also includes a main memory, such as random access memory (RAM), and can also include auxiliary memory. The auxiliary memory can include, for example, a hard disk drive, a solid state drive, and / or a removable storage drive. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner. As will be understood, the removable storage unit includes a computer-usable storage medium in which computer software and / or data are stored.
[0126] In various embodiments, the auxiliary storage may include other similar devices for allowing a computer program or other instructions to be loaded into a computer system. Such devices may include, for example, removable storage units and interfaces. Such examples may include program cartridges and cartridge interfaces (such as those found in video game devices), removable memory chips (such as erasable programmable read-only memories (EPROMs), programmable read-only memories (PROMs)), and associated sockets, or other removable storage units and interfaces that allow software and data to be transferred from the removable storage unit to the computer system.
[0127] The terms "computer program medium", "computer usable medium", and "computer readable medium" are used generally to refer to media such as removable storage drives and hard disks installed in hard disk drives. These computer program products provide software to the computer system.
[0128] The computer system may also include a communication interface. The communication interface allows software and data to be transferred between the computer system and external devices. Examples of such communication interfaces may include modems, network interfaces (such as Ethernet cards), communication ports, etc. The software and data transferred via the communication interface are in the form of signals, which may be electrical, electromagnetic, optical, or other signals that can be received by the communication interface. These signals are provided to the communication interface via a communication path (e.g., a channel). The channel carries the signals and can be implemented using wires, cables, optical fibers, telephone lines, cellular links, radio frequency (RF) links, wireless, and other communication channels.
[0129] As used herein, an "identifier" may be any suitable identifier that uniquely identifies an item. For example, the identifier may be a globally unique identifier ("GUID"). A GUID may be an identifier created and / or implemented under the Universal Unique Identifier standard. Additionally, a GUID may be stored as a 128-bit value, which may be displayed as 32 hexadecimal digits. The identifier may also include a primary number and a secondary number. Both the primary number and the secondary number may be 16-bit integers.
[0130] In various embodiments, the server may include an application server (e.g., POSTGRES PLUS ADVANCED etc.). In various embodiments, the server may include a web server (e.g., Apache, IIS, Web server, System Network Server, running on or operating system virtual machine).
[0131] A network client includes any device or software that communicates via any network, such as, for example, any device or software discussed herein. The network client can include Internet browsing software installed within a computing unit or system to conduct online transactions and / or communications. These computing units or systems can take the form of a computer or a group of computers, although other types of computing units or systems can be used, including personal computers, laptops, notebooks, tablets, smartphones, cellular phones, personal digital assistants, servers, pool servers, mainframe computers, distributed computing clusters, kiosks, terminals, point-of-sale (POS) devices or terminals, televisions, or any other device capable of receiving data over a network. The network client can include an operating system (e.g., WINDOW operating system, operating system, operating system, operating system, etc.) and various traditional support software and drivers, which are typically associated with a computer. The network client can also run INTERNET software, software, GOOGLE CHROME TM software, software, or any of the countless software packages available for browsing the Internet.
[0132] As will be understood by those skilled in the art, the network client may or may not have direct contact with a server (e.g., an application server, a web server, etc., as discussed herein). For example, the network client can access the services of a server through another server and / or hardware component, which can be directly or indirectly connected to the Internet server. For example, the network client can communicate with the server via a load balancer. In various embodiments, network client access is via a network or the Internet through a commercially available web browser software package. In this regard, the network client can be in a residential or commercial environment capable of accessing a network or the Internet. The network client can implement security protocols such as Secure Sockets Layer (SSL) and Transport Layer Security (TLS). The network client can implement several application layer protocols, including HTTP, HTTPS, FTP, and SFTP.
[0133] Various system components can be appropriately coupled to the network independently, individually, or jointly via data links, which include, for example, a connection to an Internet service provider (ISP) through a local loop, which typically communicates with a standard modem, cable modem, DISH ISDN, digital subscriber line (DSL), or various wireless communication methods are used in combination. It is worth noting that the network can be implemented as other types of networks, such as an interactive television (ITV) network. In addition, the system contemplates the use, sale, or distribution of any goods, services, or information via any network having similar functionality as described herein.
[0134] The system contemplates uses associated with network services, utility computing, pervasive and personalized computing, security and identity solutions, autonomic computing, cloud computing, commodity computing, mobile and wireless solutions, open source, biometrics, grid computing, and / or mesh computing.
[0135] Any communication, input, storage, database, or display discussed herein can be facilitated by a website having web pages. As used herein, the term "web page" does not imply a limitation on the types of documents and applications available for interaction with a user. For example, in addition to standard HTML documents, a typical website may also include various forms, applets, programs, active server pages (ASP), common gateway interface scripts (CGI), extensible markup language (XML), dynamic HTML, cascading style sheets (CSS), AJAX (Asynchronous JavaScript and XML) programs, helper applications, plug-in programs, etc. The server may include a network service that receives requests from a web server, the request including a URL and an IP address (192.168.1.1). The web server retrieves the appropriate web page and sends the data or application of the web page to the IP address. A network service is an application capable of interacting with other applications via a communication means such as the Internet. Network services are typically based on standards or protocols such as XML, SOAP, AJAX, WSDL, and UDDI. Network service methods are well known in the art and are covered in many standard texts. For example, Representational State Transfer (REST) or RESTful, network services can provide a way to achieve interoperability between applications.
[0136] The computing unit of the network client can also be equipped with an Internet browser that uses standard dial-up, cable, DSL, or any other Internet protocol known in the art to connect to the Internet or an intranet. Transactions originating from the network client may pass through a firewall to prevent unauthorized access by users from other networks. In addition, additional firewalls can be deployed between different components of the CMS to further enhance security.
[0137] Encryption can be performed by any technique now available or likely to become available in the art—for example, Twofish, RSA, El Gamal, Schorr signature, DSA, PGP, PKI, GPG (GnuPG), HPE Format-Preserving Encryption (FPE), Voltage, Triple DES, Blowfish, AES, MD5, HMAC, IDEA, RC6, and symmetric and asymmetric cryptosystems. The system and method can also incorporate SHA-family cryptographic methods, elliptic curve cryptography (e.g., ECC, ECDH, ECDSA, etc.) and / or other post-quantum cryptographic algorithms under development.
[0138] A firewall can include any hardware and / or software suitably configured to protect CMS components and / or enterprise computing resources from users of other networks. Additionally, a firewall can be configured to restrict or constrain access to the various systems and components behind the firewall from network clients connected through a web server. A firewall can reside in different configurations, including stateful inspection, proxy-based, access control lists, and packet filtering, etc. A firewall can be integrated within a web server or any other CMS component, or can alternatively reside as a separate entity. A firewall can implement Network Address Translation (“NAT”) and / or Network Address Port Translation (“NAPT”). A firewall can accommodate various tunneling protocols to facilitate secure communications, such as those used in virtual private networks. A firewall can implement a Demilitarized Zone (“DMZ”) to facilitate communications with a public network such as the Internet. A firewall can be integrated as software within an Internet server or any other application server component, reside in another computing device, or take the form of a stand-alone hardware component.
[0139] Any database discussed herein can include relational, hierarchical, graph, blockchain, object-oriented structures, and / or any other database configuration. Any database can also include a flat file structure, where data can be stored in a single file in rows and columns, without a structure for indexing and without a structural relationship between records. For example, a flat file structure can include delimited text files, CSV (comma-separated values) files, and / or any other suitable flat file structure. Common database products that can be used to implement a database include (Armonk, NY) available from Corporation (Redwood Shores, CA) of various database products, Corporation (Redmond, Washington) of MICROSOFT or MICROSOFT SQL of MySQL AB (Uppsala, Sweden) Redis, Apache the company's MapR-DB, or any other suitable database product. Additionally, any database can be organized in any suitable manner, for example, as a data table or a lookup table. Each record can be a single file, a series of files, a series of linked data fields, or any other data structure.
[0140] As used herein, big data can refer to a partially or fully structured, semi-structured, or unstructured data set that includes millions of rows and hundreds of thousands of columns. For example, a big data set can be compiled from purchase transaction histories over a period of time, from web registrations, from social media, from records of charges (ROC), from summaries of charges (SOC), from internal data, or from other suitable sources. A big data set can be compiled without descriptive metadata such as column types, counts, percentiles, or other interpretive data points.
[0141] The association of certain data can be achieved by any desired data association technique (such as those known or practiced in the art). For example, the association can be done manually or automatically. Automatic association techniques can include, for example, database searches, database merges, GREP, AGREP, SQL, using key fields in a table to speed up the search, sequential searches through all tables and files, sorting the records in a file according to a known order to simplify the lookup, etc. The association step can be done through a database merge function, for example, using "key fields" in a pre-selected database or data sector. Various database tuning steps are envisioned to optimize database performance. For example, frequently used files (such as indexes) can be placed on a separate file system to reduce input / output ("I / O") bottlenecks.
[0142] More particularly, the "keyword field" divides the database according to the high-level categories of the objects defined by the keyword field. For example, certain types of data can be designated as keyword fields in multiple related data tables, and then the data tables can be linked based on the data types in the keyword fields. The data corresponding to the keyword fields in each of the linked data tables is preferably the same or of the same type. However, data tables with similar but not identical data in the keyword fields can also be linked by using, for example, AGREP. According to one embodiment, any suitable data storage technology can be utilized to store data without a standard format. Any suitable technology can be used to store data sets, including, for example, storing a single file using the ISO / IEC 7816-4 file structure; implementing a domain so as to select a dedicated file that exposes one or more base files containing one or more data sets; utilizing data sets stored in a single file using a hierarchical archive system; data sets stored as records in a single file (including compressed, SQL-accessible, hashed via one or more keys, numeric, alphabetically ordered by a first tuple, etc.); data stored as binary large objects (BLOBs); data stored as ungrouped data elements encoded using ISO / IEC 7816-6 data elements; data stored as ungrouped data elements encoded using ISO / IEC Abstract Syntax Notation (ASN.1), as in ISO / IEC 8824 and 8825; other proprietary technologies, which may include fractal compression methods, image compression methods, etc.
[0143] In various embodiments, the ability to store multiple types of information in different formats is facilitated by storing the information as BLOBs. Thus, any binary information can be stored in the storage space associated with the data set. As discussed above, the binary information can be stored in association with the system or outside the system but attached to the system. The BLOB method can use fixed storage allocation, circular queue techniques, or best practices regarding memory management (e.g., paged memory, least recently used, etc.) to store the data set as ungrouped data elements formatted as binary blocks via a fixed memory offset. By using the BLOB method, the ability to store various data sets with different formats facilitates the storage of data in the database by multiple and unrelated owners of the data sets or the storage of data associated with the system. For example, a first data set that can be stored can be provided by a first party, a second data set that can be stored can be provided by an unrelated second party, and a third data set that can be stored can be provided by a third party unrelated to the first and second parties. Each of these three exemplary data sets can contain different information stored using different data storage formats and / or technologies. In addition, each data set can contain data subsets that may also be different from other subsets.
[0144] As stated above, in various embodiments, data can be stored without regard to a general format. However, when provided for manipulation of data in a database or system, data sets (e.g., BLOBs) can be annotated in a standard manner. The annotation can include a short header, a footer, or other suitable indicators associated with each data set, each data set being configured to convey information useful for managing the various data sets. For example, the annotation can be referred to herein as a "conditional header", "header", "footer", or "status", and can include an indication of the data set status or can include an identifier associated with a particular publisher or owner of the data. In one example, the first three bytes of each data set BLOB can be configured or capable of being configured to indicate the status of that particular data set; for example, loaded, initialized, ready, blocked, removable, or deleted. Subsequent bytes of the data can be used to indicate the identity of, for example, the publisher, user, transaction / member account identifier, etc. Each of these conditional annotations is further discussed herein.
[0145] Data set annotations can also be used for other types of status information and various other purposes. For example, data set annotations can include security information for establishing access levels. For example, the access level can be configured to allow only certain individuals, levels of employees, companies, or other entities to access the data set, or to allow access to specific data sets based on transactions, merchants, publishers, users, etc. Additionally, the security information can restrict / allow only certain actions, such as accessing, modifying, and / or deleting the data set. In one example, the data set annotation indicates that only the data set owner or user can delete the data set, that various identified users can access the data set for reading, and that other users are completely excluded from accessing the data set. However, other access restriction parameters can also be used, thus allowing various entities to access data sets with various appropriate permission levels.
[0146] Data including a header or footer can be received by a separate interactive device configured to add, delete, modify, or supplement the data based on the header or footer. Thus, in one embodiment, the header or footer is not stored on the transaction device with the data owned by the associated publisher, but rather appropriate actions can be taken by providing the user with appropriate options for actions to be taken at a separate device. The system can contemplate a data storage arrangement where the header or footer of the data, or the header history or footer history, is stored on the system, device, or transaction tool in association with the appropriate data.
[0147] Those skilled in the art will also understand that, for security reasons, any database, system, device, server, or other component of a system can include any combination thereof at a single location or multiple locations, where each database or system includes any one of a variety of suitable security features (such as firewalls, access codes, encryption, decryption, compression, decompression, etc.).
[0148] Practitioners will also realize that there are many ways to display data within a browser-based document. The data can be represented as standard text or in a fixed list, scrollable list, drop-down list, editable text field, fixed text field, pop-up window, etc. Similarly, there are many ways to modify data in a web page, such as, for example, entering free text using a keyboard, selecting menu items, check boxes, option boxes, etc.
[0149] The data can be big data processed by a distributed computing cluster. The distributed computing cluster can be, for example a software cluster that is configured to process and store large data sets, where some nodes include a distributed storage system and some nodes include a distributed processing system. In this regard, the distributed computing cluster can be configured to support the software distributed file system (HDFS) as specified by the Apache Software Foundation at www.hadoop.apache.org / docs
[0150] As used herein, the term "network" includes any cloud, cloud computing system, or electronic communication system or method that combines hardware and / or software components. Communication between parties can be accomplished through any suitable communication channel, such as, for example, a telephone network, extranet, intranet, Internet, point-of-interaction device (point-of-sale device, personal digital assistant (e.g., device, device), cellular phone, kiosk, etc.), online communication, satellite communication, offline communication, wireless communication, repeater communication, local area network (LAN), wide area network (WAN), virtual private network (VPN), networked or linked devices, keyboard, mouse, and / or any suitable form of communication or data input. Additionally, although the system is often described herein as being implemented using the TCP / IP communication protocol, the system can also be implemented using IPX, program, IP-6, NetBIOS, OSI, any tunneling protocol (such as IPsec, SSH, etc.), or any number of existing or future protocols. If the network has the nature of a public network, such as the Internet, it may be advantageous to assume that the network is insecure and open to eavesdroppers. Specific information regarding the protocols, standards, and application software used in conjunction with the Internet is generally known to those skilled in the art and thus need not be detailed here.
[0151] "Cloud" or "cloud computing" includes a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction. Cloud computing can include location-independent computing that provides resources, software, and data to computers and other devices on demand via shared servers.
[0152] As used herein, "transmission" can include sending electronic data from one system component to another via a network connection. Additionally, as used herein, "data" can include information encompassing, for example, commands, queries, files, data for storage, etc. in digital or any other form.
[0153] Any database discussed herein can include a distributed ledger maintained by multiple computing devices (e.g., nodes) on a peer-to-peer network. Each computing device maintains a copy and / or partial copy of the distributed ledger and communicates with one or more other computing devices in the network to verify data and write data to the distributed ledger. The distributed ledger can utilize features and functions of blockchain technology, including, for example, consensus-based verification, immutability, and cryptographically linked data blocks. A blockchain can include a ledger of interconnected blocks containing data. A blockchain can provide enhanced security as each block can hold individual transactions and the results of any blockchain executable. Each block can be linked to the previous block and can include a timestamp. Blocks can be linked because each block can include the hash of the previous block in the blockchain. Linked blocks form a chain where only one successor block is allowed to link to another predecessor block of a single chain. Forks can occur in cases where a divergent chain is established from a previously unified blockchain, although typically only one divergent chain will be maintained as the consensus chain. In various embodiments, a blockchain can implement smart contracts for implementing data workflows in a decentralized manner. The system can also include applications deployed on user devices, such as, for example, computers, tablets, smartphones, Internet of Things devices ("IoT" devices), etc. The applications can communicate with the blockchain (e.g., directly or via blockchain nodes) to transmit and retrieve data. In various embodiments, a governing organization or consortium can control access to data stored on the blockchain. Registration with the (one or more) governing organizations can participate in the blockchain network.
[0154] Data transfers performed via a blockchain-based system can propagate to connected peers within the blockchain network within a duration determinable by the block creation time of the particular blockchain technology implemented. For example, on a -based network, new data entries can become available within approximately 13 to 20 seconds after write completion. On a On the Fabric 1.0 platform, the duration is driven by the specific consensus algorithm selected and can be executed within seconds. In this regard, compared with existing systems, the propagation time in the system can be improved, and the implementation cost and time to market can also be significantly reduced. The system also provides higher security, at least in part due to the immutability of the data stored in the blockchain, reducing the possibility of tampering with various data inputs and outputs. In addition, the system can also provide higher data security by performing encryption processing on the data before storing it on the blockchain. Therefore, by using the system described herein to transmit, store, and access data, the security of the data is improved, which reduces the risk of harm to the computer or network.
[0155] In various embodiments, the system can also reduce database synchronization errors by providing a common data structure, thus at least partially improving the integrity of the stored data. The system also provides higher reliability and fault tolerance than traditional databases (e.g., relational databases, distributed databases, etc.) because each node operates using a complete copy of the stored data, thus at least partially reducing the downtime caused by local network outages and hardware failures. The system can also improve the reliability of data transfer in a network environment with reliable and unreliable peers because each node broadcasts messages to all connected peers, and since each block includes a link to the previous block, nodes can quickly detect lost blocks and propagate requests for the lost blocks to other nodes in the blockchain network.
[0156] The specific blockchain implementation described herein provides improvements over traditional technologies by using a decentralized database and an improved processing environment. In particular, the blockchain implementation improves computer performance by, for example, leveraging decentralized resources (e.g., lower latency). Distributed computing resources improve computer performance by, for example, reducing processing time. In addition, distributed computing resources improve computer performance by enhancing security using, for example, cryptographic protocols.
[0157] Any communication, transmission, and / or channel discussed herein can include any system or method for delivering content (e.g., data, information, metadata, etc.) and / or the content itself. The content can be presented in any form or medium, and in various embodiments, the content can be delivered electronically and / or be capable of being presented electronically. For example, a channel can include a website, a mobile application or device (e.g., GOOGLECHROMECAST TM 、 etc.) uniform resource locator (“URL”), a document (e.g., Word or EXCELTM , Portable Document Format (PDF) documents, etc., "e-books", "e-magazines", applications or micro-applications (as described herein), Short Message Service (SMS) or other types of text messages, e-mails, messages, tweets, Multimedia Messaging Service (MMS), and / or other types of communication technologies. In various embodiments, the channels may be hosted or provided by data partners. In various embodiments, the distribution channels may include at least one of a merchant website, a social media website, an affiliate or partner website, an external vendor, mobile device communication, a social media network, and / or a location-based service. The distribution channels may include at least one of a merchant website, a social media website, an affiliate or partner website, an external vendor, and mobile device communication. Examples of social media websites include
[0158]
[0159] etc. Examples of affiliate or partner websites include AMERICAN etc.
Claims
1. A method for pose detection, the method for pose detection comprising: One or more processors receive an image of a human from a sensor; The one or more processors create a bounding box on the image using a convolutional neural network (CNN), wherein the convolutional neural network is trained using gradient descent based on a hand-drawn bounding box around the image of the human, pixel data of the human in the image, thermal data around the human, and environmental data around the human; The one or more processors adjust the bounding box based on the environmental data around the human; The one or more processors obtain bounding box data from within the bounding box; The one or more processors use the convolutional neural network to determine discriminative features of the image in the bounding box based on the bounding box data, filters, and overhead thermal features of the human presence, wherein the convolutional neural network uses stochastic gradient descent to create a feature map of the discriminative features; The one or more processors determine one or more poses of the human based on the bounding box data and the discriminative features; The one or more processors determine an aggregated pose based on one or more of the poses across multiple frames over a period of time, wherein the aggregated pose smooths out changes in the one or more poses across the multiple frames over the period of time; and The one or more processors determine an event as a potential fall based on the aggregated pose changing from at least one of a standing pose or a sitting pose to a lying pose and the lying pose persisting for a certain amount of time based on a timestamp.
2. The method according to claim 1, wherein, The aggregated pose is the pose that occurs most frequently in a set of one or more of the poses.
3. The method according to claim 1, the method further comprising: The one or more processors limit the resolution of the image based on at least one of privacy concerns, sensor power consumption, data cost, data bandwidth, computational cost, or computational bandwidth.
4. The method according to claim 1, wherein, The convolutional neural network determining the discriminative features includes: The one or more processors create a vector of features of the image; The one or more processors use the convolutional neural network to learn a 2D representation of the image; The one or more processors flatten the image; The one or more processors supplement the image with positional encoding; The one or more processors use a transformer decoder to receive multiple learned positional embeddings for encoder output; and The one or more processors pass the learned positional embeddings of the encoder output to a shared feed-forward network (FFN), wherein the shared feed-forward network provides box predictions by predicting at least one of a detection class and a bounding box or a "no object" class.
5. The method according to claim 1, the method further comprising: The algorithm of the convolutional neural network is adjusted by the one or more processors based on environmental data, where the environmental data includes at least one of the following: outdoor temperature, indoor temperature, floor plan, stove, laptop, gender of the human, age of the human, height of the sensor, clothing of the human, or body weight of the human.
6. The method according to claim 1, the method further comprising: The convolutional neural network is applied by the one or more processors to the interior of the bounding box; The size and position of the bounding box are predicted by the one or more processors; and A pose prediction for the image is created by the one or more processors based on the gradient descent-based training of the convolutional neural network.
7. The method according to claim 1, wherein, Based on the confidence score, the following is performed: determining the event as a potential fall.
8. The method according to claim 1, wherein, Determining one or more of the poses is performed for a frame in the image captured by the sensor.
9. The method according to claim 1, the method further comprising post-processing by at least one of the following: Occluding, by the one or more processors, spatial locations where the event cannot reasonably occur; or Filtering, by the one or more processors, the event based on a confidence threshold for detection of the event.
10. The method according to claim 1, the method further comprising at least one of the following: Determining, by the one or more processors, whether the potential fall is located in a blacklisted spatial location; or Determining, by the one or more processors, a threshold confidence level at which the potential fall is determined to be a confirmed fall.
11. The method according to claim 1, the method further comprising: The one or more processors determine whether the potential fall is outside a blacklist spatial location; and A confirmed fall alert is sent by the one or more processors based on the potential fall being outside the blacklist location.
12. The method according to claim 1, the method further comprising: The resolution of the image is restricted by the one or more processors based on at least one of the following: privacy concerns, power consumption of the sensor, cost of the pixel data, bandwidth of the pixel data, computational cost, or computational bandwidth.
13. The method according to claim 1, the method further comprising: The one or more processors label one or more of the poses of the human in the image.
14. The method according to claim 1, wherein The image is a portion of a video clip of the human.
15. The method according to claim 1, wherein Obtaining the bounding box data includes: obtaining the bounding box data in at least one of the cases of over time or during an initial calibration session.
16. The method according to claim 1, the method further comprising: The discriminative features from the overhead thermal feature pattern of the human are analyzed by the one or more processors to determine one or more of the poses of the human.
17. The method according to claim 1, wherein One or more of the poses include at least one of the following: sitting, standing, lying down, dancing, running, or eating.
18. The method according to claim 1, the method further comprising: The temperature of the human in the space is determined by the one or more processors based on infrared energy data regarding the infrared (IR) energy from the human. The position coordinates of the human in the space are determined by the one or more processors; The position coordinates of the human are compared by the one or more processors with the position coordinates of a fixed object; and In response to the temperature of the human being within a range and in response to the position coordinates of the human being different from the position coordinates of the fixed object, the one or more processors determine that: the human is a person.
19. An article of manufacture for pose detection, the article of manufacture for pose detection comprising a non-transitory tangible computer-readable storage medium having instructions stored thereon that, in response to execution by one or more processors, cause the one or more processors to perform operations including the following: Receiving, by the one or more processors, an image of a human from a sensor; Creating, by the one or more processors, a bounding box on the image using a convolutional neural network (CNN), wherein The convolutional neural network uses gradient descent-based training based on a hand-drawn bounding box around the image of the human, pixel data of the human in the image, thermal data around the human, and environmental data around the human; The bounding box is adjusted by the one or more processors based on the environmental data around the human; The one or more processors obtain bounding box data from within the bounding box; The one or more processors use the convolutional neural network to determine discriminative features of the image in the bounding box based on the bounding box data, filters, and the overhead thermal signature of the human presence, wherein the convolutional neural network uses stochastic gradient descent to create a feature map of the discriminative features; The one or more processors determine one or more poses of the human based on the bounding box data and the discriminative features; The one or more processors determine an aggregated pose based on one or more of the poses across multiple frames over a period of time, wherein the aggregated pose smooths out variations in the one or more poses across the multiple frames over the period of time; and The one or more processors determine an event as a potential fall based on the aggregated pose changing from at least one of a standing pose or a sitting pose to a lying pose and the lying pose persisting for a certain amount of time based on a timestamp.
20. A system for pose detection, the system for pose detection comprising: One or more processors; And A tangible non-transitory memory configured to communicate with the one or more processors, The tangible non-transitory memory having instructions stored thereon that, in response to execution by the one or more processors, cause the one or more processors to perform operations including the following: The one or more processors receive an image of a human from a sensor; The one or more processors create a bounding box on the image using a convolutional neural network (CNN), wherein the convolutional neural network is trained using gradient descent based on a hand-drawn bounding box around the image of the human, pixel data of the human in the image, thermal data around the human, and environmental data around the human; The one or more processors adjust the bounding box based on the environmental data around the human; The one or more processors obtain bounding box data from within the bounding box; The one or more processors use the convolutional neural network to determine discriminative features of the image in the bounding box based on the bounding box data, filters, and the overhead thermal signature of the human presence, wherein the convolutional neural network uses stochastic gradient descent to create a feature map of the discriminative features; The one or more processors determine one or more poses of the human based on the bounding box data and the discriminative features; The one or more processors determine an aggregated pose based on one or more of the poses across multiple frames over a period of time, wherein the aggregated pose smooths out variations in the one or more poses across the multiple frames over the period of time; and Determine an event as a potential fall by the one or more processors changing from at least one of a standing posture or a sitting posture to a lying posture based on the aggregated posture and the lying posture lasting for a certain amount of time based on a timestamp.
Citation Information
Patent Citations
Monitoring human location, trajectory and behavior using thermal data
US11022495B1
Thermal data analysis for determining location, trajectory and behavior
US20210278279A1
User interface for determining location, trajectory and behavior
US20220057270A1
Systems and methods for gesture-based interaction with computer systems
CA2781511A1
Method for manipulating posture of user interface and posture correction
CN102184020A