Movement Instructions
By weighting subframes of video data, motion indicators for the scene are generated, solving the problem of accurately identifying scene movement in existing technologies and achieving more efficient motion monitoring and recognition.
Patent Information
- Application Number
- CN201980084163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-19
- Filing Date
- 2019-12-03
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2039-12-03
AI Technical Summary
Existing technologies lack effective methods for generating motion indications for a scene, especially when extracting motion data from video data, making it difficult to accurately identify and monitor motion information in the scene.
By receiving video data, motion measurements are determined in subframes of the video data, and these measurements are weighted to generate motion indicators for the scene. The weighting process can be based on historical motion data of the subframes, the position of the subframes within the scene, and the scene's visibility characteristics, and can be updated during runtime.
It improves the accuracy and efficiency of scene movement recognition, reduces data storage and processing requirements, and is suitable for non-contact sleep monitoring and other applications such as security monitoring and animal tracking.
Smart Images

Figure CN113243025B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to generating scene-specific movement instructions based on scene-based video data. Background Technology
[0002] Motion data can be extracted from video data. However, there remains a need for alternative arrangements to generate movement instructions for a scene. Summary of the Invention
[0003] In a first aspect, this specification describes an apparatus comprising: means for receiving video data for a scene; means for determining motion measurements of at least some of the subframes (e.g., subdivisions of video data frames) of the video data; means for weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and means for generating a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements.
[0004] The weighting can be binary weighting. For example, by providing binary weighting, some subframes can be ignored when generating a scene-specific motion indication from a combination of weighted motion measurements. In some embodiments, the weighting can be non-binary.
[0005] The weighting of at least some subframes may depend at least in part on one or more of the following: historical motion data for the subframe; the subframe's position within the scene; and the scene's visibility characteristics.
[0006] In some embodiments, one or more weights are updated in operation. For example, the weights can be adjusted when mobile data is received. For example, as more mobile data for a particular subframe is acquired, historical mobile data that may be associated with the weights of that subframe can be updated, so that the weights can be updated in use or in operation.
[0007] In some embodiments, the resulting motion indication can be generated offline based on motion measurement data acquired and weighted during operation (e.g., this can save storage requirements and / or reduce data processing requirements). Motion measurements for a subframe include an indication of the degree of movement between consecutive instances of the corresponding subframe.
[0008] In some embodiments, determining a motion measurement for a subframe of video data includes: means for determining an average distance measurement for a corresponding subframe of received video data; and means for determining, for one or more average distance measurements, whether the average distance measurement differs from the average distance for a corresponding subframe of a preceding (e.g., immediately preceding) video data frame instance.
[0009] If the average distance measurement between consecutive subframes is greater than a threshold, the component used to determine the motion measurement for a subframe of video data determines that motion has occurred. The threshold can be based on a determined standard deviation threshold (e.g., the threshold might be five standard deviations of the average).
[0010] Video data can include depth data for multiple pixels of an image of a scene. Furthermore, motion measurements for subframes of the video data can be determined based on changes in the received depth data for the corresponding subframe. For example, the received depth data can be determined by calculating the average distance evolution for each subframe.
[0011] The video data is received from a multi-mode camera that includes an infrared projector and sensors.
[0012] Subframes may include regular tile grids (such as square tiles); however, many other arrangements are also possible. Subframes can have many different shapes, and it is not necessary for all subframes to have the same shape for all embodiments.
[0013] The scene can be a fixed scene. For example, a scene can be captured using a camera with a fixed position and a constant orientation and focal length.
[0014] The components may include: at least one processor; and at least one memory, including computer program code, wherein the at least one memory and the computer program code are configured to cause the execution of the device together with the at least one processor.
[0015] In a second aspect, this specification describes a method comprising: receiving video data for a scene; determining motion measurements for at least some of a plurality of subframes of the video data; weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and generating a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements. The motion measurements of the subframes include an indication of the degree of movement between consecutive instances of the respective subframes.
[0016] The weighting can be binary weighting or non-binary weighting. Alternatively or additionally, the weighting of a subframe can depend at least in part on one or more of the following: historical motion data for the subframe; the subframe's position within the scene; and the scene's visibility characteristics.
[0017] In some embodiments, one or more weights are updated during operation / use.
[0018] In some embodiments, determining a motion measurement for a subframe of video data includes: determining an average distance measurement for the corresponding subframe of the received video data; and, for one or more average distance measurements, determining whether the average distance measurement differs from the average distance for the corresponding subframe of a preceding (e.g., immediately preceding) video data frame instance.
[0019] Motion measurements for subframes of video data can be determined based on changes in received depth data for the corresponding subframe. For example, received depth data can be determined by calculating the average distance evolution for each subframe.
[0020] In the third aspect, this specification describes any apparatus configured to perform any of the methods described with reference to the second aspect.
[0021] In the fourth aspect, this specification describes computer-readable instructions that, when executed by a computing device, cause the computing device to perform any of the methods described with reference to the second aspect.
[0022] In a fifth aspect, this specification describes a computer program comprising instructions for causing a device to perform at least the following: receiving video data for a scene; determining motion measurements for at least some of a plurality of subframes of the video data; weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and generating a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements.
[0023] In a sixth aspect, this specification describes a computer-readable medium (such as a non-transient computer-readable medium) including program instructions stored thereon for performing at least the following: receiving video data for a scene; determining motion measurements for at least some of a plurality of subframes of the video data; weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and generating a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements.
[0024] In a seventh aspect, this specification describes an apparatus comprising: at least one processor; and at least one memory including computer program code, which, when executed by the at least one processor, causes the apparatus to: receive video data for a scene; determine motion measurements for at least some of a plurality of subframes of the video data; weight the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and generate a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements.
[0025] In an eighth aspect, this specification describes an apparatus comprising: a first input for receiving video data for a scene and / or for acquiring video data for a scene; a first control module for determining motion measurements of at least some of the subframes of the video data; a second control module for weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframe; and a third control module for generating a motion indication for the scene from a combination (such as a sum) of some or all of the weighted motion measurements. At least some of the first, second, and third control modules may be combined. Attached Figure Description
[0026] The example will now be described in a non-limiting manner with reference to the following diagram, wherein:
[0027] Figure 1 This is a block diagram of a system according to an example embodiment;
[0028] Figure 2 This is a flowchart illustrating the algorithm according to an example embodiment;
[0029] Figure 3 Example output based on an example embodiment is shown;
[0030] Figure 4 Example output based on an example embodiment is shown;
[0031] Figure 5 This is a block diagram of a system according to an example embodiment;
[0032] Figure 6 This is a flowchart illustrating the algorithm according to an example embodiment;
[0033] Figure 7 This is a flowchart illustrating the algorithm according to an example embodiment;
[0034] Figure 8 Initialization data according to an example embodiment is shown;
[0035] Figure 9A and Figure 9B Results based on an example embodiment are shown;
[0036] Figure 10A and Figure 10B Results based on an example embodiment are shown;
[0037] Figure 11 This is a block diagram of the components of a system according to an example embodiment; and
[0038] Figure 12A and Figure 12B Tangible media for storing computer-readable code according to example embodiments are shown, namely a removable memory cell and a compact disc (CD), the computer-readable code performing operations when executed by a computer. Detailed Implementation
[0039] Sleep disorders are associated with a number of problems, including mental and medical issues. Sleep disorders are sometimes classified into several subcategories, such as intrinsic sleep disorders, extrinsic sleep disorders, and circadian rhythm-based sleep disorders.
[0040] Examples of intrinsic sleep disorders include idiopathic narcolepsy, narcolepsy, periodic limb movement disorder, restless legs syndrome, sleep apnea, and misunderstanding of sleep state.
[0041] Examples of external sleep disorders include alcohol-dependent sleep disorder, food allergy-related insomnia, and general sleep deprivation.
[0042] Examples of sleep disorders affecting circadian rhythms include advanced sleep stage syndrome, delayed sleep stage syndrome, jet lag, and shift worker sleep disorders.
[0043] A better understanding of sleep physiology and pathophysiology may help improve the care received by individuals with this difficulty. Many sleep disorders are diagnosed primarily based on self-reported complaints. The lack of objective data can hinder the understanding of cases and the care provided.
[0044] Non-contact, discrete longitudinal home monitoring may be the most suitable option for monitoring sleep quality over a period of time, and may be superior to clinic-based monitoring, for example, because sleep patterns may vary daily depending on food intake, lifestyle, and health status.
[0045] Figure 1This is a block diagram of a system according to an example embodiment, generally indicated by reference numeral 10. System 10 includes a bed 12 and a camera 14. For example, the camera may be mounted on a tripod 15. As discussed further below, the camera 14 may be a depth camera, although this is not necessary for all embodiments.
[0046] In the use of system 10, a patient (or some other subject) can sleep on bed 12, and camera 14 is used to record various aspects of sleep quality. For example, bed 12 can be in the patient's home, thus enabling home monitoring potentially over an extended period of time.
[0047] Figure 2 The flowchart illustrates the algorithm according to an example embodiment, generally indicated by reference numeral 20.
[0048] Algorithm 20 begins at operation 22, where video data is received (e.g., from camera 14 described above). In some example embodiments, camera 14 is a depth camera (e.g., a camera including an RGB-D sensor) such that the video data received in operation 22 includes depth data for a multi-pixel image of the scene. Example depth cameras may capture pixel-by-pixel depth information via infrared (IR) mediated structured light modes with stereo sensing or time-of-flight sensing to generate a depth map. Some depth cameras include 3D sensing capabilities, allowing for a synchronized image stream of depth information. Many other sensors are also possible (including video sensors that do not provide depth information).
[0049] Figure 3 An example output according to an exemplary embodiment is shown, generally indicated by reference numeral 30. Output 30 is an example of received video data in operation 22 of algorithm 20.
[0050] The received video data in operation 22 can be video data of a fixed scene, where the camera is in a fixed position and has a constant orientation and focal length. Therefore, for example, when a person is sleeping on bed 12 of the system 10 described above, changes in the video data are related to the person's movement. By providing a fixed scene, changes in the video data (e.g., due to user movement) can be monitored over time (e.g., after several hours of sleep).
[0051] The motion measurement is determined at operation 24 of algorithm 20. For example, motion between consecutive frames of the video data captured in operation 22 can be determined.
[0052] More specifically, operation 24 can determine motion measurements for each of multiple subframes of video data. For example, Figure 4An example output according to an exemplary embodiment is shown, generally indicated by reference numeral 40. Output 40 shows data from output 30 that has been divided into multiple subframes (sixteen subframes are shown in output 40). A subframe is a subdivision of a video data frame. For example, a frame can be divided into multiple subframes. Subframes can be square, but can take different forms (such as circular or irregular shapes). Furthermore, it is not necessary for all subframes to have the same shape.
[0053] Therefore, although example output 40 includes subframes in the form of a regular tile grid, this is not mandatory. Other arrangements are also possible. The size, shape, and / or arrangement of the subframes can be different and can be configurable.
[0054] Operation 24 determines motion measurements between consecutive instances of each subframe. In one example implementation, consecutive instances of subframes are separated every 5 seconds (although other intervals are possible, of course). As discussed further below, motion measurements for each subframe across multiple subframes of video data can be determined based on changes in received depth data for the respective subframe (e.g., by calculating the average distance evolution for each subframe). Even with more subframes available, separating consecutive instances of subframes into time intervals such as 5 seconds increases the likelihood of measurable motion occurring between multiple frames and reduces the system's data storage and processing requirements.
[0055] At operation 26, the determined motion measurements are weighted to generate multiple weighted motion measurements. As discussed in detail below, the weighting can be subframe-dependent. For example, the weighting for each subframe can be based at least in part on the historical motion data of that subframe. Therefore, subframes where recent motion occurred may have a higher weight than subframes where no motion occurred.
[0056] Finally, in operation 28, a motion indication based on weighted motion measurements is provided. For example, operation 28 may determine whether motion is detected in any subframe of a video image instance. The motion indication can be generated from a combination (such as a sum) of some or all weighted motion measurements.
[0057] Figure 5 This is a block diagram of a system according to an example embodiment, generally indicated by reference numeral 50. System 50 includes an imaging device 52 (such as camera 14), a data storage module 54, a data processing module 56, and a control module 58.
[0058] Data storage module 54 can store measurement data (e.g., the output of operation 24 of algorithm 20) under the control of control module 58. The stored data is processed by data processor 56 to generate an output (e.g., the output of operation 28 of operation 20). The output can be provided to control module 58. In an alternative arrangement, weighted data (e.g., the output of operation 26) can be stored by data storage module 54, and the stored (weighted) data is processed by data processor 56 to generate an output (e.g., the output of operation 28 of operation 20). A potential advantage of storing weighted data is that, in some cases, the storage requirement for weighted data may be less than the storage requirement for measurement data (e.g., especially if the weighting is binary weighting).
[0059] As discussed further below, the video data received in operation 22 is typically noisy. The weighting in operation 26 aims to improve the signal-to-noise ratio of the data in order to improve the quality of the motion indication generated in operation 28. Various weighting arrangements are described herein. These weighting arrangements can be used individually or in any combination.
[0060] The weighting in Operation 26 can be binary weighting. Therefore, when generating motion indications (in Operation 26), some subframes can be ignored. With appropriate weighting, data from some subframes that do not contain useful motion data can be ignored, thereby improving the overall signal-to-noise ratio. Data with zero weighting may not need to be stored, potentially reducing the need for data storage and processing.
[0061] The weighting of a particular subframe can depend at least in part on the historical motion data of that subframe. For example, subframes that have moved in the past (e.g., the most recent past) may have a higher weighting. In this way, the weighted sum can be biased towards subframes that have moved in the past (e.g., the most recent past). This tends to increase the signal-to-noise ratio.
[0062] The weighting of a particular subframe can depend at least in part on its position within the scene. For example, a central subframe may have a higher weighting than an outer subframe (assuming that movement is more likely to occur closer to the center of the image received in operation 22). This may make sense because the operator of a camera (such as camera 14) could potentially arrange the system such that the patient or object on bed 12 is close to the center of the image captured by the camera.
[0063] The weighting of a particular subframe can depend at least in part on the visual characteristics of the scene. For example, in the system 10 described above, a subframe that is above the horizontal plane of the bed 12 can be weighted higher than a subframe that is below the surface of the bed (assuming that movement is more likely to occur above the bed than below it).
[0064] The weighting for a specific subframe can change over time. For example, as more motion data for a particular subframe becomes available, the "historical motion data" that may be associated with the weighting of that subframe can be updated during operation.
[0065] The camera 14 of system 10 and / or the imaging device 52 of system 50 may be a multi-mode camera including a color (RGB) camera and an infrared (IR) projector and sensor. The sensor may send a near-infrared (IR) light array into the field of view of imaging device 52, where a detector receives the reflected IR, and an image sensor (e.g., a CMOS image sensor) runs computational algorithms to construct a grid-based video of real-time, three-dimensional depth values. The information acquired and processed in this way can be used, for example, to identify individuals, their movement, posture, and body attributes, and / or can be used to measure size, volume, and / or classify objects. Images from the depth camera can also be used for obstacle detection by locating floors and walls.
[0066] In one example implementation, the imaging device is implemented using a Kinect (RTM) camera provided by Microsoft. In one example, the frame rate of the imaging device 52 is 33 frames per second.
[0067] The IR projector and detector of imaging device 52 can generate a depth matrix, where each element of the matrix corresponds to the distance from an object in the image to the imaging device. The depth matrix can be converted to a grayscale image (e.g., the darker the pixel, the closer it is to the sensor). If the object is too close or too far, pixels may be set to zero, making them appear black.
[0068] Depth frames can be captured by imaging device 52 and stored in data storage module 54 (e.g., in binary file format). In the example implementation, data storage device 54 is used to record depth frames of objects whose sleep duration ranges from 5 to 8 hours. In one example implementation, the data storage requirement for one night's sleep is approximately 200GB.
[0069] The stored data can be processed by data processor 56. Data processing can be online (e.g., during data collection), offline (e.g., after data collection), or a combination of both. As discussed further below, data processor 56 can provide output.
[0070] Figure 6 This is a flowchart illustrating an algorithm according to an example embodiment, generally indicated by reference numeral 60, showing an example implementation of operation 24 of algorithm 20. As described above, operation 24 determines the movement between consecutive frames of the captured video data.
[0071] Algorithm 60 begins at operation 62, where the average distance value for each subframe is determined. Then, at operation 64, the distance between subframes (based on the distance determined in operation 62) is determined. As discussed further below, distances between subframes above a threshold can indicate movement between consecutive frames.
[0072] For example, consider the received video data 30 in the example of operation 22. The average distance between two frames of video data 30 can be determined by comparing the average distance between two consecutive frames. In one example, all pixels in the matrix are reordered in a vector of length 217088 (424*512), with each row of the depth matrix placed consecutively in the vector.
[0073] After removing all zero pixels from multiple frames, the average distance of frame number i is defined as:
[0074]
[0075] in,
[0076] D∈G\Z where
[0077] G = {the set of pixels in the frame}; and
[0078] Z = {the set of zero-value pixels in the frame}.
[0079] The average distance of consecutive frames in the dataset has been calculated (Operation 62), and the determination of movement in the relevant scene can be made (Operation 64).
[0080] One method for determining movement is as follows. First, an average value is defined based on the set of all non-zero value pixels. Next, a change in position is noted if the difference between the average distances of two consecutive frames exceeds a certain threshold θ1. This is given by: M(i) - M(i-1) ≥ θ1.
[0081] Multiple values of the standard deviation and the maximum difference are tested in order to find the optimal threshold (as discussed further below).
[0082] Now consider the received video data 40 in the example of operation 22. The average distance value can be determined by determining the average distance of each subframe.
[0083] The average distance between subframe number j in frame number i is defined as:
[0084]
[0085] in:
[0086] D∈G′\Z′ where
[0087] G′ = {the set of pixels in subframe j}; and
[0088] Z' = {the set of zero-value pixels in subframe j}.
[0089] Now, instead of detecting motion in the entire frame, we can detect motion within each subframe. If the average distance between two consecutive subframes differs from a certain threshold θ2, the algorithm will detect motion in that subframe, resulting in motion detection across the entire frame. This is given by: M(i,j) - M(i-1,j) ≥ θ2.
[0090] Figure 7 This is a flowchart illustrating the algorithm according to an example embodiment, generally indicated by reference numeral 70. Algorithm 70 begins at operation 72, which is performed during the initialization phase. The initialization phase is performed without the presence of a patient (or any other object) (therefore movement should not be detected). As described below, the initialization phase is used to determine the noise level in the data. Next, at operation 74, the data collection phase is performed. Finally, at operation 76, a subframe containing data indicating movement (the data collected in operation 74) is identified (thus implementing operation 28 of the algorithm 20 described above). In operation 76, the subframe is identified based on a threshold distance level set depending on the noise level determined in operation 72.
[0091] Figure 8 Initialization data according to an example embodiment is shown, generally indicated by reference numeral 80. Initialization data 80 includes average distance data 82, a first image 83, a second image 84, and a noise representation 85. The first image 83 corresponds to the average distance data of data point 86, and the second image 84 corresponds to the average distance data of data point 87.
[0092] Average distance data 82 shows how the determined average frame distance changes over time. Figure 8 The 18-minute time period is shown in the figure. Since data 80 was collected without movement (e.g., no patient or object on bed 12 of system 10), the variation in mean distance data 82 represents noise. Noise representation 85 expresses the noise as a Gaussian distribution. This distribution can be used to determine the standard deviation of the noise data (as indicated by the standard deviation shown in data 82). The maximum difference between two consecutive means is also plotted on output 80.
[0093] Operation 72 attempts to evaluate the inherent noise in the frames of output 80. In this way, the inherent variations in the still data can be determined so that motion detection is not obscured by noise. Operation 72 can be implemented using the principles of algorithm 20 described above. Therefore, video data (i.e., output 82) can be received. The distances in data 82 can then be determined (see operation 24 of the algorithm) and can be weighted (operation 26 of algorithm 20).
[0094] In the data collection phase 74 of Algorithm 70, the object enters a sleep state and the relevant dataset is collected. Then, in operation 76, knowledge of noise is used to determine the likelihood that changes in the distance data indicate actual movement (rather than noise). For example, if the received distance data in a subframe deviates from the previous average distance by more than five standard deviations (as determined in operation 72), movement can be considered to have occurred.
[0095] Experimental results
[0096] To verify the reproducibility of the experiment, fourteen nights of testing were conducted, involving eleven different subjects. Data was stored in video format, which was used as the real-world scenario. Real-world motion detection was manually recorded by observing the video and compared to the motion detection provided by the algorithm being tested.
[0097] Figure 9A and Figure 9B Results according to the example embodiments are shown, generally indicated by reference numerals 92 and 94, respectively. The results are based on the average distance evolution of fourteen subjects (i.e., how the average distance determination changes over time). The average percentage of good detections (true positives) across all experiments is shown together with the overestimation of movement (false positives). Results 92 and 94 are based on the above references. Figure 3 The image data is described in the form of frames (i.e., frames are not divided into subframes).
[0098] Results 92 show the number of correct moving markers (true positives) and incorrect moving markers (false positives) for different standard deviation threshold levels. At a moving threshold of two standard deviations (ensuring that at least two standard deviations of average distance change are detected), 45% of true positive measurements and 81% of false positive measurements were detected. As the threshold level increases, the number of false positives decreases (reaching zero at five standard deviations). However, at five standard deviations, the true positive rate drops to 26%. Therefore, Results 92 demonstrates poor performance.
[0099] Result 102 shows the number of correct move markers (true positives) and incorrect move markers (false positives) when using the maximum distance difference as the threshold. Performance is also poor.
[0100] Figure 10A and Figure 10B Results according to the example embodiments are shown, generally indicated by reference numerals 102 and 104, respectively. The results are based on the average distance evolution of fourteen subjects. The average percentage of good detections (true positives) across all experiments is shown together with the overestimation of movement (false positives). Results 102 and 104 are based on the above references. Figure 4 The image data is described in the form of frames (i.e., the frame is divided into subframes and each subframe is considered separately).
[0101] Results 102 show the number of correct motion markers (true positives) and incorrect motion markers (false positives) for different standard deviation threshold levels. At a motion threshold of five standard deviations, 70% of true positive measurements and 3% of false positive measurements were detected. Therefore, the subframe arrangement results are significantly better than the reference above. Figure 9A and Figure 9B The results describe the full-frame layout. When using the maximum distance difference as a threshold, result 104 shows similarly good performance.
[0102] The examples above generally involve using images from a depth camera when generating motion indications for a scene. This is not mandatory for all embodiments. For example, changes in consecutive video image subframes can be used to determine motion in a scene.
[0103] The examples described above generally relate to general sleep monitoring, and more specifically to the detection of data related to sleep disorders. This is not essential for all embodiments. For example, the principles discussed herein can be used to determine movement in a scene for other purposes, such as tracking movement in a video stream (i.e., a "tracking camera" for hunting or conservation) for the purpose of detecting the presence of moving animals, or detecting moving people for security monitoring purposes.
[0104] For the sake of completeness, Figure 11 This is a schematic diagram of components of one or more example embodiments previously described, collectively referred to below as processing system 300. Processing system 300 may have a processor 302, a memory 304 tightly coupled to the processor and including RAM 314 and ROM 312, and optionally, a user input 310 and a display 318. Processing system 300 may include one or more network / device interfaces 308 for connection to a network / device, such as a wired or wireless modem. Interface 308 may also operate as a connection to other devices, such as devices that are not network-side devices. Therefore, direct connections between devices / devices without network involvement are possible.
[0105] The processor 302 is connected to each of the other components in order to control their operation.
[0106] Memory 304 may include non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD). The ROM 312 of memory 314 stores the operating system 315 and may also store software applications 316. The RAM 314 of memory 304 is used by the processor 302 to temporarily store data. The operating system 315 may include code that implements aspects of the aforementioned algorithms 20, 60, and 70 when executed by the processor. Note that in the case of small devices, memory is best suited for small size applications, i.e., hard disk drives (HDDs) or solid-state drives (SSDs) are not always used.
[0107] The processor 302 can take any suitable form. For example, it can be a microcontroller, multiple microcontrollers, a processor, or multiple processors.
[0108] The processing system 300 can be a standalone computer, server, console, or its network. The processing system 300 and the necessary structural components can all be within a device, such as an Internet of Things (IoT) device, i.e., embedded in a very small size.
[0109] In some example embodiments, the processing system 300 may also be associated with external software applications. These may be applications stored on a remote server device / device and may run partially or exclusively on the remote server device / device. These applications may be referred to as cloud-hosted applications. The processing system 300 may communicate with the remote server device / device to utilize the software applications stored there.
[0110] Figure 12A and Figure 12B Tangible media are shown, namely a removable memory unit 365 and a compact disc (CD) 368, which store computer-readable code that, when run by a computer, can execute the methods according to the example embodiments described above. The removable memory unit 365 may be a memory stick, such as a USB memory stick, having internal memory 366 for storing computer-readable code. The computer system can access memory 366 via connector 367. CD 368 may be a CD-ROM or DVD, etc. Other forms of tangible storage media may be used. Tangible media can be any device / apparatus capable of storing data / information, wherein the data / information can be exchanged between devices / apparatus / networks.
[0111] Embodiments of the present invention can be implemented in software, hardware, application logic, or a combination of software, hardware, and application logic. The software, application logic, and / or hardware can reside on memory or any computer medium. In exemplary embodiments, the application logic, software, or instruction set is maintained on any of a variety of conventional computer-readable media. In the context of the present invention, "memory" or "computer-readable medium" can be any non-transitory medium or device that can contain, store, communicate, propagate, or transmit instructions for use in or associated with instruction execution, such as a computer.
[0112] In the relevant context, references to "computer-readable storage medium," "computer program product," "tangible embodiment of a computer program," or "processor" or "processing circuitry" should be understood to include not only those with different architectures, such as single / multiprocessor architectures and sequencer / parallel architectures, but also special-purpose circuits, such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processing devices / apparatus, and other devices / apparatus. References to computer programs, instructions, code, etc., should be understood to express software used for programmable processor firmware, such as programmable content of hardware devices / apparatus as processor instructions or configuration or configuration settings of fixed-function devices / apparatus, for gate arrays, programmable logic devices / apparatus, etc.
[0113] In this application, the term "circuit" means all of the following: (a) a purely hardware circuit implementation (e.g., an implementation in analog and / or digital circuits only) and (b) a combination of circuits and software (and / or firmware), such as (if applicable): (i) a combination of (multiple) processors or (ii) (multiple) processors / software (including (multiple) digital signal processors), software, and (multiple) memories that work together to enable a device (e.g., a server) to perform various functions and (c) circuits, such as (multiple) microprocessors or a portion thereof, which require software or firmware for operation, even if the software or firmware does not exist physically.
[0114] If necessary, the different functions discussed here can be executed in different orders and / or simultaneously with each other. Furthermore, if necessary, one or more of the above functions can be optional or can be combined. Similarly, it will be understood that... Figure 2 , Figure 6 and Figure 7 The flowchart is merely an example and the various operations described therein can be omitted, reordered, and / or combined.
[0115] It should be understood that the above-described exemplary embodiments are purely illustrative and do not limit the scope of the invention. Other variations and modifications will become apparent to those skilled in the art upon reading this specification.
[0116] Furthermore, the disclosure of this application should be understood to include any novel feature or any novel combination of features or any generalization thereof explicitly or implicitly disclosed herein, and new claims may be formulated during the implementation of this application or any application derived therefrom to cover any such feature and / or combination of such features.
Claims
1. A device for determining scene movement, comprising: A component for receiving video data for a fixed scene, wherein the video data includes depth data of a plurality of pixels of an image of the scene; Components for determining motion measurements for at least some of a plurality of subframes of the video data, wherein each of the plurality of subframes is a subdivision of a frame of the video data; Components for weighting the motion measurements to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframes, and the weighting for at least some of the subframes depends at least in part on: The visual characteristics of the scene, and One or more of the following: historical movement data of the subframe, and the position of the subframe within the scene; as well as A component for generating a motion indication for the scenario from some or all of the weighted motion measurements. The motion measurement for the subframe includes an indication of the degree of motion between consecutive instances of the corresponding subframe.
2. The apparatus of claim 1, wherein the weighting is binary weighting.
3. The apparatus of claim 1, wherein one or more of the weights are updated during operation.
4. The apparatus of claim 1, wherein the component for determining the motion measurement value for the subframe of the video data comprises: A component for determining the average distance measurement for a corresponding subframe of the received video data. Among the components, for one or more of the average distance measurements, it is determined whether the average distance measurement differs from the average distance for the corresponding subframe of the preceding video data frame instance, in order to determine whether movement has occurred.
5. The apparatus of claim 1, wherein if the average distance measurement between consecutive subframes is greater than a threshold, the component for determining the motion measurement for the subframe of the video data determines that motion has occurred.
6. The apparatus of claim 1, wherein the motion measurement for the subframe of the video data is determined based on changes in received depth data for the corresponding subframe.
7. The apparatus of claim 1, wherein the video data is received from a multi-mode camera including an infrared projector and sensors.
8. The apparatus of claim 1, wherein the subframe comprises a regular tile grid.
9. The apparatus according to any one of claims 1 to 8, wherein the component comprises: At least one processor; as well as At least one memory, including computer program code, the at least one memory and the computer program code being configured to cause the execution of the device together with the at least one processor.
10. A method for determining scene movement, comprising: Receive video data for a fixed scene, wherein the video data includes depth data of multiple pixels of an image of the scene; Determine motion measurements for at least some of a plurality of subframes of the video data, wherein each of the plurality of subframes is a subdivision of a frame of the video data; The motion measurements are weighted to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframes, and the weighting for at least some of the subframes depends at least in part on: The visual characteristics of the scene, and One or more of the following: historical movement data of the subframe, and the position of the subframe within the scene; as well as A motion indication for the scenario is generated from some or all of the weighted motion measurements. The motion measurement for the subframe includes an indication of the degree of motion between consecutive instances of the corresponding subframe.
11. A computer program product comprising instructions for causing a device to execute at least the following: Receive video data for a fixed scene, wherein the video data includes depth data of multiple pixels of an image of the scene; Determine motion measurements for at least some of a plurality of subframes of the video data, wherein each of the plurality of subframes is a subdivision of a frame of the video data; The motion measurements are weighted to generate a plurality of weighted motion measurements, wherein the weighting depends on the subframes, and the weighting for at least some of the subframes depends at least in part on: The visual characteristics of the scene, and One or more of the following: historical movement data of the subframe, and the position of the subframe within the scene; as well as A motion indication for the scenario is generated from some or all of the weighted motion measurements. The motion measurement for the subframe includes an indication of the degree of motion between consecutive instances of the corresponding subframe.
Citation Information
Patent Citations
Generating a breathing alert
US20180064369A1