Directional facing
By integrating a lighting system to receive sound and using an artificial intelligence model to estimate orientation, privacy and compliance issues are resolved, enabling non-invasive orientation detection suitable for applications such as fall detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SIGNIFY HOLDING BV
- Filing Date
- 2024-12-23
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies have privacy and compliance issues when estimating a person's facing direction, especially in fall detection where wearable devices are difficult to utilize effectively.
By integrating a lighting system, using first and second illuminators to receive the sound emitted by a person, estimating the facing direction based on the spatial directional characteristics of the sound and an artificial intelligence model, and combining a thermopile sensor and a hidden Markov model to track the person's position, non-invasive facing direction detection is achieved.
It enables efficient estimation of human orientation without infringing on privacy, making it suitable for applications such as fall detection and reducing the need for compliance with wearable devices.
Smart Images

Figure CN122497891A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to lighting systems that integrate orientation detection, and more particularly to estimating a person's orientation based on human-generated sounds. Background Technology
[0002] The direction a person is facing is a crucial factor in lighting control and other applications, including smart home systems. Furthermore, orientation information can help understand human-to-human interactions. In retail applications, orientation information can indicate which shelf and / or product a person is facing. Orientation detection is also relevant in surveillance systems, computer games, driver alertness monitoring in automobiles, assistive technologies, elderly care, and suspicious behavior detection. Optical cameras are frequently used to detect the direction people are facing. Cameras can also be used to detect if a person has fallen. However, privacy concerns may limit the use of cameras in certain applications. In some cases, wearable devices are used for fall detection. However, compliance with wearable devices, such as that of older adults, can pose challenges to their use in fall detection. Therefore, a solution may be needed to estimate a person's orientation with minimal or no privacy and compliance concerns, and to use this orientation information for purposes such as fall detection. Summary of the Invention
[0003] This disclosure generally relates to lighting systems for integrated orientation detection, and more particularly, to estimating a person's orientation based on sounds emitted by a person. In an example embodiment, a method for orientation detection using a lighting system includes receiving sounds emitted by a person in a region by a first illuminator and a second illuminator. The method also includes determining spatial directivity features of the sound based on sound levels received by at least the first and second illuminators, wherein the sound levels are at at least two frequencies of the sound. The method further includes estimating a person's orientation based on the spatial directivity features of the sound by executing an artificial intelligence (AI) model trained using spatial directivity feature data labeled with orientation.
[0004] In another example embodiment, the lighting system for orientation detection includes a first illuminator comprising a first processor and a first microphone, wherein the first illuminator is configured to receive sound generated by a person in a region via the first microphone. The lighting system is configured to determine a first spatial directionality feature of the sound based on a first sound level at at least two frequencies of the sound received by the first illuminator. The system also includes a second illuminator comprising a second processor and a second microphone, wherein the second illuminator is configured to receive sound generated by a person via the second microphone. The lighting system is configured to determine a second spatial directionality feature of the sound based on a second sound level at at least two frequencies of the sound received by the second illuminator. The lighting system is configured to execute an AI model to estimate the person's orientation based at least on the first and second spatial directionality features.
[0005] These and other aspects, objects, features, and embodiments will become apparent from the following description and the appended claims. Attached Figure Description
[0006] The accompanying drawings will now be consulted. These drawings are not necessarily to scale, and in them: Figure 1 An orientation-detection lighting system according to an example embodiment is shown; Figure 2 The example embodiment is shown by Figure 1 A graph of the horizontal surface sound level of the frequency components of a person's voice; Figure 3 The example embodiment is shown by Figure 1 A graph showing the vertical plane sound level of the frequency components of a person's voice; Figure 4 A horizontal plane according to an example embodiment is shown. Figure 1 A top view of the area facing the person; Figure 5 The following is illustrated in the vertical plane according to an example embodiment. Figure 1 A side view of the area in which a person is facing; Figure 6 An example embodiment is shown in Figure 1 Fall incidents in the area Figure 1 The direction a person is facing; Figure 7 The example embodiment is shown, corresponding to Figure 1 A block diagram of the lighting system's illuminators; and Figure 8 A method for orientation detection according to an example embodiment is shown.
[0007] The accompanying drawings illustrate only exemplary embodiments and should therefore not be considered as limiting the scope. Elements and features shown in the drawings are not necessarily drawn to scale; rather, the emphasis is on clearly illustrating the principles of the exemplary embodiments. Furthermore, certain dimensions or positions may be exaggerated to aid in visually conveying these principles. In the drawings, the same reference numerals used in different figures may denote similar or corresponding, but not necessarily identical, elements. Detailed Implementation
[0008] In the following paragraphs, exemplary embodiments will be described in more detail with reference to the accompanying drawings. Well-known components, methods, and / or processing techniques are omitted or briefly described in the description. Furthermore, reference to various features(s) of the embodiments does not imply that all embodiments must include all of the mentioned features(s).
[0009] Figure 1 A directional detection lighting system 100 according to an example embodiment is shown. In some example embodiments, system 100 includes luminaires 102, 104, 106, 108, and 110. Luminaires 102-110 may be in area 114, such as in a room. For example, area 114 may be a room in a nursing home, a shop, etc. Luminaires 102-108 may be recessed lights, for example, mounted in the ceiling 116 of area 114. Luminaire 110 may be, for example, a table lamp located on a piece of furniture 118 or another type of lighting. For example, luminaire 110 may be a floor light on the floor 124 of area 114 instead of on furniture 118.
[0010] In some example embodiments, each of the illuminators 102-110 may include a processor, a microphone, and an occupancy sensor. The optical module of each of the illuminators 102-110 may emit light provided by the specific illuminator. Each of the illuminators 102-110 may receive sound via its respective microphone, and the processor of each of the illuminators 102-110 may process the sound received via the corresponding microphone.
[0011] In some example embodiments, system 100 can detect one or more people, such as person 120 in area 114. For example, one or more of illuminators 102-110 may include occupancy sensors. For illustration, illuminators 102-110 may each include a passive infrared (PIR) sensor and / or a thermopile (e.g., a single-pixel thermopile) that can be used to detect person 120 in area 114, as will be readily understood by those skilled in the art who benefit from this disclosure. For example, a processor of illuminator 102 may detect person 120 based on detection signals from the occupancy sensors of illuminator 102. Illuminators 102-110 may also estimate the position of person 120. For example, illuminators 102-110 may use a corresponding thermopile to estimate the position of person in area 114, wherein the sensed temperature of person 120 depends on the position of person 120 relative to illuminators 102-110. For example, since the locations of illuminators 102-110 in area 114 can be known, different sensed temperature values of person 120 sensed by the corresponding thermopile of illuminators 102-110 can be used to estimate the location of person 120 in area 114. To illustrate, one of the illuminators 102-110 or server 112 of system 100 can receive information from other illuminators and estimate the location of person 120. Alternatively or additionally, server 112 of system 100 can receive information from illuminators 102-110 and estimate the location of person 120. After estimating the location of person 120, a Hidden Markov Model (HMM) can be applied to track person 120 as they move around in area 114. For example, the processor of one of the illuminators 102-110 or server 112 can execute an HMM module to track person 120 in area 114. Server 112 can operate as an edge device interacting with a cloud server.
[0012] In some example embodiments, system 100 can determine the facing direction of person 120. The facing direction of person 120, as used herein, refers to the face of person 120. For example, system 100 can determine the facing direction of person 120 based on one or more sounds emitted by person 120. For illustration, illuminators 102-110 can determine the spatial directional characteristics of the sound 130 emitted by person 120, and the facing direction of person 120 can be determined by one or more of illuminators 102-110 and / or by server 112 based on the spatial directional characteristics determined by two or more of illuminators 102-110. The facing direction of person 120 can consider both horizontal and vertical directions. For example, person 120 can have facing directions such as face up, face forward, face down, right face up, etc., where, for example, front, right, left, and back are designated as horizontal regions forming a 360-degree radius around person 120, and where up, front, and down are designated as vertical regions relative to the face of person 120.
[0013] Typically, the facing direction of person 120 can be user-defined. For example, the facing direction of person 120 can be defined relative to the structure of a general space (e.g., a room) that may represent other spaces of interest. For illustration, the facing direction of person 120 could be the front wall 126 of area 114. As another example, when the front of person 120 faces the front wall 126 and the face of person 120 faces the ceiling 116, the facing direction of person 120 could be both the front wall 126 and the ceiling 116. As yet another example, when the front of person 120 faces the front wall 126 and the face of person 120 faces the floor 124, the facing direction of person 120 could be both the front wall 126 and the floor 124 of area 114. As yet another example, person 120 could face the rear wall 128 and the floor 124, such that the facing direction of person 120 is both the rear wall 128 and the floor 124. In some cases, the facing direction of person 120 can be defined relative to an object such as shelf 132, where, for example, the possible facing directions of person 120 could be the top shelf, middle shelf, and bottom shelf of shelf 132.
[0014] In some example embodiments, the facing direction of person 120 may be defined more generally using azimuth and vertical angles (i.e., elevation and depression) or a range of azimuth and vertical angles. Typically, the complete azimuth range is from 0 to 360 degrees; for example, 0 degrees corresponds to north, 90 degrees to east, 180 degrees to south, and 270 degrees to west. The vertical angle ranges from 90 degrees to -90 degrees relative to the horizontal plane.
[0015] In some example embodiments, illuminators 102-110 can determine the spatial directionality characteristics of sound, such as sound 130 emitted by person 120. The spatial directionality characteristics of sound 130 determined by a particular one of illuminators 102-110 can be defined by, for example, the relationship between the sound levels (i.e., audio signal strength or signal power) of sound 130 at two or more frequency components of sound 130 received by the particular illuminator. For example, the spatial directionality characteristics of sound 130 determined by illuminator 102 can be defined by the relationship between the sound levels of sound 130 at two frequency components of sound 130 received by illuminator 102. The relationship between sound levels can be expressed as, for example, the difference between two sound levels, the ratio of a smaller sound level to a higher sound level, or another expression readily understood by those skilled in the art who benefit from this disclosure. As another example, the spatial directionality characteristics of the sound 130 determined by the illuminator 102 can be defined by the relationship between the sound levels of the sound 130 at the first and second frequency components of the sound 130 received by the illuminator 102, and also by, for example, the relationship between the sound levels of the first, second, and third frequency components of the sound 130 when the illuminator 102 receives the sound 130.
[0016] Figure 2 The example embodiment is shown by Figure 1 Graph 200 shows the horizontal sound level curve of the frequency components of a sound emitted by person 120 (such as sound 130). (Reference) Figure 1 and Figure 2 In some example embodiments, different frequency components of sound 130 have different horizontal spatial directivity. For example, the 1 kHz and 8 kHz frequency components of sound 130 have different horizontal spatial directivity, depending on the horizontal orientation of person 120. To illustrate, when person 120 faces... Figure 2 When the horizontal plane is marked as "forward," the sound level of the 8 kHz frequency component of sound 130 located directly in front of person 120 is significantly higher than the sound level located behind person 120. In other words, a microphone located in front of person 120 can easily receive the 8 kHz frequency component of sound 130, while the 8 kHz frequency component may be too weak to be received by a microphone located behind person 120, for example, in... Figure 2 The middle mark indicates the horizontal direction behind. In contrast to the 8 kHz frequency component, the microphones in front of and behind the person 120 can easily receive the 1 kHz frequency component of sound 130, although they may be at different sound levels.
[0017] In some example embodiments, the similarity and difference in the horizontal spatial directionality of the 1 kHz and 8 kHz frequency components of sound 130 relative to microphones at different locations in region 114 can define a spatial directionality feature of sound 130, which can be used to estimate or otherwise determine the horizontal orientation of person 120. That is, the spatial directionality feature of sound 130 defined based on the similarity and difference in the horizontal spatial directionality of the 1 kHz and 8 kHz frequency components of sound 130 can be used to determine the horizontal orientation of person 120. As another example, the spatial directionality feature of sound 130 defined based on the similarity and difference in the horizontal spatial directionality of the 1 kHz and 4 kHz frequency components of sound 130 can be used to determine the horizontal orientation of person 120. As yet another example, the spatial directionality feature of sound 130 defined based on the similarity and difference in the horizontal spatial directionality of the 1 kHz and 4 kHz frequency components of sound 130, and based on the similarity and difference in the horizontal spatial directionality of the 1 kHz and 8 kHz frequency components of sound 130, can be used to estimate or otherwise determine the horizontal orientation of person 120. Generally speaking, the similarity and difference in the horizontal spatial directionality of different frequency components of sound 130 can be used to estimate or otherwise determine the horizontal orientation of person 120.
[0018] Figure 3 The example embodiment is shown by Figure 1 Graph 300 shows the vertical plane sound level of the frequency components of a human-generated sound (such as sound 130). (Reference) Figure 1 and Figure 3 In some example embodiments, different frequency components of sound 130 have different vertical spatial directivity. For example, the 1 kHz and 8 kHz frequency components of sound 130 have different vertical spatial directivity, depending on the vertical orientation of person 120. To illustrate, when person 120 is... Figure 3 When the vertical plane is oriented at zero degrees, the sound level of the 8 kHz frequency component of sound 130 at a position in front of person 120 (e.g., at 0 degrees in the vertical plane) is significantly higher than at a position behind person 120 (e.g., at 180 degrees in the vertical plane). That is, a microphone located in front of person 120 can easily receive the 8 kHz frequency component of sound 130, while the 8 kHz frequency component may be too weak to be adequately received by a microphone located behind person 120 (e.g., at 180 degrees in the vertical plane). Conversely, microphones in front of and behind person 120 can easily receive the 1 kHz frequency component of sound 130, although they may be at different sound levels.
[0019] In some example embodiments, the similarity and difference in the vertical spatial directionality of the 1 kHz and 8 kHz frequency components of sound 130 relative to the microphone at different locations in region 114 can define a spatial directionality feature of sound 130, which can be used to estimate or otherwise determine the facing direction of person 120. That is, the spatial directionality feature of sound 130 defined based on the similarity and difference in the vertical spatial directionality of the 1 kHz and 8 kHz frequency components of sound 130 can be used to determine the vertical facing direction of person 120. As another example, the spatial directionality feature of sound 130 defined based on the similarity and difference in the vertical spatial directionality of the 1 kHz and 4 kHz frequency components of sound 130 can be used to determine the vertical facing direction of person 120. As another example, the spatial directionality features of sound 130, defined based on the similarity and difference in the vertical spatial directionality of its 1 kHz and 4 kHz frequency components, and based on the similarity and difference in the vertical spatial directionality of its 1 kHz and 8 kHz frequency components, can be used to estimate or otherwise determine the facing direction of person 120. In general, the similarity and difference in the vertical spatial directionality of different frequency components of sound 130 can be used to estimate or otherwise determine the vertical facing direction of person 120.
[0020] In some example embodiments, a frequency domain analysis of sound 130 can be performed to determine the spatial directional characteristics of sound 130. For illustration, the time domain representation of sound 130 is shown in equation (1).
[0021] Equation (1) In equation (1), k refers to a specific microphone, and y k (t) is the intensity (e.g., amplitude) of the audio signal 130 received by microphone k, such as microphone k is Figure 1 The microphone is one of the illuminators 102-110 shown. k It is channel state information, and includes audio fading due to distance. k θ(t) and φ(t) are the angles of arrival of sound 130 at microphone k, and can correspond to the azimuth (horizontal) and elevation or depression (vertical) angles of the direction in which person 120 is facing, respectively.
[0022] Equation (2) provides the frequency domain representation of sound 130, which is the Fast Fourier Transform (FFT) of the time domain representation of sound 130 given in Equation (1), where f m This represents the various frequencies of a sound at 130 Hz.
[0023] Equation (2) For example, f0 (i.e., m=0) can be a base or reference frequency of 100 Hz. Figure 1 Each of the illuminators 102-110 shown may include one or more microphones and a processor, and may determine the audio signal intensity at frequencies of interest (such as 100 Hz, 1 kHz, 2 kHz, 4 kHz, and 8 kHz), each frequency corresponding to a corresponding value of m in equation (2). As mentioned above, each illuminator may also determine one or more spatial directional characteristics of sound 130 based on the audio signal intensities of two or more of the frequencies of interest.
[0024] In some example embodiments, to minimize the impact of sound level variations emitted by different people or even the same person in region 114, the audio signal intensity of the sounds received by two or more of the illuminators 102-110 can be "normalized" based on the corresponding audio signal intensity of the sounds at a reference frequency f0. For illustration, the audio signal intensity of the sounds 130 received by two or more of the illuminators 102-110 at frequencies of interest (such as 1 kHz, 2 kHz, 4 kHz, and / or 8 kHz) can be "normalized" using the audio signal intensity of the sounds 130 at a reference frequency f0 (such as, for example, 100 Hz). Alternatively, the audio signal intensity of another frequency of the sounds 130 can be used instead of 100 Hz. As used herein, “normalization” refers to the audio signal strength of sound 130 (or other sound) at a specific frequency (such as, for example, 1 kHz, 2 kHz, 4 kHz and 8 kHz) as expressed by equation (2) divided by the audio signal strength of sound 130 (or other corresponding sound) at a reference frequency f0 (e.g., 100 Hz), as shown in equation (3).
[0025] Equation (3) In equation (3), Indicates that in the case marked f m The "normalized" audio signal intensity at a frequency of 130, where different m values represent different frequencies. For example, f m The frequencies can be 1 kHz, 2 kHz, 4 kHz, 8 kHz, where each frequency corresponds to a different value of m in equation (3). Each illuminator of the illuminators 102-110 that receive the sound 130 can determine one or more spatial directional characteristics of the sound 130 based on the “normalized” audio signal strength of the sound 130, for example at 1 kHz and 4 kHz, 1 kHz and 8 kHz and / or other frequency combinations as described above.
[0026] In some example embodiments, one or more of the illuminators 102-110 and / or server 112 may execute an artificial intelligence (AI) model to estimate the facing direction of person 120 based on sound 130. Specifically, one or more of the illuminators 102-110 and / or server 112 may execute an AI model to estimate the facing direction of person 120 based on two or more spatial directional features from the illuminators 102-110 receiving sound 130. For illustration, audio signal intensity data can be used to train the AI model. For example, training data may include the audio signal intensity of sound received by a microphone (i.e., the microphone of the illuminator) in an area (e.g., a room), where the audio signal intensity is relative to multiple frequencies. The number of frequencies may be M, where M is 2 or greater. For example, the frequencies may be 1 kHz and 4 kHz; 1 kHz and 8 kHz; 1 kHz, 4 kHz and 8 kHz; or 1 kHz, 2 kHz, 4 kHz and 8 kHz.
[0027] To illustrate, the audio signal intensity at two or more frequencies of a sound emitted by a person, and the audio signal intensity received by the corresponding microphones of multiple illuminators (e.g., 2, 4, or 5 illuminators) in a region, can be labeled with the facing direction of a particular person. Alternatively or additionally, facing direction labels can be applied to spatial directional features derived from the audio signal intensity of the sound received by the corresponding microphones of multiple illuminators (e.g., 2, 3, 4, or 5 illuminators). For example, the spatial directional features derived from each of the multiple illuminators in a region can be labeled with the facing direction of the person emitting the sound. As another example, the spatial directional features derived from multiple illuminators in a region can be labeled as a set with the facing direction of the person emitting the sound. The facing direction of the person used for labeling purposes can be obtained from a camera, a multi-pixel thermopile sensor, personal observation, etc. Multiple such labeled signal intensity data and / or labeled spatial directional features can be included in training data obtained based on sounds emitted by one or more people at different locations and with different facing directions in a region, and can be used to train an AI model.
[0028] Typically, the labeled dataset of audio signal strength in the training data can correspond to the data shown in equation (2). , where M ranges from, for example, 1 to M, and where M is 2 or greater, and where k represents the number of microphones in the illuminator of the region. For example, by and The audio signal intensity of the sound, where f1 is, for example, 1 kHz and f2 is 8 kHz, can be labeled using the facing direction of the person emitting the sound. That is, the individual audio signal intensity at two or more frequencies, or a group of audio signal intensities at two or more frequencies, can be considered a spatial directional feature and can be labeled using facing direction. Alternatively or additionally, at each illuminator including a microphone for receiving sound used to generate training data, [the following is included]. and Or from and Derived spatial orientation features (e.g., The sound can be labeled using the facing direction of the person emitting the sound. Audio signal intensity data or spatial directionality features of such labels from multiple illuminators can be used to train an AI model. For example, spatial directionality features derived from multiple illuminators based on sound can be labeled as units. Typically, training data can include the audio signal intensity and / or spatial directionality features of multiple such labels for multiple sounds emitted by one or more people and captured by microphones from multiple illuminators.
[0029] In some alternative embodiments, the training data may include or may be obtained from the “normalized” audio signal intensity (and / or spatial directionality features derived therefrom) of the sound at multiple frequencies of each sound, wherein the “normalized” audio signal intensity of each sound at multiple frequencies is labeled with the facing direction of the person emitting the sound. For illustration, the training data may be based on the equation (3) The normalized audio signal strength, where m ranges from, for example, 1 to M, where M is 2 or greater.
[0030] In some alternative embodiments, the angles of arrival θ(t) and φ(t) of the sound 130 received by a particular illuminator can be considered as spatial directivity features of the sound 130. For example, the angles of arrival θ(t) and φ(t) can be determined or otherwise approximated by assuming initial values for the angles of arrival θ(t) and φ(t) in equation (1) and iteratively updating these values by comparing the audio signal intensity calculated at multiple frequencies of the sound 130 with the actual audio signal intensity determined from the sound 130 received by the illuminator (e.g., illuminator 102) receiving the sound 130. Spatial directivity features determined in this way can be used to train AI models and, in a manner described herein, relative to models based on, for example, Figure 2 and Figure 3 Spatial directional characteristics are determined by the directional differences of some frequency components (e.g., 1 kHz and 8 kHz) of the sound 130 shown to estimate the facing direction of the person 120.
[0031] Figure 4A horizontal plane according to an example embodiment is shown. Figure 1 A top view of the facing area of a person at angle 120. (Reference) Figure 1-4 It can indicate a person's horizontal orientation based on a specified area, such as, for example, Figure 4 The front, right, back, and left areas are shown. That is, person 120 can face the front, right, back, or left area, regardless of their position. For example, the front area can be between boundary lines 402 and 408, which can form, for example, a right angle. The right area can be between boundary lines 402 and 404, which can form, for example, a right angle. The back area can be between boundary lines 404 and 406, which can form, for example, a right angle. The left area can be between boundary lines 406 and 408, which can form, for example, a right angle. The horizontal facing areas (front, right, back, and left) can be arranged such that person 120, directly facing north, has a line of sight extending through the center of the front area (as shown by the dashed arrow).
[0032] Figure 5 The following is illustrated in the vertical plane according to an example embodiment. Figure 1 A side view of the facing direction area of person 120. (Reference) Figure 1-5 It can indicate a person's vertical orientation based on a specified area, such as, for example, Figure 5 The diagram illustrates the upward, forward, and downward directions. That is, regardless of the position of person 120, person 120 can face an upward, forward, or downward direction. For example, the upward direction can be between boundary line 502 and vertical line V, forming a 60-degree angle, with vertical line V extending through person 120. The forward direction can be between boundary lines 502 and 504, forming a 60-degree angle. The downward direction can be between boundary line 502 and vertical line V, forming a 60-degree angle. The vertical facing direction can extend 360 degrees around person 120.
[0033] In some example embodiments, the horizontal facing direction regions (front, right, back, and left) and the vertical facing direction regions (up, front, and down) can specify the facing direction of person 120. For example, the facing direction of person 120 can be front-facing, front-up, front-down, right-facing-front, right-facing-up, right-facing-down, back-facing-forward, back-facing-up, etc. Such facing direction labels can be used in training data for training an AI model, which is executed to estimate the facing direction of person 120. For illustration, one or more of illuminators 102-110 and / or server 112 can execute a trained AI model to estimate the facing direction of person 120. Figure 1 The person 120 shown has a left-facing, forward-facing orientation.
[0034] In some alternative embodiments, without departing from the scope of this disclosure, the methods may differ. Figure 4 and Figure 5 The diagram illustrates how facing direction can be defined. For example, a combination of azimuth and elevation or depression angles can be used as labels for training data to indicate the facing direction of the person emitting the sound. In this way, the execution of a trained AI model can provide facing direction in terms of both azimuth and elevation / depression angles.
[0035] In some example embodiments, training data for the AI model can be obtained based on sounds emitted by people in a space (e.g., a room in an elderly residence) with the same lighting layout as one or more rooms where the AI model will be used. For example, Figure 1 Region 114 shown can be used to obtain training data, and illuminators 102-110 can then be used with the trained AI model to estimate the facing direction of a person, such as person 120. In this case, facing direction labels, such as different wall labels (e.g., front wall, right wall, etc.), can be used as training data labels. As another example, different object labels, such as shelf #1, etc., can be used as facing direction labels and thus as the facing direction estimated by performing the AI model. In some alternative embodiments, the facing direction can be labeled as regions I and J, where I refers to the horizontal plane and J refers to the vertical plane. In some alternative embodiments, the AI model can be trained such that the AI model is penalized for relying on latent space representations when estimating facing direction based on sound.
[0036] In some alternative embodiments, self-learning can be used to train the AI model. For example, a self-learning method can also be applied where acoustic directionality features based on at least two different audio frequencies are recorded at two or more illuminators. A self-supervised method then clusters audio events based on their acoustic directionality features. An autoregressive predictive coding (APC) algorithm uses signals from k-1 microphones (i.e., k-1 illuminators) to predict the signal from the kth illuminator. Since the speaker's facing direction plays a crucial role in determining the signal strength distribution among the k microphones in the room, the APC encoder learns a latent representation of the spatial distribution of the temporal audio signal, primarily based on the facing direction and the audio arrival angle. Subsequently, the AI model is tuned to the facing direction, such as the facing direction region, using a small amount of ground-based real-world data.
[0037] Return to reference Figure 1In some example embodiments, two or more of the illuminators 102-110 may receive sound 130, and the spatial directionality characteristics of sound 130 may be determined, for example, based on the 1 kHz and 8 kHz audio signal intensities of sound 130 received at the particular two or more illuminators. Each of the two or more illuminators 102-110 may be referenced above, for example. Figure 1-3 The described method determines the corresponding spatial directionality features of sound 130. Server 112 can receive spatial directionality features from two or more specific illuminators and execute a trained AI model to estimate the facing direction of person 120 based on the spatial directionality features. For example, server 112 can estimate the facing direction of person 120, for example, according to a reference... Figure 4 and 5 The described facing direction indicates left-facing. Alternatively, if the AI model is trained using other facing direction specifications such as azimuth and elevation / depression, server 112 can estimate the facing direction of person 120 accordingly. In some alternative embodiments, one or more of illuminators 102-110 may receive spatial directionality features from one or more other illuminators in system 100 and estimate the facing direction of person 120 in the manner described with respect to server 112.
[0038] Figure 6 An example embodiment is shown in Figure 1 The facing direction of person 120 during the fall event in area 114. (Reference) Figure 1 and 6 In some example embodiments, person 120 may emit sounds 602, 604, 606, 608, 610 immediately before and during the fall event. Figure 6 In the diagram, the arrows extending through sounds 602-610 are intended to indicate the facing direction of person 120 at different positions P1, P2, P3, P4, and P5. At position P1, Figure 1 System 100 can detect person 120, estimate person 120's position, and begin tracking person 120. At position P1, person 120 may emit sound 602 immediately before starting to slip on floor 124. System 100 can, for example, detect person 120 in area 114, estimate person 120's position, and track person 120. System 100 can also use the AI model described above to estimate person 120's facing direction based on the spatial directionality features of sound 602 determined in the manner described above by two or more of the illuminators 102-110.
[0039] In some example embodiments, when person 120 begins to slip at position P2, person 120 may emit sound 604. The facing direction of person 120 at position P2 may also be different from the facing direction of person 120 at position P1. When sound 604 is received by two or more of illuminators 102-110, system 100 can use the AI model described above to determine the facing direction of person 120 at position P2 based on sound 604. As person 120 falls further to position P3, person 120 may still have a different facing direction and may emit sound 606. When sound 606 is received by two or more of illuminators 102-110, system 100 can determine the facing direction of person 120 at position P3 based on sound 606. At position P4, person 120 may be mostly on floor 124 and may emit sound 608. When sound 608 is received by two or more of the illuminators 102-110, system 100 can determine the facing direction of person 120 at position P4 based on sound 608. At position P5, person 120 may attempt to sit up or stand up and may emit sound 610. When sound 610 is received by two or more of the illuminators 102-110, system 100 can determine the facing direction of person 120 at position P5 based on sound 610.
[0040] In some example embodiments, system 100 (e.g., server 112) can process the orientation determined based on sounds 602-610 to perform fall detection by determining whether the sequence of orientations matches a fall event or accident. That is, system 100 can determine the sequence of orientations based on a time series of spatial directional features determined from sounds 602-610, and can detect whether person 120 has experienced a fall event / accident based on the sequence of orientations. Typically, system 100 may have or be able to access a database of multiple sets of orientation sequences corresponding to a person's fall events, and system 100 can compare the orientation sequence determined based on sounds emitted by person 120 with multiple fall orientation sequences to perform fall detection. By detecting and tracking person 120, system 100 can ensure that sounds 602-610 were emitted by the same person.
[0041] In some alternative embodiments, sound 602-610 may be a segment of a single sound made by person 120. For clarity, Figure 6 The order of facing directions shown is an example, and person 120 may have different facing directions during other fall events / accidents.
[0042] In some example embodiments, system 100 (e.g., server 112 of system 100) may process sounds 602-610 captured by microphones of one or more of the illuminators 102-110 to determine whether one or more of the sounds 602-610 indicate a fall event / accident or are otherwise associated with a fall event / accident. For example, server 112 may compare each of the sounds 602-610 with a sound database (e.g., sounds of panic, surprise, etc.) created by people immediately before, during, and immediately after a fall event / accident. By comparing the sounds 602-610 with a database of fall-related sounds, server 112 may confirm whether system 100 has detected a fall event / accident based on the facing direction of person 120. The database of fall-related sounds may be stored on server 112 or may be accessed by server 112 in other ways, such as from a cloud server.
[0043] In some example embodiments, system 100 (e.g., server 112 of system 100) can use information from one or more radar devices of illuminators 102-110 to negate, for example... Figure 6 The result of fall detection (e.g., slipping or tripping to the ground) of person 120 performed in the facing direction of person 120. For example, because people typically experience relatively rapid movement during a fall event / accident, if no change in the speed of person 120 is detected based on radar signals from one or more radar devices of illuminators 102-110, server 112 can base its actions on... Figure 6 The orientation of person 120 is used to negate / invalidate, for example, the result of a fall detection performed by server 112. To illustrate, if no change in the speed of person 120 is detected, or if the speed change is inconsistent with the fall event / accident during a time frame overlapping with the detection of the fall event / accident (e.g., too slow), server 112 can negate / invalidate the fall event / accident detection determined based on the orientation of person 120. If the speed of person 120 is consistent with a fall event or accident, processor 702 can confirm the result of the performed fall detection determined based on the orientation.
[0044] although Figure 6The sequence / time series of orientation during the fall event / accident is shown, but system 100 can use the sequence / time series of orientation of person 120 for other purposes, such as detecting the activity of person 120. For example, system 100 can be used to estimate the degree of interest of person 120 in one or more products on a shelf. As another example, system 100 can be used to detect the alertness of person 120 based on changes in the orientation of person 120. As yet another example, system 100 can be used to prevent shoplifting by detecting restlessness in a store.
[0045] Generally, system 100 can distinguish between a person 120 who has experienced a fall and a person 120 engaged in other activities. For example, system 100 can distinguish between a person 120 corresponding to walking and talking, followed by lying down and talking, and a sequence / time series of facing directions corresponding to walking and talking, followed by a fall and shouting for help. As another example, system 100 can distinguish between a sequence / time series of facing directions corresponding to a person 120 singing while changing body posture and a sequence / time series of facing directions corresponding to a person 120 walking and talking followed by a fall and shouting for help. As described above, system 100 can compare the sequence / time series of facing directions of person 120 with a database of facing direction sequences corresponding to fall events and other activities.
[0046] Figure 7 The example embodiment is shown in the corresponding Figure 1 A block diagram of illuminators 700 in system 100, comprising illuminators 102-110. Typically, illuminator 700 may represent each of illuminators 102-110 and may include illuminators in other systems. Figure 1 Other lighting fixtures in system 100. (See reference) Figure 1-7In some example embodiments, illuminator 700 may include processor 702 (e.g., microprocessor), optical module 704, one or more microphones 706 (e.g., microphone array), radar device 708, memory device 710 (e.g., non-volatile memory device such as flash memory), communication interface unit 712, and occupancy sensor 716. Processor 702 may execute software code stored in memory device 710 to perform the operations described herein concerning illuminators 102-110. For example, an AI model may be stored in memory device 710. Communication interface unit 712 may send and receive signals conforming to one or more wireless communication standards. For example, communication interface unit 712 may send and receive Bluetooth Low Energy (BLE) signals, Wi-Fi signals, ZigBee signals, and / or other wireless signals. Processor 702 may use communication interface unit 712 to send information to, for example, other illuminators and / or servers such as server 112. Processor 702 may also receive information, such as information from other illuminators, through communication interface unit 712. In some example embodiments, the optical module 704 can provide light provided by the illuminator 700. For example, the illuminator 700 can be an embedded illuminator, a hanging illuminator, a floor illuminator, a table illuminator, etc. The processor 702 can control the operation of the optical module 704, which may include, for example, one or more light sources, such as light-emitting diodes (LEDs).
[0047] In some example embodiments, one or more microphones 706 may capture sounds emitted by a person, such as person 120. Processor 702 may process audio data derived from the sounds captured by microphone 706. For example, processor 702 may filter out noise from the captured sounds (e.g., sound 130) and may determine the audio signal strength (i.e., spectral power distribution) of the sounds at multiple frequencies (e.g., 1 kHz, 2 kHz, 4 kHz, and / or 8 kHz). Processor 702 may determine the spatial directionality characteristics of the sounds captured by microphone 706 based on the audio signal strength at at least two or more of the sound frequencies. Processor 702 may transmit the audio signal strength and / or spatial directionality characteristics to, for example, another illuminator of system 100, such as server 112, via communication interface unit 712. Alternatively or additionally, processor 702 may execute one or more AI models 714 to perform operations such as estimating the facing direction of a person based on two or more spatial directionality characteristics determined by the illuminators of system 100.
[0048] In some example embodiments, the processor 702 may be based on, for example, reference Figure 6The description of a person's orientation sequence is used to perform fall detection. For example, processor 702 can estimate the orientation sequence of person 120 based on the time series of spatial orientation features, and then determine whether person 120 has experienced a fall event based on the orientation sequence.
[0049] In some example embodiments, radar device 708 can be used to confirm a fall (e.g., slip or trip to the ground) of a person (such as person 120) detected based on an orientation determined from sounds emitted by the person. For example, if a change in the person's speed determined based on radar signals from radar device 708 is consistent with or inconsistent with a fall event / accident, processor 702 can confirm the result of fall detection performed based on an orientation determined using spatial directional features from multiple illuminators. If the change in the person's speed is inconsistent with a fall event / accident, processor 702 can reject the result of fall detection performed based on an orientation determined using spatial directional features from multiple illuminators.
[0050] In some example embodiments, processor 702 can process the sound captured by microphone 706 to determine the type of sound emitted by a person. For example, a sound of fear or surprise may indicate a fall. Processor 702 can compare the sound received via microphone 706 with a database of sounds made by a person immediately before, during, and after a fall / accident. Thus, based on the type of sound used to determine the person's orientation, processor 702 can confirm the results of fall detection performed based on an orientation sequence. As mentioned above, the orientation sequence can be determined using spatial directionality features, for example, based on sound captured by microphones from multiple illuminators.
[0051] In some example embodiments, the occupancy sensor 716 may include a PIR sensor, a thermopile sensor, and / or another type of sensor that can be used to detect a person and even estimate the person's location. For example, a multi-pixel thermopile may be used to estimate the location of a person, for example, in region 114. The occupancy sensor 716 may operate in conjunction with the processor 702 to detect and estimate the location of a person. The processor 702 may also use information from the occupancy sensor 716 to track people. For example, the processor 702 may execute an HMM software module to track person 120 in region 114 based on information from the occupancy sensor 716.
[0052] In some alternative embodiments, server 112 may perform some of the operations described herein with respect to processor 702. For example, server 112 may receive audio data derived from sound captured by microphone 706 and may determine the audio signal strength (i.e., spectral power) of two or more frequencies of sound. As another example, server 112 may determine the audio signal strength (i.e., spectral power) of sounds at two or more frequencies. Figure 1 The spatial directionality characteristics of the sound captured by each of the illuminators 102-110 in system 100. As another example, server 112 can be based on the spatial directionality characteristics of the sound captured by each of the illuminators 102-110 in system 100. Figure 1 Based on the spatial orientation features of one or more of the illuminators 102-110 in system 100 or based on the spatial orientation features determined by server 112 based on frequency data, one or more AI models are executed to estimate the facing direction of a person (e.g., person 120).
[0053] In some example embodiments, the luminaire 700 may include a driver that can, for example, receive AC power from a municipal power line and provide DC power to the components of the luminaire 700, as will be readily understood by those skilled in the art who benefit from this disclosure. In some alternative embodiments, the luminaire 700 may include a driver that can receive AC power from a municipal power line and provide DC power to the components of the luminaire 700. Figure 7 The illuminator 700 may have more or fewer components without departing from the scope of this disclosure. In some alternative embodiments, some components of the illuminator 700 may be integrated into a single component without departing from the scope of this disclosure. In some alternative embodiments, radio frequency (RF) signals may be used instead of radar signals or as a supplement to radar signals without departing from the scope of this disclosure.
[0054] Figure 8 A method 800 for orientation detection according to an example embodiment is shown. (Reference) Figure 1-8 In some example embodiments, method 800 includes, in step 802, detecting a person 120 in area 114 and estimating the position of the person 120. For example, one or more of illuminators 102-110 can detect and estimate the position of the person 120. Alternatively, server 112 can detect and estimate the position of the person 120 based on occupancy sensor information from one or more of illuminators 102-110. In step 804, method 800 may include tracking the person 120 in area 114. For example, server 112 may execute an HMM software module to track the person 120 in area 114 based on occupancy information from illuminators 102-110.
[0055] In some example embodiments, at step 806, method 800 may include receiving sound 130 (and / or one or more of sounds 602-610) generated by person 120 in area 114. For example, two or more of illuminators 102-110 may receive sound 130 or one or more of sounds 602-610. In some example embodiments, at step 808, method 800 may include filtering the sound, for example, removing noise from noise. Filtering may also be performed by implementing bandpass filters and / or lowpass filters to remove some frequency components of the sound. Each illuminator in illuminators 102-110 that receives sound may also filter the sound received by one or more microphones of that particular illuminator. In some alternative embodiments, each illuminator that receives sound may send audio data generated from the sound to server 112, and server 112 may process (e.g., filter) the audio data representing the sound received by that particular illuminator.
[0056] In some example embodiments, in step 810, method 800 may include determining spatial directional characteristics of one or more of sound 130 and / or sound 602-610. For example, the spatial directional characteristics of sound 130 determined by illuminator 102 may be defined by the relationship between the sound levels (i.e., spectral power or audio signal intensity) of sound 130 at two frequency components (e.g., 1 kHz and 8 kHz) of sound 130 received by illuminator 102. For example, the spatial directional characteristics of sound 130 received by illuminator 102 may be the audio signal level (i.e., audio signal intensity) at a first frequency (e.g., 1 kHz) and the audio signal level at a second frequency (e.g., 4 kHz or 8 kHz). As another example, the spatial directional characteristics of sound 130 received by illuminator 102 may be the difference between the sound level at the first frequency (e.g., 1 kHz) and the audio signal intensity at the second frequency (e.g., 4 kHz or 8 kHz). As described above, the sound levels at different frequencies of sound 130 can be "normalized" by the sound levels at a reference frequency (e.g., 100 Hz) of sound 130, and the "normalized" sound levels can be used to determine the spatial directivity characteristics of sound 130. Because the spectral power distribution (i.e., the signal power distribution at different frequencies) of sound 130 received by a specific illuminator of system 100 can depend on the position of person 120 relative to the specific illuminator, in some example embodiments, the position of person 120 can be used to determine the spatial directivity characteristics of the specific illuminator. In some alternative embodiments, server 112 can determine the spatial directivity characteristics of sound 130 with respect to each of the illuminators 102-110 receiving sound 130, and provide server 112 with audio data of sound 130 or spectral power distribution information of sound 130.
[0057] In some example embodiments, in step 812, method 800 may include estimating the facing direction of person 120 based on the spatial directionality characteristics of sound 130. For example, if Figure 1 If the illuminator 102 is positioned in front of the person 120, then the sound 130 captured by the illuminator 102 can have similar sound levels at frequencies of 1 kHz and 8 kHz. In contrast, the sound 130 captured by the illuminator 104, located behind the person 120, may have significantly different sound levels at frequencies of 1 kHz and 8 kHz, which can be based on... Figure 2 and 3 The curves 200 and 300 shown are used to determine this. For example, server 112 can execute the AI model described above to estimate the facing direction of person 120 based on the spatial directional features of sound 130, where the spatial directional features of sound 130 are determined based on sound 130 received by illuminators 102 and 104. Server 112 can also use one or more spatial directional features from one or more other illuminators (e.g., illuminator 106) from system 100 (such as illuminators 106-110). As described above, the spatial directional features of sound 130 with respect to multiple illuminators can be determined by the illuminators. Alternatively, the spatial directional features of sound 130 relative to multiple illuminators can be determined by server 112 based on audio data from the illuminators, generated by sound 130 received by the illuminators, or based on spectral power distribution information of sound 130 received from the illuminators.
[0058] In some example embodiments, in step 814, method 800 may include performing fall detection of person 120 based on a sequence of orientation of person 120, the sequence of orientation being estimated based on a time series of spatial directional features of sounds emitted by person 120 (such as sounds 602-610). For example, server 112 may be based on the manner described above, for example, with respect to step 812. Figure 6 The orientation of person 120 is estimated using the sounds 602-610 shown. The orientation of person 120 can be estimated based on time-series data of spatial directional characteristics obtained from or otherwise derived from the sounds 602-610. Server 112 can determine whether the orientation sequence determined based on the sounds 602-610 corresponds to or otherwise indicates a fall or incident, for example, by comparing the orientation sequence with an orientation sequence known to correspond to a fall or incident. If server 112 determines that person 120 has fallen, server 112 can send a notification (e.g., a text message or email) to entities such as elderly care service operators, nurses, family members, etc.
[0059] By using the voice emitted by a person, system 100 can estimate the direction a person is facing. Using the voice of a person instead of a camera to estimate their facing direction may be desirable, for example, because using voice provides better privacy. By using the facing direction sequence estimated using the voice, system 100 can detect fall events / accidents. Performing fall detection while providing improved privacy may be desirable in homes, shops, etc., for the elderly.
[0060] In some alternative embodiments, Figure 1 System 100 may include more or fewer illuminators than shown without departing from the scope of this disclosure. In some alternative embodiments, illuminators 102-110 may be installed in a different configuration than shown without departing from the scope of this disclosure. In some alternative embodiments, illuminators 102-108 may be pendant illuminators, wall lamps, etc. In some exemplary embodiments, illuminator 110 may be a table lamp or another type of illuminator without departing from the scope of this disclosure. In some alternative embodiments, region 114 may have a different shape than shown without departing from the scope of this disclosure. In some alternative embodiments, multiple persons may be in region 114 without departing from the scope of this disclosure.
[0061] In some alternative embodiments, method 800 may include more or fewer steps than shown without departing from the scope of this disclosure. In some alternative embodiments, the steps of method 800 may be performed in a different order than shown without departing from the scope of this disclosure.
[0062] Although specific embodiments have been described in detail herein, these descriptions are by way of example. The features of the exemplary embodiments described herein are representative, and in alternative embodiments, certain features, elements, and / or steps may be added or omitted. Furthermore, those skilled in the art can make modifications to aspects of the exemplary embodiments described herein without departing from the scope of the appended claims, which are intended to be interpreted in the broadest possible sense to include modifications and equivalent structures.
Claims
1. A method (800) for detecting a person's facing direction using an illumination system (100), the method comprising: The sound (130) produced by the person (120) in the area (114) is received (806) by the first illuminator (102) and the second illuminator (104); (810) The spatial directionality characteristics of the sound are determined based on the sound levels received by at least the first and second illuminators, wherein the sound levels are at least two frequencies of the sound; and By executing an artificial intelligence (AI) model (714), the orientation of a person is estimated (812) based on the spatial directionality features of sound. The AI model (714) is trained using spatial directionality feature data labeled with the orientation.
2. The method of claim 1 further includes performing (814) fall detection of a person (120) based on the person's facing direction sequence, the person's facing direction sequence including the person's facing direction, wherein the facing direction sequence is determined based on a time series of spatial directional features of one or more sounds (130, 602-610) emitted by the person (120).
3. The method of claim 2, further comprising detecting (802) a person in the area and estimating the location of the person.
4. The method of claim 3, further comprising tracking (804) people in the area.
5. The method of claim 2, further comprising confirming the result of the person fall detection based on the type of the one or more sounds.
6. The method of claim 2, further comprising negating the results of a person fall detection based on the speed of a person in the area determined using radar-based and / or radio frequency (RF) sensing.
7. The method of claim 1 further includes filtering (808) the sound to remove noise from the sound.
8. The method according to claim 1, wherein, The direction of a person's orientation is further estimated based on additional spatial directional features, which are determined based on the additional sound level of the sound (130) received by the third illuminator (106, 108, 110), wherein the additional sound level of the sound is at at least two frequencies of the sound.
9. The method of claim 1, wherein the spatial directionality feature of the sound comprises a first spatial directionality feature of the sound and a second spatial directionality feature of the sound, wherein the first spatial directionality feature is determined by a first illuminator (102) or a server (112) based on a first sound level of the sound at at least two frequencies of the sound received by the first illuminator (102), and wherein the second spatial directionality feature of the sound is determined by a second illuminator (104) or a server (112) based on a second sound level of the sound at at least two frequencies of the sound received by the second illuminator (104).
10. The method of claim 1, wherein estimating the person’s facing direction is performed by a first illuminator (102), a second illuminator (104), or a server (112) configured to communicate with the first and second illuminators.
11. A lighting system (100) for detecting the orientation of a person, the system comprising: A first illuminator (102, 700) includes a first processor (702) and a first microphone (706), wherein the first illuminator is configured to receive sound (130) generated by a person (120) in area (114) via the first microphone, and wherein the illumination system is configured to determine a first spatial directionality feature of the sound based on a first sound level at at least two frequencies of the sound received by the first illuminator; and The second illuminator (104, 700) includes a second processor (702) and a second microphone (706), wherein the second illuminator is configured to receive a sound generated by a person via the second microphone, wherein the lighting system is configured to determine a second spatial directionality feature of the sound based on a second sound level at at least two frequencies of the sound received by the second illuminator, and wherein the lighting system is configured to execute an artificial intelligence (AI) model (714) to estimate the person's facing direction based at least on the first spatial directionality feature and the second spatial directionality feature.
12. The lighting system (100) of claim 11, wherein the lighting system is further configured to perform fall detection of the person (120) based on a sequence of the person's (120) orientation, and wherein the sequence of orientation is determined based on a time series of spatial directional features of one or more sounds emitted by the person.
13. The lighting system (100) according to claim 11 further includes a third illuminator (106, 108, 110), wherein, The system is configured to further estimate the person's facing direction based on the third spatial directionality characteristics of the sound determined by the third illuminator, the third spatial directionality characteristics of the sound being determined based on the third sound level of the sound at at least two frequencies received by the third illuminator.
14. The lighting system (100) of claim 11, wherein the first processor (702) is configured to execute the AI model to estimate the facing direction of a person (120) based at least on a first spatial directional feature of the sound and a second spatial directional feature of the sound.
15. The lighting system (100) of claim 11, wherein the server (112) is configured to execute an AI model (714) to estimate the person’s facing direction based at least on a first spatial directional feature of the sound and a second spatial directional feature of the sound.