Determination of location and orientation of an entity in an area
Patent Information
- Application Number
- PCT/US2025/011435
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-13
- Publication Date
- 2025-08-21
AI Technical Summary
Existing smart home and building technologies face challenges in integrating sound localization to determine both the location and orientation of sound sources, as well as in effectively communicating between devices using different network protocols, leading to inefficiencies and fragmented control systems.
A system and method for determining the location and orientation of sound sources using triangulation with synchronized sound receivers integrated into household devices, employing machine learning for sound recognition and command execution, and utilizing a standard communication protocol like Matter for unified device control.
Enables precise determination of sound source location and orientation, allowing for intelligent device control based on user direction and position, enhancing security and automation in smart environments while overcoming protocol fragmentation.
Smart Images

Figure US2025011435_21082025_PF_FP_ABST
Abstract
Description
DETERMINATION OF LOCATION AND ORIENTATION OF AN ENTITY IN AN AREACROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present disclosure claims priority to Provisional Application No. 63 / 620,664, entitled DETERMINATION OF LOCATION AND ORIENTATION OF AN ENTITY IN AN AREA, and filed 12 January 2024, the content of which is hereby incorporated by reference herein in its entirety.FIELD OF THE INVENTION
[0002] The present disclosure relates to novel and advantageous systems and methods for determining location and orientation of an entity or sound source in an area. Particularly, the present disclosure relates to novel and advantageous systems and methods for determining location and orientation of orientation of people and animals in an area. More particularly, the present disclosure relates to novel and advantageous systems and methods for determining location and orientation of people and animals in an area, and use of such systems and methods in a smart building environment for identifying devices that may be targeted to receive commands, and executing such commands on the appropriate devices.BACKGROUND OF THE INVENTION
[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.
[0004] The Internet of Things (loT) refers to the network of connected objects that are able to collect and exchange data in real time using embedded sensors. Thermostats, cars, lights, refrigerators, appliances, and other devices can be connected to the loT. The rapid growth of the loT has led to the possibility of smart communities, smart buildings, and smart homes. More specifically, the loT can be used to make communities, buildings, and homes (as examples) smarter and more efficient. The smart communities, buildings, and homes in turn enable assessment of conditions quantitatively. Such smart communities, buildings, and homes can use state-of-the-art technology and real-time data analytics to improve environmentalsustainability, reduce the digital divide, and enhance people’s lives with smarter, personalized, and more intuitive services and experiences
[0005] A smart building is a structure that uses automated processes to automatically control and monitor the building’s operations, including HVAC, lighting, security, and others, maximizing user comfort while minimizing energy consumption. A smart home refers to a habitation having a communication network, sensors, household devices, and appliances that can be remotely identified, accessed, monitored, and controlled.
[0006] Smart communities, buildings, and homes can make life and work easier in the following ways:• Comfort for occupants with controlled lighting, temperature, and humidity;• Automated control of a building’s HVAC, lighting, electrical, shading, access, and security systems;• Cost optimization with analyzing building usage patterns and adjusting availability or automation;• Reduced environmental impact by analyzing indoor and outdoor environmental conditions, occupants’ behavior, and other data that can optimize energy and water consumption;• Integration capabilities with the ability to be embedded into older structures;• Preventative maintenance by analyzing real-time and historical equipment data; and• Enhanced health and well-being with access to control systems and improving indoor air quality through efficient HVAC operation.
[0007] Ongoing issues with smart home technology are thought to have slowed the pace of adoption thereof and development of smart communities, buildings, and homes. The technology currently used in smart buildings is piecemeal, relying on a range of connected devices to perform different functions. These devices can control different aspects of a building or home, such as the thermostat, garage doors, or security cameras, but are not generally integrated in a meaningful manner.
[0008] One pervasive issue with creating an integrated smart community, building, or home that uses loT devices is that devices today use a variety of different network protocols. These protocols may include, for example, wifi, Bluetooth, Zigbee, and Z-wave. This means that the devices cannot all connect or communicate with each other. Any new device may require a new app and, in the aggregate, lead to a digitized version of the “too many remotes” dilemma.It is thought that a standard communication protocol, such as Matter, could address many issues associated with the number of network protocols currently used.
[0009] Sound localization is the process of identifying spatial coordinate of a sound source based on the sound signal received by an array of microphones. Sound localization is often based on a triangulation of the sound signal to determine the location of the sound source. Little effort has been given to both identifying the location of the sound source and the directionality of the sound source. For example, while it may be possible to determine where someone is positioned in a room, it has been challenging to determine in what direction that person is facing in that position in a room.
[0010] Further, minimal effort has been given to integrating sound localization into a smart home or building. Sound localization may be integrated with a smart building or home to provide security functionality, to enable commands, and other.BRIEF SUMMARY OF THE INVENTION[OH] The following presents a simplified summary of one or more embodiments of the present disclosure in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments, and is intended to neither identify key or critical elements of all embodiments, nor delineate the scope of any or all embodiments.
[0012] The present disclosure, in one or more embodiments, relates to a method for determining location and orientation of an entity in an area is provided. The method may comprise mapping a plurality of sound receivers in the area, detecting a sound from the entity, triangulating an x-y position of the entity based on the detected sound, and determining orientation of the entity based on the detected sound. The sound receivers may be incorporated into household devices including at least one of an electrical outlet, a light switch, a light fixture, a television, a speaker, a microwave, an occupancy sensor, a smoke detector, a thermostat, an air quality monitoring device, and a contact sensor.
[0013] Mapping the plurality of sound receivers may comprise mapping the plurality of sound receiver with respect to the area and with respect to one another and may further comprise mapping locations of structures within the area. Detecting a sound may comprise detecting an audio stream and identifying an audio frame in the audio stream or identifying a plurality of audio frames in the audio stream.
[0014] In some embodiments, at least two sound receivers are provided on different vertical planes from one another. Using the at least two sound receivers, the method may further comprise determining a z position of the entity.
[0015] In some embodiments, the method further comprises synchronizing the sound receivers to a same absolute time base. Synchronizing the sound receivers may be done by outputting a synchronization sound signal, receiving the synchronization sound signal by the plurality of sound receivers, and processing the synchronization sound signal. In various embodiments, the synchronization sound signal may be in a frequency spectrum not audible by humans or pets.
[0016] The method may further comprise identifying a room in which the entity is positioned and discarding the sound if the room is not the area.
[0017] The method may further comprise identifying the entity based on the detected sound and, in some embodiments, when the entity issues a command, the system may cross check whether the entity has permissions for the command.
[0018] In other embodiments, a method for localization of a sound source in an area is provided comprising mapping a plurality of sound receivers in the area, detecting a sound from the sound source, identifying the sound, determining whether the sound is threat related, executing security measures if the sound is threat related. For example, when the sound is identified as breaking glass, the sound is determined as threat related, and taking security measure comprises calling authorities. Further, when the sound is identified as screaming, they method may include coordinating a call to the area to evaluate safety.
[0019] The present disclosure, in one or more embodiments, additionally relates to a system for determining location and orientation of an entity in an area is provided, in accordance with some embodiments. The system may include a plurality of sound receivers for receiving a sound, a memory for storing an audio stream of the sound, and a processor capable of triangulating an x-y position of the entity using the audio stream and determining orientation of the entity using the audio stream. The system may further include a sound emitter for outputting a synchronization sound signal. In some embodiments, the system may be integrated into a smart building.
[0020] While multiple embodiments are disclosed, still other embodiments of the present disclosure will become apparent to those skilled in the art from the following detailed description, which shows and describes illustrative embodiments of the invention. As will be realized, the various embodiments of the present disclosure are capable of modifications in various obvious aspects, all without departing from the spirit and scope of the presentdisclosure. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] While the specification concludes with claims particularly pointing out and distinctly claiming the subject matter that is regarded as forming the various embodiments of the present disclosure, it is believed that the invention will be better understood from the following description taken in conjunction with the accompanying Figures, in which:
[0022] Figure 1 illustrates a flow diagram of a method for localizing an entity in a room based on sound generated by that entity, in accordance with one embodiment.
[0023] Figure 2 illustrates a device with a sound receiver provided therein, in accordance with one embodiment.
[0024] Figure 3 illustrates a top view of a smart room with localization capabilities, in accordance with one embodiment.
[0025] Figure 4a illustrates the room of Figure 3 with sound receivers provided on the same plane, in accordance with one embodiment.
[0026] Figure 4b illustrates a room with sound receivers provided on the same plane, in accordance with one embodiment.
[0027] Figure 5 illustrates the room of Figure 3 with sound receivers provided on different planes, in accordance with one embodiment.DETAILED DESCRIPTION
[0028] The present disclosure relates to novel and advantageous systems and methods for determining location and orientation of an entity or sound source in an area (also referred to interchangeably as a room herein). Particularly, the present disclosure relates to novel and advantageous systems and methods for determining location and orientation of people and animals in an area. More particularly, the present disclosure relates to novel and advantageous systems and methods for determining location and orientation of people and animals in an area and use of such systems in a broader smart building environment for identification of devices that are targeted for commands, and executing such commands on the devices.
[0029] The system and method use triangulation to identify a location of a sound source (or entity) in an area (such as a room). The system and method may use artificial intelligence, built on machine learning models, to analyze and identify audio signals emitted by the sound source and, in some instances, to execute actions based on an identified audio signal. In variousembodiments, the sound source may be a person or an animal. In some embodiments, the system and method identifies whether the person is likely a child or an adult. The system and method may further identify the orientation, or directionality, of the entity. In the context of a person, orientation / directionality refers to the direction the person is facing. The system may be integrated into or configured to work with a smart home such that the position and orientation of a person, and in some cases the identity of the person, can be used in operation of devices in the smart home.
[0030] The system for determining location and / or orientation in an area can be integrated into a smart community, building, or home. For example, by determining the position, and / or the orientation of a person in a room, commands given by that person may be interpreted by the system and then executed by a specific smart device. More specifically, by identifying what direction a person is facing, the system may interpret what device is being targeted for commands. For example, if a person says “open blinds,” the system can identify what direction the person is facing and the blinds towards which the person is looking may be commanded to open. Additionally and / or alternatively, by determining the position of a person or animal in a room, certain devices may automatically be triggered on or off. The system may further be configured to take certain actions upon detection and identification of a specific sound. For example, if a security-related sound is identified, the system may call police.
[0031] Figure 1 illustrates a flow diagram of a method for determining location and orientation of a sound source, such as an entity, in a room based on sound generated by the sound source. The method uses one or more sound receivers located within a space. The sound receivers may be, for example, microphones. The microphones may be incorporated into devices, including smart devices, within the area. For example, the microphones may be incorporated into light sockets, lights switches, thermostats, etc.
[0032] The sound receivers are mapped 12 with respect to one another and with respect to location in the space. In general, localization may be improved with at least four sound receivers being provided in a given area. However, in various embodiments, fewer than four or more than four sound receivers may be used.
[0033] The sound receivers detect a sound 14 from a sound source or entity. The sound may be, for example, words or other vocalizations, crying, coughs, sneezes, footsteps, etc. The sound receivers may detect sounds useful for security monitoring, such as ingress and egress noises, door slamming, glass breaking, screaming (not detected as speech), an alarm, a siren, and other emergency-type noises. Other noises that may be detected include, for example,laughter, barking, meowing, a bird chirping, a car, an engine, music playing, a thunderstorm, etc.
[0034] The system triangulates a position 16, 18 of the sound source based on the position of the sound source relative the microphones. Triangulating the origin of a sound can depend on a time difference of arrival (TDOA) calculation and locating the sound source using geometric relationships between microphones and the sound source. The calculated position may include x, y, and z coordinates and orientation / directionality. The x-y position reflects the position of the entity in a 2-dimensional area of the room. The z position reflects the height position of the sound origin (which could be, for example, feet or mouth). The system may determine directionality 20 of the sound source. Orientation / directionality reflects the directionality of the sound origin or, more specifically, what direction the sound source is facing. In some embodiments, directionality may not be determined.
[0035] In some embodiments, the system and method may have sound recognition and / or speech recognition capabilities for identifying the sound source. Based on identification of the sound, the system may execute various actions. Accordingly, the system may execute, or cause to be executed, one or more commands 22 based on the position and / or orientation of the sound source.
[0036] Placement and Mapping of Sound Receivers
[0037] In some embodiments, during design of a smart building, desired positions of sound receivers, or microphones, may be determined and the building constructed with sounds receivers in the desired positions. In other embodiments, sound receivers may be placed at preset locations common to certain types of buildings. In yet other embodiments, a building may be retrofit with sound receivers wherein desired positions of the sound receivers are determined and the sound receivers put in place at those locations.
[0038] Figure 2 illustrates a device 30 with a sound receiver 32 provided associated therewith, in accordance with one embodiment. The sound receiver 32 may be integral to the device, may be placed on or in the device, or may, in some embodiments, be placed without a device. The device 30 shown in Figure 2 is an electrical outlet. In alternative embodiments, the sound receiver may be included in any other suitable device including, for example, an electrical outlet, a light switch, a light fixture, a television, a speaker, a microwave, occupancy sensor, smoke detector, a thermostat, and air quality monitoring device, a contact sensor, etc. Any combination of devices housing sound receivers may be used and it is not necessary to use the same type of device in any given area or room and, in some embodiments, the sound receiver may be provided without association with a particular device. In some embodiments, it may beuseful to provide sound receivers at different heights such as in different types of devices that are typically found at different heights - for example, an outlet, a smoke detector, and a thermostat.
[0039] In the embodiment of Figure 2, the device 30 comprises an outlet having two sockets for receiving an electrical plug 34, a USB charging socket 36, a USB-C charging socket 38, and a sound receiver or microphone 32. In the embodiment shown, the sound receiver is generally centrally placed on the outlet. However, the sound receiver may be provided at any location on or within the outlet. In general, the device holding the sound receiver is in an unmovable position such that the exact position of the sound receiver is known and can be used in triangulating the position of a sound source.
[0040] Placement of the sound receivers may be done such that triangulation of location and orientation / directionality of a sound source is possible. The area and the location of each sound receiver in the area is known - space, position in space, position to each other, etc. - via mapping and may be used to triangulate a position of a sound source in the area. When mapping the sound receivers, the size of the area and structures (such as chairs or tables) in the area may also be mapped such that echoing may be accounted for.
[0041] Figure 3 illustrates a top view of a smart room 40 with four sound receivers 42 provided therein. The sound receivers 42 may be provided at any suitable location in the room, such as on or within (e.g. flush mounted) the walls. The location of the sound receivers 42 on the walls may accommodate doors, windows, etc. and need not be uniform one wall to another. For ease of reference, sound receivers are discussed herein as provided associated with a wall but this is not intended to be limiting and, so long as mapped accordingly, one or more sound receivers may not be provided on a wall and may be associated with a ceiling, a bookcase, a mantle, or other. In some embodiments, the sound receivers may be provided as part of a power socket or outlet such as shown in Figure 2. The sound receivers may be provided at the same height or at different heights. For example, if a sound receiver is provided in a socket, standard sockets are provided 12-16” off of the floor (though may be provided at a different height). In some embodiments, the smart room may be designed such that at least one sound receiver is located proximate a height of a sound source, such as at head level for a person of average height. Providing sound receivers at different heights may be advantageous for determining a height of a sound source.
[0042] Figures 4a and 4b illustrate views of a room 44 having sound receivers 46 provided on the same plane, in accordance with various embodiments. Figure 4a illustrates two walls and a floor of the room. Figure 4b illustrates three walls of a room. It is to be appreciated that thesame principals apply to other walls of the room. In the embodiments shown, sound receivers 46 are shown provided one on each wall of the room. It is to be appreciated that more or fewer sound receivers may be provided. For example, in some embodiments no sound receivers may be provided on a given wall or more than one sound receiver may be provided on a given wall.
[0043] Figure 5 illustrates a view of a room 50 having sound receivers 52, 54 provided on different planes, in accordance with one embodiment. Figure 5 illustrates two walls and a floor of the room. As shown, one sound receiver 52 is provided near the floor, such as at a typical outlet location and one sound receiver is 54 provided near the head height of an average human, such as a thermostat. It is to be appreciated that the same principals apply to other walls of the room.
[0044] Detect Sound from a Sound Source
[0045] The sound receivers are used to detect a sound signal from a sound source. In various embodiments, the sound receivers may be always-on listening devices, may be devices that are on at preset intervals, may be manually turned on, or other. In some embodiments, triangulation of the position of a sound source may only occur after a sound receiver hears a wake word, upon an occurrence of an event (such as a person entering through a front door, or other), etc .
[0046] In general, a sound is detected by all sound receivers in the area of the sound. For example, if a person in a room says something, all sound receivers in the room will typically detect that sound. Further, in some instances, sound receivers in other areas (such as an adjacent room) may detect the sound. The system and method may be programmed with a cutoff based on noise level or volume.
[0047] Room to room identification, or identification of the room in which a sound source is located if sound receivers in different rooms pick up a noise, can be built into the localization algorithm based on the known architectures of rooms. More specifically, for any given audio stream there are aspects to the audio stream that will indicate if a sound source is in the same room. For example, when sounds move through a wall, frequency of the sound is modulated (typically lowered). In some embodiments, the system and method may be comprehensive such that sounds detected in one room may be compared with sounds detected in another room and the system may discard redundant but modulated recordings.
[0048] When a sound is detected, one or more pieces of data about the sound and / or sound source may be identified. These may include, for example, volume of the sound, words in the sound, whether the sound source appears to be moving, etc.
[0049] In instances where many people are talking, the system may be unable to parse out and identify individual voices. However, other sounds - such as a dog barking, glass breaking, ora baby crying - may be identifiable over the combined other voices. The system thus may be configured to continue recording and attempting to identify sounds.
[0050] The system and method may be configured to record and analyze any suitable length of audio stream. In some embodiments, the system and method use 200, 300, or 500 millisecond audio streams. Each recorded sound stream may comprise one or more sound frames.
[0051] Triangulating Location of Sound Source
[0052] Based upon the detected sound, a location of the sound source is determined via triangulation. The received sound and detail about the sound signal are used to determine the position of the sound source. This may include not only position but also vertical distance above ground and / or directionality of the sound source.
[0053] Challenges exist with triangulating a position of a sound source using a plurality of sound receivers. In general, accuracy of triangulation is dependent upon how in synch the recording of the sound signal by each sound receiver is. Accordingly, in some embodiments, it may be useful to account for variable latency or individual machine delay. Various methods can be used for accounting for individual machine delay among sound receivers.
[0054] To enhance the efficiency and accuracy of triangulation, the system and method may be configured such that the sound recordings have the same absolute time base (rather than the same relative time base). An absolute time base, or value, coordinates the recordings with a specific real-world time. In one embodiment, the system and method synchronizes the recordings from each sound receiver during processing of the recordings.
[0055] The more accurate the time synchronization, the faster and more accurately triangulation can be performed. In general, it may be useful for the synchronization method to provide a same-time base for recordings on the order of l / 100thof a millisecond. For sound receivers to have the same time base, one or more devices within an area (of which the location is known (absolute or in relation to the devices supplied with the sound receivers)) may be supplied with means to generate a synchronization sound signal. The device for outputting a synchronization sound may be a device with a sound receiver or may be a different device. The synchronization sound signal may be in a frequency spectrum not audible by humans or pets, but recognizable by the system sound receivers. As all locations of all devices (the device generating the synchronization sound signal and the sound receivers) are known, the time of each device can be synchronized. The synchronization sound signal can also have the absolute time encoded, or parts of the absolute time encoded (e.g. seconds in the minute, seconds is the hour, etc.). A synchronization device may be provided including a sound emitter for outputtingthe synchronization sound signal. Alternatively, one of the devices including a sound receiver may also include a sound emitter for outputting the synchronization sound signal.
[0056] In some embodiments, the synchronization sound signal may be emitted from the device periodically or at specific times. The device for outputting a synchronization sound may emit a piezo-tone or piezo-noise. The piezo-tone may be, for example, a very high or very low frequency timing signal that is generally inaudible to humans. Receipt of the piezo-tone by each sound receiver can be used to calibrate ensuing triangulation calculations. The synchronization sound thus time tags a recording such that while the recordings may themselves not be synchronized (because of machine lag or other), synchronization may be done after recording independent from the sound receivers.
[0057] Using a synchronization tone, the algorithms associated with preprogrammed knowledge of location of sound receivers may be fine-tuned for timing purposes.
[0058] Alternatively, in one embodiment another method of ensuring the same time base is to use real time ethernet or other transmission technology having a deterministic protocol. Deterministic protocol minimizes delay variation, providing consistent latency, and having a predictable amount of time for data packets to travel across the network. Using a method having deterministic protocol ensures that the data is transmitted and received not just quickly, but consistently and predictably.
[0059] X-Y Position
[0060] The position of a sound source is determined based on the time of emission (ToE) of a sound emitted by the sound source, a time of arrival (To A) of the sound at a sound receiver, and the corresponding time of flight (ToF) of the sound between the sound source and each sound receiver. A first determined location of the sound source may be the x-y position of the sound source, or the 2-dimensional position of a sound source along an x-y plane.
[0061] To triangulate a position of an entity or sound source, an audio (or sound) frame is analyzed. The system and method can calculate a localization value of an audio frame having a duration of at as little as 1 / 10thof a millisecond. In some embodiments, the system and method may be configured to calculate a localization value only of audio frames having a duration of at least 2x that at the time of synchronization. In some implementations, a plurality of different sound frames may be used to enhance the accuracy of localization.
[0062] By triangulation of the signals (different run-times from sound source to sound receiver) and evaluation of the strength of the signal, the position of the noise source can be determined. Due to the specific frequency response, persons / animals can be distinguished from each other - or the number of different persons / animals can be determined. It is to be notedthat there are different modulations to sound that happen as sound travels, for example, reflection. Sound travels predictably and, based on the known variables in a given area, the algorithm of the system and method may be programmed to account for sound modulation that would be expected of sound traveling through the known space with the known factors.
[0063] For triangulation, the different signal propagation times of the individual audio signals are determined. More specifically, once a sound is emitted by a sound source, that sound propagates through air in an area and reaches the sound receivers. The sound takes a different amount of time to reach each sound receiver in an area due to the different distances to be travelled by the sound waves. The system triangulates the position of the sound source based on the ToE of the sound and the ToA of the sound at the various sound receivers, described more fully below.
[0064] In some embodiments, triangulation is performed on any sounds having minimum quality metrics, including duration and noise level. Factors contributing the quality of the sound include the quality of the sound sensor, the frequency band of the sound sensor, and the accuracy of the model of a selected machine learning (ML) algorithm for the system (discussed below).
[0065] Triangulation may be done by determining one or more audio frames that occur in an audio stream (the propagation of a sound wave from the sound source to the sound receiver) of a sound receiver (e.g. a microphone with the highest level, or a microphone that reported the existence of an audio stream last in a room) in the signals of the other sound sources.
[0066] Assuming that all of the sound frames have the same time base and that the position in space of each sound receiver to one another is known for the time of recording, there is a signal propagation time difference for each of the sound receivers. The propagation time difference of a sound frame is equivalent to the difference in the distance to the sound source. To determine the position of the sound source in space, the distance to the sound source of each sound receiver may be displayed as a sphere. If all radii are extended simultaneously until all spheres separate or affect every other sphere at least once, the resulting intersection of all spheres results in the probability space of the location of the sound source. The equation to be solved is a nonlinear equalization equation, which can be solved, for example, with the Levenberg-Marquardt algorithm.
[0067] In some embodiments, localization, or position determination is performed with 10Hz or 5Hz. Because of the small time frame necessary for triangulation, the system and method are able to detect movement of a sound source.
[0068] Triangulation data can be improved or plausibly checked by further location information from other system (e.g WiFi / Bluetooth) interference measurements. Additionally, advanced synchronization techniques can be employed to ensure that all sound receivers have the same absolute time base, significantly improving triangulation accuracy. This can be achieved by emitting a synchronization sound signal, inaudible to humans, which is recognized by the system's sound receivers. The synchronization sound signal can also encode absolute time, allowing for precise time alignment across devices. Furthermore, machine learning models can be utilized to analyze and identify audio signals, filtering out noise and enhancing the precision of triangulation. By incorporating these methods, the reliability and accuracy of triangulation data can be greatly enhanced.
[0069] Z Position
[0070] A second determined location of the sound source may be the z position of the source, or the vertical distance of the sound source above ground. More specifically, the vertical distance of the sound source (such as the mouth of a speaker) can be determined. This can be used to extrapolate, for a human sound source, a height of the human and, based on this, whether the human is an adult or a child. Determination of vertical position of a sound source may be enhanced by providing sound receivers at multiple planes, or heights.
[0071] Determining Directionality of Sound Source
[0072] Directionality
[0073] The system and method can further be used to determine the direction that a sound source is facing, also referred to as viewing direction. For example, for a human speaker, the system and method can determine what direction the speaker is facing and, based on that, extrapolate to what device a command should be directed.
[0074] More specifically, the orientation of the head can be determined in humans and thus conclusions can be drawn about the direction of view (same as speech). This information can also be used to create context in speech recognition, e.g. with the command, "Open windows" can be recognized which window is meant.
[0075] In one embodiment, the viewing direction may be determined by the frequency level of the sound receivers. To compare the levels with each other, one or more audio frames is identified in all audio streams, as in the location determination. The difference in levels at one or more frequencies or an averaged level over several or all frequencies can then be calculated.
[0076] Determining Information based on Changes to Ambient Noise
[0077] Oftentimes, there are persistent audio streams in a room that outside of the frequency range audible to the human ear. These are referred to herein as ambient noises. A personwalking through a room will modulate the ambient noises by their presence. Accordingly, the system and method may be programmed such that changes to an ambient noise may be detected and the system can identify instances where such change is due to a person being present in a room. Ambient noises may include HVAC sounds, pet sounds, sounds external to the living space, and other artificial and natural sounds that would occur in the environment.
[0078] Determining Information about Sound Source
[0079] Based on information gained by the sound receivers, detail about the sound source may be extrapolated. For example, as discussed above, based on the z-position of the sound source, height may be extrapolated and a determination made of whether a human sound source is an adult or child.
[0080] In some embodiments, the system and method may make determinations about the sound source based on the received sound, in some cases without reference to the location of the sound source. This may include identification of a person, specifically or as an “unknown person.”
[0081] In some embodiments, the system and method may execute commands on a smart home based upon determined information about a sound source. For example, if a child is detected, certain items in a room may be turned off or deactivated - such as an oven, power sockets, or other items that could pose a danger to a child.
[0082] Command Control
[0083] In a smart community, building, or house, voice commands may be given to operate various devices. These may include, for example, lights, speakers, windows, blinds, sinks, thermostat, and others.
[0084] Using the system and method described herein, control of devices may be enhanced. For example, based on directionality of a sound source saying “Close blinds,” the system can identify what blinds the sound source is facing and can close those blinds. Further, in some embodiments, the system and method may identify the identity of the speaker and may be trained to identify that, when Speaker A asks for blinds to be closed, she prefer them closed 75%. While when Speaker B asks for blinds to be closed, he prefers them closed 100%.
[0085] In some embodiments, the system may be programmed with permissions for certain commands. For example, the ability to turn off a security system or to open a safe may be associated with certain permissions. Upon receipt of commands requiring such permissions, the system may identify the sound source and evaluate whether the sound source has the required permissions.
[0086] Artificial Intelligence
[0087] The system and method may use artificial intelligence to identify sound sources and further to evaluate and execute commands.
[0088] Capabilities of the system and method may be enhanced using machine learning. More specifically, machine learning may be used to provide the system with capabilities to identify certain sounds, identify specific speakers, and the like. Sounds the identification of which may be enhanced by machine learning may include any of those identified above, such as human voice, pet sounds, weather sounds (rain, thunder), security -related sounds (door opening, glass breaking, alarms, sirens), etc. The system may further incorporate artificial intelligence to continually evolve sound identification, to take actions upon sound identification, to connect to other smart devices, etc.
[0089] Sound identification, including voice identification, may be enabled by developing a training set for the sound(s) to be identified. A plurality of samples of each of the sounds (including, for example, of a specific individual’s voice) are collected and labelled as a dataset from a specific source. Relevant features of the sound, such as pitch formants, and other acoustic properties, are extracted. A training model is trained on each dataset by associating the extracted features with the labelled source. Any suitable machine learning model may be used including, for example, Gaussian Mixture Models (GMMs), Hidden Markov Models (HMMs), Convolutional Neural Networks (CNNs), and Recurrent Neural Networks (RNNs). The trained model may then be integrated into software used in system and method.
[0090] In addition to training the system to identify specific sounds, the system may be designed to trigger action upon detection of a type of sound or a type of sound at a certain location.
[0091] In some embodiments, the system may be configured to alert a homeowner, security, the police, or other upon identification of a security-related sound. The system may be trained with what actions to take upon detection of what sounds and, if applicable, at what locations. For example, if the system detects glass breaking and triangulates that the glass breaking occurred at a window, the system may alert authorities. In contrast, if the system detects glass breaking and triangulates that glass breaking occurring at a kitchen sink, the system may not take action. In some embodiments, the system may be configured to check upon safety of someone inside the area upon detection of a possible threat related sound. For example, if a scream is detected, the system may be configured to cause a security company to call the building in which the area is located to confirm safety of residents.
[0092] In other embodiments, the system may be configured to coordinate with smart devices in the a smart building based on receipt of a command. For example, a speaker may say “turnon lights”. The system can identify where the speaker is and what direction the speaker is facing. Based on this, the system can identify what lights to turn on and can send a command to those lights. Similar actions can be taken for turning devices (tv, music speakers, faucets, lights) on and off.
[0093] In yet a further embodiment, the system can follow commands only if given by a specific person. For example, if the system receives a command “unlock safe,” the system may be configured to cross check whether the speaker of the command has permissions to unlock the safe.
[0094] The system can be figured for detecting unknown people in the room. If the system has been trained on known voices, the system can identify when an unknown person is in the room.
[0095] The system can be configured to constantly learn about the people in the smart building or smart house.
[0096] For purposes of this disclosure, any system described herein may include any instrumentality or aggregate of instrumentalities operable to compute, calculate, determine, classify, process, transmit, receive, retrieve, originate, switch, store, display, communicate, manifest, detect, record, reproduce, handle, or utilize any form of information, intelligence, or data for business, scientific, control, or other purposes. For example, a system or any portion thereof may be a minicomputer, mainframe computer, personal computer (e.g., desktop or laptop), tablet computer, embedded computer, mobile device (e.g., personal digital assistant (PDA) or smart phone) or other hand-held computing device, server (e.g., blade server or rack server), a network storage device, or any other suitable device or combination of devices and may vary in size, shape, performance, functionality, and price. A system may include volatile memory (e.g., random access memory (RAM)), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, ROM, and / or other types of nonvolatile memory (e.g., EPROM, EEPROM, etc.). A basic input / output system (BIOS) can be stored in the non-volatile memory (e.g., ROM), and may include basic routines facilitating communication of data and signals between components within the system. The volatile memory may additionally include a high-speed RAM, such as static RAM for caching data.
[0097] Additional components of a system may include one or more disk drives or one or more mass storage devices, one or more network ports for communicating with external devices as well as various input and output (VO) devices, such as digital and analog general purpose VO, a keyboard, a mouse, touchscreen and / or a video display. Mass storage devices may include, but are not limited to, a hard disk drive, floppy disk drive, CD-ROM drive, smart drive, flashdrive, or other types of non-volatile data storage, a plurality of storage devices, a storage subsystem, or any combination of storage devices. A storage interface may be provided for interfacing with mass storage devices, for example, a storage subsystem. The storage interface may include any suitable interface technology, such as EIDE, ATA, SATA, and IEEE 1394. A system may include what is referred to as a user interface for interacting with the system, which may generally include a display, mouse or other cursor control device, keyboard, button, touchpad, touch screen, stylus, remote control (such as an infrared remote control), microphone, camera, video recorder, gesture systems (e.g., eye movement, head movement, etc.), speaker, LED, light, joystick, game pad, switch, buzzer, bell, and / or other user input / output device for communicating with one or more users or for entering information into the system. These and other devices for interacting with the system may be connected to the system through I / O device interface(s) via a system bus, but can be connected by other interfaces such as a parallel port, IEEE 1394 serial port, a game port, a USB port, an IR interface, etc. Output devices may include any type of device for presenting information to a user, including but not limited to, a computer monitor, flat-screen display, or other visual display, a printer, and / or speakers or any other device for providing information in audio form, such as a telephone, a plurality of output devices, or any combination of output devices.
[0098] A system may also include one or more buses operable to transmit communications between the various hardware components. A system bus may be any of several types of bus structure that can further interconnect, for example, to a memory bus (with or without a memory controller) and / or a peripheral bus (e.g., PCI, PCIe, AGP, LPC, I2C, SPI, USB, etc.) using any of a variety of commercially available bus architectures.
[0099] One or more programs or applications, such as a web browser and / or other executable applications, may be stored in one or more of the system data storage devices. Generally, programs may include routines, methods, data structures, other software components, etc., that perform particular tasks or implement particular abstract data types. Programs or applications may be loaded in part or in whole into a main memory or processor during execution by the processor. One or more processors may execute applications or programs to run systems or methods of the present disclosure, or portions thereof, stored as executable programs or program code in the memory, or received from the Internet or other network. Any commercial or freeware web browser or other application capable of retrieving content from a network and displaying pages or screens may be used. In some embodiments, a customized application may be used to access, display, and update information. A user may interact with the system,programs, and data stored thereon or accessible thereto using any one or more of the input and output devices described above.
[0100] A system of the present disclosure can operate in a networked environment using logical connections via a wired and / or wireless communications subsystem to one or more networks and / or other computers. Other computers can include, but are not limited to, workstations, servers, routers, personal computers, microprocessor-based entertainment appliances, peer devices, or other common network nodes, and may generally include many or all of the elements described above. Logical connections may include wired and / or wireless connectivity to a local area network (LAN), a wide area network (WAN), hotspot, a global communications network, such as the Internet, and so on. The system may be operable to communicate with wired and / or wireless devices or other processing entities using, for example, radio technologies, such as the IEEE 802. xx family of standards, and includes at least Wi-Fi (wireless fidelity), WiMax, and Bluetooth wireless technologies. Communications can be made via a predefined structure as with a conventional network or via an ad hoc communication between at least two devices.
[0101] Hardware and software components of the present disclosure, as discussed herein, may be integral portions of a single computer, server, controller, or message sign, or may be connected parts of a computer network. The hardware and software components may be located within a single location or, in other embodiments, portions of the hardware and software components may be divided among a plurality of locations and connected directly or through a global computer information network, such as the Internet. Accordingly, aspects of the various embodiments of the present disclosure can be practiced in distributed computing environments where certain tasks are performed by remote processing devices that are linked through a communications network. In such a distributed computing environment, program modules may be located in local and / or remote storage and / or memory systems.
[0102] As will be appreciated by one of skill in the art, the various embodiments of the present disclosure may be embodied as a method (including, for example, a computer-implemented process, a business process, and / or any other process), apparatus (including, for example, a system, machine, device, computer program product, and / or the like), or a combination of the foregoing. Accordingly, embodiments of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, middleware, microcode, hardware description languages, etc.), or an embodiment combining software and hardware aspects. Furthermore, embodiments of the present disclosure may take the form of a computer program product on a computer-readable medium or computer-readable storagemedium, having computer-executable program code embodied in the medium, that define processes or methods described herein. A processor or processors may perform the necessary tasks defined by the computer-executable program code. Computer-executable program code for carrying out operations of embodiments of the present disclosure may be written in an object oriented, scripted or unscripted programming language such as Java, Perl, PHP, Visual Basic, Smalltalk, C++, or the like. However, the computer program code for carrying out operations of embodiments of the present disclosure may also be written in conventional procedural programming languages, such as the C programming language or similar programming languages. A code segment may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, an object, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0103] In the context of this document, a computer readable medium may be any medium that can contain, store, communicate, or transport the program for use by or in connection with the systems disclosed herein. The computer-executable program code may be transmitted using any appropriate medium, including but not limited to the Internet, optical fiber cable, radio frequency (RF) signals or other wireless signals, or other mediums. The computer readable medium may be, for example but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples of suitable computer readable medium include, but are not limited to, an electrical connection having one or more wires or a tangible storage medium such as a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a compact disc readonly memory (CD-ROM), or other optical or magnetic storage device. Computer-readable media includes, but is not to be confused with, computer-readable storage medium, which is intended to cover all physical, non-transitory, or similar embodiments of computer-readable media.
[0104] Various embodiments of the present disclosure may be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products. It is understood that each block of the flowchart illustrations and / or block diagrams, and / or combinations of blocks in the flowchart illustrations and / or block diagrams,can be implemented by computer-executable program code portions. These computerexecutable program code portions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a particular machine, such that the code portions, which execute via the processor of the computer or other programmable data processing apparatus, create mechanisms for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Alternatively, computer program implemented steps or acts may be combined with operator or human implemented steps or acts in order to carry out an embodiment of the invention.
[0105] Additionally, although a flowchart or block diagram may illustrate a method as comprising sequential steps or a process as having a particular order of operations, many of the steps or operations in the flowchart(s) or block diagram(s) illustrated herein can be performed in parallel or concurrently, and the flowchart(s) or block diagram(s) should be read in the context of the various embodiments of the present disclosure. In addition, the order of the method steps or process operations illustrated in a flowchart or block diagram may be rearranged for some embodiments. Similarly, a method or process illustrated in a flow chart or block diagram could have additional steps or operations not included therein or fewer steps or operations than those shown. Moreover, a method step may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.
[0106] As used herein, the terms “substantially” or “generally” refer to the complete or nearly complete extent or degree of an action, characteristic, property, state, structure, item, or result. For example, an object that is “substantially” or “generally” enclosed would mean that the object is either completely enclosed or nearly completely enclosed. The exact allowable degree of deviation from absolute completeness may in some cases depend on the specific context. However, generally speaking, the nearness of completion will be so as to have generally the same overall result as if absolute and total completion were obtained. The use of “substantially” or “generally” is equally applicable when used in a negative connotation to refer to the complete or near complete lack of an action, characteristic, property, state, structure, item, or result. For example, an element, combination, embodiment, or composition that is “substantially free of’ or “generally free of’ an element may still actually contain such element as long as there is generally no significant effect thereof.
[0107] To aid the Patent Office and any readers of any patent issued on this application in interpreting the claims appended hereto, applicants wish to note that they do not intend any ofthe appended claims or claim elements to invoke 35 U.S.C. § 112(f) unless the words “means for” or “step for” are explicitly used in the particular claim.
[0108] Additionally, as used herein, the phrase “at least one of [X] and [Y],” where X and Y are different components that may be included in an embodiment of the present disclosure, means that the embodiment could include component X without component Y, the embodiment could include the component Y without component X, or the embodiment could include both components X and Y. Similarly, when used with respect to three or more components, such as “at least one of [X], [Y], and [Z],” the phrase means that the embodiment could include any one of the three or more components, any combination or sub-combination of any of the components, or all of the components.
[0109] In the foregoing description various embodiments of the present disclosure have been presented for the purpose of illustration and description. They are not intended to be exhaustive or to limit the invention to the precise form disclosed. Obvious modifications or variations are possible in light of the above teachings. The various embodiments were chosen and described to provide the best illustration of the principals of the disclosure and their practical application, and to enable one of ordinary skill in the art to utilize the various embodiments with various modifications as are suited to the particular use contemplated. All such modifications and variations are within the scope of the present disclosure as determined by the appended claims when interpreted in accordance with the breadth they are fairly, legally, and equitably entitled.
Claims
ClaimsWhat is claimed is:
1. A method for determining location and orientation of an entity in an area, the method comprising: mapping a plurality of sound receivers in the area; detecting a sound from the entity; triangulating an x-y position of the entity based on the detected sound; and determining orientation of the entity based on the detected sound.
2. The method of claim 1, further comprising synchronizing the sound receivers to a same absolute time base.
3. The method of claim 2, wherein synchronizing the sound receivers comprises outputting a synchronization sound signal, receiving the synchronization sound signal by the plurality of sound receivers, and processing the synchronization sound signal.
4. The method of claim 3, wherein the synchronization sound signal is in a frequency spectrum not audible by humans or pets.
5. The method of claim 1, wherein detecting a sound comprises detecting an audio stream and identifying an audio frame in the audio stream.
6. The method of claim 1, wherein detecting a sound comprises detecting an audio stream and identifying a plurality of audio frames in the audio stream.
7. The method of claim 1, wherein mapping the plurality of sound receivers comprises mapping the plurality of sound receiver with respect to the area and with respect to one another.
8. The method of claim 7, further comprising mapping locations of structures within the area.
9. The method of claim 1, wherein at least two sound receivers are provided on different vertical planes from one another.
10. The method of claim 9, further comprising determining a z position of the entity.
11. The method of claim 1, further comprising identifying a room in which the entity is positioned and discarding the sound if the room is not the area.
12. The method of claim 1, wherein the sound receivers are incorporated into household devices including at least one of an electrical outlet, a light switch, a light fixture, a television, a speaker, a microwave, an occupancy sensor, a smoke detector, a thermostat, an air quality monitoring device, and a contact sensor.
13. The method of claim 1, further comprising identifying the entity based on the detected sound.
14. The method of claim 13, wherein when the entity issues a command, the system cross checks whether the entity has permissions for the command.
15. A method for localization of a sound source in an area, the method comprising: mapping a plurality of sound receivers in the area; detecting a sound from the sound source; identifying the sound; determining whether the sound is threat related; and executing security measures if the sound is threat related.
16. The method of claim 15, wherein when the sound is identified as breaking glass, the sound is determined as threat related, and taking security measure comprises calling authorities.
17. The method of claim 15, wherein when the sound is identified as screaming, a call to the area is made to evaluate safety.
18. A system for determining location and orientation of an entity in an area, the system comprising: a plurality of sound receivers for receiving a sound; a memory for storing an audio stream of the sound; a processor capable of: triangulating an x-y position of the entity using the audio stream; and determining orientation of the entity using the audio stream.
19. The system of claim 18, further comprising a sound emitter for outputting a synchronization sound signal.
20. The system of claim 18, wherein the system is integrated into a smart building.
Citation Information
Patent Citations
Audio recognition method and audio recognition device
EP4258264A1
Positioning system for determining a location of an object
US20210141050A1
Devices and methods for 3D position determination
US20230324497A1