Body language detection and microphone control
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- UNIVERSAL CITY STUDIOS LLC
- Filing Date
- 2023-03-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional show attractions struggle to provide accurate and consistent interactions between guests and show elements due to ambient noise interference, limiting the flexibility and effectiveness of guest interactions.
A system incorporating a gimbal-mounted shotgun microphone, a camera, and a processor that identifies the primary human speaker through image or video feed analysis and directs the microphone to focus on that speaker, thereby enhancing interaction by reducing ambient noise.
The system improves interaction quality by isolating the primary speaker's voice from ambient noise, allowing for more accurate command recognition and control of show elements, thereby enhancing the immersive experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] This section is intended to introduce the reader to various aspects of art that may be related to various aspects of the present disclosure, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background to facilitate a better understanding of the various aspects of the present disclosure. As such, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0002] Entertainment venues, such as theme parks, amusement parks, theaters, cinemas, stadiums, and concert halls, have been installed to provide a variety of immersive experiences to an audience of guests. These entertainment venues may include show attractions (e.g., movies, plays, rides, games) that provide an immersive experience to guests. For example, traditional show attractions may allow guests to interact with various show elements of the traditional show attraction. However, it is currently recognized that traditional show attractions are not adequately designed to allow for precise and / or consistent interaction of a particular guest in an audience with various show elements of the traditional show attraction. For example, the range or flexibility of interaction of a particular guest with the show elements of the traditional show attraction may be reduced by ambient noise from the audience. Thus, it is currently recognized that improved interactive show attractions are desirable. Summary of the Invention
[0003] The following summarizes certain embodiments common in scope to the subject matter of the original claims. These embodiments are not intended to limit the scope of the disclosure, but rather to provide merely a brief summary of some disclosed embodiments. Indeed, the disclosure may include a variety of forms that may be similar to or different from the embodiments set forth below.
[0004] In one embodiment, a system includes a gimbal, a shotgun microphone coupled to the gimbal, a camera, and at least one processor configured to receive data indicative of an image or video feed from the camera, determine a location of a primary human speaker based on the image video feed, and actuate the gimbal to point the shotgun microphone toward the location of the primary human speaker.
[0005] In one embodiment, a system includes a microphone assembly, a camera, and at least one processor. The at least one processor is configured to receive data indicative of an image or video feed from the camera and determine a primary human speaker within a group of people and a location of the primary human speaker based on the data indicative of the image or video feed. The at least one processor is also configured to control the microphone assembly based on the location of the primary human speaker.
[0006] In an embodiment, one or more tangible, non-transitory computer readable media include instructions that, when executed by at least one processor, cause the at least one processor to perform various operations. The operations include receiving data from a camera indicative of an image or video feed capturing a group of humans. The operations also include determining a primary human speaker within the group of humans via a body language detection algorithm that receives the data indicative of the image or video feed. The operations also include controlling a motorized gimbal to aim a shotgun microphone at the primary human speaker within the group of humans. The operations also include receiving data via the shotgun microphone indicative of sounds captured by the shotgun microphone. The operations also include determining a command issued by the primary human speaker based on the data indicative of sounds captured by the shotgun microphone. The operations also include controlling show elements based on the command.
[0007] These and other features, aspects, and advantages of the present disclosure will be better understood from the following detailed description when read in conjunction with the accompanying drawings, in which like parts are designated with like numerals throughout. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a schematic diagram of a show attraction including a control assembly configured to identify a primary human speaker in a human group and control a shotgun microphone to point toward the location of the primary human speaker, according to an aspect of the present disclosure. [Diagram 2] 2 is a schematic diagram of the show attraction of FIG. 1 in which the control assembly is configured to identify a new primary human speaker in a human group and control the shotgun microphone to point toward a new location of the new primary human speaker, according to an embodiment of the present disclosure. [Diagram 3] 2 is a schematic perspective view of a portion of the show attraction of FIG. 1, in which a control assembly is configured to control aspects of the show attraction based on commands received from a primary human speaker via a shotgun microphone, in accordance with an embodiment of the present disclosure. [Figure 4] 2 is a schematic diagram of a body language detection algorithm employed in the control assembly of the show attraction of FIG. 1 to identify the primary human speaker, according to an embodiment of the present disclosure. [Diagram 5] FIG. 2 is a process flow diagram illustrating a method for controlling the show attraction of FIG. 1 to identify a primary human speaker, point a shotgun microphone toward the location of the primary human speaker, and control show elements of the show attraction, according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of a show attraction including a control assembly configured to identify a primary human speaker in a human group and select a microphone from a microphone array that corresponds to the primary human speaker, according to an aspect of the present disclosure. [Figure 7] FIG. 7 is a process flow diagram illustrating a method for controlling the show attraction of FIG. 6 to identify a primary human speaker, select a microphone from the microphone array corresponding to the primary human speaker, and control a show element of the show attraction, according to an embodiment of the present disclosure. [Figure 8]FIG. 1 is a schematic diagram of a teleconferencing system including a control assembly configured to identify a primary human speaker in a human group and select a microphone from a microphone array that corresponds to the primary human speaker, in accordance with an aspect of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] One or more specific embodiments are described below. In order to describe these embodiments concisely, not all of the features of the implementations are described herein. It is to be understood that in the development of any such implementation, as in any engineering or design project, numerous implementation-specific decisions must be made to achieve the developer's particular objectives, such as adhering to system-related and business-related constraints that may vary from implementation to implementation. Moreover, it is to be understood that such a development effort may be complex and time-consuming, but would be a routine undertaking of design, fabrication and manufacture for those of ordinary skill in the art having the benefit of this disclosure.
[0010] When introducing elements of various embodiments of the disclosure, the articles "a," "an," and "the" are intended to mean that there are one, two, or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements. It should also be understood that references to "one embodiment" or "an embodiment" of the disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also contain the recited features.
[0011] The present disclosure relates generally to a show attraction system configured to capture commands from a primary human speaker within a group of humans and control show elements of the show attraction based on the commands from the primary human speaker. Specifically, the present disclosure relates to a control assembly that identifies a primary human speaker within a group of humans (e.g., via processing an image or video feed of the group of humans), aims a shotgun microphone at the identified primary human speaker, and controls one or more show elements of the show attraction based on the commands received by the shotgun microphone from the identified primary human speaker.
[0012] An entertainment venue may include various show attractions (e.g., movies, plays, rides, games) that enable interaction between guests participating in the show attraction and various show elements of the show attraction. For example, a guest may issue a command that is received by a microphone of the show attraction, and a control assembly of the show attraction may effect modification of one or more show elements (e.g., physical show props, electronic screens, show lights, etc.) based on the command. Unfortunately, the use of voice commands may be problematic due to ambient noise interfering with the voice commands. For example, a group of guests may participate in a show attraction, and a subset of the group, such as one guest, may be tasked with issuing commands that are captured by the microphone and executed by the show attraction. Because ambient noise is generated by other guests and / or other sources, it may be difficult to separate the commands from the ambient noise.
[0013] According to the present disclosure, a show attraction may employ a camera to capture an image or video feed of a group of guests, hereafter referred to as the human group. A controller may receive data indicative of the image or video feed from the camera and identify a primary human speaker from the human group based on the data indicative of the image or video feed. For example, the human group may converse or otherwise make noise throughout the duration of the show attraction, while one or more humans from the human group may be tasked at various intervals of the show attraction with issuing commands to be performed by the show attraction. The controller may receive data indicative of the image or video feed and execute a body language detection algorithm that identifies a primary human speaker (e.g., a speaker issuing commands) from the human group based on the data. In practice, the body language detection algorithm may be used to detect various body languages indicative of a primary (e.g., active, controlling, or dominant) human speaker, such as facial expressions, hand gestures, head gestures, body postures, body movements, body orientations, and body positions. That is, certain types of body language can indicate a primary (e.g., active, controlling, or dominant) speaker, while other types of body language can indicate a secondary (e.g., passive, subdued, or deferential) speaker or non-speaker.
[0014] In some embodiments, the show attraction may invoke (or otherwise incite) specific body language indicative of a designated primary human speaker to issue a command. For example, the show attraction may instruct or otherwise ask guests to wave or nod, and detect these via a body language detection algorithm executed by the controller based on data indicative of the image or video feed. Additionally or alternatively, the show attraction may instruct or otherwise ask guests to wave a wand or other prop, and detect the prop waving action via a body language detection algorithm executed by the controller based on data indicative of the image or video feed to identify the primary human speaker from a crowd of humans. In addition to identifying the primary human speaker, the controller may also identify the location of the primary human speaker based on data indicative of the image or video feed. For example, various components of the show attraction (e.g., cameras) may include static positions or origins known to the controller and may be configured to enable the controller to determine or infer the location of the primary human speaker within the show attraction.
[0015] After identifying the primary human speaker and the location of the primary human speaker, the controller can control motorized gimbals (or other motion platforms) coupled to various show elements of the show attraction and / or a shotgun microphone to aim the shotgun microphone at the primary human speaker. The controller can initially focus on various show elements of the show attraction and control electronic displays, physical show props, and / or show lights, etc. based on the identification of the primary human speaker and its location. The electronic displays can be controlled to present a digital avatar having certain characteristics, such as, for example, a particular shape, color, size, brightness, or directionality, corresponding to the primary human speaker. The physical show props can be controlled to a position adjacent to the location of the primary human speaker, for example. The show lights can be controlled to shine a spotlight toward the location of the primary human speaker, for example. Other show elements controlled based on the identification of the primary human speaker and the location of the primary human speaker can also be employed in accordance with the present disclosure.
[0016] As described above, the controller also controls a motorized gimbal (or other motion platform) coupled to the shotgun microphone to aim the shotgun microphone at the primary human speaker (also referred to as the primary speaker). As will be appreciated by those skilled in the art, a "shotgun microphone" as used in this disclosure refers to a relatively highly directional type of microphone configured to capture sound within a limited area and / or within a limited directional range. A shotgun microphone can be contrasted with other types of microphones, such as omnidirectional microphones configured to capture sound from multiple directions. Other types of microphones capture sound from multiple directions with relatively similar strength, and therefore tend to capture ambient noise that interferes with the target sound (e.g., commands from the main speaker) that is being captured. Thus, a shotgun microphone can provide the advantage of rejecting ambient noise to better capture the target sound (e.g., commands from the main speaker). Furthermore, acoustic signals captured via other types of microphones may require relatively intensive and expensive processing software to filter ambient noise from the target sound. Thus, according to the present embodiment, signal processing of the acoustic signals captured by the shotgun microphone can be relatively more cost-effective and efficient.
[0017] As a non-limiting example, a shotgun microphone may include a relatively long tube (e.g., a tube having a length of about 6 inches [15.2 centimeters] to about 30 inches [76.2 centimeters]), a group of phase cancelling holes or slits through the tube, and a diaphragm disposed inside the tube. The diaphragm may receive sound from the phase cancelling holes or slits and, in some embodiments, the open front end of the tube. Generally, sounds originating from where the shotgun microphone is pointed are in phase and are captured to a greater extent than sounds originating from other locations that are out of phase and are captured less (or not at all). Thus, shotgun microphones are generally better at rejecting ambient noise than other types of microphones, such as omnidirectional microphones.
[0018] The shotgun microphone may be aimed at the identified primary human speaker, and then capture sounds made by the primary human speaker and communicate data indicative of the captured sounds to the controller. The controller may determine commands issued by the primary human speaker based on the data indicative of the captured sounds. For example, the controller may execute a voice recognition algorithm that receives the data indicative of the captured sounds and determines commands issued by the primary human speaker. Based on the commands, the controller may control one or more show elements of the show attraction, such as electronic screens, physical show props, show lights, or other show elements. As non-limiting examples, the primary human speaker may issue commands to change the color of a digital avatar presented on an electronic screen, to change the position of a physical show prop, or to change the direction, intensity, or color of show lights. Other commands related to control of show elements may also be employed in accordance with the present disclosure.
[0019] In some embodiments, the control of show elements based on the identity of the primary human speaker (or its location) may be combined or sequentially performed with the control of show elements based on commands issued by the primary human speaker. As a non-limiting example, the controller may cause the digital avatar presented on the electronic display to include a shape (e.g., a triangle) corresponding to the identity of the primary human speaker, and cause the digital avatar presented on the electronic display to adjust another characteristic (e.g., change to a larger size) corresponding to a command issued by the primary human speaker. Thus, after execution of a command issued by the primary human speaker, the digital avatar presented on the electronic screen may include a larger sized triangle. The present disclosure may also employ other control features, as described in detail below with reference to the drawings. In general, embodiments of the present disclosure operate to improve interaction between guests participating in a show attraction and show elements of the show attraction, reduce costs or processing complexity associated with this interaction, and / or improve the guest experience through this interaction, as compared to conventional embodiments. These and other features are described in detail below with reference to the drawings.
[0020] FIG. 1 is a schematic diagram of an embodiment of a show attraction 10 including a control assembly 11 configured to identify a primary human speaker within a human group and control a shotgun microphone 12 to orient toward the location of the primary human speaker. In the illustrated embodiment, the control assembly 11 includes, among other features, a shotgun microphone 12 coupled to a gimbal 14 having a motor 16. As will be appreciated in light of the following description, in accordance with the present disclosure, the shotgun microphone 12 is employed to improve the capture of sound from a particular guest within a guest audience by rejecting ambient noise associated with the guest audience or other sources. In general, the shotgun microphone 12 is a relatively highly directional type of microphone configured to capture sound within a limited area and / or within a limited directional range. The shotgun microphone 12 can be contrasted with other types of microphones, such as omnidirectional microphones, that are configured to capture sound from multiple directions. Because other types of microphones, such as omnidirectional microphones, capture sound from multiple directions, they are more likely to capture ambient noise that interferes with the target sound that is being captured. Thus, the shotgun microphone 12 can provide the advantage of rejecting ambient noise to better capture target sounds and not requiring intensive and expensive signal processing compared to other types of microphones that capture high levels of ambient noise.
[0021] As a non-limiting example, the shotgun microphone 12 may include a tube 13 (e.g., a tube having a length of about 6 inches [15.2 centimeters] to about 30 inches [76.2 centimeters]), a group of phase-cancelling holes or slits 15 extending through the tube 13, and a diaphragm 17 disposed inside the tube 13. The diaphragm 17 may receive sound from the phase-cancelling holes or slits 15 and, in some embodiments, from the open front end 19 of the tube 13. Generally, sounds originating from where the shotgun microphone 12 is pointed are in phase and are captured to a greater extent than sounds originating from other locations that are out of phase and are captured less (or not at all). Thus, the shotgun microphone 12 is generally better at rejecting ambient noise than other types of microphones, such as omnidirectional microphones. As described in more detail below, the show attraction 10 includes a control assembly 11 configured to identify a primary human speaker within a human group, direct the shotgun microphone 12 to the primary human speaker, and execute various control features of the show elements of the show attraction 10. Thus, the shotgun microphone 12 captures the sounds produced by the primary human speaker without high levels of ambient noise compared to other types of microphones.
[0022] For example, the show attraction 10 includes a controller 18 having a processing circuit 20, a memory circuit 22, and a communication circuit 24. The memory circuit 22 stores instructions that, when executed by the processing circuit 20, cause the processing circuit 20 to perform various operations, which are described in more detail below. The communication circuit 24 is configured to enable communication between the controller 18 and various other aspects of the show attraction 10. For example, the communication circuit 24 may enable communication between the controller 18 and the motors 16 of the gimbal 14, the shotgun microphone 12 coupled to the gimbal 14, the camera 26 of the control assembly 11, the show lights 28, an additional microphone 30 (e.g., an omnidirectional microphone) of the control assembly 11, a show prop motor 32 of a physical show prop 34 (e.g., a train, trolley, or animated figure), an electronic screen 36, or any combination thereof.
[0023] According to the present disclosure, the camera 26 is configured to provide data indicative of an image or video feed of a staging area 38 of the show attraction 10 to the controller 18. Guests (e.g., a first person 40, a second person 42, a third person 44, a fourth person 46, and a fifth person 48) may be located in the staging area 38 and may be captured in an image or video feed by the camera 26. The first person 40, the second person 42, the third person 44, the fourth person 46, and the fifth person 48 may be collectively referred to as a human group 50. The controller 18 may identify a primary human speaker in the human group 50 based on the data indicative of the image or video feed received from the camera 26. For example, the controller 18 may execute a body language detection algorithm that receives the data indicative of the image or video feed, and the body language detection algorithm identifies a primary human speaker in the human group 50. As will be appreciated in light of the ensuing description, a primary human speaker may be identified through execution of a body language detection algorithm based on detected facial expressions indicative of a primary (or active) human speaker, hand gestures indicative of a primary (or active) human speaker, head gestures indicative of a primary (or active) human speaker, body postures indicative of a primary (or active) human speaker, body movements indicative of a primary (or active) human speaker, body orientations indicative of a primary (or active) human speaker, etc. In practice, certain types of body language may necessarily indicate a primary (e.g., active, controlling, or dominant) speaker, whereas other types of body language (or lack of detected body language) may indicate a secondary (e.g., passive, inhibited, or respectful) speaker or a non-speaker.
[0024] Additionally, in some embodiments, the show attraction 10 may solicit (or otherwise incite) certain body language indicative of the primary human speaker. For example, the show attraction 10 may instruct or otherwise solicit the primary human speaker to wave, wave a wand or other prop, nod, etc., and detect these via body language detection algorithms executed by the controller 18 based on data indicative of the images or video feeds received from the camera 26. Additionally, in some embodiments, the controller 18 may analyze audio captured by the additional microphone 30 (e.g., an omnidirectional microphone) in addition to data indicative of the images or video feeds to facilitate identification of the primary human speaker. That is, while audio captured by the additional microphone 30 may not be utilized to determine commands issued by the primary human speaker, audio captured by the additional microphone 30 may be utilized in combination with data indicative of the images or video feeds to facilitate identification of the primary human speaker.
[0025] In the illustrated embodiment, the controller 18 determines that a third person 44 in the group of people 50 is the primary human speaker based on a hand gesture (e.g., a wave) and locates the location of the primary human speaker relative to the shotgun microphone 12 (or the origin of the shotgun microphone 12). In practice, the camera 26 may be positioned at a stationary location such that the location of the identified primary human speaker relative to the camera 26 is known. Additionally, the location of the shotgun microphone 12 (or the origin of the shotgun microphone 12) may also be known such that the location of the primary human speaker relative to the shotgun microphone 12 (or the origin of the shotgun microphone 12) may be determined, inferred, or interpolated based on the location of the primary human speaker relative to the camera 26. Other location techniques may also be used in accordance with the present disclosure. For example, the performance area 38 may include various unique indicators 39 having known locations distributed around the perimeter of the performance area 38 and captured in an image or video feed by the camera 26. Based on the proximity of a primary human speaker, such as third human speaker 44, to one of these unique indicators 39, the location of the primary human speaker can be determined.
[0026] In accordance with the present disclosure, the controller 18 controls the shotgun microphone 12 to point toward the location of the third person 44 based on the identification of the third person 44 as the primary human speaker and his / her location. The shotgun microphone 12 can be controllable via the gimbal 14 and corresponding motor 16 to rotate in a first circumferential direction 41 about a first axis 43, in a second circumferential direction 45 about a second axis 47, and in a third circumferential direction 49 about a third axis 51. The shotgun microphone 12 can also be controllable via the gimbal 14 and corresponding motor 16 to translate along the first axis 43, the second axis 47, and the third axis 51.
[0027] In some embodiments, multiple instances of motors 16 may be employed to enable movement of shotgun microphone 12 via gimbal 14 as described above. Additionally, gimbal 14 as used herein may refer to any structural mounting structure that may enable translation and / or rotation of shotgun microphone 12 (e.g., via actuation by one or more motors 16). Additionally, shotgun microphone 12, gimbal 14, and motors 16 may be collectively referred to as microphone assembly 53. In general, rotation and / or translation of shotgun microphone 12 via gimbal 14 and corresponding motors 16 as described above may enable shotgun microphone 12 to be aimed anywhere or essentially anywhere within performance area 38. In the illustrated embodiment, shotgun microphone 12 is aimed at third human 44 based on third human 44 being identified as the primary human speaker. Thus, the sounds captured by the shotgun microphone 12 include sounds produced by the third person 44 and exclude (or to a lesser extent include) ambient noise produced by the first person 40, the second person 42, the fourth person 46 and the fifth person 48, and / or other ambient noise sources.
[0028] Before continuing with the discussion of the shotgun microphone 12 capturing commands issued by a primary human speaker (e.g., third human 44 in the illustrated embodiment), it is noted that show elements of the show attraction 10 may be controlled based on the identification of the primary human speaker and its location. For example, the controller 18 may control the electronic screen 36 to present a digital avatar 54 corresponding to the third human 44 being the primary human speaker. In the illustrated embodiment, the triangular shape of the digital avatar 54 corresponds to the third human 44, however, other shapes of the digital avatar 54 may correspond to the first human 40 (e.g., circle), the second human 42 (e.g., rectangle), the fourth human 46 (e.g., star), and the fifth human 58 (e.g., pentagon). Additionally or alternatively, other characteristics of the digital avatar 54 may be controlled to correspond to the identification of the primary human speaker. For example, in some embodiments, the orientation of the digital avatar 54 may be controlled (e.g., so that the digital avatar 54 faces towards the primary human speaker) based on the identity of the primary human speaker. Further, in some embodiments, the digital avatar 54 may include a digital representation of a human, robot, animal or some other living being, which may be presented on the electronic screen 36 to face the location of the primary human speaker.
[0029] Further show elements may be controlled based on the identification of the primary human speaker and its location. For example, the controller 18 may control the show lights 28 to correspond to the identification of the third human speaker 44 as the primary human speaker, such as controlling the show lights 28 so that a spotlight is directed toward the third human speaker 44. Additionally, the controller 18 may control a physical show prop 34 (e.g., a train, trolley, or animated figure) to move the physical show prop 34 (e.g., along a track 52) to a position adjacent the identified primary human speaker as shown.
[0030] 2 illustrates an example in which show elements are controlled differently based on the deviation of the primary human speaker from a third human 44. For example, in FIG. 2, a fourth human 46 from a group of humans 50 is identified as the primary human speaker based on a facial expression, such as a smile, that indicates the primary human speaker. In FIG. 2, based on the identification of the fourth human 46 as the primary human speaker, the controller 18 controls a physical show prop 34 (e.g., a train, trolley, or animated figure) to a position adjacent to the fourth human 46 along a track 52, etc. The controller 18 also controls the electronic screen 36 to display a digital avatar 54 (e.g., having a star shape) corresponding to the fourth human 46 and controls the show lights 28 to spotlight the fourth human 46 based on the identification of the fourth human 46 as the primary human speaker in FIG. 2.
[0031] As shown in both FIG. 1 and FIG. 2, the shotgun microphone 12 is directed to the location of the identified primary human speaker (e.g., the third human 44 in FIG. 1 and the fourth human 46 in FIG. 2). The controller 18 can receive data from the shotgun microphone 12 indicative of the sounds captured by the shotgun microphone 12. As described above, due to the directionality of the shotgun microphone 12, a relatively large portion of the sounds captured by the shotgun microphone 12 can originate from the primary human speaker at which the shotgun microphone 12 is directed, rather than from ambient noise sources (e.g., other people in the human group 50). It is noted that in some embodiments, the directionality of the shotgun microphone 12 can be adjusted based at least in part on the sounds captured by the shotgun microphone 12. For example, the shotgun microphone 12 can be controlled initially based on the identification of the primary human speaker and its location. If the sounds captured by the shotgun microphone 12 include ambient noise with an intensity above the ambient noise threshold, the directionality of the shotgun microphone 12 can be controlled (e.g., adjusted) in stages to identify a location where the ambient noise falls below the ambient noise threshold. As will be appreciated in light of the following description, the controller 18 is capable of determining a command issued by a primary human speaker based on data indicative of sounds captured by the shotgun microphone 12 and then executing that command.
[0032] For example, Figure 3 is a schematic diagram of an embodiment of a portion of the show attraction 10 of Figure 1 in which the control assembly 11 is configured to control aspects of the show attraction 10 based on commands received from a primary human speaker. As described above, the controller 18 of the control assembly 11 can receive a data input 70 indicative of sounds captured by the shotgun microphone 12 of Figures 1 and 2 and execute a voice recognition algorithm to determine or infer a command based on the data input 70. The command can include an instruction to control one of the show elements of the show attraction 10. For example, the command can include an instruction to control a characteristic of the digital avatar 54 presented on the electronic screen 36, such as the size, shape, color, brightness, or directionality of the digital avatar 54, the position or other characteristic of the physical show props 34 (e.g., a train, trolley, or animated figure), and the direction, color, intensity, brightness, or other characteristic of the show lights 28.
[0033] In some embodiments, one or more of the show elements may be controlled based on a combination or sequence of both the identity of the primary human speaker (or its location) and commands issued by the primary human speaker. For example, as described above, the digital avatar 54 presented on the electronic screen 36 of FIG. 1 may include a triangle corresponding to the identity of the third human speaker 44 as the primary human speaker. The commands issued by the primary human speaker (e.g., the third human speaker 44) may include instructions to increase the size of the digital avatar 54 such that the electronic screen 36 presents a larger triangle size 72 corresponding to the digital avatar 54, or instructions to move the position of the digital avatar 54 such that the electronic screen 36 presents a differently positioned triangle 74 corresponding to the digital avatar 54.
[0034] Other modifications (e.g., shape, orientation, directivity, etc.) of the digital avatar 54 based on commands from the primary human speaker may also be employed in accordance with the present disclosure. As another example, the digital avatar 54 may include a digital human, robot, or animal that may be controlled to face the third human 44 based on the identification of the third human 44 as the primary human speaker, and may be controlled to emote or react (e.g., smile, wave, jump, laugh) in response to commands issued by the third human 44 as the primary human speaker. Other show elements of the show attraction 10 may also be controlled based on the identification of the primary human speaker and the commands issued by the primary human speaker. For example, a physical show prop 34 (e.g., a train, trolley, or animated figure) may be moved (e.g., along the track 52) via the show prop motor 32 to a different show prop position 76 based on commands, and a spotlight direction 78 generated by the show lights 28 may be changed based on commands. The combination of control of show elements based on both identification of the primary human speaker (and its location) and commands issued by the primary human speaker can enhance the immersive experience of the human population 50 in the show attraction 10 as compared to conventional systems and methods.
[0035] As discussed above, a primary human speaker may be identified from the human group 50 based on an analysis of the image or video feed via a body language detection algorithm employed by the controller 18. Figure 4 is a schematic diagram of an embodiment of a body language detection algorithm 90 employed in the show attraction 10 of Figure 1 to identify a primary human speaker. In the illustrated embodiment, the body language detection algorithm 90 includes the processing steps of receiving a data input indicative of an image or video feed (block 92) and determining body language characteristics of the human group captured in the image or video feed based on the data input 92 (block 94). For example, the body language characteristics may include each human's facial expression 96, each human's hand gestures 98, each human's head gestures 100, each human's body posture 102, each human's body movement 104, each human's body orientation 106, each human's body position 108, and so forth 110. Body language characteristics can be determined from the data input 70 based on computer vision techniques such as, for example, three-dimensional (3D) pose analysis or estimation, motion analysis or estimation (e.g., tracking and / or optical flow), shape recognition, face recognition and feature extraction.
[0036] The body language detection algorithm 90 then compares the detected body language characteristics as described above with reference body language characteristics (block 95). In the illustrated embodiment, the reference body language characteristics are stored in the database 112. The reference body language characteristics may include only characteristics indicative of a primary (e.g., active, controlling, or dominant) human speaker. Alternatively, the reference body language characteristics may include a first subset of characteristics indicative of a primary human speaker and a second subset of characteristics indicative of a secondary (e.g., passive, inhibited, or respectful) human speaker or non-speaker.
[0037] The body language detection algorithm 90 then identifies the primary human speaker based on a comparison of the detected body language characteristics to the reference body language characteristics (block 114). In embodiments where the reference body language characteristics include characteristics indicative of a primary human speaker and characteristics indicative of a secondary human speaker (or non-speaker), the body language detection algorithm 90 may operate to exclude members within the human population identified as secondary human speakers (or non-speakers). Additionally, the detection algorithm 90 may exclude humans who do not exhibit detectable body language characteristics from identification as a primary human speaker. The body language detection algorithm 90 may also utilize a correspondence (or match) between the detected body language characteristics and the reference body language characteristics indicative of a primary human speaker to identify the primary human speaker.
[0038] The body language detection algorithm 90 may also include aspects of machine learning or artificial intelligence. For example, the body language detection algorithm 90 may include an operation of verifying whether the determination of the primary human speaker was accurate (block 116). In an embodiment, upon detection of a verifiable command issued by the identified primary human speaker, the identification of the primary human speaker may be verified as accurate. Additionally or alternatively, upon failure to detect a verifiable command issued by the identified primary human speaker, the identification of the primary human speaker may be verified as inaccurate. Other verification techniques may also be employed in accordance with the present disclosure. Once it has been verified whether the identification of the primary human speaker was accurate, certain aspects employed in the body language detection algorithm 90 may be updated. For example, any of the data processing, comparison, or analysis in blocks 94, 95, and 114 may be updated. As an example, certain types of body language characteristics detected in block 94 may be added or removed in future iterations of the body language detection algorithm 90. Additionally or alternatively, the database 112 may be updated based on the validation performed in block 116, which stores various reference body language characteristics for comparison with the detected body language characteristics. Updating the body language detection algorithm 90 based on the validation step in block 116 may allow tuning of the body language detection algorithm 90. It should be noted that body language characteristics indicative of a primary human speaker may deviate based on specific cultural or regional expressions, the evolution of body language over time, etc. By allowing updating or modification of the body language detection algorithm 90 in light of the validation step in block 116, the body language detection algorithm 90 may be enhanced and / or adjusted over time to address changes or differences in body language for the reasons discussed above.
[0039] In general, the body language detection algorithm 90 may be employed as part of a broader process employed in the show attraction 10 of FIG. 1 to identify a primary human speaker and control various show elements of the show attraction 10 based at least in part on commands issued by the primary human speaker. For example, FIG. 5 is a process flow diagram illustrating a method 150 of controlling the show attraction 10 of FIG. 1 to identify a primary human speaker, point a shotgun microphone at the location of the primary human speaker, and control show elements of the show attraction. The method 150 illustrated in FIG. 5 includes receiving data indicative of an image or video feed captured of a group of people from a camera via a controller (block 152). For example, the group of people may be present in a presentation area of the show attraction, and the camera may be configured to capture an image or video feed of the presentation area.
[0040] The method 150 also includes determining a primary human speaker (and its location) in the human group via a body language detection algorithm executed by the controller to receive data indicative of the image or video feed (block 154). For example, the controller may execute the body language detection algorithm 90 shown in FIG. 4. In general, the body language detection algorithm is configured to detect various body language characteristics, such as facial expressions, hand gestures, head gestures, body posture, body movements, body orientation, and body position of each member of the human group. As discussed above, certain types of body language may indicate a primary (e.g., active, controlling, or dominant) speaker, while other types of body language may indicate a secondary (e.g., passive, inhibited, or respectful) speaker or a non-speaker. The body language detection algorithm may be configured to compare the detected body language characteristics to baseline body language characteristics to identify the primary human speaker. As discussed above, the controller may also verify the location of the primary human speaker in addition to detecting the primary human speaker.
[0041] The method 150 also includes controlling (block 156) via the controller a show element configuration of the show attraction based on the identity of the primary human speaker, the location of the primary human speaker, or both. For example, the controller may control a particular show element to many different show element configurations. The controller may select a show element configuration from many different show element configurations based on the identity of the primary human speaker, the location of the primary human speaker, or both. As described above, the controller may control electronic displays, physical show props (e.g., trains, trolleys, or animated figures), show lights, or combinations thereof, based on the identification of the primary human speaker and / or its location. In some embodiments, the electronic displays may be controlled to present a digital avatar having characteristics (e.g., size, shape, color, orientation, etc.) corresponding to the primary human speaker or its location.
[0042] As an example, the digital avatar may be controlled to include a first size based on a first guest being identified as the primary human speaker and a second size different from the first size based on a second guest being identified as the primary human speaker. Additionally or alternatively, the controller may control physical show props to a position corresponding to the primary human speaker or its location. Further, the controller may control show lights to shine a spotlight on the primary human speaker. Other controls may be employed in accordance with the present disclosure.
[0043] The method 150 also includes controlling the motorized gimbal via the controller to point the shotgun microphone toward the location of the primary human speaker (block 158). For example, the motorized gimbal can be controlled to rotate and / or translate the shotgun microphone so that the shotgun microphone points toward the location of the primary human speaker. As described above, the shotgun microphone is configured to reject ambient noise while capturing sounds emitted by the primary human speaker. Compared to other types of microphones, such as omnidirectional microphones, the shotgun microphone is good at rejecting ambient noise (e.g., from other members of a human group). By rejecting ambient noise well, the shotgun microphone can transmit an acoustic signal to the controller that does not require intensive and expensive signal processing techniques associated with other types of microphones, such as omnidirectional microphones.
[0044] Method 150 also includes receiving data via the shotgun microphone indicative of sounds captured by the shotgun microphone (block 160). As described above, the sounds captured by the shotgun microphone can correspond primarily to sounds produced by the primary human speaker while filtering out ambient noise and / or being recorded at a relatively low intensity. Method 150 also includes determining, via the controller, a command issued by the primary human speaker based on the data indicative of the sounds captured by the shotgun microphone (block 162). As described above, the controller can execute a voice recognition algorithm to determine a command issued by the primary human speaker based on the data indicative of the sounds captured by the shotgun microphone. In one embodiment, the voice recognition algorithm can detect a language associated with the command. In another embodiment, the language can be known and the voice recognition algorithm can be directed to the known language.
[0045] The method 150 also includes controlling show elements based on the commands via the controller (block 164). For example, electronic screens, physical show props (e.g., a train, trolley, or animated figure), show lights, another show element, or any combination thereof, may be controlled based on commands issued by the primary human speaker. In some embodiments, the control of show elements based on the identity of the primary human speaker or its location (e.g., block 156) and the control of show elements based on commands issued by the primary human speaker (e.g., block 164) may be performed in combination or sequence. For example, the controller may cause a digital avatar presented on an electronic display to include a shape (e.g., a triangle) that corresponds to the identity of the primary human speaker, and cause a digital avatar presented on an electronic display to include another characteristic (e.g., a larger triangle size) that corresponds to commands issued by the primary human speaker. Alternatively, the show element control based on commands issued by the primary human speaker may be completely separate from the show element control based on the identity of the primary human speaker. For example, electronic screens can be controlled to present a digital avatar with certain characteristics based on the identity of the primary human speaker (or its location), while physical show props can be controlled to positions based on commands issued by the primary human speaker.
[0046] 6 is a schematic diagram of an embodiment of a show attraction 10 including a control assembly 11 configured to identify a primary human speaker (e.g., a third human 44) within a group of humans 50 (e.g., a first human 40, a second human 42, a third human 44, a fourth human 44, and a fifth human 48). Additionally, the show attraction 10 of FIG. 6 includes a microphone array 200, referred to in the particular case of this disclosure as a microphone assembly, employing a number of microphones, such as a first microphone 202, a second microphone 204, a third microphone 206, a fourth microphone 208, and a fifth microphone 210, and a controller 18 of the control assembly 11 selects one (or a subset) of the microphones 202, 204, 206, 208, 210 within the microphone array 200 based on the identification of the primary human speaker and / or its location.
[0047] For example, as described above, the controller 18 may determine that the third human 44 is the primary human speaker based on a body language detection algorithm that processes data indicative of an image or video feed received from the camera 26. Various body language such as facial expressions, hand gestures, head gestures, body posture, body movements, body orientation, body position, etc. may be detected and may be indicative of a primary human speaker. Based on the detection of the third human 44 as the primary human speaker and the location of the third human 44, the controller 18 may select the third microphone 206 from the microphone array 200 based on the proximity, position or orientation of the third microphone 206 relative to the third human 44 (or its location). For example, a first microphone 202 may correspond to the location of a first person 40, a second microphone 204 may correspond to the location of a second person 42, a third microphone 206 may correspond to the location of a fourth person 44, a fifth microphone 208 may correspond to the location of a fifth person 46, and a sixth microphone 210 may correspond to the location of a sixth person 48. However, the number of microphones in the microphone array 200 may also differ from the number of people in a human group 50 located in the performance area 38.
[0048] In general, the microphones in the microphone array 200 are selected based on the primary human speaker and / or the location of the primary human speaker. The controller 18 can then receive data from the selected microphones 206 and determine the commands issued by the primary human speaker (e.g., the third human 44) based on the data. In some embodiments, the non-selected microphones 202, 204, 208, 210 can be deactivated, stopped, or otherwise controlled such that the controller 18 does not receive and / or consider data from the non-selected microphones 202, 204, 208, 210. In this manner, ambient or other noise (e.g., noise from the first human 40, the second human 42, the fourth human 46, and the fifth human 48) can be prevented from interfering with the target audio captured by the third microphone 206. Note that in some embodiments, a subset of the microphones 202, 204, 206, 208, 210 can also be selected based on the identity of the primary human speaker. For example, in response to determining that the third human 44 is the primary human speaker, the controller 18 can receive data from the second microphone 204, the third microphone 206, and the fourth microphone 208 while rejecting data from the first microphone 202 and the fifth microphone 210. Additionally, the microphones 202, 204, 206, 208, 210 of the microphone array 200 can be directional or shotgun microphones. Additionally, the microphones 202, 204, 206, 208, 210 can be overhanging or front facing so as to be directed toward the performance area 38 without interfering with the show attraction 10 and the corresponding guest experience. Other aspects of the illustrated show attraction 10 (e.g., the show lights 28, the additional microphones 30, the show props 34 and tracks 52, the electronic screens 36, and the unique indicators 39) can operate the same or similarly to the embodiment shown in FIGS. 1, 2, and 3.
[0049] FIG. 7 is a process flow diagram illustrating an embodiment of a method 250 for controlling the show attraction of FIG. 6 to identify a primary human speaker, select a microphone from the microphone array corresponding to the primary human speaker, and control show elements of the show attraction. In the illustrated embodiment, the method 250 includes receiving data indicative of an image or video feed captured of the human group from a camera via a controller (block 252). Block 252 can be the same as or similar to block 152 of method 150 shown in FIG. 5 and described in detail above. The method 250 also includes determining a primary human speaker (and its location) within the human group via a body language detection algorithm executed by the controller and receiving the data indicative of the image or video feed (block 254). Block 254 can be the same as or similar to block 154 of method 150 shown in FIG. 5 and described in detail above. The method 250 also includes controlling, via the controller, a show element configuration of the show attraction based on the identity of the primary human speaker, the location of the primary human speaker, or both (block 256). Block 256 may be the same as or similar to block 156 of method 150 shown in FIG. 5 and described in detail above.
[0050] Method 250 also includes selecting, via the controller, from the microphone array a microphone (or a subset of microphones) corresponding to the detected primary human speaker (block 258). For example, as described above, the microphone array may have many microphones pointing toward various known locations in a performance area in which a human group is present. The controller may select the microphone corresponding to the primary human speaker based on an identification of the primary human speaker (and its location) within the human group. Other microphones (e.g., non-selected microphones) in the microphone array may be deactivated or otherwise excluded. Thus, the controller receives and processes audio captured by the selected microphone corresponding to the primary human speaker (or its location), which may capture a relatively small amount of ambient noise (e.g., compared to other non-selected microphones in the microphone array).
[0051] Method 250 also includes receiving, at the controller, data from the selected microphone indicative of sounds captured by the selected microphone (block 260). As described above, the sounds captured by the selected microphone can correspond primarily to sounds produced by the primary human speaker while filtering out ambient noise and / or being recorded at a relatively low intensity. Method 250 also includes determining, via the controller, a command issued by the primary human speaker based on the data indicative of the sounds captured by the selected microphone (block 262). As described above, the controller can execute a voice recognition algorithm to determine a command issued by the primary human speaker based on the data indicative of the sounds captured by the selected microphone. Method 250 also includes controlling, via the controller, a show element based on the command (block 264). Block 264 can be the same as or similar to block 164 of method 150 shown in FIG. 5 and described in detail above.
[0052] 8 is a schematic diagram of an embodiment of a teleconferencing system 300 deployed in a venue 302, such as a conference room, meeting room, auditorium, or other venue for holding a conference call or similar event. The teleconferencing system 300 includes a control assembly 304 configured to identify a primary human speaker within a group of people 306 and select a microphone 310 corresponding to the primary human speaker from a microphone array 308. In some instances of the present disclosure, the microphone array 308 may be referred to as a microphone assembly. The control assembly 304 also includes the microphone array 308, a camera assembly 312 (e.g., including one or more cameras) configured to capture video or images within the venue 302, a controllable media device 314 (e.g., a television, a projector, a computer, a speaker, or any combination thereof), and a controller 316.
[0053] The controller 316 includes a processor 318 and a memory 320. The memory 320 includes instructions that, when executed by the processor 318, cause the processor 318 to perform various functions. For example, the controller 316 is configured to receive data from the camera assembly 312 indicative of images or videos captured by the camera assembly 312. Additionally, the controller 316 is configured to execute a body language detection algorithm, as described in detail above, to determine a primary human speaker from the human group 306 within the venue 302. In the illustrated embodiment, the human group 306 includes thirteen humans 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346. A seventh human 334 has been identified as the primary human speaker based on a waving gesture that the controller 316 (e.g., via the body language detection algorithm) has identified as indicative of a primary human speaker.
[0054] In response to determining that the seventh person 334 is the primary human speaker, the controller 316 can select a microphone 310 from the microphone array 308 that corresponds to the seventh person 334 (or its location). In fact, the microphone array 308 in the illustrated embodiment includes twelve microphones 310, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368 positioned at different positions and / or orientations. The controller 316 can select a microphone 310 from the microphone array 308 based on the proximity of the microphone 310 to the primary human speaker (e.g., the seventh person 334) and / or the directionality of the microphone 310 with respect to the primary human speaker. Thus, controller 316 can receive data indicative of audio captured by microphone 310, while rejecting data indicative of audio captured by non-selected microphones 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368. In this manner, data indicative of audio received by controller 316 (e.g., from microphone 310) can emphasize sounds produced by the primary human speaker (e.g., seventh human 334) in contrast to ambient noise produced by other members of human group 306 and / or other sound sources within venue 302.
[0055] In the illustrated embodiment, the number of microphones in the microphone array 308 is different from the number of people in the human group 306. However, the number of microphones in the microphone array 308 can be the same as the number of people in the human group 306. Furthermore, in accordance with the present disclosure, each of the microphones 310, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368 in the microphone array 308 can also be part of a respective larger device, such as a computer, cell phone, or laptop, communicatively coupled to the controller 316. For example, in one embodiment, the selected microphone 310 can be part of a laptop corresponding to the seventh person 334 identified as the primary human speaker. Furthermore, the camera assembly 312 can include a number of cameras associated with each larger device (e.g., computer, cell phone, laptop) corresponding to the human group 306. In this manner, the control assembly 304 can include devices corresponding to the human group 306, each device including a respective microphone of the microphone array 308 and a camera of the camera assembly 312. Additionally, in some embodiments, the teleconferencing system 300 may be used to broadcast a conference call to various devices (e.g., computers, laptops) over a network. Thus, audio captured by a selected microphone 310 (or a subset of the microphones) may be broadcast over the network to various destination devices.
[0056] According to the present disclosure, the controller 316 may be communicatively coupled to the media device 314 and configured to control the media device 314 based on the audio captured by the selected microphone 310. For example, the media device 314 may include a speaker configured to amplify the audio captured by the microphone 310 and received by the controller 316. Additionally, the media device 314 may correspond to a computer or laptop connected to the teleconferencing system 300 from a location remote from the venue 302 (e.g., via a network) and may play the audio captured by the selected microphone 310 through the speaker of the media device 314 at the remote location. Additionally or alternatively, the controller 316 may determine a command issued by a primary human speaker based on the audio captured by the microphone 310. As discussed above, a voice recognition algorithm may be employed to determine the command. In some embodiments, the command may correspond to a desired control of the media device 314. For example, the media device 314 may include a projector or a television on which a presentation is presented. The controller 316 may control the media devices 314 to change slides or pages of the presentation based on commands issued by the primary human speaker, initiate video or other graphics associated with the presentation, or otherwise control the presentation.
[0057] The above-described systems and methods may provide technical advantages over conventional systems and methods. For example, the disclosed show attraction systems and methods may enable improved interaction between a guest audience (e.g., a particular guest within a guest audience) and show elements of a show attraction as compared to conventional systems and methods. Additionally, the disclosed systems and methods may enable improved audio capture of commands from one or more guests within an audience, thereby reducing inaccurate interaction through inaccurate show element control, reducing processing resources (e.g., complexity and cost) required to determine commands issued by one or more guests within an audience (thus improving associated computational operations), and improving the guest experience.
[0058] While only certain features have been illustrated and described herein, many modifications and changes will occur to those skilled in the art and it is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the present disclosure.
[0059] The technology presented and claimed herein refers to and applies to tangible objects and specific examples of a practical nature that positively improve the art and are therefore not abstract, intangible, or purely theoretical. Moreover, if any claim appended at the end of this specification contains one or more elements designated as "means for [performing] ... [function]" or "step for [performing] ... [function]," such elements are to be construed pursuant to 35 U.S.C. 112(f). On the other hand, for any claim containing an element designated in any other manner, such elements are not to be construed pursuant to 35 U.S.C. 112(f). [Explanation of symbols]
[0060] 10. Show Attractions 11 Control Assembly 12 Shotgun Mike 13 Tube 14 Gimbal 15 Holes or slits 16 Motor 17 Diaphragm 18 Controller 19 Front end 20 Processing circuit 22 Memory Circuit 24 Communication Circuits 26 Camera 28 Showlight 30 More Microphones 32 Show Props Motor 34 Physical Show Props 36 Electronic Screen 38 Performance Area 39 Unique Indicators 40 The First Man 41 First circumferential direction 42 The Second Man 43 First Axis 44 The Third Person 45 Second Circumferential Direction 46 The Fourth Man 47 Second Axis 48 The Fifth Man 49 Third Circumferential Direction 50 Human Population 51 The third axis 52 orbit 53 Microphone Assembly 54 Digital Avatar
Claims
1. It is a system, Gimbal and, A shotgun microphone coupled to the gimbal, Camera and, At least one processor, Equipped with, The aforementioned at least one processor is The camera receives first data indicating an image or video feed. Based on the first data showing the image or video feed, the primary human speaker within the human group and the location of the primary human speaker are determined. After determining the location of the primary human speaker, the gimbal is controlled to point the shotgun microphone towards the location of the primary human speaker. The system receives second data indicating the sound captured by the shotgun microphone via the aforementioned shotgun microphone. Based on the second data representing the sound captured by the shotgun microphone, the command uttered by the primary human speaker is determined. The command is executed by controlling the show elements of the system. It is configured in such a way. system.
2. Equipped with an electronic screen configured to output a digital avatar, The aforementioned at least one processor is Based on the identity of the primary human speaker, the location of the primary human speaker, or both, the characteristics of the digital avatar are determined. Control the electronic screen to output the digital avatar having the aforementioned characteristics. It is configured in such a way. The system according to claim 1.
3. The show element comprises the show element which is operable between multiple show element positions, The aforementioned at least one processor is Based on the identity of the primary human speaker, the location of the primary human speaker, or both, a show element location among the plurality of show element locations is determined. The command is executed by controlling the show element to move it to the position of the show element. It is configured in such a way. The system according to claim 1.
4. The system according to claim 1, wherein the at least one processor is configured to execute a body language detection algorithm to identify a hand gesture indicating the primary human speaker based on the first data representing the image or video feed.
5. The system according to claim 1, wherein the at least one processor is configured to execute a body language detection algorithm to identify a facial expression indicating the primary human speaker based on the first data representing the image or video feed.
6. The system according to claim 1, wherein the at least one processor is configured to execute the command by controlling the characteristics of a digital avatar displayed on an electronic screen corresponding to the show element of the system.
7. The system according to claim 1, wherein the show elements include an electronic screen, physical show props, lights, or a combination thereof.
8. The system according to claim 1, wherein the show element includes an electronic screen, and the at least one processor is configured to execute the command by controlling the electronic screen to change the color of a digital avatar presented on the electronic screen.
9. The system according to claim 1, wherein the show element includes a physical show prop, and the at least one processor is configured to execute the command by controlling the physical show prop to change its position.
10. It is a system, Microphone assembly and Camera and, At least one processor, Equipped with, The aforementioned at least one processor is The camera receives first data indicating an image or video feed. Based on the first data showing the image or video feed, the primary human speaker within the human group and the location of the primary human speaker are determined. After determining the location of the primary human speaker, the microphone assembly is controlled to point its microphone towards the location of the primary human speaker. The microphone receives second data indicating the sound captured by the microphone. Based on the second data representing the sound captured by the microphone, the command uttered by the primary human speaker is determined. The command is executed by controlling the show elements of the system. It is configured in such a way. system.
11. The microphone assembly includes a gimbal, The at least one processor is configured to control the microphone assembly so as to point the microphone at the location of the primary human speaker by controlling the gimbal. The system according to claim 10.
12. The gimbal includes a motor, The at least one processor is configured to control the microphone so that it points the microphone towards the location of the primary human speaker by controlling the motor of the gimbal. The system according to claim 11.
13. The microphone assembly includes a microphone array, The aforementioned at least one processor is From the microphone array, identify a first microphone that does not correspond to the location of the primary human speaker. Deactivate the first microphone. It is configured in such a way. The system according to claim 10.
14. The aforementioned at least one processor is From the microphone array, identify the second microphone corresponding to the location of the primary human speaker. The second microphone receives third data indicating further sound captured by the second microphone. It is configured in such a way. The system according to claim 13.
15. Equipped with a speaker, The at least one processor is configured to control the speaker to output audio corresponding to the sound captured by the microphone. The system according to claim 10.
16. Equipped with an electronic screen, The at least one processor is configured to control the visual presentation displayed on the electronic screen based on the second data. The system according to claim 10.
17. One or more tangible, non-temporary computer-readable media containing instructions, wherein the instructions, when executed by at least one processor, Receiving first data from the camera that shows an image or video feed containing a group of people, The first data, which represents the image or video feed, is received via a body language detection algorithm to determine the primary human speaker and the location of the primary human speaker. Controlling the motorized gimbal to point the shotgun microphone at the location of the primary human speaker, The system receives second data indicating the sound captured by the shotgun microphone via the aforementioned shotgun microphone, Based on the second data representing the sound captured by the shotgun microphone, the command uttered by the primary human speaker is determined. Controlling the show elements based on the aforementioned command, One or more tangible, non-temporary computer-readable media that cause at least one processor to perform the above.
18. The instruction causes the at least one processor to control the show element or further show elements based on the identity of the primary human speaker, the location of the primary human speaker, or both, when executed by the at least one processor, one or more tangible non-temporary computer-readable media according to claim 17.
19. When the instruction is executed by the at least one processor, Controlling the first characteristics of the show element based on the identity of the primary human speaker, the location of the primary human speaker, or both thereof, Based on the command, the second characteristic of the show element is controlled so that the show element simultaneously includes the first characteristic and the second characteristic. The one or more tangible non-temporary computer-readable media according to claim 18, wherein the at least one processor performs the following.