AUDIO OUTPUT BASED ON DYNAMIC AUDIO FRAME SELECTION
The data management system synchronizes and selects high-quality audio frames based on metadata to address unsynchronized audio issues in conference systems, enhancing audio quality and usability.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Conference systems face inefficiencies due to overlapping and unsynchronized audio inputs from multiple devices, leading to poor audio quality and user experience, especially when devices are not synchronized or located in the same environment, resulting in feedback, background noise, and echo.
A data management system aligns in-room computing devices to synchronize audio frames and dynamically selects higher-quality frames based on metadata, generating an output that suppresses noise and echo, ensuring clear audio for both in-room and remote participants.
The solution enhances audio quality and usability by selecting and processing audio frames to produce a clear and synchronized output, improving the overall conference experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
GENERAL STATE OF THE ART
[0001] Conference systems (e.g., videoconferencing systems) allow entities to communicate in real time over a network. Such systems can receive audio data from sensors, process the data to generate output, and transmit that same output to output devices. For example, the systems can receive audio data from one device and, based on that audio data, transmit the same output to every other device connected to the system. Because multiple devices can be located in the same place, receiving and transmitting data output in this way can increase resource consumption and negatively impact usability. Furthermore, using such a process can be time-consuming and inefficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Reference numerals may be reused in the drawings to indicate the correspondence between the designated elements. The drawings are provided to illustrate exemplary embodiments described in this document and are not intended to limit the scope of protection of this disclosure. Fig. Figure 1A is a block diagram illustrating an exemplary environment in which a data management system can process data received from and / or to be forwarded to computing devices. Fig. 1B is a block diagram illustrating an example environment in which a data manager of the data management system from Fig. 1A can process data that has been received from and / or is to be forwarded to computing devices. Fig. 2 depicts a general architecture of a computing device or computing system that provides a data management system capable of processing data received from and / or forwarded to computing devices and managing the assignment of roles to entities. Fig. Figure 3 is a pictorial diagram of a sample environment that includes a variety of computing devices, sensors, and output devices. Fig. Figure 4 illustrates a routine for generating output based on a dynamic audio frame selection. DETAILED DESCRIPTION
[0003] In general terms, aspects of this disclosure relate to the efficient processing of audio received by shared devices as part of a consolidated conference where multiple users are in the same environment. Such a conference system may output overlapping audio based on the audio received from the shared devices. In such scenarios, the audio output by a conference system may be ineffective (e.g., unintelligible) due to the overlapping audio received from the shared devices. One approach to audio processing in such scenarios is to deploy fully centralized devices within the environment. However, such centralized devices can be relatively expensive and complex.
[0004] In some cases, such a conference system may output audio received by the room's computer(s) regardless of the quality, quantity, type, etc., of the audio captured by those computer(s). Thus, the output may be based on data received by the computer(s) regardless of whether a speaker is located closer to another computer. Furthermore, the output may be based on data received by one computer regardless of whether that computer captures more background noise, interference, etc., compared to another. In such cases, the output may contain feedback, background noise, echo, interference, crackling, etc., leading to inconsistent meetings in shared conference environments.
[0005] Another problem of particular concern is that the computing devices providing the audio may not be synchronized, and / or the audio itself may not be synchronized. For example, the computing devices may provide audio at different times, not for the same period, and so on. In such cases, such a conference system can be inefficient and / or ineffective.
[0006] Providing such output can cause significant problems for the conference system. For example, while the audio may be complete, intelligible, and / or clear for participants inside the room, the output may be incomplete, unintelligible, and / or unclear for participants outside the room. Such problems can be cascading (cumulative) and / or can reduce the functionality of the conference system (e.g., by making the output incomprehensible to users), leading to an undesirable user experience. For example, the problems can increase cumulatively if additional computing devices are added to the one or more initial computing devices located in the same space (e.g., in-room computing devices for a conference).
[0007] Embodiments of the present disclosure address these problems by providing efficient audio handling in such scenarios without requiring complex, centralized audio processing equipment. In particular, embodiments of the present disclosure enable a data management system to match audio from a variety of computing devices (e.g., in-room computing devices) and dynamically select audio frames (e.g., a user-defined set of audio frames, a variable set of audio frames at runtime, etc.) from the matched audio data to produce an output without including the unselected audio frames in the output. For example, a data management system can dynamically select an audio source from a variety of in-room audio sources and the associated audio for each period of a variety of time periods.
[0008] To enable the dynamic selection of an audio frame, the data management system can align in-room computing devices (e.g., devices located within the room itself). The data management system can align the in-room computing devices by providing an input to each device (e.g., forcing or commanding an alignment or synchronization of the in-room computing devices). For example, the data management system can align the in-room computing devices by providing an input specifying a time period for synchronization, a command to perform synchronization, and so on.
[0009] By aligning the room's internal computing devices, the data management system can obtain audio data from all or some of the internal computing devices, allowing the data management system to select a specific internal computing device (and associated audio data) for inclusion in an output. For example, the data management system can obtain audio data that includes a first audio frame associated with a first computing device (e.g., captured by it), a second audio frame associated with a second computing device, a third audio frame associated with a third computing device, and so on, all associated with the same time period.
[0010] To identify audio frames that belong to the same time period and are eligible for selection within that period, the data management system can align the audio data (e.g., time-align). For example, the data management system can time-align audio frames from all or some of the multiple in-room computing devices (e.g., align the audio frames according to a specific time period). In some cases, to align the audio data for all or some of the multiple in-room computing devices, the data management system can identify timestamps of the audio data received by each device and align the audio data using those timestamps. In some cases, the data management system can filter (e.g., block, ignore, etc.) audio data that belongs to a different timestamp.
[0011] Using the matched audio data, the data management system can compare audio frames assigned to the same time period and select one audio frame from those frames to generate output. For example, the data management system can select a first audio frame, assigned to a first computing device, from the audio data for a first time period, and a second audio frame, assigned to a second computing device, from the audio data for a second time period (e.g., based on the movement of a loudspeaker, or because the first computing device picks up feedback and / or background noise during the second time period, etc.).
[0012] To select the audio frame, the data management system can receive metadata (e.g., generated by the data management system, a variety of in-room computing devices, a separate system, etc.) that is associated with the audio data and can specify the quality of the audio data. It can then use this metadata to select the audio frame. For example, the audio data might be represented as a waveform, and the metadata might specify an amplitude, rate, gap, etc., associated with the waveform. In another example, the metadata might include one or more parameters (e.g., location, volume, pitch, speed, timing, intensity, etc.) associated with the data.
[0013] Because metadata can specify the quality of the audio data (e.g., amplitude, rate, gap, etc.), the data management system can use this metadata to identify and dynamically select a higher-quality audio frame compared to other audio frames from the same time period. For example, the data management system can use the metadata to identify and dynamically select an audio frame predicted to be representative of a given environment over a specific time period. In some cases, the data management system can use the metadata to identify and dynamically select an audio frame with the highest amplitude, rate, gap, etc., compared to other audio frames from the same time period.
[0014] Using the audio frames selected for all or part of the multitude of time periods, the data management system can generate audio data for external computing devices, providing an in-room experience. Specifically, the data management system can nest the selected audio frames to create an audio stream (e.g., a continuous audio stream) that includes, for each time period, one audio frame of higher quality compared to other audio frames of the same time period. The data management system can then provide audio-based output to the external computing devices (e.g., one or more remote computing devices).
[0015] To improve the quality and / or usability of the audio data, which includes dynamically selected audio frames for each time period, the data management system can further process the audio data to generate the output and / or additional outputs. For example, the data management system can identify background noise and / or an echo and suppress and / or cancel it out to produce the output. In another example, the data management system can identify a powered loudspeaker associated with the output and provide an identifier for that loudspeaker.
[0016] As a person skilled in the art will recognize in light of the present disclosure, the embodiments disclosed herein improve the ability of computing systems to enable and / or facilitate a conference (e.g., a conference call, a teleconference, a meeting, a call, a conference event, etc.) between multiple computing devices. Furthermore, the embodiments disclosed herein address technical problems associated with computing systems; in particular, the difficulties in adapting data transmitted between computing devices during a conference. These technical problems are addressed by the various technical solutions described herein, including the use of metadata for the dynamic selection of audio frames. The dynamic selection of audio frames can enable elastic audio (e.g., elastic audio input, elastic audio output, etc.).) by enabling an output and / or input to include the dynamic (e.g., elastic) selection of audio frames. Thus, the present disclosure represents an improvement over existing computing systems in general.
[0017] The foregoing aspects and many of the associated benefits of this revelation will be even more appreciated when they are better understood by reference to the following description in conjunction with the accompanying drawings.
[0018] Fig. Figure 1A is a block diagram of an illustrative operating environment 100A in which first computing devices 101, second computing devices 103, one or more output devices 120, one or more sensors 122, and a media management system 140 can interact with a computing system 110 via a physical connection or a network 104. The computing system 110 can include and / or utilize a variety of connections (e.g., a variety of channels) with the first computing devices 101, the second computing devices 103, the one or more output devices 120, the one or more sensors 122, and / or the media management system 140. For example, the computing system 110 can use a different connection for all or some of the first computing devices 101.
[0019] For illustration, various examples of first computing devices 101 and second computing devices 103 communicating with the computing system 110 are shown, including a desktop computer, a laptop, and a mobile phone. In general, the first computing devices 101 and the second computing devices 103 can be any computing device, such as a desktop, laptop, or tablet computer, a personal computer, a portable computer, a server, a personal digital assistant (PDA), a hybrid PDA / mobile phone, a mobile phone, an electronic book reader, a set-top box, a voice control device, a camera, a digital media playback device, and the like.
[0020] The computer system 110 can equip the first computing devices 101 and the second computing devices 103 with one or more user interfaces, command-line interfaces (CLI), application programming interfaces (API), and / or other programmatic interfaces for participating in a conference. The interfaces can, for example, include an identifier for an active loudspeaker assigned by the computer system 110.
[0021] In some cases, all or some of the first computing devices 101 may be located at a first location (e.g., hardwired, mounted, placed, etc.). The first location may, for example, be a room (e.g., a conference room), and the first computing devices 101 may be hardwired to a table located in that room. All or some of the first computing devices 101 may receive sensor data (e.g., audio data) associated with the first location and / or provide an output.
[0022] In some cases, all or some of the second computing devices 103 may be located at one or more second locations that differ from the first location (e.g., hardwired, attached, placed, etc.). The second computing devices 103 may, for example, be portable user computing devices with a dynamic location (e.g., the location of the second computing devices 103 may change over time). All or some of the second computing devices 103 may receive sensor data associated with the one or more second locations and / or provide an output.
[0023] In some cases, the second computing devices 103 can be located remotely from the first computing devices 101. For example, the first computing devices 101 can be computing devices of entities within the space (e.g., participants within the space), and the second computing devices 103 can be computing devices of entities outside the space (e.g., participants outside the space).
[0024] In some cases, the first computing devices 101 and / or the second computing devices 103 can be dynamic, so that the number of first computing devices 101 and / or the number of second computing devices 103 can change over time (e.g., based on additional entities). For example, the number of first computing devices 101 can change depending on the number of entities within the room for a conference, and the number of second computing devices 103 can change depending on the number of entities remote from the conference.
[0025] In some cases, the Computing System 110 can identify one or more primary locations associated with a conference (e.g., main conference location, central conference location, conference origin, personal conference location, in-room location, etc.) and one or more secondary locations associated with the conference (e.g., remote locations). For example, the Computing System 110 can identify the one or more primary locations based on user input, parsing of a conference agenda, a number of computing devices connecting to the conference via the one or more primary locations (as opposed to one or more secondary locations), and so on.In another example, the computing system 110 can determine that more computing devices are connecting to the conference via a first location compared to computing devices connecting to the conference via one or more second locations, and the computing system 110 can identify the first location as an internal location. Based on the identification of the one or more first locations, the computing system 110 can classify one or more computing devices as first computing devices 101 (computing devices connecting via the one or more first locations) or second computing devices 103 (computing devices connecting via one or more second locations).
[0026] The Media Management System 140 can be any computing system used to implement a communication platform, enabling connections between computing devices. The communication platform can provide computing devices with an interface (e.g., an application programming interface) that allows them to exchange data. For example, the communication platform can allow computing devices to exchange audio data, text data, video data, and so on. In some cases, the communication platform can enable real-time communication.
[0027] The first computing devices 101, the second computing devices 103, the media management system 140, and the computing system 110 can communicate via a network 104, which can include any wired network, wireless network, or a combination thereof. For example, the network 104 can be a personal network, a local area network, a wide area network, a broadcast network (e.g., for radio or television), a cable network, a satellite network, a mobile phone network, or a combination thereof. As another example, the network 104 can be a publicly accessible network of interconnected networks, possibly operated by several different parties, such as the Internet. In some embodiments, the network 104 can be a private or semi-private network, such as a company or university intranet.Network 104 can comprise one or more wireless networks, such as a Global System for Mobile Communications (GSM) network, a Code Division Multiple Access (CDMA) network, a Long Term Evolution (LTE) network, or any other type of wireless network. Network 104 can use protocols and components for communicating over the Internet or any of the other types of networks mentioned above. For example, the protocols used by Network 104 can include Hypertext Transfer Protocol (HTTP), HTTP Secure (HTTPS), Message Queue Telemetry Transport (MQTT), Constrained Application Protocol (CoAP), and the like. Protocols and components for communicating over the Internet or any of the other aforementioned types of communication networks are well known to those skilled in the art and are therefore not described in detail in this document.
[0028] The one or more output devices 120 can include audio output devices (e.g., loudspeakers), image output devices (e.g., displays), etc. The one or more sensors 122 can include audio data sensors (e.g., microphones), image data sensors (e.g., cameras), etc. The one or more output devices 120 and the one or more sensors 122 can be located at a first location. For example, the one or more output devices 120 and the one or more sensors 122 can be located at the same location as the first computing devices 101.
[0029] In some cases, the one or more output devices 120 and the one or more sensors 122 can be located at a specific location (e.g., a first location) (e.g., hardwired, attached, placed, etc.) to obtain sensor data associated with that location. For example, the location can be a room (e.g., a conference room), and the one or more output devices 120 and the one or more sensors 122 can be attached to a wall of the room, to a ceiling of the room, to a table in the room, etc.
[0030] In some cases, the one or more output devices 120 and / or the one or more sensors 122 can be part of the first computing devices 101. For example, the one or more output devices 120 and / or the one or more sensors 122 can be hardware components of the first computing devices 101. In another example, the first computing devices 101 can include the one or more output devices 120, the one or more sensors 122, and / or the computing system 110.
[0031] In Fig. 1A The first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 can provide sensor data (e.g., image data, audio data, etc.) to the computing system 110. The first computing devices 101, the second computing devices 103, and / or the output devices 120 can receive sensor data from the computing system 110. The sensor data can specify an environment of the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122. The sensor data can, for example, specify one or more entities within the environment.
[0032] In some cases, the second computing devices 103 can provide sensor data to the media management system 140, and the media management system can then provide the sensor data to the computing system 110, while the first computing devices 101 and the one or more sensors 122 can provide sensor data directly to the computing system 110. For example, the first computing devices 101 and the one or more sensors 122 can provide sensor data directly to the computing system 110, and the computing system 110 can generate output based on the sensor data and provide the output to the media management system 140 for transmission to the second computing devices 103. In another example, the second computing devices 103 can provide sensor data to the media management system 140, which can then forward an output based on the sensor data to the computing system 110.The computing system 110 can process the output and provide the processed output to the first computing devices 101 and / or the output devices 120.
[0033] In some cases, all or some of the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 may include a respective buffer (e.g., a respective data buffer). All or some of the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 may write data to the buffer and write data from the buffer to a data storage device (e.g., a respective data storage device, a shared data storage device, etc.) and / or the computing system 110. For example, all or some of the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 may write data periodically or aperiodically to the buffer and / or write data from the buffer to the data storage device and / or the computing system 110 (e.g., based on a time period, a size of the data, etc.).
[0034] To enable the processing of sensor data, the computing system 110 includes a data management system 130. In some cases, the computing system 110 may be separate and / or located remotely from the data management system 130. The data management system 130 includes an audio data storage device 132, a metadata storage device 136, and a data manager 131. The data manager 131 can receive and / or store data from all or part of the audio data storage device 132 and / or the metadata storage device 136. The audio data storage device 132 stores audio data 133A and audio data 133B, and the metadata storage device 136 stores metadata 137.
[0035] The audio data 133A and / or the audio data 133B may include a portion of the sensor data provided by the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122. For example, the audio data 133A and / or the audio data 133B may include audio data generated by one or more audio sensors. In some cases, the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 may provide the audio data 133A, and the data management system 130 may generate the audio data 133B (e.g., based on the audio data 133A). In some cases, the audio data 133B may include a portion of the audio data 133A. For example, the data management system 130 can filter the audio data 133A and identify the audio data 133B based on filtering the audio data 133A.
[0036] The metadata 137 can contain data associated with the audio data 133A and / or the audio data 133B. The metadata 137 can be based on the audio data 133A and / or the audio data 133B (e.g., it can be generated based on them). For example, the metadata 137 can be generated based on the shape of a waveform from the audio data 133A. As discussed in this document, the metadata 137 can contain and / or specify one or more parameters, one or more field-value pairs, one or more measurements, one or more ratings, etc.
[0037] In some cases, the metadata may include a measurement of a rate (e.g., a zero-crossing rate). For example, the metadata may specify a rate at which a signal (e.g., the audio data 133A) changes from positive to zero to negative or from negative to zero to positive. The rate may specify a pitch and / or tone associated with the audio data 133A, and the computing system 110 may use the rate to identify a portion of the audio data 133A that corresponds to a particular pitch and / or tone. In some cases, the computing system 110 may perform speech recognition by using the rate to identify a portion of the audio data 133A that corresponds to speech. For example, the computing system 110 may use the rate to identify a portion of the audio data 133A that is predicted to represent an environment (e.g., speech in the environment).
[0038] In some cases, the metadata may include an amplitude (e.g., mean square amplitude, maximum absolute amplitude, etc.). For example, the mean square amplitude may indicate an average amplitude of the audio data 133A, and the maximum absolute amplitude may indicate a maximum amplitude of the audio data 133A. The computing system 110 can use the amplitude to identify a portion of the audio data 133A that is associated with a largest amplitude and / or a largest average amplitude compared to other portions of the audio data 133A. For example, the computing system 110 can use the amplitude to identify a portion of the audio data 133A that is predicted to represent an environment (e.g., speech in the environment).
[0039] In some cases, the metadata may include a gap (e.g., a maximum gap at zero crossing). For example, the maximum gap at zero crossing may specify a difference between a minimum amplitude and a minimum amplitude of audio data 133A. The computing system 110 can use the gap to identify a portion of audio data 133A that corresponds to a gap satisfying a certain threshold. For example, the computing system 110 can use the gap to identify a portion of audio data 133A that has a gap within a threshold range or below a threshold value.In another example, the computing system 110 can use the gap and the amplitude to identify a part of the audio data 133A that has a gap within a threshold range or below a threshold and has an amplitude that is greater than the amplitudes of other parts of the audio data 133A that have a gap within the threshold range or below the threshold.
[0040] In some cases, the data management system 130 can receive the metadata 137 from the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122. For example, a first computing device of the first computing devices 101 can receive audio data 133A, generate metadata 137 based on the audio data 133A, and provide the audio data 133A and the metadata 137 to the data management system 130. In some cases, all or some of the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 can provide a respective part (e.g., a set) of the audio data 133A, respective image data, and / or the metadata 137 to the data management system 130.
[0041] In some cases, the data management system 130 (or a separate system) can generate the metadata 137. For example, the data management system 130 can receive the audio data 133A, analyze the audio data 133A, and generate the metadata 137.
[0042] In some cases, the data management system 130 may include an image data store 134. In some cases, the data management system 130 may not include an image data store 134. The data manager 131 can receive data from and / or store data in the image data store 134. The image data store 134 can store image data 135. In some cases, the metadata 137 may include data associated with the image data 135 (e.g., the metadata 137 may be generated based on the image data 135).
[0043] The image data 135 can include a portion of the sensor data provided by the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122. For example, the image data 135 can include image data generated by one or more image sensors. The image data 135 can include one or more individual images. For example, the image data 135 can include a single image, a sequence of images, a video, etc. In another example, the image data 135 can include one or more images superimposed with audio data.
[0044] In some cases, the data management system 130 can identify the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122. For example, the data management system 130 can identify that the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 are providing data to the data management system 130 (e.g., data that meets a threshold). In another example, the data management system 130 can identify the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122 based on an input (e.g., a registration, a login, etc.) from the first computing devices 101, the second computing devices 103, and / or the one or more sensors 122.In another example, the data management system 130 can determine one or more entities that are associated with the first computing devices 101, the second computing devices 103 and / or the one or more sensors 122 (e.g., that an entity is within a field of view of an associated image sensor) and can identify the first computing devices 101, the second computing devices 103 and / or the one or more sensors 122 based on the determination.
[0045] Since all or some of the first computing devices 101 can capture audio data associated with the same environment, but which may not be synchronized (e.g., operating at different times) to obtain and compare audio frames for the same period captured by all or some of the first computing devices 101, the data manager 131 can identify and synchronize (e.g., synchronize) the first computing devices 101. For example, the data manager 131 can synchronize the first computing devices 101 by instructing (e.g., causing) the first computing devices 101 to flush a respective buffer synchronously (e.g., flush a respective buffer at the same time, simultaneously, concurrently, etc.). In some cases, the data manager 131 can instruct the first computing devices 101 to flush a respective buffer synchronously by providing the first computing devices 101 with an input (e.g.,a buffer instruction and / or time data). For example, the buffer instruction could be an instruction to clear a buffer of the respective computing device.
[0046] In some cases, all or some of the first computing devices 101 may empty a respective buffer out of sequence based on the input. For example, all or some of the first computing devices 101 may empty a respective buffer based on the size of the data stored in that buffer, a time period, etc., and all or some of the first computing devices 101 may empty a respective buffer based on the input before the data size meets a threshold (e.g., exceeds, reaches, etc.) and / or before a time period meets a threshold. In some cases, the input may force all or some of the first computing devices 101 to empty a respective buffer.
[0047] In some cases, the data manager 131 can determine time data for synchronizing the first computing devices 101 and provide the time data to the first computing devices 101. For example, the data manager 131 can generate or receive a timestamp and provide the timestamp to the first computing devices 101. Based on the timestamp, all or some of the first computing devices 101 can timestamp audio data (e.g., using the timestamp) and provide the audio data (e.g., timestamped audio data) to the data manager 131.
[0048] In some cases, the data manager 131 can issue a buffer instruction to the first computing devices 101 to align them. For example, the buffer instruction might include an instruction to flush a corresponding buffer. In response to the buffer instruction, all or some of the first computing devices 101 can flush a corresponding buffer and provide audio data to the data manager 131.
[0049] As discussed in this document, the data manager 131 can obtain audio data 133A from the first computing devices 101 based on the alignment of the first computing devices 101. In some cases, the data manager 131 can obtain metadata from the first computing devices 101 (e.g., with or separately from the audio data 133A).
[0050] Based on the received audio data 133A, the data manager 131 can adjust the audio data 133A (e.g., temporally) and filter a portion of the audio data 133A. The data manager 131 can dynamically select portions of the audio data 133A (e.g., the adjusted audio data) based on the metadata 137 and generate audio data 133B based on this dynamic selection of audio data portions. The data manager 131 can use the metadata 137 to identify, for all or part of a multitude of time periods, a portion of the audio data 133A that exhibits higher quality compared to other portions of the audio data 133A associated with the same time period (e.g., less feedback, higher amplitude indicating a voice, less background noise, etc.).For example, the data manager 131 can use the metadata 137 to identify and dynamically select a portion of the audio data 133A with the largest amplitude compared to other portions of the audio data 133A assigned to the same time period. In another example, the data manager 131 can use the metadata 137 to identify a portion of the audio data 133A that has a gap within or below a threshold range and an amplitude greater than the amplitudes of other portions of the audio data 133A assigned to the same time period that have a gap within or below the threshold range. The data manager 131 can then generate output based on the audio data 133B and provide the output to the output devices 120, the secondary computing devices 103, and / or the media management system 140.
[0051] Fig. Figure 1B is a block diagram of an illustrative operating environment 100B in which a data manager 131 is operated (e.g., the one described in this document with reference to Fig. 1A discussed data administrators 131). As in Fig. As shown in Figure 1B, the data manager 131 comprises and implements a matching component 172, a selection component 174, a multiplexer 176, an active loudspeaker component 178, and a processing component 180. It is understood that the data manager 131 may include more, fewer, or different components. For example, the data manager 131 may not include an active loudspeaker component. In some cases, a separate system or component (e.g., a separate component of the data management system) may include and / or implement one or more of the matching component 172, the selection component 174, the multiplexer 176, the active loudspeaker component 178, and / or the processing component 180.
[0052] The matching component 172 can receive and match the audio data (e.g., based on the matching of the first computing devices 101). To match the audio data, the matching component 172 can determine whether all or part of the audio data meets a threshold. For example, the matching component 172 can determine whether a time value associated with the audio data corresponds to the time data, whether a size value in the audio data corresponds to a size threshold, and so on. Based on determining that part of the audio data does not meet a threshold (e.g., a time value associated with part of the audio data does not correspond to the time data), the matching component 172 can filter out that part of the audio data to produce matched audio data (e.g., filtered audio data). The matching component 172 can then provide the matched audio data to the selection component 174.
[0053] Selection component 174 can dynamically select an audio frame from the audio data (e.g., the matched audio data) for each period within a plurality of time periods. For example, the audio data can contain a plurality of audio frames for each period within the plurality of time periods, and selection component 174 can select an audio frame from the respective plurality of audio frames for each period. Selection component 174 can dynamically select an audio frame by choosing any audio frame for a given period from the plurality of audio frames assigned to that period. Furthermore, each audio frame from the plurality of audio frames can be dynamically selectable for the respective period. Each audio frame within the plurality of audio frames can be assigned to a specific computing device of the first computing devices 101.For example, a given set of audio frames can include a first audio frame, which is assigned to a first time period and is received by a first computing device, a second audio frame, which is assigned to a second time period and is received by a second computing device, and so on.
[0054] Selection component 174 can dynamically select an audio frame based on the respective metadata associated with each of the multiple audio frames. For example, selection component 174 can select an audio frame with the highest amplitude (e.g., as identified by the metadata) compared to other audio frames. In another example, selection component 174 can select an audio frame with the highest amplitude compared to other audio frames, even if a threshold is not met (e.g., the amplitude does not meet a threshold, a gap in the audio frame does not meet the threshold, etc.). In some cases, selection component 174 can select an audio frame based on a comparison of a first set of metadata associated with a first audio frame with a second set of metadata associated with a second audio frame.In some cases, the selection component 174 can select an audio frame based on a comparison of a set of metadata associated with an audio frame against a threshold value.
[0055] Based on the dynamic selection of an audio frame for each period of a multitude of time periods, the selection component 174 can identify and / or generate audio data (e.g., a set of audio frames). For example, the selection component 174 can generate audio data by filtering the audio data (e.g., by filtering out unselected audio frames). The selection component 174 can provide the audio data to the active loudspeaker component 178 and the multiplexer 176.
[0056] The multiplexer 176 can receive the audio data and generate an audio stream based on that data. For example, the multiplexer 176 can generate a continuous audio stream. In some cases, the multiplexer 176 can interweave the audio data within the audio stream to generate it (e.g., by mixing, inserting, etc.). The multiplexer 176 can provide the audio stream to the processing component 180. In some cases, the multiplexer 176 can provide the audio stream as output to a second processing unit.
[0057] In some cases, the processing component 180 can receive the audio stream, process it, and provide an output. For example, the processing component 180 can perform noise reduction (e.g., background noise reduction), gain control (e.g., automatic gain control), and / or echo cancellation. Since the first computing devices 101 may be located in the same environment (e.g., the same room), the audio stream may contain an echo, noise, a weak signal, a strong signal, etc. Therefore, the processing component 180 can process the sensor data by performing noise reduction, gain control, echo cancellation, etc., to improve the user experience. The processing component 180 can provide the processed audio stream as output to a second computing device.
[0058] In some cases, the processing component 180 can process the audio stream based on audio data from all or some of the first computing devices 101. For example, the processing component 180 can use audio data not included within the audio stream to identify an echo, gain, and / or noise, and process the audio stream to suppress the noise, adjust the audio stream according to the gain, and / or cancel the echo. While the audio data may not be included within the audio stream (for example, because a speaker is located closer to another computing device), the audio data not included within the audio stream can indicate an echo, gain, and / or noise, and the processing component 180 can use the audio data to process the audio stream.
[0059] In some cases, the processing component 180 can provide the output and / or indication of the processed audio stream (e.g., an indicator of echo, gain, background noise, etc.) to all or some of the first computing devices 101.
[0060] The active loudspeaker component 178 can receive the metadata and / or audio data. For example, the audio data can include all or part of the audio data provided to the multiplexer 176. In another example, the audio data can include additional audio data received from the first computing devices 101 (e.g., different audio data than that provided to the matching component 172, the selection component 174, and / or the multiplexer 176). In yet another example, the audio data can include audio data received from one or more third computing devices.
[0061] Based on the metadata and / or audio data, the active loudspeaker component 178 can perform active loudspeaker identification to identify an entity (and role) associated with the sensor data. For example, the active loudspeaker component 178 can provide the sensor data to a machine learning model (e.g., implemented by the active loudspeaker component 178) trained to output an identifier of a specific entity and / or computing device associated with the sensor data. The computing system can provide an identifier (e.g., a watermark, border, etc., in relation to image data) of the active loudspeaker to one or more third-party computing devices. For example, the computing system can cause the identifier to be displayed.
[0062] Fig. Figure 2 shows a general architecture of a computer system (referred to as data management system 130) that is operated to manage (e.g., forward) sensor data associated with a conference. The general architecture of the system shown in Fig. The data management system 130, as illustrated in Figure 2, comprises an arrangement of computer hardware modules and software modules that can be used to implement aspects of the present disclosure. For example, aspects of the present disclosure can be implemented by computer hardware modules (e.g., a processor, a processing device, a calculating device, etc.) or by software modules. In some cases, one or more first aspects of the present disclosure can be implemented by computer hardware modules and one or more second aspects of the present disclosure can be implemented by software modules. The hardware modules can be implemented with physical electronic devices. The data management system 130 can include many more (or fewer) elements than those shown in Figure 2. Fig. The elements shown in Figure 2 include those shown. However, it is not necessary for all of these commonly used elements to be shown in order to provide an enabling revelation. Furthermore, the elements shown in Figure 2 may include those shown in Figure 2. Fig. 2 illustrated general architectures can be used to represent one or more of the others in Fig. 1A illustrated components to be implemented.
[0063] As illustrated, the data management system 130 includes a processing unit 290, a network interface 292, a drive 294 for computer-readable media, and an input / output device interface 296, all of which can communicate with each other via a communication bus. The network interface 292 can provide a connection to one or more networks or computing systems. The processing unit 290 can thus receive information and instructions from other computing systems or services via the network 104. The processing unit 290 can also communicate to or from the memory 280 and, furthermore, provide information for an optional display (not shown) via the input / output device interface 296. The input / output device interface 296 can also accept input from an optional input device (not shown).
[0064] The memory 280 can contain computer program instructions (grouped in some embodiments as units) that the processing unit 290 executes to implement one or more aspects of the present disclosure, together with data used to facilitate or support such execution. Although in Fig. While shown as a single set of memory 280, in practice the memory 280 can be divided into levels, such as primary and secondary memory, where these levels may include (among others) RAM, 3D XPOINT memory, flash memory, magnetic storage, and the like. For example, it can be assumed that, for the purposes of this description, primary memory represents the main working memory of the data management system 130, which has a higher speed but a lower overall capacity than secondary, tertiary, and so on.
[0065] Memory 280 can store an operating system 284 that provides computer program instructions for use by the processing unit 290 in the general administration and operation of the data management system 130. Memory 280 can also contain computer program instructions and other information for implementing aspects of this disclosure. For example, in one embodiment, memory 280 includes a data manager 131 for managing the data as described above. Memory 280 also includes audio data 133 and metadata 137. In some cases, memory 280 can include image data 135. The audio data 133, image data 135, and / or metadata 137 can be cached locally in the data management system 130, for example, in the form of a file mapped to memory.For example, the data management system 130 can receive the audio data 133, the image data 135 and / or the metadata 137 and store the audio data 133, the image data 135 and / or the metadata 137 in memory 280.
[0066] The data management system 130 from Fig. Figure 2 is an illustrative configuration of such a device, of which others are possible. For example, although the data management system 130 is shown as a single device, in some embodiments it may be implemented as a logical device hosted by multiple physical host devices. In other embodiments, the data management system 130 may be implemented as one or more virtual devices running on a physical computing device. While in Fig. 2 described as a data management system 130, similar components can be used in some embodiments to manage other components in the environment 100. Fig. 1A to implement the devices shown.
[0067] As discussed above, the first computing devices 101, the second computing devices 103, the output devices 120, and / or the sensors 122 can be deployed to conduct a conference. The first computing devices 101, the output devices 120, and / or the sensors 122 can be located at a first location (e.g., in a conference room) associated with the conference (e.g., an origin point, a home base, a central point, etc., of the conference), and the second computing devices 103 can be located at one or more second locations (e.g., located some distance from the first location). In some cases, the first location can be the location of the computing system 110.
[0068] To illustrate how the first computing devices 101, the output devices 120 and / or the sensors 122 can be located at the first location, is Fig. 3 A pictorial diagram of an example environment 300, which includes a plurality of devices for receiving sensor data and providing an output (e.g., to an entity). The plurality of devices may be similar to the first computing devices 101, the output devices 120, and / or the sensors 122, as described above with reference to Fig. 1A discussed. The multitude of devices can be modular, allowing devices to be removed, added, or modified in real time as needed.
[0069] The multitude of devices located in the vicinity of 300 may include computing devices 302A, 302B, 302C, 302D, 302E, and 302F, computing systems 304A and 304B, output devices 306, and sensors 308 and 310. In some cases, all or some of the computing devices 302A, 302B, 302C, 302D, 302E, and 302F and / or the computing systems 304A and 304B may include one or more output devices and / or sensors. Furthermore, the output device 306 may include one or more sensors, and the sensors 308 and 310 may include one or more output devices.
[0070] As discussed above, all or some of the computing devices 302A, 302B, 302C, 302D, 302E and 302F, the computing systems 304A and 304B, the output devices 306 and the sensors 308 and 310 may be located within the environment 300. For example, the environment 300 may be a conference room. All or part of the computing devices 302A, 302B, 302C, 302D, 302E and 302F, the computing systems 304A and 304B, the output device 306 and the sensors 308 and 310 may be attached to the environment 300, attached to an object (e.g. a wall, a table, a ceiling, etc.) in the environment 300, placed in the environment 300, hardwired within the environment 300, mounted within the environment 300, etc.In some cases, all or part of the computing devices 302A, 302B, 302C, 302D, 302E and 302F, the computing systems 304A and 304B, the output device 306 and the sensors 308 and 310 may be mounted within the environment 300 using one or more hardware brackets, stands, etc.
[0071] As discussed above, all or some of the computing devices 302A, 302B, 302C, 302D, 302E, and 302F can be user computing devices (e.g., tablets). All or some of the computing devices 302A, 302B, 302C, 302D, 302E, and 302F can include one or more sensors (e.g., image sensors, audio sensors, etc.) and / or output devices (e.g., microphones, displays, etc.) to receive sensor data. The environment 300 can include more, fewer, or different computing devices.
[0072] The output device 306 can provide an output. For example, the output device 306 can include an audio output device (e.g., a microphone) for outputting audio signals and / or an image output device (e.g., a display) for outputting images (e.g., still images, videos, etc.). The output device 306 can receive data from one or more of the computing systems 304A and 304B and provide the output. The environment 300 can include more, fewer, or different output devices.
[0073] Sensors 308 and 310 can receive sensor data. For example, sensors 308 and 310 can include an audio sensor to receive audio data and / or an image sensor to receive image data. Sensors 308 and 310 can forward the sensor data to one or more of the computing systems 304A and 304B. Environment 300 can contain more, fewer, or different sensors.
[0074] The 304A and 304B computer systems can be compared with the above by reference to Fig. The discussed computing system 110 may be similar to, and / or include or implement, the computing system 110. In some cases, computing system 304A may be a primary computing system and computing system 304B may be a backup computing system. In some cases, computing system 304A and computing system 304B may perform different functions. For example, computing system 304A can receive, process, and forward data assigned to a first subset of computing devices 302A, 302B, 302C, 302D, 302E, and 302F, output device 306, and / or sensors 308 and 310, and computing system 304B can receive, process, and forward data assigned to a second subset of computing devices 302A, 302B, 302C, 302D, 302E, and 302F, output device 306, and / or sensors 308 and 310. In some cases, computing systems 304A and 304B can access computing resources (e.g.,The environment 300 shares memory to receive, process, and forward data assigned to the computing devices 302A, 302B, 302C, 302D, 302E, and 302F, the output device 306, and / or the sensors 308 and 310. The environment 300 may contain more, fewer, or different computing systems.
[0075] All or some of the computing devices 302A, 302B, 302C, 302D, 302E, and 302F can receive output from the computing systems 304A and 304B and forward sensor data to them. For example, all or some of the computing devices 302A, 302B, 302C, 302D, 302E, and 302F can have a hard-wired connection to the computing systems 304A and 304B. In some cases, the hard-wired connection between the computing devices 302A, 302B, 302C, 302D, 302E and 302F and the computing systems 304A and 304B, the power connections for the computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or the network connections for the computing devices 302A, 302B, 302C, 302D, 302E and 302F can be routed within mounting brackets for the computing devices 302A, 302B, 302C, 302D, 302E and 302F in such a way that cables for the computing devices 302A, 302B, 302C, 302D, 302E and 302F are not clearly visible to the human eye.Therefore, all or some of the computing devices 302A, 302B, 302C, 302D, 302E and 302F can forward sensor data to the computing systems 304A and 304B.
[0076] In some cases, the 302A, 302B, 302C, 302D, 302E, and 302F computing units can be connected to the 304A and 304B computing systems via a hub (e.g., a Universal Serial Bus hub ("USB" hub)). Furthermore, the 302A, 302B, 302C, 302D, 302E, and 302F computing units can be connected to a switch (e.g., a network switch).
[0077] As discussed above, the computing systems 304A and 304B can synchronize the computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or the sensors 308 and 310 and / or the data to be forwarded to the computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or the output device 306. Based on the synchronization, the computing systems 304A and 304B can receive sensor data from the computing devices 302A, 302B, 302C, 302D, 302E, and 302F and / or the sensors 308 and 310, and / or forward an output to the computing devices 302A, 302B, 302C, 302D, 302E, and 302F and / or the output device 306. In some cases, the computing systems 304A and 304B can receive sensor data from one or more sensors of the computing systems 304A and 304B (e.g., from one or more audio sensors of the computing systems 304A and 304B).For example, the sensor data can specify audio within environment 300, an image of environment 300, audio output through computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or output device 306, an image displayed via a display of computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or output device 306 (e.g., screen sharing data).
[0078] The computing systems 304A and 304B can process the sensor data received from the computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or the sensors 308 and 310, and / or the data to be forwarded to the computing devices 302A, 302B, 302C, 302D, 302E and 302F and / or the output device 306. To process the sensor data, the computing systems can filter, normalize, transform, adapt, encode (e.g., using an encoder from the 304A and 304B computing systems), synchronize, and so on. For example, the 304A and 304B computing systems can process the sensor data by filtering it (e.g., based on selecting one audio frame for each time period), performing echo cancellation, automatic gain control, and background noise removal.
[0079] Based on the processing of sensor data, the 304A and 304B computing systems can generate an output (e.g., processed sensor data). In some cases, the output may include a continuous audio stream. The 304A and 304B computing systems can generate the output and provide it to a media management system for transmission to one or more devices (e.g., one or more computing devices, one or more sensors, one or more output devices, etc.). In some embodiments, the 304A and 304B computing systems can provide the output directly to one or more devices (e.g., remote devices). In some cases, the 304A and 304B computing systems cannot provide output to the 302A, 302B, 302C, 302D, 302E, and 302F computing devices and / or the 308 and 310 sensors. The output can be an image output (e.g., an output of a single image, an output of a video, etc.), an audio output, etc.include those assigned to external conference participants.
[0080] In some cases, the output may include instructions for displaying image data and / or outputting audio data. For example, the 304A and 304B computer systems can instruct (e.g., cause) the display of image data and / or the output of audio data.
[0081] As discussed above, a computing system can generate output based on a dynamic selection of audio frames. In some cases, the computing system can dynamically select an audio frame for a given period from a large number of audio frames received from a large number of computing devices and can ignore the unselected audio frames. Referring to Fig. Section 4 describes an illustrative routine 400 for generating output based on dynamic audio frame selection. Routine 400 can be executed, for example, by the computer system 110. Fig. 1A is implemented (which may include, for example, a computing device, data processing hardware, memory, etc.). In some cases, routine 400 can be implemented by a processor. Routine 400 begins in block 402, where the computing system provides a buffer instruction (e.g., to a multitude of first computing devices). For example, the computing system can provide the buffer instruction to a first computing device and a second computing device.
[0082] In some cases, before providing the buffer instruction to the plurality of first computing devices, the computing system can identify a plurality of first computing devices (e.g., including one or more microphones for receiving audio data, one or more displays, one or more speakers, one or more image sensors, etc.). For example, the computing system can identify the plurality of first computing devices as being located within a first environment. In another example, the computing system can identify the plurality of first computing devices as providing data (e.g., audio data) to the computing system. In yet another example, the computing system can identify a plurality of first computing devices that have associated metadata that meets a threshold (e.g., the metadata specifies a rate that is equal to or exceeds a threshold).Based on the identification of the multitude of first computing devices, the computing system can provide the buffer instruction to the multitude of first computing devices.
[0083] In some cases, the computing system can determine initial time data (e.g., a timestamp) and provide the buffer instruction based on this initial time data. For example, the initial time data could specify a time to execute the buffer instruction, a timestamp for stamping audio data, etc. In some cases, the computing system can determine the initial time data (e.g., a timestamp) and provide the initial time data.
[0084] In some cases, in response to receiving the buffer command and / or the initial timing data, all or some of the plurality of first computing devices can flush a respective buffer (e.g., an audio buffer) associated with that particular computing device (e.g., a buffer located on or attached to that particular computing device) based on the buffer command (e.g., flush synchronously). For example, one first computing device of the plurality of first computing devices can flush a first buffer based on the buffer command, and a second computing device of the plurality of first computing devices can flush a second buffer based on the buffer command.
[0085] Based on the provision of the buffer instruction, the computing system receives audio frames and metadata at block 404. In some cases, the computing system may receive the audio frames and metadata as a plurality of sets of audio and a plurality of sets of metadata. For example, the computing system may receive a plurality of sets of audio (e.g., a plurality of first sets of audio) and a plurality of sets of metadata. The computing system may receive a respective set of audio and a respective set of metadata from all or some of the plurality of first computing devices (e.g., in response to the provision of the buffer instruction to all or some of the plurality of first computing devices). For example, the computing system may receive a first set of audio (e.g.,a first set of audio frames) and a first set of metadata are obtained from a first computing device, and a second set of audio data (e.g., a second set of audio frames) and a second set of metadata are obtained from a second computing device.
[0086] All or part of the multitude of audio data sets can contain a multitude of audio frames. Furthermore, all or part of the multitude of audio data sets can contain a respective audio frame for each period of a multitude of time periods.
[0087] The metadata can identify and / or specify one or more rates, gaps, or amplitudes associated with a corresponding audio frame. For example, the metadata can specify the amplitude of the audio frames. In some cases, the metadata can be based on the shape of one or more associated audio frames. In some cases, the computing system can generate the metadata (e.g., based on the received audio data). In other cases, multiple computing devices or a separate system can generate the metadata.
[0088] In some cases, the computing system can determine that the audio data is aligned (e.g., time-aligned within a specific threshold). For example, the computing system can verify the alignment of one or more audio frames associated with a first computing device with one or more audio frames associated with a second computing device. In some cases, the computing devices can determine that one or more audio frames (e.g., one or more audio frames associated with a fourth computing device) are not aligned with one or more other audio frames (e.g., one or more audio frames associated with a first computing device, one or more audio frames associated with a second computing device, etc.).Based on the determination that one or more audio frames are not aligned, the computing system can filter one or more audio frames (e.g., discard, ignore, etc.) (e.g., discard, ignore, etc.).
[0089] Using the received audio frames (e.g., the filtered audio frames), the computing system at block 406 generates audio data based on the metadata. The computing system can generate the audio data based on the audio frames received from the multitude of first computing devices (e.g., a first set of audio frames provided by a first computing device and a second set of audio frames provided by a second computing device).
[0090] In some cases, the computing system used to generate the audio data can, based on the metadata, identify and dynamically select a specific audio frame from a multitude of audio frames for each period within the multitude of time periods. This frame is then assigned to the respective period and to all or part of the multitude of first computing devices. The computing system can use the metadata to identify and dynamically select a specific audio frame for each period within the multitude of time periods that is predicted to have higher quality compared to a multitude of audio frames assigned to the same period (e.g., less feedback, less echo, adjusted gain, less background noise, indication of a voice, etc.).
[0091] In some cases, the computing system used to generate the audio data can compare the metadata (e.g., with other metadata, with a threshold, etc.) to identify and dynamically select a specific audio frame. For example, the computing system can compare a first set of metadata associated with a first audio frame, obtained by a first computing device from a multitude of first computing devices, with a second set of metadata associated with a second audio frame, obtained by a second computing device from the multitude of first computing devices. To compare the metadata, the computing system can compare a first metadata value associated with a first audio frame and a second metadata value associated with a second audio frame to determine which value is greater (e.g.,which indicates a larger amplitude of the first value or the second value).
[0092] In an example of metadata comparison, the computing system can determine that the volume of a first audio frame, associated with a first time period, exceeds the volume of a second audio frame, which is also associated with the first time period. Based on this determination, the computing system can identify and dynamically select the first audio frame and not select the second.
[0093] In another example of metadata comparison, the computing system can determine that a gap associated with the first audio frame (e.g., a difference between two or more values associated with the first audio frame) meets a threshold (e.g., is equal to the threshold, exceeds it, etc.). Based on this determination, the computing system can identify and dynamically select the second audio frame and exclude the first audio frame (which, for example, might be associated with a volume level higher than that of the second audio frame).
[0094] The computing system can determine a set of audio frames (e.g., a third set of audio frames) based on the identification and dynamic selection of audio frames. The set of audio frames can include audio frames provided by two or more computing devices (e.g., the set of audio frames can include one audio frame from a set of audio frames provided by a first computing device and one audio frame from a set of audio frames provided by a second computing device). The computing system can determine the audio data (e.g., a second set of audio data) based on the selected set of audio frames.
[0095] The audio data can contain at least one audio frame received from all or some of the multiple first computing devices. For example, the audio data can contain a first audio frame for a first period received from a first computing device, a second audio frame for a second period received from a second computing device, and so on.
[0096] To provide a representation of the audio data, the computing system at block 408 forwards an output based on the audio data. The computing system can forward the output to one or more secondary computing devices (e.g., to one or more secondary computing devices located away from the multitude of primary computing devices). For example, the computing system can instruct the output of the output device.
[0097] In some cases, the computing system can generate a continuous audio stream based on the audio data and produce output based on that continuous audio stream. For example, the computing system can interlace the audio frames of the audio data into a continuous audio stream and produce output that includes the continuous audio stream. In some cases, the computing system can generate the output using a multiplexer.
[0098] In some cases, the computing system can process one or more audio data points or the continuous audio stream to generate output. For example, the computing system can perform one or more noise reduction, automatic gain control, or echo cancellation operations on the audio data and / or the continuous audio stream and generate the output.
[0099] In some cases, the computing system can identify an active loudspeaker based on audio data (e.g., audio data generated by the computing system, audio data and / or additional data obtained from multiple first computing devices, additional audio data obtained from a third computing device, etc.). The computing system can determine an identifier for the active loudspeaker and forward the identifier to one or more second computing devices. Routine 400 then terminates at block 410.
[0100] In various embodiments, the routine can contain 400 more, fewer, different, or other combinations of blocks than those in Fig. The four illustrated examples include: For instance, in some embodiments, routine 400 may not include forwarding an output based on the audio data. As another example, block 402 may be omitted in some cases, so the data management system does not provide the buffer instruction. The examples in Fig. The illustrated routine 400 is therefore to be understood as illustrative and not restrictive.
[0101] It is understood that not all objects or advantages can necessarily be achieved in accordance with any particular embodiment described in this document. Thus, the person skilled in the art will recognize, for example, that certain embodiments may be configured to be operated in a manner that achieves or optimizes one advantage or group of advantages as taught in this document, without necessarily achieving other objectives or advantages as may be taught or suggested in this document.
[0102] All processes described in this document can be embodied in software code modules and thereby fully automated, including one or more specific computer-executable instructions that are carried out by a computing system. The computing system can include one or more computers or processors. The code modules can be stored on any type of non-transferable, computer-readable medium or other computer storage device. Some or all of the procedures can be embodied in specialized computer hardware.
[0103] Many variations beyond those described herein are evident from this disclosure. For example, depending on the embodiment, certain actions, events, or functions of one of the algorithms described herein may be performed in a different order, added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the execution of the algorithms). Furthermore, in certain embodiments, actions or events may be performed not sequentially but simultaneously, for example, through multithreading, interrupt handling, or multiple processors or processor cores, or on other parallel architectures. Additionally, different tasks or processes may be performed by different machines and / or computing systems that can work together.
[0104] The various illustrative logic blocks and modules described in connection with the embodiments disclosed in this document are implemented or executed by a machine such as a processing unit or processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described in this document. A processor may be a microprocessor; alternatively, the processor may be a controller, a microcontroller, a state machine, combinations thereof, or the like. A processor may include electrical circuits configured to process computer-executable instructions.In another embodiment, a processor includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although primarily described in this document in relation to digital technology, a processor can also primarily include analog components. A computing environment can include any type of computer system, including, but not limited to, a microprocessor-based computer system, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a calculating machine within a device.
[0105] Conditional language, such as "can," "could," "would," or "would like," should, unless explicitly stated otherwise, be understood in the context in which it is generally used to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not. Thus, such conditional language should not generally imply that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for deciding, with or without user input or prompts, whether these features, elements, and / or steps are included or performed in a particular embodiment.
[0106] Disjunctive language, such as the expression "at least one of X, Y, or Z," is to be understood, unless explicitly stated otherwise, in the context in which it is generally used to indicate that an object, concept, etc., can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to express that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z.
[0107] All process descriptions, elements, or blocks in the flowcharts described in this document and / or depicted in the accompanying figures are to be understood as potential representations of modules, segments, or code sections containing one or more executable instructions for implementing specific logical functions or elements in the process. The embodiments described in this document include alternative implementations in which elements or functions can be deleted, executed out of sequence, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved, as would be known to a person skilled in the art.
[0108] Unless explicitly stated otherwise, articles such as "a" or "an" should generally be interpreted as including one or more of the described elements. Similarly, expressions such as "a device configured for this purpose" should include one or more of the named devices. Such one or more named devices may also be configured together to perform the specified recitations. For example, "a processor configured to perform recitation A, B, and C" may include a first processor configured to perform recitation A, in conjunction with a second processor configured to perform recitations B and C.
[0109] Several exemplary embodiments of the disclosure can be described by the following sentences: Sentence 1: A system, encompassing: Data processing hardware; and Memory in communication with the data processing hardware, wherein the memory stores instructions which, when executed on the data processing hardware, cause the data processing hardware to: Identifying a plurality of first computing devices as being located within a first environment; Providing a buffer instruction to the plurality of first computing devices, wherein each computing device of the plurality of first computing devices is configured to synchronously flush a respective buffer based on the buffer instruction; Received from the multitude of first computing devices, a multitude of first sets of audio data and associated metadata based on the buffer instruction, wherein each first set of audio data of the multitude of first sets of audio data comprises a respective audio frame for each period of a multitude of time periods; dynamic selection, for each period of the multitude of time periods, of a respective audio frame from the multitude of first sets of audio data based on the associated metadata; Determining a second set of audio data based on dynamic selection, for each period of the plurality of periods, of the respective audio frame, wherein the second set of audio data comprises at least one respective audio frame obtained from each first computing device of the plurality of first computing devices; Generating a continuous audio stream based on the second set of audio data; Performing one or more noise reduction, automatic gain control, or echo cancellation processes on the continuous audio stream and subsequently generating an output; and Forwarding the output to one or more secondary computing devices located within a secondary environment. Sentence 2: The system according to Sentence 1, wherein the execution of the instructions on the data processing hardware further causes the data processing hardware to do the following: Determining an active loudspeaker based on one or more of the plurality of first sets of audio data or a plurality of third sets of audio data obtained from the plurality of first computing devices; and Forwarding an identifier of the active loudspeaker to one or more secondary computing devices. Sentence 3: The system according to Sentence 1 or Sentence 2, wherein each first computing device of the plurality of first computing devices comprises the following: a microphone to obtain a first set of audio data. Sentence 4: The system according to one of Sentences 1 to 3, wherein the associated metadata specifies one or more of a rate, gap or amplitude that are associated with a respective audio frame of the plurality of first sets of audio data. Sentence 5: A procedure comprising: Providing a buffer instruction to a first computing device and a second computing device to cause the first computing device to empty a first buffer based on the buffer instruction, and to cause the second computing device to empty a second buffer based on the buffer instruction; Received from the first computing device, a first set of audio frames and a first set of metadata following the provision of the buffer instruction to the first computing device, wherein the first set of audio frames comprises, for each period of a plurality of time periods, a respective audio frame of the first set of audio frames; Received from the second computing device, a second set of audio frames and a second set of metadata following the provision of the buffer instruction to the second computing device, wherein the second set of audio frames comprises a respective audio frame of the second set of audio frames for each period of the plurality of time periods; Identify, for each period of the multitude of time periods, a respective audio frame from the first set of audio frames and the second set of audio frames based on the first set of metadata and the second set of metadata; Determining a third set of audio frames based on identifying, for each period of the multitude of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, wherein the third set of audio frames comprises a specific audio frame from the first set of audio frames and a specific audio frame from the second set of audio frames; Generating a continuous audio stream based on the third set of audio frames; and Forwarding an output based on the continuous audio stream to one or more of the first computing device, the second computing device, or a third computing device. Sentence 6: The procedure according to sentence 5, further comprising: Comparing the first set of metadata and the second set of metadata, whereby the identification, for each period of the multitude of time periods, of the respective audio frame from the first set of audio frames and the second set of audio frames is based on the comparison of the first set of metadata and the second set of metadata. Sentence 7: The procedure according to sentence 5 or sentence 6, further comprising: Comparing the first set of metadata and the second set of metadata; and Determine that a volume associated with a first audio frame of the first set of audio frames exceeds a volume associated with a second audio frame of the second set of audio frames, based on comparing the first set of metadata and the second set of metadata, including identifying, for each period of the multitude of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, and identifying the first audio frame based on determining that the volume associated with the first audio frame exceeds the volume associated with the second audio frame. Sentence 8: The procedure according to one of sentences 5 to 7, further comprising: Comparing the first set of metadata and the second set of metadata; Determine that a volume level assigned to a first audio frame of the first set of audio frames exceeds a volume level assigned to a second audio frame of the second set of audio frames, based on a comparison of the first set of metadata and the second set of metadata; and Determining that a gap associated with the first audio frame satisfies a threshold, comprising identifying, for each period of the plurality of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, and identifying the second audio frame based on determining that the gap satisfies the threshold. Sentence 9: The procedure according to one of Sentences 5 to 8, wherein the first set of metadata is based on a form of the first set of audio frames and wherein the second set of metadata is based on a form of the second set of audio frames. Sentence 10: The method according to any of sentences 5 to 9, wherein generating the continuous audio stream includes generating the continuous audio stream using a multiplexer. Sentence 11: The procedure according to one of sentences 5 to 10, further comprising: Determine that the first set of audio frames and the second set of audio frames are aligned. Sentence 12: The procedure according to one of sentences 5 to 11, further comprising: Received from a fourth computing device, a fourth set of audio frames, and a third set of metadata, wherein the fourth set of audio frames comprises, for each period of the plurality of time periods, a respective audio frame of the second set of audio frames; and Determine that the fourth set of audio frames is not aligned with one or more of the first set of audio frames or the second set of audio frames, wherein the generation of the continuous audio stream is further based on determining that the fourth set of audio frames is not aligned with one or more of the first set of audio frames or the second set of audio frames. Sentence 13: The procedure according to one of sentences 5 to 12, further comprising: Determining an active loudspeaker based on the first set of audio frames, the first set of metadata, the second set of audio frames, and the second set of metadata; and Forwarding an identifier of the active loudspeaker to the third computing device. Sentence 14: The procedure according to one of sentences 5 to 13, further comprising: Receiving a fourth set of audio frames from a fourth computing device; Determining an active loudspeaker based on the first set of audio frames, the first set of metadata, the second set of audio frames, and the second set of metadata; and Forwarding an identifier of the active loudspeaker to the third computing device. Sentence 15: Non-transient computer-readable storage medium containing computer-executable instructions which, when executed by a processor, cause the processor to: Providing a buffer instruction to a first computing device and a second computing device; Received from the first computing device, a first audio frame and a first set of metadata following the provision of the buffer instruction to the first computing device; Received from the second computing device, a second audio frame and a second set of metadata following the provision of the buffer instruction to the second computing device; Generating audio data from the first audio frame and the second audio frame, based on the first set of metadata and the second set of metadata, wherein the audio data comprises either the first audio frame or the second audio frame; and Forwarding an output based on the audio data to one or more of the first computing device, the second computing device, or a third computing device. Sentence 16: The non-transient computer-readable medium according to Sentence 15, wherein the first computing device clears a buffer of the first computing device in response to the buffer instruction and wherein the second computing device clears a buffer of the second computing device in response to the buffer instruction. Sentence 17: The non-transient computer-readable medium according to Sentence 15 or Sentence 16, wherein the audio data comprises a continuous audio stream. Sentence 18: The non-transient computer-readable medium according to any one of Sentences 15 to 17, wherein the execution of the computer-executable instructions by the processor further causes the processor to: Receiving a third audio frame from the first computing device; and Receiving a fourth audio frame from the second computing device, wherein the audio data comprise the first audio frame and the fourth audio frame. Sentence 19: The non-transient computer-readable medium according to any one of Sentences 15 to 18, wherein the execution of the computer-executable instructions by the processor further causes the processor to: Receiving a third audio frame from the first computing device; and Receiving a fourth audio frame from the second computing device, wherein, in order to generate the audio data, the execution of the computer-executable instructions by the processor further causes the processor to do the following: Nesting the first audio frame and the fourth audio frame within a continuous audio stream, with the audio data encompassing the continuous audio stream. Sentence 20: The non-transient computer-readable medium according to any of Sentences 15 to 19, wherein the execution of the computer-executable instructions by the processor further causes the processor to: Performing one or more noise reduction, automatic gain control, or echo cancellation processes on the audio data; and Generating the output based on performing one or more of the noise reduction, automatic gain control, or echo cancellation operations on the audio data.
Claims
[1] Procedure, encompassing: Providing a buffer instruction to a first computing device and a second computing device to cause the first computing device to empty a first buffer based on the buffer instruction, and to cause the second computing device to empty a second buffer based on the buffer instruction; Received from the first computing device, a first set of audio frames and a first set of metadata following the provision of the buffer instruction to the first computing device, wherein the first set of audio frames comprises, for each period of a plurality of time periods, a respective audio frame of the first set of audio frames; Received from the second computing device, a second set of audio frames and a second set of metadata following the provision of the buffer instruction to the second computing device, wherein the second set of audio frames comprises a respective audio frame of the second set of audio frames for each period of the plurality of time periods; Identify, for each period of the multitude of time periods, a respective audio frame from the first set of audio frames and the second set of audio frames based on the first set of metadata and the second set of metadata; Determining a third set of audio frames based on identifying, for each period of the multitude of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, wherein the third set of audio frames comprises a specific audio frame from the first set of audio frames and a specific audio frame from the second set of audio frames; Generating a continuous audio stream based on the third set of audio frames; and Forwarding an output based on the continuous audio stream to one or more of the first computing device, the second computing device, or a third computing device. [2] Method according to claim 1, further comprising: Comparing the first set of metadata and the second set of metadata, whereby the identification, for each period of the multitude of time periods, of the respective audio frame from the first set of audio frames and the second set of audio frames is based on the comparison of the first set of metadata and the second set of metadata. [3] The method of claim 2, further comprising: Determine that a volume associated with a first audio frame of the first set of audio frames exceeds a volume associated with a second audio frame of the second set of audio frames, based on comparing the first set of metadata and the second set of metadata, comprising identifying, for each period of the plurality of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, and further identifying the first audio frame based on determining that the volume associated with the first audio frame exceeds the volume associated with the second audio frame. [4] The method of claim 3, further comprising: Determining that a gap associated with the first audio frame satisfies a threshold, comprising identifying, for each period of the plurality of time periods, the respective audio frame from the first set of audio frames and the second set of audio frames, and further identifying the second audio frame based on determining that the gap satisfies the threshold. [5] Method according to any one of claims 1 to 4, wherein the first set of metadata is based on a form of the first set of audio frames and wherein the second set of metadata is based on a form of the second set of audio frames. [6] Method according to any one of claims 1 to 5, wherein generating the continuous audio stream comprises generating the continuous audio stream using a multiplexer. [7] Method according to any one of claims 1 to 6, further comprising: Determine that the first set of audio frames and the second set of audio frames are aligned. [8] Method according to any one of claims 1 to 7, further comprising: Received from a fourth computing device, a fourth set of audio frames, and a third set of metadata, wherein the fourth set of audio frames comprises, for each period of the plurality of time periods, a respective audio frame of the second set of audio frames; and Determine that the fourth set of audio frames is not aligned with one or more of the first set of audio frames or the second set of audio frames, wherein the generation of the continuous audio stream is further based on determining that the fourth set of audio frames is not aligned with one or more of the first set of audio frames or the second set of audio frames. [9] The method of claim 8, further comprising: Selecting an active loudspeaker based on the fourth set of audio frames; and Forwarding an identifier of the active loudspeaker to the third computing device. [10] Method according to any one of claims 1 to 8, further comprising: Determining an active loudspeaker based on the first set of audio frames, the first set of metadata, the second set of audio frames, and the second set of metadata; and Forwarding an identifier of the active loudspeaker to the third computing device. [11] Non-transient computer-readable storage medium containing computer-executable instructions which, when executed by a processor, cause the processor to do the following: Providing a buffer instruction to a first computing device and a second computing device; Received from the first computing device, a first audio frame and a first set of metadata following the provision of the buffer instruction to the first computing device; Received from the second computing device, a second audio frame and a second set of metadata following the provision of the buffer instruction to the second computing device; Generating audio data from the first audio frame and the second audio frame, based on the first set of metadata and the second set of metadata, wherein the audio data comprises either the first audio frame or the second audio frame; and Forwarding an output based on the audio data to one or more of the first computing device, the second computing device, or a third computing device. [12] Non-transient computer-readable medium according to claim 11, wherein the first computing device clears a buffer of the first computing device in response to the buffer command and wherein the second computing device clears a buffer of the second computing device in response to the buffer command. [13] Non-transient computer-readable medium according to claim 11 or claim 12, wherein the execution of the computer-executable instructions by the processor further causes the processor to: Receiving a third audio frame from the first computing device; and Receiving a fourth audio frame from the second computing device, where the audio data includes the first audio frame and the fourth audio frame. [14] Non-volatile computer-readable media according to claim 13, wherein, in order to generate the audio data, the execution of the computer-executable instructions by the processor further causes the processor to do the following: Nesting the first audio frame and the fourth audio frame within a continuous audio stream, with the audio data encompassing the continuous audio stream. [15] Non-transient computer-readable medium according to any one of claims 11 to 14, wherein the execution of the computer-executable instructions by the processor further causes the processor to: Performing one or more noise reduction, automatic gain control, or echo cancellation processes on the audio data; and Generating the output based on performing one or more of the noise reduction, automatic gain control, or echo cancellation operations on the audio data.