Audio device and method of operation thereof

JP2024542696A5Pending Publication Date: 2025-10-01KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024532472
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-02
Filing Date
2022-11-25
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Existing audio distribution systems in decentralized environments, such as medical settings, disaster recovery, and remote meetings, face challenges in providing efficient, flexible, and reliable audio delivery that supports diverse user preferences and maintains privacy, while requiring complex and sub-optimal solutions.

Method used

An audio device that captures audio beams, steers them to different sound sources, analyzes speech characteristics, and adjusts output signals based on speaker categories and user characteristics, allowing selective inclusion or exclusion of audio segments to enhance user experience and privacy.

Benefits of technology

The system provides improved audio distribution by dynamically adjusting audio signals to prioritize relevant speakers and exclude inappropriate content, enhancing user experience and privacy in decentralized scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The audio device includes an audio capture device 201 that forms multiple audio beams and generates an audio capture signal for each audio beam of the multiple audio beams. A beam steering device 203 directs each audio beam toward a different sound source. An analyzer 211 analyzes at least a first audio capture signal to determine speech characteristics of the audio of the first audio capture signal. A classifier 213 determines a speaker category of a plurality of speaker categories for the first audio capture signal according to the speech characteristics. An audio generator 205 generates an audio output signal by combining audio capture signals including the first audio capture signal. An adapter 215 adjusts the first audio output signal according to the first speaker category. For example, some audio of some speaker categories may be fully or partially muted in the audio output signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus and method for generating one or more audio output signals, and in particular, but not exclusively, to generating audio signals for distribution to a remote device. [Background technology]

[0002] Currently, there is an increasing trend to perform many tasks and interactions that were traditionally performed directly by groups of humans in the same location in a more decentralized way. Many of these new decentralized approaches are fundamentally based on the distribution of audio signals between different locations, allowing various participants in different locations to interact and communicate.

[0003] For example, in healthcare environments such as ambulances, wards, operating theatres, etc., multiple healthcare professionals may wish to interact to improve services. In traditional healthcare environments, all participants need to be in the same location to collaborate efficiently. However, increasingly, such scenarios are being realised or facilitated using voice distribution, as a number of people can be located remotely from the patient. For example, in a real operating theatre there are a number of people including the attending physician, nurses, patient etc., but in addition to that, a number of specialists can be located remotely, sometimes far away. For example, technical experts for equipment can be located remotely, or a consultant doctor with very specific expertise who is also located remotely can be involved.

[0004] Another example is in the field of disaster recovery or service / maintenance of various equipment, including computer equipment, where there may be day-to-day users and possibly field service engineers on-site, and various engineers in remote locations who specialize in different aspects of the equipment.

[0005] Thus, many hands-on service sessions may involve a large number of participants, either on-site or at remote locations. These participants may include: Field Service Engineer Local technical staff Local clinical staff Non-technical or non-clinical staff and specialists, administrative staff and / or support staff In some cases, the patient (or family) ·others

[0006] As another example, many conferences and even trials are increasingly being conducted using video conferencing that also includes audio conferencing. In such cases, often many participants (e.g., judges, lawyers, court staff) are in a main conference room or courtroom, while other participants (e.g., defendants or witnesses) may join via appropriate audiovisual conferencing links.

[0007] While such an approach tends to be very advantageous in many situations, it relies heavily on efficient voice delivery. Furthermore, it also raises many new issues and challenges, such as how to support maintaining privacy, efficient targeted communication, practical and preferably low complexity implementation, etc. The priorities and requirements of the voice signal provided to the external source may vary greatly from embodiment to embodiment. There is a great demand to be able to provide voice services and signals that improve the overall experience and can support distributed locations.

[0008] Currently, such scenarios are typically supported by simple audio conferencing solutions where ambient / room audio is picked up by one or more microphones and transmitted to a remote source. However, such basic approaches tend to be suboptimal, and it would be desirable to use more efficient and capable audio distribution, including audio distribution that can provide additional services and / or features that can further enhance or support the distributed user experience.

[0009] Improved approaches would be advantageous in many scenarios, particularly approaches that allow for better operation, more flexibility, less complexity, easier implementation, improved user experience, more reliable or robust operation, reduced computational burden, greater applicability, easier operation, better support for flexible scenarios, better adaptability, and / or improved performance and / or operation. Summary of the Invention

[0010] SUMMARY OF THE DISCLOSURE Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination.

[0011] According to one aspect of the present invention, an audio device is provided, the audio device including: an audio capture device for capturing audio in an environment, the audio capture device forming a plurality of audio beams and generating an audio capture signal for each audio beam of the plurality of audio beams; a beam steering device for directing each audio beam of the plurality of audio beams towards a different audio source; an analyzer for analyzing at least a first audio capture signal to determine speech characteristics of the audio of the first audio capture signal; a classifier for determining a first speaker category from a plurality of speaker categories for a first audio source of the first audio capture signal in response to the speech characteristics; a speech generator for generating a first audio output signal for a first user by combining audio capture signals including the first audio capture signal; and an adapter for adjusting the first audio output signal in response to the first speaker category and the characteristics of the first user.

[0012] The present invention provides an improved audio distribution system and may provide improved support for or even enable many audio-based applications and services. This approach may provide improved support, for example, for including remote participants interacting with people locally present in an environment, specifically, for example, in a room. This approach may provide improved support and / or facilitate many practical applications and services that allow people in different locations to effectively participate in the same activity, and in some cases allow all participants to interact and communicate with each other efficiently. This approach may often provide a user experience that is closer to what would be experienced in a situation where all participants are in the same location. This approach may often provide an improved and differentiated experience, for example, for remote participants.

[0013] This approach may result in adjustment and improved control of the audio provided via the output audio signal in many scenarios: For example, the generated audio may be adjusted to filter out some speakers or parts of the voice that are not appropriate for the remote participant.

[0014] The apparatus may be configured to transmit the audio output signal to a remote device, for example over a network. At least a portion of the speech characteristics may be indicative of cognitive content of the speech of the first audio capture signal.

[0015] In some embodiments, the apparatus may include an output processor for outputting a first audio output signal. The output processor may be configured to provide the first audio signal to a remote device. The output processor may include an audio encoder for encoding the first audio signal and a communication unit for communicating the encoded first audio signal to the remote device. The communication unit may be configured to transmit the encoded first audio signal over a communication channel, such as a communication channel provided by a network.

[0016] According to an optional feature of the invention, the voice generator further generates a different second voice output signal for the second user by combining the voice capture signal including the first voice capture signal, and the adapter individually adjusts the second voice output signal according to the first speaker category and characteristics of the second user.

[0017] This may result in improved performance, operation, and / or user experience in many embodiments and scenarios.

[0018] In some embodiments, the apparatus may include an output processor for outputting a second audio output signal. The output processor may be configured to provide the second signal to a remote device. The output processor may include an audio encoder for encoding the second audio signal and a communication unit for communicating the encoded second audio signal to the remote device. The communication unit may be configured to transmit the encoded first audio signal over a communication channel, such as a communication channel provided by a network.

[0019] According to an optional feature of the invention, the analyzer detects words in the first audio capture signal and determines at least a first one of the speech characteristics in response to the detected words.

[0020] This may result in improved performance, operation, and / or user experience in many embodiments and scenarios, and may result in practical and / or typically more accurate determination of appropriate speaker categories.

[0021] According to an optional feature of the invention, the analyzer determines a first one of the speech characteristics in response to natural language processing (NLP) of the detected words.

[0022] This may improve operation and / or performance in many embodiments and cases.

[0023] According to an optional feature of the invention, the characteristic of the first and / or second user may be an access rights characteristic indicating a degree of permitted access of the user to at least one information category.

[0024] According to an optional feature of the invention, the adapter selects, in response to a first speaker category, which of the plurality of audio capture signals to include in combination to generate a first audio output signal.

[0025] This may improve applications and / or services in many embodiments and cases. This approach can automatically adjust and customize the audio output signal to exclude some speaker categories.

[0026] According to an optional feature of the invention, the audio device further includes a content analyzer for analyzing a segment of the first audio capture signal to determine a content category of the segment from a plurality of content categories, and the adapter adjusts the first audio output signal in response to the content category.

[0027] This may improve applications and / or services in many embodiments and cases. This approach may automatically adjust and customize the audio output signal to filter out some content, e.g., certain sentences / utterances of some speakers, from the audio in the environment. For example, personal information may be filtered out except from speakers authorized to disclose personal information.

[0028] According to an optional feature of the invention, the adapter attenuates segments of the first audio capture signal for at least one combination of the content category and the first speaker category and does not attenuate segments of the first audio capture signal for at least one other combination of the content category and the first speaker category.

[0029] This may improve applications and / or services in many embodiments and cases. This approach allows, for example, the system to attenuate or mute audio content that is desired not to be communicated to the remote participants.

[0030] The adapter may be specifically configured to mute segments of the first audio capture signal for at least one combination of the content category and the first speaker category, and to not mute segments of the first audio capture signal for at least one other combination of the content category and the first speaker category.

[0031] According to an optional feature of the invention, there is a user interface for presenting a representation of the attenuated segments.

[0032] This can result in improved performance and user experience in many embodiments.

[0033] According to an optional feature of the invention, the classifier includes a signature generator for generating a signature of the sound source in response to a frequency distribution of an audio capture signal of the sound source, and a memory unit for storing the signature of the sound source linked to a speaker category determined for the sound source, wherein the signature generator generates a first signature of the first sound source in response to the first sound source being detected, and the classifier matches the first signature with the signatures stored in the memory unit and determines a first speaker category of the first sound source in response to the speaker category linked to the stored signature.

[0034] This may result in improved performance and operation in many embodiments, and in particular may allow for faster and / or more accurate speaker classification.

[0035] According to an optional feature of the invention, the sound capture device detects a new sound source, and the beam steering device, in response to detecting the new sound source, switches the sound beam from being directed towards a previous sound source to being directed towards the new sound source, and selects the previous sound source from the multiple sound sources to which the beam is directed in response to a speaker category of the previous sound source.

[0036] This may result in improved performance and / or operation in many embodiments, in particular improved and / or faster adaptation to changes in active speakers.

[0037] According to an optional feature of the invention, the audio device further includes a detector for detecting an active voice capture signal including a currently active speech signal, and a user interface for presenting an indication of a speaker category assigned to a source of the active voice capture signal.

[0038] This can often result in improved performance and / or user experience.

[0039] According to an optional feature of the invention, the speech generator adjusts at least one combining weight of the speech capture signals in response to the first speaker category.

[0040] This can often result in improved performance and / or user experience.

[0041] According to an optional feature of the invention, the sound capture device generates a variable sound beam, and the beam steering device varies the variable sound beam to detect potential new sound sources and determines whether there is a match between the potential new sound source and a sound source toward which any of the multiple sound beams is directed, the determination of whether there is a match being made in response to a comparison of at least one of characteristics of the variable sound beam and characteristics of a sound beam of the multiple sound beams to characteristics of a sound capture signal of the variable sound beam and characteristics of the sound capture signal of the multiple sound beams, and if no match is detected, steers the sound beam from the previous sound source to the potential new sound source.

[0042] This can often result in improved performance and / or user experience.

[0043] According to one aspect of the present invention, there is provided a method of operating an audio device, the method comprising the steps of capturing sound in an environment by forming a plurality of sound beams and generating a sound capture signal for each sound beam of the plurality of sound beams; directing each sound beam of the plurality of sound beams towards a different sound source; analysing at least a first sound capture signal to determine speech characteristics of the sound of the first sound capture signal; determining a first speaker category from a plurality of speaker categories for a first sound source of the first sound capture signal in response to the speech characteristics; generating a first audio output signal for a first user by combining audio capture signals including the first audio capture signal; and adjusting the first audio output signal in response to the first speaker category and characteristics of the first user.

[0044] These and other aspects, features and advantages of the invention will be elucidated and elucidated with reference to the embodiment(s) described hereinafter. [Brief description of the drawings]

[0045] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Figure 1] Figure 1 shows an example of a usage scenario with audio delivery. [Diagram 2] FIG. 2 illustrates example elements of an apparatus according to some embodiments of the present invention. [Diagram 3] FIG. 3 illustrates an example of a knowledge graph that may be used in an apparatus according to some embodiments of the present invention. [Figure 4] FIG. 4 shows example elements of a classifier for an apparatus according to some embodiments of the present invention. [Diagram 5] FIG. 5 illustrates example elements of a beamformer for an apparatus according to some embodiments of the present invention. [Figure 6] FIG. 6 shows an example of the evolution of a conversation context. [Figure 7] FIG. 7 shows an example approach for detecting changes in conversational context. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0046] FIG. 1 illustrates an example of a configuration and application in which a distributed interaction between people is supported by an audio distribution / communication system. In this example, a group of various people 101 are in a room 103, and an audio distribution / communication apparatus 105 supports communication with one or more remote devices / participants 107. In this example, the audio communication apparatus 105 is configured to communicate with multiple remote audio devices 107, each supporting one or more remote participants. In some cases, the remote audio device 107 may simply have the capability to render and play received audio data, and the capability to capture and encode audio as audio data suitable for transmission to the audio communication apparatus 105. For example, the remote audio device 107 may be a relatively low-complexity conferencing device. In other embodiments, more complex devices may be used, including devices with capabilities such as supporting video communication, presenting data to a user, etc.

[0047] As an example scenario, the room 103 may be an operating room or examination room where a number of different people may be present. These people may have different roles and functions and may include, for example, a patient, a surgeon, one or more nurses, the patient's relatives, technical support staff, etc. The current activity may be further assisted by remote participants, such as, for example, medical professionals, consultant doctors, technical support staff, patient associates, etc.

[0048] Another example scenario could include a remote service application where a number of people may be present, for example to maintain or repair technical equipment in room 103. The people present could include workers performing the day-to-day work, field technical engineers, supervisors, etc. The scenario could be assisted by people in remote locations, for example one or more technical experts for different pieces of equipment, operational support engineers, etc.

[0049] Yet another example could be a courtroom with audio-based links to participants such as the defendant or witnesses.

[0050] In these and similar scenarios, multiple people come together in the same location to collaborate with remote participants, work collaboratively as a group, and often perform complex and sometimes critical tasks that require input and participation from a variety of people with different roles and expertise, aided by audio distribution and communication systems that are key to efficient working and interaction.

[0051] Such scenarios can often be supported by low-complexity traditional audio conferencing systems, but tend to provide suboptimal user experiences and tasks in many situations. Below, approaches are described that typically provide approaches that can improve tasks and user experiences, and typically provide approaches that can improve the resolution and completion of ongoing tasks.

[0052] FIG. 2 illustrates an example of an audio device that may specifically correspond to audio communication device 105 of the example of FIG.

[0053] The audio communication device 105 of FIG. 2 includes an audio capture device 201 configured to capture audio in an environment, in this example, in a room 103 .

[0054] The sound capture device 201 is configured to form multiple sound beams and typically generate a sound capture signal for each sound beam.

[0055] It will be appreciated that various approaches and algorithms are known for generating an audio beam that provides directional capture of audio. For example, typically, directional audio capture may be achieved using a microphone array that includes multiple microphones / audio capture elements, for example, arranged in a line. As is well known to those skilled in the art, directional audio capture may be formed by applying appropriate phase shifts to multiple different microphone signals and combining them. Each capture signal corresponds to an audio beam formed from the microphone array. In this example, the audio capture device 201 may be configured to generate multiple such combined signals, each combined signal being an audio capture signal corresponding to a beam. The audio beam may be dynamically changed by dynamically changing the weight (phase) of each microphone.

[0056] It should be appreciated that other approaches for directional sound capture may be used in other embodiments, for example, in some embodiments multiple directional microphones may each generate a sound capture signal, and mechanical motors may be used to dynamically change the direction of the corresponding beams of the directional microphones.

[0057] The audio communication device 103 further includes a beam steering device 203 configured to control the direction of the audio beams formed by the audio capture device 201. This device is specifically configured to direct each audio beam towards an audio source, typically a different audio source.

[0058] Various techniques are known for steering a sound beam towards a sound source, including both techniques for detecting the sound source and for tracking the sound source after detection. The beam steering unit 203 may use any suitable algorithm, and in particular, in many embodiments, an algorithm that dynamically adjusts the weights of the beamforming combination of captured signals from the microphone array.

[0059] Thus, the sound capture device 201 and the beam steering device 203 may implement the functionality of capturing multiple sound capture signals in multiple sound beams that are directed toward sound sources in the environment, such that the sound capture signal for a particular sound beam typically represents the sound source captured by the beam.

[0060] The sound sources may correspond specifically to people in the room, or may include other sound sources that may possibly be present in the room, for example. In some embodiments, additional circuitry may be included to distinguish and select sound sources that correspond to human speech, for example using complex speech detection, or using less complex detection, for example based on simpler characteristics (e.g., whether the frequency distribution matches that expected from speech, etc.). In other embodiments, the beams may simply be directed, for example, to stronger sound sources, which may be deemed to capture the talker in the room (and it may be acceptable or even advantageous for an application for one or more beams to pick up sound sources other than the talker).

[0061] The audio capture devices 201 are coupled to an audio generator 205 configured to combine multiple audio capture signals from the audio capture devices 201 to generate an audio output signal. In some embodiments, the audio capture signals may be combined by being downmixed to a number of (sub)signals or channels less than the number of audio capture signals. For example, the audio capture signals may be combined to a single mono or stereo signal. The audio generator 205 may then generate an encoded combined signal by encoding the downmix signal. In some embodiments, the downmix signal may be accompanied by parametric upmix data to restore the individual audio capture signals.

[0062] In some embodiments, a combined encoded signal may be generated that includes the separately encoded audio capture signals and in which the encoded data is combined, for example into a single bitstream.

[0063] In this example, the audio generator 205 is coupled to a communication unit 207 configured to communicate with a remote device. The communication unit 207 may use any suitable communication approach to establish a communication link with the remote device 107. In some embodiments, the communication link may be a direct (usually bidirectional) link, such as a direct wireless communication link or a wired link. However, in most embodiments, the link may be formed over a network, which may specifically be a general-purpose network, and in many embodiments may be formed over the Internet. Thus, in many embodiments, the communication unit 207 may include a network interface, e.g., an interface specifically to the Internet, that enables communication with the remote device 107 using a suitable network technology / standard.

[0064] In this example, the audio generator 205 generates an encoded audio output signal and sends it to the communication unit 207, which can transmit the signal over the Internet to one or more remote devices 107, so that the captured audio can be rendered by the remote device 107 and the remote participants can hear the audio from the room 103.

[0065] In this example, the remote devices 107 have the capability to capture local audio and transmit it to the audio communication device 103. For example, each remote device 107 may have a microphone to capture audio from the corresponding participant, encode the audio, and transmit the encoded audio data to the audio communication device 103. The communication unit 207 receives the encoded audio from the remote devices 107 and transmits the encoded audio to a renderer 209 that may be capable of rendering the received audio signal. The rendered audio may be sent to a local sound transducer, such as a speaker, that emits the audio into the room 103.

[0066] Thus, in some embodiments, the system may include the ability to receive and render audio from remote devices / participants in the main environment / room. Thus, two-way audio distribution and sharing is possible. However, it will be appreciated that in other embodiments, one or more of the remote devices 107 may not record or provide audio, and indeed, in some embodiments, the audio communication unit 103 may not have the ability to receive audio from remote devices or render audio. Nevertheless, this approach may be very useful, for example, in cases where remote participants communicate by other means, such as entering text that can be displayed on a display in the room. In other examples, no communication from the remote devices / participants to the room is provided, but the remote participants may, for example, perform actions that aid in the activity in the room (e.g., selectively switching equipment on / off, providing test inputs, or taking control (e.g., controlling a remote load based on audio received from the room 103)).

[0067] In some embodiments and scenarios, the audio broadcast may be complemented by other communications and interactions, such as an accompanying video broadcast that allows remote participants to see what is happening in the room.

[0068] However, in this example, audio communication device 103 is configured to do more than just capture and transmit audio to remote device 107. Audio communication device 103 also has the ability to selectively adjust the output audio, and in particular, the audio that is transmitted to remote device 107.

[0069] The voice capture device 201 is included within the voice communication device 103 of Figure 2 and is coupled to an analyzer 211. The analyzer 211 is configured to analyze one or more voice capture signals to determine voice speech characteristics of the corresponding voice capture signals.

[0070] For example, analyzer 211 may be configured to detect speech content of the voice capture signals, specifically to analyze each voice capture signal and determine speech content characteristics of each voice capture signal. Analyzer 211 may be specifically configured to perform voice recognition to detect spoken words in one or more voice capture signals.

[0071] The analyzer 211 is coupled to a classifier 213 configured to determine a speaker category of the source of the audio capture signal. Typically, the classifier 213 is configured to determine the speaker category for all audio capture signals, or for all audio capture signals that are considered to be speech signals, for example. For example, the classifier 213 may include speech detection criteria that are evaluated to determine whether an individual audio capture signal is considered to correspond to audio from an audio source in the form of a speaker or a non-speaker. Various techniques for detecting speech are known, and any suitable approach may be used. As a low complexity example, the classifier 213 may simply determine whether the audio source captured by a given audio capture signal is speech or not based on the number of words detected per time unit by the speech recognition process of the analyzer 211. The classifier 213 then determines the speech category of each audio capture signal that is considered to have captured speech. Thus, the classifier 213 may have a set of predefined categories. The classifier 213 may select a category from a set of predefined categories for each of one, more, or all of the audio capture signals. Thus, each audio capture signal believed to have captured an audio signal may be assigned a category from a (typically pre-determined) set of categories.

[0072] The classifier 213 is coupled to an adapter 215, which is further coupled to the audio generator 205. The adapter 215 is configured to adjust the audio output signal based on the speaker category determined for the audio capture signal (typically based on the category determined for all audio capture signals). Correspondingly, the audio communication device 103 may be configured to associate speaker categories with the audio capture signals resulting from audio beamforming towards the sound sources in the room. The audio transmitted to the remote device can then be adjusted based on these speaker categories. Thus, the audio provided to the remote participant can be adjusted / modified based on the speaker category.

[0073] As an example, the audio generator 205 can selectively attenuate, and in particular mute, individual audio capture signals or portions thereof based on speech characteristics of the captured audio, in particular based on cognitive speech content determined from detected words, etc., and based on the speaker category determined for that audio beam.

[0074] In some embodiments, the adapter 215 may be configured to select which of the audio capture signals generated by the audio capture device 201 to include in the combination to generate an audio output signal based on the detected speaker category.

[0075] The audio generator 205 may be configured, for example, to exclude some speaker categories from the combined audio signal. Thus, the system can form beams that track individual speakers in the room and, based on the determined speaker categories, generate a combined output audio signal in which the audio capture signals of the beams are included or excluded based on the speaker categories. Thus, the system can automatically adjust the provided audio to include only some speakers and exclude others, such as based on their role in the activity. For example, a remote participant may hear a medical professional in the examination room, but not a patient or a relative.

[0076] In many embodiments, the analyzer 211 may be configured to detect words in the audio capture signal. The analyzer 211 may generate speech characteristics based on these words, such as, for example, the number of words detected, the length of the words, the category into which the words are classified (e.g., by matching the detected words with words stored in memory with associated word categories), etc. In many embodiments, the detected words may be directly used as the speech characteristics provided to the classifier 213, and classification may be performed based on the detected words.

[0077] Many techniques and algorithms are known for speech recognition for determining words and phrases in captured speech, and it will be appreciated that the analyzer 211 may use any suitable approach.

[0078] In many embodiments, the audio communication device 103 may be configured to perform voice recognition, speech-to-text conversion, and subsequent processing based on the generated text. In some cases, modification of the audio capture signal may also be performed based on the text. For example, after the text has been processed and adjusted, a text-to-speech conversion operation may be performed that may generate an audio signal that may specifically be considered to be a modified version of the corresponding audio capture signal. Adjustment and modification of the text, and therefore the corresponding audio capture signal, may be performed based on the speaker category detected for that audio capture signal.

[0079] In many embodiments, the determination of the speaker category based on the detected words may be based on natural language processing (NLP) of the detected words. For example, the words (and phrases) detected by the analyzer 211 may be converted to text and then processed by an NLP process that can determine characteristics of the speech that can be used to determine the speaker category.

[0080] As a specific example, the detected speech from the audio capture signal may be provided to an NLP module of the classifier 213, which may process to determine a speaker category based on an NLP-based classification algorithm. Each speaker category may correspond to a role in the activity, e.g. consultant doctor, surgeon, patient, nurse, relative, technical support, lab technician, etc. in the medical example, or user, service engineer, expert, operator, etc. in the technical service example.

[0081] An NLP-based classification model can be constructed to determine the speaker category, specifically the role, of each participant.

[0082] For example, to create such a model, a base model may be trained using appropriate data representing all possible participants and roles. Preparation / generation of training data may be implemented using a training dataset that includes data points specific to each particular role (e.g., if the profiles are patients and families, a guide / rule may be used that they will never use technical words / phrases when in a procedure / lab, but will mostly use words and phrases related to the disease (sentiments derived from those words or phrases may also be used as additional parameters)). Once such profile-specific data is prepared, a classification model can be built using such data as a training dataset.

[0083] Participant-specific data may be used to train the base model. Participant-specific data may include specific words or phrases (as may be used during conversations in a diagnostic laboratory) that represent categories of participants. Such words, phrases, and sentences are processed using NLP techniques (such as stemming, part-of-speech (POS) tagging, word embeddings, etc.) to extract unique feature vectors. These feature vectors can be used to train machine learning models for participant classification, and neural network architectures may be used to create the models.

[0084] After the model is generated (and appropriately trained / tuned), the voice capture signal data (e.g., in the form of speech features corresponding to detected words / phrases / sentences) is used as input to the model. Typically, the input to the model is in the form of text, where the detected speech may have been converted into appropriate text by the analyzer 211. The model may then classify the source of the voice capture signal into one of the categories / roles.

[0085] For example, the classifier 213 can determine the speaker category by extracting a feature vector from the text (converted from the speech) based on a trained model, which can be used as an input to a model for identifying the participant. If the participant is a patient's family member, the feature vector contains unique values ​​specific to him / her, and the classifier model is trained to identify such unique feature values.

[0086] In some embodiments, the audio communication device 103 may be configured to select the audio capture signals, and therefore the speakers, to be included in the generated output audio signal, as described above. This selection is dynamic and may be modified and changed depending on the current circumstances.

[0087] In some embodiments, the audio communication device 103 has a content analyzer 217 configured to analyze segments of the audio capture signal and determine content categories for those segments. Thus, in addition to the speaker categories being determined for individual beams / audio capture signals / sources / speakers, the audio communication device 103 can further classify individual segments of audio / speech according to appropriate content classes.

[0088] The adapter 215 may be further configured to adjust the audio output signal based on the determined content category.

[0089] In many embodiments, the adapter 215 may be configured to adjust the level of an audio segment, in particular to attenuate the segment, depending on the content category determined for the segment. In many embodiments, the adapter 215 may be configured to control the audio generator 205 to attenuate segments assigned to one content category while not attenuating segments assigned to another category.

[0090] Thus, in such an embodiment, the audio communication device 103 may be configured to determine, for example, particular audio segments that correspond to particular content categories and particular speaker categories, and to attenuate, often mute, such segments entirely.

[0091] For example, in a remote service application, the audio communication device 103 may detect that the remote service engineer may mention certain detailed information that may be sensitive and that the remote service engineer may not be authorized to convey to a third party. However, a supported local client may be able to validly disclose such information, e.g., the client may have authorization to disclose the information. The audio communication device 103 may detect segments classified as corresponding to such restricted disclosure and remove those segments from the audio output signal if they are present in the audio capture signal associated with a speaker category that is not authorized to disclose such information. However, the segments are not removed if the speaker category corresponds to a speaker category that is authorized to disclose the information.

[0092] As another example, in a medical scenario, one of the content categories may be associated with personal information and thus this content category may be assigned to segments representing names, addresses, emails, account numbers, dates (birthdays), etc. The adapter 215 may control the speech generator 205 to remove such personal information from the output stream. For example, a patient may be asked to provide personal data as part of an examination (e.g., dementia examination), but such information is removed from the speech output signal. However, if a specialist or doctor includes a name or phone number (e.g., in a conversation with a relative), this information is included in the speech output signal.

[0093] In many embodiments, a speaker category may correspond to multiple different authority levels for disclosing information. A content category may be associated with multiple different confidentiality levels. Thus, in many embodiments, a speaker category may be associated with a set of different content categories that a speaker of that speaker category has disclosure authority. The adapter 215 may be configured to mute segments assigned to content categories for which the speaker category of the audio capture signal / speaker does not have disclosure authority.

[0094] Various algorithms and approaches for determining content information and for identifying content (e.g., of spoken words and sentences) are well known, and it will be appreciated that one skilled in the art can use any suitable such known approach.

[0095] In particular, in many embodiments, NLP processing can be used to determine the content category, and in fact many embodiments may use the same NLP modules and processing that are used for speaker classification. For example, speech recognition may be performed by converting the detected words to text and applying NLP processing to the resulting text. Examples of suitable NLP processing techniques include tokenization, stemming, word embedding, context embedding, etc.

[0096] More specifically, the NLP module may first perform free speech anonymization. For example, text anonymization techniques may be applied to remove parts of phrases that contain sensitive information, e.g., references to the patient's name, age, etc. This anonymization can have different levels of depth, ranging from removing all identifiers to anonymizing parts of the data. This approach allows implementing different permission levels depending on the application's preferences.

[0097] Detecting words / phrases for de-identification can be achieved by running Named Entity Recongition (NER) on the text / sentence and classifying named entities / words into predefined categories such as people, organizations, places, etc.

[0098] Such NER methods may not be ideally trained and thus may misclassify important data (e.g., utterances that are important and related to diagnostics or services, etc., and should be conveyed unmuted). In some embodiments, this problem may be mitigated by using a knowledge graph (including relationships between nodes) specific to the technical conversation between the local and remote participants. An example of such a graph is shown in FIG. 3. After NER identification, the identified words of the segments to be excluded may be further checked using such a knowledge graph and, if found as nodes in the graph, the NER identification may be modified. This may reduce the risk of important data / information being lost during the de-identification process.

[0099] After anonymization, i.e., after segments to be removed have been identified, the audio capture signal segments (e.g., audio signals / waveforms) for the segments to be included in the audio output signal may be assembled and combined to form a continuous audio stream. The resulting audio output signal may be transmitted to a remote device / participant 107.

[0100] For example, speech recognition can detect and divide the voice capture signal into voice fragments, which are represented by text sent to the NLP module. Each segment can be assigned a speaker category and a content category. Each segment can be associated with, for example, a timestamp. When sorted by timestamp, the phrases form a chat room-like dialogue captured in the voice phrases. The segments can be, for example, individual words, or individual words within a segment can be indicated by timestamps identifying the beginning and / or end of the word. The generation of the output signal by combining the voice fragments / segments, possibly removing some fragments / segments, can be based on the timestamps. Essentially, all voice fragments corresponding to segments that are intended to be included (e.g., not containing sensitive information or from participants with appropriate authorization) are included in the combination. However, for time instances corresponding to segments identified as not to be included, the corresponding voice fragments are not included in the combination. For example, in the time intervals corresponding to these segments, the voice capture signal of the corresponding beam can be replaced by another voice signal or simply set to a zero signal (e.g., a silence signal with zero amplitude). In some embodiments, the audio capture signal of a muted segment may be replaced with, for example, a (e.g., predetermined) default audio clip / signal. For example, a white noise or tone signal may be used to replace the audio capture signal during a segment in which the audio capture signal contains content that should not be included in the audio output signal.

[0101] In some embodiments, the audio communication device 103 may be configured to include only the strongest audio capture signal (e.g., currently highest level / amplitude) in the output audio signal. The audio generator 205 may receive the audio capture signal from the targeted beam along with the beam's identity and the beam's speaker category. The audio generator 205 may then determine the speaker category of the strongest audio capture signal, and therefore specifically the role of the currently strongest sound source / speaker. Depending on the category, the adapter 215 may control the audio generator 205 to include this audio in the audio output signal or whether to exclude this audio from transmission to the remote participant.

[0102] In many embodiments, the audio generator 205 may be configured to generate multiple different audio output signals. In particular, the audio generator 205 may generate different audio output signals for different remote devices / participants 107. In particular, the audio generator 205 may be configured to generate an output audio signal for a first remote participant and a second different output audio signal for a second participant. The signals may then be transmitted to the corresponding different remote devices 107. The generation of the different audio output signals may be performed, for example, by obtaining two output audio signals by including different audio capture signals, in particular different segments / audio fragments, in the downmix.

[0103] The adapter 215 may be configured to adjust the audio output signal for different users / remote participants individually. This adjustment may be made depending on the speaker categories and the characteristics of the users. For example, a level of authority or permission may be associated with the users, and the audio output signal provided to each user may depend on this level. For example, a participant may have a high level of permission to hear all conversations in a room (e.g., a court reporter in a room at a court), while another participant may only have permission to hear basic information only (e.g., a member of the public attending a trial out of general interest). The adapter 215 may, for example, allow one remote participant to hear audio from all participants, and therefore include all audio capture signals in the output audio signal. However, for another user with a lower level of permission, only some speaker categories may be allowed, and audio capture signals of speakers identified as belonging to another speaker category may be muted. For example, a court reporter may be able to hear all audio, while a member of the public may only hear some speakers, and not, for example, a witness under protection.

[0104] In many embodiments, the characteristic of the user used to adjust the audio output signal may be an access privilege characteristic that indicates the degree of access the user is permitted to at least one information category.

[0105] In some embodiments, the adapter 215 may be further configured to individually adjust the audio output signal depending on the content category of the audio segments. For example, some remote participants may have permission to hear personal information while other remote participants may not have permission to hear such information. In this case, audio segments containing personal information may be removed from the audio output signal for remote participants in the second category but not from the audio output signal for remote participants in the first category.

[0106] It will be appreciated that the adapter 215 may implement the adjustments in any suitable manner. For example, in some embodiments, a fixed rule-based approach may be used, where the rules may define which content category and speaker category combinations should be included in the audio output signal, and which combinations should be excluded from the audio output signal. There may be separate rules for each characteristic of a user / remote participant. For example, a set of user / remote participant categories may be defined, and a specific set of rules may be implemented for each category. A remote participant may be associated with a particular category (e.g., set by the user or by an operator / controller of the session), and the rules for that particular category may be used by the adapter 215 when generating an audio output signal for that user.

[0107] In many embodiments, the audio communication device 103 may be configured to employ a variety of techniques to dynamically adjust and modify the audio source, and specifically the speaker, that is being tracked.

[0108] In many embodiments, the audio communication device 103 may be capable of detecting audio sources that are not the source of an audio capture signal / beam, i.e., may detect new audio sources that are not currently tracked by any beam. Specifically, the audio capture device 201 may detect new audio sources that are not currently associated with any audio capture signal or formed beam.

[0109] As an example, the sound capture device 201 may generate a variable sound beam that can be controlled to change direction under the control of the beam steering device 203 to look for new sound sources. Such a sound beam may also be referred to as a search beam or a variable beam, while a sound beam that tracks a sound source may also be referred to as a targeted beam.

[0110] The sound capture device 201 may generate a rotating or moving beam that may pause if a strong sound signal is detected. In such a case, the sound capture signal of the rotating beam may be correlated with the sound capture signal of the beam that is currently tracking the sound source. If the correlation is high enough, e.g., the beam direction is close enough to the beam direction of the existing sound capture signal, the sound source is deemed to be an already tracked source rather than a new source. However, if the correlation is low or the direction is significantly different, a new sound source is deemed to have been detected.

[0111] Thus, it may be determined whether the detected potential new sound source matches any sound beams that the sound beam is directed at, i.e. tracked by the targeted beam. The match determination is based on comparing the characteristics of the variable sound beam with the characteristics of the targeted beam and / or comparing the characteristics of the sound capture signal of the variable sound beam with the characteristics of the sound capture signal of the targeted sound beam. For example, if the sound signals from the beams are sufficiently correlated with each other and the directions of the sound beams are sufficiently close, a match determination can be considered to have been made. If there is a match, it is determined that the detected sound source is an already tracked sound source, otherwise it is determined that a new sound source has been detected.

[0112] When a new sound source is detected, the beam steering device 203 may be configured to switch the audio beam to be directed from the previous sound source to the new sound source. Typically, the number of beams that can be implemented to accurately track a sound source is very limited, e.g., only five beams may be generated simultaneously. This limits the number of different sound sources that can be tracked simultaneously to a small number. The voice communication device 103 may be configured to dynamically switch between that limited number of beams to track the most appropriate sound source, e.g., primarily the loudest sound source, or e.g., the most active sound source.

[0113] However, in the approach of the voice communication device 103 of Figure 2, the selection of which sound source to track may also depend on the determined speaker category. In particular, when switching one of the sound beams to a newly detected sound source, the beam steering device 203 may select the beam to assign to the newly detected sound based on the speaker category, and thus the previous sound source / beam / voice capture signal to be dropped to change the beam assignment.

[0114] For example, the speaker categories may be ranked / ordered relative to one another, with each speaker category assigned a priority. For example, in an operating room, the surgeon speaker category may be assigned the highest priority, followed by the medical professional speaker category, then the patient speaker category, then the medical support staff speaker category, and finally the relative speaker category. In a court of law, the judge speaker category may be assigned the highest priority, followed by the lawyer speaker category, then the witness speaker category, and finally the speaker category including all other roles.

[0115] Specifically, the beam steering device 203 can select the beam to be reassigned starting from the lowest priority, such that the beam assigned to the speaker / sound source having the lowest priority among the currently assigned beams is selected.

[0116] Such an approach allows for improved adaptation of the available beams to evolving and changing scenarios, typically more quickly, and in particular more efficiently to changes in speakers present or active in the room, while still ensuring that the most important information is provided to remote participants.

[0117] In some embodiments, the audio communication device 103 may further include functionality for facilitating and / or improving adaptation to new audio sources. In particular, the audio communication device 103 may include means for storing data related to currently tracked / detected audio sources, the stored data including an indication of the assigned speaker category. When a new audio source / signal is detected, the audio communication device 103 may search the stored data to evaluate whether the audio source was previously tracked. If so, the stored speaker category may be extracted and used, for example, as an initial speaker category for the new audio source.

[0118] As an example, as shown in FIG. 4, the classifier 213 further includes a signature generator 403 and a signature store 405 in addition to a main classification processor 401 that performs speaker classification as described above.

[0119] The signature generator 403 can generate a signature of the sound source based on the frequency distribution of the sound source's audio capture signal. The sound source signature can be a distinctive mark, feature, or characteristic that can be generated from the audio capture signal. The signatures of the audio capture signals of different sound sources tend to be different.

[0120] As an example, for a voice capture signal that is detected to belong to a certain speaker category, the signature generator 403 may determine a frequency distribution by repeatedly performing an FFT on the voice capture signal. The resulting frequency spectrum may be averaged and a signature generated from the averaged frequency spectrum. In some cases, the frequency spectrum may be used directly as the signature. In other embodiments, some processing / analysis of the frequency spectrum may be performed. For example, the smallest frequency interval that constitutes 70% of the total energy may be determined and used as the signature.

[0121] The signature generator 403 may be configured to store the signature determined for the audio source in the signature store 405. In addition to the signature, further information about the audio source may also be stored. In particular, the speaker category determined for the audio source may be stored and linked to the signature. Thus, after some time, the signature store 405 may contain various signatures and associated speaker categories.

[0122] The classifier 213 may be configured to determine the speaker category of an audio source based on the stored signature. Specifically, when a new audio source is detected, the signature generator 403 may generate a signature of a new audio capture signal for that audio source. The signature generator 403 may specifically compare the new signature with the signatures stored in the signature store. If a match is found (according to appropriate matching / similarity criteria), the signature generator 403 may extract the linked speaker category and assign this to the new audio source.

[0123] In some cases, the determination of the speaker category of a new audio source may be made by simply assigning it to the speaker category stored for the matching signature. This may be applied, for example, when the match is very high. In other embodiments, or when the match is lower, the stored signature may be used as the initial speaker category, or as an initial candidate speaker category. Such an approach still requires a classification process to be performed, but typically requires less speech data to be analyzed and evaluated before a speaker category can be assigned.

[0124] In many practical examples, the audio capture device 201 may include multiple adaptive beamformers, such as the beamformer described in WO2017EP84679A. The beamformers may be based on block processing. For a 16 kHz audio signal, typically frames of 256 samples may be used. For each frame, all beamformers may calculate their outputs to provide an audio capture signal corresponding to the frame. Furthermore, the audio communication device 103 may determine for each frame which beams are active and whether a new beam needs to be formed for a new sound source.

[0125] A free-running (variable) sound beam may be formed by an adaptive beamformer, for example as described in US 7146012 or US 7602926. Based on the sound capture signal generated for the free-running beam, potential new sound sources can be detected.

[0126] In some cases, in each frame, only one of the multiple audio capture signals is selected. For example, classification and, for example, speech recognition and / or NLP processing may be performed only on the strongest signal in each frame. The following operations may be used: 1. If the strongest sound source is in one of the targeted beams that is tracking the sound source or is fixed in one direction, for example, it can be checked whether the distance between the free running beam and the targeted beam is small. If so, the signal-to-interference ratio is likely high enough for the free running beam to capture the sound source, and one of the targeted beams can be updated to track this new sound source. The strongest sound capture signal, or possibly all of the sound capture signals, can be provided to the analyzer 211 and classifier 213, and possibly also to the content analyzer 217. These can then classify the speaker and / or the content. 2. If the free-running beam has the strongest sound source, it may be determined whether there is a targeted beam near the free-running beam (e.g., using the distance determination approach disclosed in WO2017EP83680A). a. If no overlapping beams are found, a new beam is created by copying the coefficients of the swing-running variable beam to a new targeted beam. If the number of targeted beams has already reached the maximum number, targeted beams must be removed first. The selection of beams to be removed can be based on various criteria. In particular, as mentioned above, the speaker category / role of the speakers in the beam can be taken into account (e.g., technical staff and clinical staff are prioritized over clerical / office staff, patients, visitors / family, etc.), and in a secondary step when multiple participants have the same role, the amount of activity during the most recent period, the energy in the beam when active, the distance between the beams, and any kind of combination can be considered. b. If overlapping beams are found at short distances, in some embodiments no action is taken. In this case it may be assumed that the targeted beams can automatically adjust themselves to a suitable solution. For longer distances, beamforming may be reinitialized using the coefficients of the free-running beam.

[0127] Figure 5 shows a typical example of how a targeted beam is generated. This approach includes at least an adaptive beamformer that generates a desired signal and a noise reference. A second adaptive filter may be used to cancel coherent noise or other sources. In addition, nonlinear post-processing may be applied to further clean up the desired signal.

[0128] In many embodiments, the structure and implementation of the beamformer for the targeted beam / dedicated beam / tracking beam / targeted beam is the same as that for the free-running variable beam, e.g., using the same filter length, etc. However, the adjustment control may be different. The free-running beamformer may always adjust to form a beam toward the strongest voice signal, while the targeted beamformer may use a more selective adjustment approach. For example, the targeted beamformer may only adjust if the signal-to-noise ratio is high enough, if speech is detected, etc. For example, the approach disclosed in WO2018EP50045A may be used to provide a robust approach.

[0129] In many embodiments, the audio capture device 201 may include a user interface 219, specifically an output user interface 219, that can present information to participants in the room. In many embodiments, the audio capture device 201 may include a display or display interface that can be used to present visual information to a user.

[0130] In many embodiments, the sound capture device 201 may be configured to present information to the participant regarding the speaker categories that have been assigned to the various sound sources and beams.

[0131] In particular, in many embodiments, the audio communication device 103 may include a detector 221 configured to detect that the active audio capture signals include a currently active speech signal. For example, the detector 221 may continually evaluate all audio capture signals to detect whether speech is currently present in the signal and thus whether the tracked audio source is a speaker and whether that speaker is currently speaking.

[0132] It will be appreciated that various approaches can be used for speech detection by detector 221. In some embodiments, speech detection may be performed simply by detecting whether the audio capture signal represents audio above a given threshold level. This may be appropriate, for example, when the only sound expected is human speech, and therefore all captured sounds can be considered as speech. In other embodiments, more complex algorithms may be used, including, for example, evaluation of frequency spectrum (e.g. distribution, presence of harmonics, etc.), dynamic changes (e.g. transients), etc. In some embodiments, detection may be performed based on the results of speech recognition (although this may be too late in many cases). It will be appreciated that various algorithms for speech detection are known and any suitable approach may be used.

[0133] The user interface 219 may receive notification of which audio capture signals / audio beams / audio sources are currently considered to be active speakers. An indication of the speaker categories assigned to these capture signals / audio beams / audio sources may then be presented / output. For example, in some embodiments, the display may show a list of all the roles of people who are currently actively speaking.

[0134] In some embodiments, the detector 221 may be configured not only to detect which audio capture signals currently contain speech and therefore who is currently speaking, but also to determine a single dominant speaker. For example, the audio capture signal from which the strongest signal was detected may be determined and deemed to be the dominant speaker. Such an approach tends to provide a high degree of confidence in cases where only one person speaks at a time, such as in a courtroom.

[0135] In such cases, the user interface 219 may be configured to present, for example, an indication of the speaker category of a single speaker rather than multiple speakers. In such a scenario, an indication may be displayed to the user indicating the current speaker role, which may change dynamically as different participants speak.

[0136] In some embodiments, the audio generator 205 may be configured to select a single audio capture signal / audio beam / audio source / speaker based on speaker category. For example, if multiple speakers are detected to be currently active, and in particular, multiple audio capture signals are detected to currently contain audio signals, a single speaker / signal may be selected based on speaker category. For example, as described above, relative priorities may be associated with the speaker categories, and a single audio capture signal / speaker may be selected as the audio capture signal associated with the highest priority group. In some embodiments, an indication of the speaker category of one speaker rather than multiple speakers may be presented. Accordingly, in such a scenario, an indication may be displayed to the user indicating the role of the currently active speaker with the highest priority.

[0137] In many embodiments, the audio communication unit 103 may be configured to transmit metadata to a remote device / participant. The metadata may include an indication of at least one speaker category. For example, instead of displaying an indicator of the speaker category on a display as described above, data indicating the speaker category of the currently active speaker may be included in the bitstream. The remote device may include functionality to extract the metadata and present speaker category information (e.g., speaker role) to the remote participant.

[0138] In some embodiments, the user interface may be configured to alternatively or additionally present an indication of segments of the audio capture signal that are attenuated or specifically muted. For example, a low volume tone may be generated during the time that one or more audio segments are muted and not included in the output audio signal. In other embodiments, an indication may be provided on the display when a segment is attenuated. In some embodiments, a simple binary indication may be provided, while in other embodiments, more detailed information may be provided. For example, an indication of the content category of the muted audio segment may be presented to the participants in the room. Such an approach may provide advantageous feedback to the participants. For example, if a segment is muted because it contains sensitive information, such as personal information, an indication may be presented on the display alerting the speaker that they are currently disclosing such personal information. This may, for example, reduce the risk of unintended disclosure by a participant.

[0139] In some embodiments, the combination of output audio signals for a given user may depend on the speaker categories assigned to the audio capture signals, and in particular, the adapter 215 may be configured to control the audio generator 205 to adjust one or more combination weights of the audio output signals based on the assigned category.

[0140] The combining weight of an audio output signal may be the weight of that audio capture signal relative to the weight of other audio output signals when combining the audio output signal into an output audio signal. Specifically, the combining weight of a given audio output signal may correspond to the relative gain of the audio output signal when combined / mixed with other audio capture signals.

[0141] For example, in some embodiments, speaker categories corresponding to particular speaker categories, such as judges, surgeons, or field service engineers, may be deemed particularly important and, accordingly, the gain of audio capture signals associated with such speaker categories may be set to a higher level when mixed with other audio capture signals, thereby producing an audio output signal in which speakers of the particular category are more audible.

[0142] In some embodiments, as described above, speaker categories may be associated with different priorities. In such cases, the gain / combination weights of the audio capture signal may be set according to the priorities of the speaker categories associated with the audio capture signal. For example, the higher the priority, the higher the gain / combination weights may be set.

[0143] Thus, in some embodiments, the voice capture signal may prioritize signals in the audio output signal by setting relative gains / combination weights according to speaker categories / priorities. This may be useful, for example, when multiple speakers are speaking simultaneously and all of them are included in the audio output signal. Based on the speaker classification, the voice capture signal may decide to give a more important speaker (e.g., the chairperson) a slightly higher gain in the output audio signal than the other speakers.

[0144] In some embodiments, speaker classification may take into account the context of the ongoing conversation.

[0145] For example, a methodology for determining the speaker category / role of a participant captured by a targeted beam is to consider the context of the ongoing conversation between the remote party and the participants in the room. Figure 6 illustrates a scenario in which a remote service engineer (RSE) is conversing with a lab technician from a remote location regarding a reported issue. In this example, there is a certain context to the conversation. During the conversation, whenever a new sentence is received from one of the participants (RSE or technician), the context of that sentence may be matched with the context of sentences in the ongoing conversation to determine whether the new sentence is part of the current conversation and whether it is from the correct participant. If there is a sentence captured by a targeted beam from another participant (e.g., a doctor talking to a patient / family member containing confidential information), that sentence does not match the context of the ongoing conversation. In this way, the role of the participant can be determined and the sentence can be discarded.

[0146] The main steps of such an approach may include: When a new sentence / text is detected from one of the audio beams of an ongoing conversation, the sentence / text of the previous conversation and the new sentence / text are passed through a context extraction module. Such a module may include the following operations: - Sentence Tokenizer: First, as a preprocessing step, the sentence / text is tokenized into a list of tokens. - Stemming and stop word removal: The next step is stemming or lemmatization of words to their base forms. Different forms of a word have the same context, so the base forms are used in natural language processing. A pre-processing step of removing stop words (such as "is", "and", "or" etc.) may also be performed. - Word Embeddings: Word embeddings represent words in a vector space where similar words are mapped close to each other based on the context of their usage. - Context Embedding: Context embedding represents the temporal interactions or context of words in a vector space. The output of this module may be two context vectors representing the context of both types of sentences / texts. In the next step, a similarity score can be output by calculating the similarity between these two context vectors using an appropriate similarity function, for example: Score context =Similarity(Context_Vector prev ,Context_Vector new The similarity score can then be compared to a threshold to determine if the new sentence / text obtained from the output of the targeted beam is part of the conversation or if it comes from another person who is in the room but not participating in this conversation.

[0147] FIG. 7 shows an example of an overall methodology that can be used to derive / calculate the contextual similarity of new sentences / text detected in the audio beam with respect to an ongoing conversation.

[0148] For clarity, the above description describes embodiments of the invention with reference to multiple different functional circuits, units, and processors. However, it will be appreciated that functionality may be suitably distributed among different functional circuits, units, or processors without detracting from the invention. For example, functionality described as being performed by multiple separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits should not be construed as indicative of a strict logical or physical structure or organization, but rather as references to suitable means for providing the described functionality.

[0149] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented as a single unit or may be physically and functionally distributed between several different units, circuits and processors.

[0150] Although the present invention has been described in relation to several embodiments, the present invention is not limited to the specific forms described in the specification. The scope of the present invention is limited only by the appended claims. Moreover, even if a feature appears to be described in relation to a specific embodiment, a person skilled in the art will recognize that various features of the above-mentioned embodiments can be combined in accordance with the present invention. In the claims, the terms "comprise", "include" and the like do not exclude the presence of other elements or steps.

[0151] Furthermore, although individually recited, a plurality of means, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. Moreover, although individual features are included in different claims, they may be combined as appropriate, and the inclusion of such features in different claims does not imply that the combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one claim category does not necessarily limit the feature to this category, but the feature may be equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features should be acted upon, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in that order. The steps may be performed in any suitable order. Furthermore, the singular does not exclude the plural; thus, the use of "first", "second", etc. does not exclude a plural. Reference signs in the claims are merely examples for clarity and do not limit the scope of the claims in any way.

[0152] The following embodiments generally illustrate examples of audio devices, methods of operating audio devices, and computer program products implementing the methods.

[0153] Embodiment 1. An audio device, comprising: an audio capture device (201) for capturing audio in an environment, the audio capture device (201) forming a plurality of audio beams and generating an audio capture signal for each audio beam of the plurality of audio beams; a beam steering device (203) for directing each of the plurality of sound beams toward a different sound source; an analyzer (211) for analyzing at least the first audio capture signal to determine audio speech characteristics of the first audio capture signal; a classifier (213) for determining a first speaker category from among a plurality of speaker categories for a first source of a first audio capture signal in response to speech characteristics; an audio generator (205) for generating a first audio output signal for a first user by combining audio capture signals including the first audio capture signal, and for generating a different second audio output signal for a second user by combining audio capture signals including the first audio capture signal; an adapter (215) for adjusting the first audio output signal according to the first speaker category; The adapter (215) individually adjusts the first audio output signal according to the first speaker category and characteristics of the first user, and individually adjusts the second audio output signal according to the first speaker category and characteristics of the second user.

[0154] 1. An audio device, comprising: an audio capture device (201) for capturing audio in an environment, the audio capture device (201) forming a plurality of audio beams and generating an audio capture signal for each audio beam of the plurality of audio beams; a beam steering device (203) for directing each of the plurality of sound beams toward a different sound source; an analyzer (211) for analyzing at least the first audio capture signal to determine audio speech characteristics of the first audio capture signal; a classifier (213) for determining a first speaker category from among a plurality of speaker categories for a first source of a first audio capture signal in response to speech characteristics; an audio generator (205) for generating an audio output signal by combining audio capture signals including the first audio capture signal; and an adapter (215) for adjusting an audio output signal according to a first speaker category.

[0155] 2. The audio device of claim 1, wherein the analyzer (211) detects words in the first audio capture signal and determines at least a first one of the speech characteristics in response to the detected words.

[0156] 3. The speech device of claim 2, wherein the analyzer (211) determines a first one of the speech characteristics in response to natural language processing (NLP) of the detected words.

[0157] 4. The audio device of any one of claims 1 to 3, wherein the audio generator (205) generates a first audio output signal for a first user and a different second audio output signal for a second user, and the adapter (215) individually adjusts the first audio output signal according to a first speaker category and characteristics of the first user, and adjusts the second audio output signal according to a second speaker category and characteristics of the second user.

[0158] 5. The audio device of any one of claims 1 to 4, wherein the adapter (215) selects which of the plurality of audio capture signals to include in combination to generate a first audio output signal in response to a first speaker category.

[0159] 6. The audio device of any one of claims 1 to 5, further comprising a content analyzer (217) for analyzing a segment of the first audio capture signal to determine a content category of the segment from a plurality of content categories, and the adapter (215) adjusts the first audio output signal depending on the content category.

[0160] 7. The audio device of claim 6, wherein the adapter (215) attenuates segments of the first audio capture signal for a combination of at least one content category and the first speaker category and does not attenuate segments of the first audio capture signal for a combination of at least one other content category and the first speaker category.

[0161] 8. An audio device as claimed in claim 7, comprising a user interface (219) for presenting an indication of the attenuated segments.

[0162] 9. The classifier (213) a signature generator (403) for generating a signature of the sound source based on a frequency distribution of the sound capture signal of the sound source; a memory unit (405) for storing a signature of the audio source linked to the speaker category determined for the audio source, The signature generator (403) generates a first signature of the first sound source in response to the first sound source being detected; 9. The audio device of claim 1, wherein the classifier (213) matches the first signature with signatures stored in the memory unit (405) and determines a first speaker category of the first audio source depending on a speaker category linked to the stored signature.

[0163] 10. The audio device of any one of claims 1 to 9, wherein the audio capture device (201) detects a new audio source, and the beam steering device (203) switches the audio beam from being directed towards a previous audio source to being directed towards the new audio source in response to the detection of the new audio source, and selects the previous audio source in response to a speaker category of the previous audio source.

[0164] 11. The audio device further comprises: a detector (221) for detecting an active voice capture signal, including a currently active speech signal; and a user interface (219) presenting an indication of speaker categories assigned to sources of active audio capture signals.

[0165] 12. The audio device according to any one of claims 1 to 11, wherein the audio generator (205) adjusts a combination weight of at least one of the audio capture signals depending on the first speaker category.

[0166] 13. The sound capture device (201) generates a variable sound beam and the beam steering device (203) Vary the variable sound beam to detect potential new sound sources, determining whether the potential new sound source matches a sound source to which any of the plurality of sound beams is directed, the determination of whether or not the potential new sound source matches in response to a comparison of at least one of the characteristics of the variable sound beam and the characteristics of the sound beam of the plurality of sound beams to the characteristics of the sound capture signal of the variable sound beam and the characteristics of the sound capture signal of the plurality of sound beams; 13. An audio device according to any one of claims 1 to 12, wherein if no match is found, the audio beam is steered from the previous sound source towards the potential new sound source.

[0167] 14. A method of operating an audio device, the method comprising: capturing sound within the environment by forming a plurality of sound beams and generating a sound capture signal for each sound beam of the plurality of sound beams; directing each sound beam of the plurality of sound beams towards a different sound source; analyzing at least a first audio capture signal to determine audio speech characteristics of the first audio capture signal; determining a first speaker category from among a plurality of speaker categories for a first source of a first speech capture signal in response to speech characteristics; generating a first audio output signal by combining the audio capture signals with the first audio capture signal; and adjusting the first audio output signal according to the first speaker category.

[0168] 15. A computer program product comprising computer program code means for performing all the steps recited in claim 14 when the program is run on a computer.

Claims

1. an audio capture device for capturing audio in an environment, the audio capture device forming a plurality of audio beams and generating an audio capture signal for each audio beam of the plurality of audio beams; a beam steering device for steering each of the plurality of sound beams toward a different sound source; an analyzer that analyzes at least a first audio capture signal to determine speech characteristics of the audio of the first audio capture signal; a classifier for determining a first speaker category from among a plurality of speaker categories representing speaker roles for a first sound source of the first captured audio signal in response to the speech characteristics; an audio generator that generates a first audio output signal for a first user by combining audio capture signals including the first audio capture signal; an adapter that adjusts the first audio output signal in response to the first speaker category and an access privilege characteristic indicative of a degree of permitted access of the first user to at least one information category; An audio device comprising:

2. 2. The audio device of claim 1, wherein the audio generator further generates a different second audio output signal for a second user by combining audio capture signals including the first audio capture signal, and the adapter individually adjusts the second audio output signal according to the first speaker category and characteristics of the second user.

3. 3. The audio device of claim 1, wherein the analyzer detects words in the first captured audio signal and determines at least a first one of the speech characteristics in response to the detected words.

4. The speech device of claim 3 , wherein the analyzer determines the first one of the speech characteristics in response to natural language processing of the detected words.

5. 5. The audio device of claim 1, wherein the adapter selects which of the plurality of audio capture signals to include in the combination to generate the first audio output signal in response to the first speaker category.

6. 6. The audio device of claim 1, further comprising a content analyzer configured to analyze a segment of the first audio capture signal to determine a content category of the segment from a plurality of content categories, and wherein the adapter adjusts the first audio output signal according to the content category.

7. 7. The audio device of claim 6, wherein the adapter attenuates segments of the first audio capture signal for at least one combination of content category and the first speaker category and does not attenuate segments of the first audio capture signal for at least one other combination of content category and the first speaker category.

8. 8. The audio device of claim 7, comprising a user interface for presenting an indication of the attenuated segments.

9. The classifier: a signature generator for generating a signature of a sound source based on a frequency distribution of an audio capture signal of the sound source; a memory unit for storing a signature of the audio source linked to the speaker category determined for said audio source; Including, the signature generator generates a first signature of the first sound source in response to the first sound source being detected; 9. The audio device of claim 1, wherein the classifier matches the first signature with signatures stored in the memory and determines the first speaker category of the first audio source according to a speaker category linked to the stored signature.

10. 10. The audio device of claim 1, wherein the audio capture device detects a new audio source, and the beam steering device switches an audio beam from being directed toward a previous audio source to being directed toward the new audio source in response to detecting the new audio source, and selects the previous audio source from a plurality of audio sources to direct a beam toward in response to a speaker category of the previous audio source.

11. the audio device further comprising: a detector for detecting an active voice capture signal including a currently active speech signal; a user interface presenting an indication of speaker categories assigned to sources of the active speech capture signal; and 11. An audio device according to any one of claims 1 to 10, comprising:

12. 12. The audio device of claim 1, wherein the audio generator adjusts a combining weight of at least one of the audio capture signals depending on the first speaker category.

13. The sound capture device generates a variable sound beam, and the beam steering device Varying the variable sound beam to detect potential new sound sources; determining whether the potential new sound source matches a sound source at which any one of the plurality of beams is directed, said determination of whether or not a match is made in response to a comparison of at least one of a characteristic of the variable sound beam and a characteristic of a sound beam of the plurality of sound beams with a characteristic of a sound capture signal of the variable sound beam and a characteristic of a sound capture signal of the plurality of sound beams; 13. An audio device according to any one of claims 1 to 12, wherein if no match is found, an audio beam is steered from a previous sound source towards the potential new sound source.

14. 1. A method of operation of an audio device, said method comprising: capturing sound in the environment by forming a plurality of sound beams and generating a sound capture signal for each sound beam of the plurality of sound beams; directing each sound beam of the plurality of sound beams towards a different sound source; analyzing at least a first audio capture signal to determine speech characteristics of the audio of the first audio capture signal; determining a first speaker category from a plurality of speaker categories representing speaker roles for a first sound source of the first captured audio signal in response to the speech characteristics; generating a first audio output signal for a first user by combining audio capture signals including the first audio capture signal; adjusting the first audio output signal in response to the first speaker category and an access privilege characteristic indicative of a degree of permitted access of the first user to at least one information category; A method comprising:

15. A computer program comprising computer program code means for performing all the steps of the method according to claim 14 when said computer program is run on a computer.