Computer-implemented method for assisting a user in a video conference, and vehicle
A method using a machine learning model to separate participant and computer-generated speech in vehicles improves road safety by minimizing driver distraction during video conferences through a distinct acoustic field and voice command functionality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MERCEDES BENZ GROUP AG
- Filing Date
- 2025-11-06
- Publication Date
- 2026-06-04
AI Technical Summary
Existing video conference systems in vehicles distract drivers by overlapping computer-generated speech with participant speech, increasing cognitive load and reducing road safety.
A computer-implemented method using a machine learning model to analyze conference speech, outputting explanatory information through a second virtual sound source positioned differently from participants, creating a distinct acoustic field to utilize the cocktail party effect, and allowing voice commands without continuous microphone activation.
Reduces driver distraction by clearly distinguishing participant and computer-generated speech, enhancing road safety through reduced cognitive load and enabling safe vehicle operation during video conferences.
Smart Images

Figure EP2025082178_04062026_PF_FP_ABST
Abstract
Description
[0001] Mercedes-Benz Group AG
[0002] Computer-implemented method for assisting a user in a video conference and vehicle
[0003] The invention relates to a computer-implemented method for assisting a user in a video conference according to the type defined in more detail in the preamble of claim 1, and to a vehicle for carrying out the method.
[0004] Modern technology allows people to stay in constant contact. In many regions of the world, the mobile network is comparatively well-developed, enabling seamless internet connectivity for mobile devices such as smartphones. This allows users to send and receive messages like emails, text messages from messenger services, and the like, as well as to make Voice over IP (VoIP) calls and hold video conferences. Even in remote areas, mobile network coverage can be provided using communication satellites. This is achieved primarily through satellite constellations, comprising numerous microsatellites orbiting the Earth in a low Earth orbit. Furthermore, modern vehicles are equipped with integrated telecommunications units, allowing users to access internet-related services from within the vehicle.
[0005] This makes it possible to participate in a video conference while driving. Using one or more microphones located inside the vehicle, the voices of the occupants can be picked up and transmitted to the respective conference participants. The vehicle's cameras can also capture a video image of the interior, showing the occupants. The audio from the video conference can be output through the speakers of the vehicle's integrated sound system. Modern vehicles typically have one or more displays located inside the vehicle, which also allows for the display of a video stream from the video conference. This video stream can, for example, show the other conference participants or a shared screen.
[0006] While operating a vehicle, the driver should focus their attention on the road. Even having a conversation can be distracting. If additional visual content is displayed on a screen, this can increase distraction even further. Therefore, there is a need to provide solutions that allow participation in video conferences while driving, allowing the driver to maintain their focus on the road.
[0007] US Patent 2020 / 0403817 A1 describes the generation of customized meeting insights based on user interactions and meeting media. The document details a specially trained machine learning model for increasing the efficiency of meeting participants, along with a corresponding training method. Based on methods of computational linguistics, also known as natural language processing (NLP), the machine learning model can analyze conversation content and identify relevant information. Additionally, the model can access and process meeting-relevant media content to evaluate its content and relevant context. The machine learning model can be connected to external sources, such as a meeting participant's digital calendar, to manipulate them selectively.The machine learning model is used to generate insights related to meetings, such as a meeting summary, highlights, meeting-descriptive metrics, action items, and the like. Meetings can be attended via various devices, such as a vehicle's integrated navigation system. The machine learning model disclosed in this publication thus enables real-time meeting assistance.
[0008] Furthermore, US patent 2019 / 0332680 A1 discloses a multilingual virtual personal assistant. This system recognizes speech in a recorded audio signal and identifies the language of the speech. Subsequently, at least a portion of the speech is translated into another language. A trained machine learning model then identifies and corrects translation errors in the translated portion, taking into account the inherent semantics of the spoken language.
[0009] Furthermore, DE 102016 103 331 A1 discloses a device and a method for playing back audio signals in a motor vehicle. Participation in a video conference is possible in the motor vehicle, whereby the other participants can be placed at different positions as virtual sound objects in an acoustic field generated in the vehicle interior.
[0010] The present invention is based on the objective of providing an improved computer-implemented method for assisting a user in a video conference, which allows simultaneous participation in a video conference and the control of a vehicle while maintaining road safety.
[0011] According to the invention, this problem is solved by a computer-implemented method for assisting a user in a video conference with the features of claim 1. Advantageous embodiments and further developments, as well as a vehicle for carrying out the method, are described in the dependent claims.
[0012] A generic computer-implemented method for assisting a user in a video conference, wherein at least two participants take part in the video conference, the user participates in the video conference from a vehicle, a machine learning model analyzes the participants' speech using computational linguistics methods to identify semantic content, and upon recognizing semantic content relevant to the user, causes the output of information explaining the semantic content to the user, furthermore provides that the explanatory information is output acoustically in the vehicle by means of computer-generated speech, wherein, according to the invention, the participants' speech is output via a first virtual sound source and the explanatory information via a second virtual sound source.which are positioned at different locations within an acoustic field in the vehicle, and the computer-generated speech output in the vehicle is filtered from an audio signal generated by a microphone that captures the acoustics in the vehicle interior. In layman's terms, the method according to the invention relates to a conference assistant. The information explaining the semantic content can be understood as a hint from the conference assistant. The conference assistant helps the driver with the cognitive processing of information related to the video conference. For this purpose, relevant content can be summarized or repeated. For example, the conference assistant can describe media content relevant to the video conference, obtain strategic information, ask topic-related questions, and provide recommendations or hints, such as: "The meeting ends in 5 minutes,"You actually wanted to talk about topic XY, which hasn't come up yet," and things like that.
[0013] However, the output of explanatory information in the vehicle in the form of computer-generated speech can also distract the user from driving. In particular, if the speech of the conference participants and the computer-generated speech come from the same direction, this creates a high cognitive load for the user. There is a particular risk that the user will mistake the computer-generated speech for that of a conference participant.
[0014] According to the invention, the participants' speech and the computer-generated speech are emanating from two different directions. The conference participants and the source of the computer-generated speech are positioned at different locations within an acoustic field created inside the vehicle. This allows the user to benefit from the so-called cocktail party effect. This effect describes the phenomenon where participants engaged in a conversation can follow it even when surrounded by a crowd of other people speaking simultaneously. Thus, even when the participants' speech is overlaid by the simultaneous output of the computer-generated speech, the user is still able to easily and reliably grasp the explanatory information both acoustically and cognitively.This reduces the user's mental load, allowing them to reliably focus their attention on driving. This improves road safety. Even while driving, the user can still participate in the video conference and mentally follow its content. If a traffic situation arises that requires increased mental attention, such as sudden braking or swerving due to an obstacle, or navigating an unfamiliar area, the user can refocus on driving. Once the situation is resolved, the conference assistant can summarize the missed content of the video conference, ensuring the user doesn't miss any relevant information.
[0015] The vehicle used by the user is equipped with a sound system suitable for providing the acoustic field. Such a sound system is characterized by several speakers arranged at different positions within the vehicle interior, which are individually controlled. Object-based surround sound formats such as Dolby Atmos are used for this purpose. The acoustic field generated by the sound system can also extend partially outside the vehicle interior.
[0016] The machine learning model for processing information associated with the video conference is executed locally in the vehicle. Thanks to prior training, the machine learning model is able to identify semantic content within the relevant information. For this purpose, at least the spoken content of the participants during the video conference is processed and analyzed by the machine learning model.
[0017] As previously described, the computer-generated speech output in the vehicle is filtered out from an audio signal produced by a microphone that captures the acoustics inside the vehicle. This reduces the risk of other video conference participants being disturbed or distracted by the computer-generated speech output. Because the computer-generated speech is created locally on a processing unit in the vehicle, it can be filtered out of the recorded audio signal particularly easily. The computer-generated speech can thus be subtracted from the corresponding audio signal.
[0018] An advantageous further development of the method according to the invention provides that the user specifies the location for positioning the second virtual sound source in the vehicle. This allows the user to try out different positions for the second virtual sound source. The user can then select the position that makes it easiest for them to differentiate between the spoken words of the participants and the computer-generated speech of the conference assistant. This allows the user to focus as much as possible on essential aspects of both the video conference and the operation of the vehicle.
[0019] According to a further advantageous embodiment of the method according to the invention, the second virtual sound source is positioned above the user's left and / or right shoulder or above the vehicle's steering wheel. The second virtual sound source can thus not only be freely positionable but also placed in predefined locations. The user can choose from several predefined target positions, or the placement of the second virtual sound source within the vehicle interior can follow fixed specifications. The second virtual sound source is preferably positioned above the user's left and / or right shoulder or above the vehicle's steering wheel, which has a particularly positive effect on achieving the cocktail party effect. The second virtual sound source can be positioned only above the left or only above the right shoulder to produce a corresponding mono output.However, two second virtual sound sources can also be provided, each assigned to the user's left and right ear above the respective shoulder. This enables stereo output of the computer-generated speech.
[0020] In particular, further acoustic effects can be used to enhance the cocktail party effect, such as reducing the volume of the spoken words while simultaneously outputting the computer-generated speech of the conference assistant and / or increasing the volume of the conference assistant.
[0021] A further advantageous embodiment of the method according to the invention provides that the machine learning model analyzes the content of a video stream from the video conference and takes the content of the video stream into account in addition to determining the relevant semantic content. The machine learning model was also trained accordingly for this purpose. For example, the facial expressions and / or gestures of the conference participants can be analyzed. For instance, it can be recognized that at least one of the participants has a questioning look. Such a reaction might occur, for example, in response to a statement made by the user. The conference assistant can then recognize this and inform the user. For example, the conference assistant could output the following information to the user: "Karl seems not to have understood the matter. It is recommended that he repeat the matter in other words."This allows the user to react appropriately to constantly changing situations during the video conference, without requiring them to look at a display device to view the video stream. This allows the user to keep their eyes on the road, further improving road safety.
[0022] The video stream can not only show other conference participants, but also shared screen content or dedicated presentation content such as slides. Documents, texts, tables, diagrams, graphics, and the like can therefore be included in the video stream. This media can be captured, analyzed, and evaluated by the machine learning model. This also allows the identification of semantic content related to the media and the output of this information in the form of computer-generated speech within the vehicle.
[0023] According to a further advantageous embodiment of the method according to the invention, the machine learning model analyzes conference media associated with the video conference and stored in a data storage device, and takes this media into account to determine the relevant semantic content. This allows media relevant to the conference to be analyzed in advance, before they are displayed in the video stream. Conference media that are not included or displayed in the video stream may also be relevant to the conversation taking place during the video conference. This allows the user to be informed about relevant content in a more comprehensive way. The data storage device can also be located locally in the vehicle. However, it is also conceivable that the machine learning model accesses an external data storage device, for example, via the internet.For example, the data storage can be integrated into a central computing facility such as a server or server cluster. The machine learning model can also access the data storage of a mobile device belonging to one of the conference participants, particularly after data sharing has been authorized. A further advantageous embodiment of the method according to the invention provides that the machine learning model transcribes the spoken words and stores the resulting transcript and / or recognized semantic content in the data storage. The corresponding transcript and / or the content artificially generated by the machine learning model can thus be supplemented in the data storage for later use. This can facilitate the subsequent acquisition of relevant information for all conference participants. In particular, the machine learning model can use the information it has itself stored in the data storage for re-access.This allows the machine learning model to access past relevant information in order to use it to generate currently relevant information. For example, the context of previously spoken words or already recognized semantic content can be important for currently relevant semantic content.
[0024] A further advantageous embodiment of the method according to the invention provides that a microphone capturing the acoustics in the vehicle interior is muted for the video conference and activated to receive a voice command for the machine learning model when the user's intention to issue the voice command is recognized. Typically, the microphone(s) intended for recording the user's speech are activated during the video conference. Thus, it is unknown when the user will speak during the video conference. If the microphone is continuously activated, the user can contribute verbally at any time, and this contribution is then relayed to the other conference participants. This increases user convenience, as they do not have to press a push-to-talk button each time.It is also conceivable that the microphone will be automatically activated for the video conference once a certain volume level is exceeded. However, this carries the risk that the user might speak too quietly, which would then not be audible to the other participants in the video conference.
[0025] As mentioned earlier, the user can interact with the conference assistant. This can be done by giving the assistant voice commands, for example, to recite a prepared reminder, to read the meeting agenda, to summarize what the conference participants have said in the last 5 minutes, or similar tasks. Voice commands are particularly suitable for interacting with the conference assistant because they allow the user to keep their eyes on the road. This minimizes the mental strain on the user. However, commands for the conference assistant can be disruptive or confusing for other conference participants, as they don't fit into the context of the ongoing video conference conversation.
[0026] Advantageously, the microphone is muted for the video conference and activated to receive the voice command when the user wants to give a corresponding voice command.
[0027] There are various ways in which a user's intention can be recognized. For example, the user can be tracked using a variety of sensors, such as cameras, biometric sensors, and the like, which allows patterns in user behavior to be identified. The machine learning model can then be trained to link previously issued voice commands with corresponding patterns and automatically switch the microphone in relevant situations.
[0028] However, users particularly prefer to use a dedicated control to communicate their intention to issue a voice command. Such a control can be a physical device, such as a button, switch, touch-sensitive surface, or similar. The control can be integrated into the center console or steering wheel, for example. This makes the control particularly easy and quick to use. Furthermore, a dedicated control ensures that the user's intention is always correctly detected. This reduces the risk of a voice command being incorrectly recognized, which could inadvertently activate the microphone.
[0029] A further advantageous embodiment of the method according to the invention provides that the machine learning model is further trained based on the user's behavior in the video conference and / or the user's reaction to an action performed by the machine learning model. This allows the machine learning model to be adapted to specific users. As already mentioned, the user may exhibit specific behavior. For example, the user may want to hear a short summary of the discussion content so far at regular intervals. In another situation, the user could command the automatic generation of media content such as diagrams, graphs, summaries, or the like, and their corresponding storage in a data repository.If such tasks for the conference assistant become more frequent, especially in response to certain events, such as the passage of a certain period of time, the recognition of characteristic phrases, sentences, terms or the like, the machine learning model can, through appropriate training, recognize corresponding situations and become active proactively, without a voice command issued by the user.
[0030] The machine learning model can perform a wide variety of actions, such as creating entries in a user's digital calendar, sending an invitation to a follow-up appointment for the video conference, automatically retrieving information from the internet, saving a transcript generated during the video conference, and the like.
[0031] A vehicle of this type, comprising a computing unit for executing a machine learning model and means for conducting a video conference, and further comprising a sound system configured to generate a three-dimensional acoustic field, is further developed according to the invention in that the computing unit and the means for conducting the video conference are configured to execute a method described above. In the vehicle according to the invention, it is thus ensured that users can participate in a video conference simultaneously while controlling the vehicle, and that the vehicle can be operated safely. The vehicle can be any road vehicle such as a car, truck, van, bus, or the like. It is also generally conceivable that it could be a rail vehicle, watercraft, or aircraft.
[0032] The equipment required for conducting a video conference includes, in particular, one or more cameras that capture the vehicle's interior. This allows for the creation of a video stream showing the user and displaying it to the other conference participants. The equipment also includes one or more microphones for capturing the acoustics within the vehicle. The video stream of the video conference, or the video streams showing the individual participants, can be displayed on one or more display devices located in the vehicle, for example, a touchscreen. The vehicle also has a suitable sound system for outputting a three-dimensional acoustic field. Sound sources can be positioned at various locations within this acoustic field. Finally, the vehicle is equipped with a telecommunications unit to establish a connection with the other participants in the video conference.The corresponding data stream can be routed partially or entirely through a central computing facility, such as a server or server network. A wide variety of video conferencing programs can be used.
[0033] Further advantageous embodiments of the computer-implemented method according to the invention for assisting a user in a video conference and of the vehicle according to the invention also result from the exemplary embodiment, which is described in more detail below with reference to the figure.
[0034] Figure 1 shows a schematic top view of a vehicle according to the invention while driving, with the person driving the vehicle participating in a video conference.
[0035] A vehicle 2 according to the invention, shown in Figure 1, allows a user 1, or the person driving the vehicle, to participate in a video conference while the vehicle 2 is in operation. An inventive method for assisting the user 1 is implemented, which minimizes the distraction of the user 1 caused by conducting the video conference. This allows the user 1 to concentrate more on driving, ultimately increasing road safety.
[0036] Vehicle 2 includes the means to conduct a video conference. This includes one or more cameras 8 for capturing the vehicle interior. This allows the user 1 to be recorded and the corresponding video stream to be displayed to the other participants in the video conference. In addition, one or more microphones 6 are provided to capture the acoustics within the vehicle interior. Vehicle 2 also has a sound system capable of generating a three-dimensional acoustic field. For this purpose, the sound system comprises several loudspeakers 9 arranged at various locations within Vehicle 2. These loudspeakers 9 can be integrated, for example, into the A-pillar, B-pillar, C-pillar, a headrest, the dashboard, the headliner, or similar locations. The sound system is capable of processing an object-based surround sound format, such as Dolby Atmos.Furthermore, the vehicle 2 includes at least one display device 10 for outputting a video stream of the video conference. Such a video stream can show the other participants of the video conference as well as additional content such as a shared screen, media content, and the like.
[0037] The vehicle 2 further comprises a processing unit 7 for processing signals generated by the camera 8 and the microphone 6 and for controlling the loudspeakers 9 and the display device 10. In addition, the processing unit 7 is in communication with the respective terminal devices of the other participants in the video conference via a telecommunications unit 11, either directly or indirectly through a central computing unit (not shown in detail). The processing unit 7 has at least read access to a computer-readable storage medium containing a computer program product, which in turn comprises machine-interpretable instructions. When executed by a processor of the processing unit 7, these instructions cause the processor to execute a computer-implemented method according to the invention for assisting the user 1 in the video conference. In other words, a so-called conference assistant for vehicles 2 is provided according to the invention.In this process, processing unit 7 executes a specially trained machine learning model that uses computational linguistics methods to analyze the speech of video conference participants in order to identify semantic content. Upon recognizing semantic content relevant to user 1, the machine learning model outputs a text explaining that semantic content in the form of computer-generated speech. In addition to speech, the machine learning model can also analyze the video stream and / or conference media stored on a data storage device. Furthermore, the machine learning model, i.e., the conference assistant, can receive voice commands from user 1, for example, to repeat a brief summary of the discussion from the last 5 minutes.
[0038] According to the invention, the spoken words of the participants are output via a first virtual sound source 3, and the explanatory information is output via a second virtual sound source 4, both positioned at different locations within the acoustic field of the vehicle 2. The acoustic field can be extended beyond the vehicle interior, as shown in Figure 1. For example, the first virtual sound source 3 can be located above the hood of the vehicle 2 or on the passenger seat. Any conceivable position within the boundaries of the vehicle 2, other than that of the second virtual sound source 4, is possible. Positioning the second virtual sound source 4 above the left and / or right shoulder of the user 1 or above the steering wheel 5 is particularly preferred.
[0039] Thus, the participants of the video conference and the computer-generated speech can be clearly distinguished by User 1. Due to the different placement, the so-called cocktail party effect can be utilized for User 1, allowing for particularly reliable differentiation between what the participants are saying and the computer-generated speech. This makes it easier for User 1 to concentrate on driving, as they are less distracted by the conference assistant than with state-of-the-art solutions. To further improve the distinction between what the participants are saying and the computer-generated speech, the volume of the computer-generated speech can be increased and the volume of what the participants are saying can be decreased. Additionally, a characteristic voice tone or...The sound for the computer-generated speech will be chosen to be particularly easy for user 1 to hear.
Claims
Mercedes-Benz Group AG Patent claims 1. A computer-implemented method for assisting a user (1) in a video conference, wherein at least two participants take part in the video conference, the user (1) participates in the video conference from a vehicle (2), a machine learning model analyzes the spoken words of the participants using methods of computational linguistics in order to identify semantic content, and upon recognizing semantic content relevant to the user (1), causes the output of information explaining the semantic content to the user, and wherein the explanatory information is output acoustically in the vehicle (2) by means of computer-generated speech, characterized in that the spoken words of the participants are output via a first virtual sound source (3) and the explanatory information via a second virtual sound source (4), which are positioned at different positions of an acoustic field in the vehicle (2);and the computer-generated speech output in the vehicle (2) is filtered out from an audio signal generated by means of a microphone (6) that detects the acoustics in the vehicle interior.
2. Method according to claim 1, characterized in that the location for positioning the second virtual sound source (4) in the vehicle (2) is specified by the user (1).
3. Method according to claim 1 or 2, characterized in that the second virtual sound source (4) is placed above the left and / or right shoulder of the user (1) or above the steering wheel (5) of the vehicle (2).
4. Method according to one of claims 1 to 3, characterized in that the machine learning model analyzes the content of a video stream of the video conference and takes the content of the video stream into account in addition to determining the relevant semantic content.
5. Method according to one of claims 1 to 4, characterized in that the machine learning model analyzes conference media associated with the video conference and stored in a data storage device and takes the conference media into account to determine the relevant semantic content.
6. Method according to claim 5, characterized in that the machine learning model transcribes the spoken word and stores the resulting transcript and / or recognized semantic content in the data storage.
7. Method according to one of claims 1 to 6, characterized in that a microphone (6) detecting the acoustics in the vehicle interior is muted for the video conference and activated to receive a voice command for the machine learning model when the intention of the user (1) to give the voice command is recognized.
8. Method according to claim 7, characterized in that the user (1) operates a dedicated control device to communicate the intention to issue the voice command.
9. Method according to any one of claims 1 to 8, characterized in that the machine learning model is further trained based on the behavior of the user (1) in the video conference and / or the response of the user (1) to an action performed by the machine learning model.
10. Vehicle (2) comprising a computing unit (7) for executing a machine learning model and means for conducting a video conference, in turn comprising a sound system configured to generate a three-dimensional acoustic field, characterized in that the computing unit (7) and the means for conducting the video conference are configured to execute a method according to one of claims 1 to 9.