Facilitation and optimization of enterprise personnel communications using multimodal artificial intelligence

The multimodal AI meeting assistive engine addresses communication challenges by analyzing and modifying user behaviors, enhancing performance and satisfaction through personalized support for enterprise personnel.

WO2026019426A1PCT designated stage Publication Date: 2026-01-22EQUIFAX INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/038332
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Enterprise personnel face challenges in managing time demands, communicating effectively, and overcoming personal discomfort, leading to isolation and reduced performance, which can negatively impact job satisfaction and overall enterprise performance.

Method used

A multimodal AI meeting assistive engine that analyzes user physical characteristics through various sensors, modifies abnormal behaviors or speech to appear normal, and provides assistance for improved communication and task completion.

Benefits of technology

Enhances communication effectiveness, improves personnel performance and satisfaction, and fosters enterprise growth by automating workarounds and providing personalized support for meetings and tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038332_22012026_PF_FP_ABST
    Figure US2024038332_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for facilitating and optimizing enterprise personnel communications using a multimodal artificial intelligence (AI) meeting assistive engine. The multimodal AI meeting assistive engine can receive or otherwise obtain information relative to user communications (e.g., meetings) and other user functions and responsibilities from a number of sources. These sources can include but are not limited to sensors, which may take the form of hardware devices such as cameras and microphones that are commonplace in a typical enterprise setting. From the sensors, the multimodal AI meeting assistive engine can receive signals that are representative of one or more physical characteristics of a user, and can determine if a user physical characteristic, such as a facial expression or a user voice, are abnormal. Abnormal physical characteristics may be modified by the multimodal AI meeting assistive engine so that the abnormal physical characteristics appear normal to virtual third parties.
Need to check novelty before this filing date? Find Prior Art

Description

FACILITATION AND OPTIMIZATION OF ENTERPRISE PERSONNEL COMMUNICATIONS USING MULTIMODAL ARTIFICIAL INTELLIGENCETechnical Field

[0001] The present disclosure relates generally to facilitating and optimizing enterprise communications, and more particularly although not exclusively, to facilitating and optimizing communications among enterprise personnel using a multimodal artificial intelligence (Al) assistive engine.Background

[0002] Enterprise personnel, especially personnel employed by large enterprises, are commonly faced with a variety of time demands and responsibilities. In addition to core job duties, these time demands and responsibilities may include, for example, ancillary tasks such as scheduling, attending, or participating in meetings, learning new technologies, resolving technical problems, and networking with both coworkers and non-coworkers, among potentially numerous other things.

[0003] Personnel may often find it difficult to address all the demands on their time. At least some personnel may also be disinclined to ask questions or raise concerns, or to pose comments or suggestions, whether relative to their particular job function or about enterprise operations in general. And while human or technology resources may be available to facilitate the resolution of required tasks and supply required information, personnel may not understand how to effectively locate and use such resources or, for example, may not be comfortable using technology with which they are unacquainted or communicating with unfamiliar human resources. Likewise, for various reasons, some personnel may be uncomfortable or awkward with respect to personal communications in general, which can make such personnel feel isolated and may also exacerbate difficulties with fulfilling assigned tasks.

[0004] Such issue can be detrimental to personnel performance, job satisfaction, and even personal well-being. Likewise, if enough personnelexperience such issues, overall enterprise performance may be detrimentally affected.Summary

[0005] Various aspects of the present disclosure provide systems and methods for facilitating and optimizing enterprise communications using multimodal artificial intelligence (Al). According to one example, a system can include a plurality of sensors configured to detect a plurality of physical characteristic of a user, and to transmit via one or more communication mediums, signals representative of the plurality of physical characteristics of the user such that the plurality of physical characteristics of the user are observable by at least one third party. The system may also include a computing device comprising a multimodal Al meeting assistive engine configured to receive or otherwise obtain the signals from the plurality of sensors; a processor and a memory, such as a non-transitory computer-readable medium, which includes instructions that are executable by the processor to cause the processor to perform various operations. According to aspects of the present disclosure, the operations can include evaluating, by the multimodal Al meeting assistive engine, the plurality of physical characteristics of the user by analyzing the signals from the plurality of sensors; determining, by the multimodal Al meeting assistive engine, based on analyzing the signals, that at least one of the plurality of physical characteristic of the user is abnormal; and in response to determining that the at least one of the physical characteristics of the user is abnormal, modifying, by the multimodal Al meeting assistive engine, the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

[0006] According to another example of the present disclosure, a non- transitory computer readable medium may contain instructions that are executable by a processor to cause the processor to perform operations. According to aspects of the present disclosure, the operations can include causing a multimodal Al meeting assistive engine to receive or otherwise obtain, from a plurality of sensors, signals representing a plurality of physicalcharacteristics of a user, where the signals are transmitted by the plurality of sensors via one or more communication mediums such that the plurality of physical characteristics of the user are observable by at least one third party. According to aspects of the present disclosure, the operations can also include evaluating, by the multimodal Al meeting assistive engine, the plurality of physical characteristics of the user by analyzing the signals from the plurality of sensors; determining, by the multimodal Al meeting assistive engine, based on analyzing the signals, that at least one physical characteristic of the plurality of physical characteristics of the user is abnormal; and in response to determining that the at least one physical characteristic of the user is abnormal, modifying, by the multimodal Al meeting assistive engine, the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

[0007] According to an additional example of the present disclosure, a method is provided. The method may include receiving or otherwise obtaining, by a multimodal Al meeting assistive engine of a computing device, from a plurality of sensors, signals representing a plurality of physical characteristics of a user, the signals transmitted by the plurality of sensors via one or more communication mediums such that the plurality of physical characteristics of the user are observable by at least one third party. The method may also include causing, by a processor of the computing device, the multimodal Al meeting assistive engine to evaluate the plurality of physical characteristics of the user by analyzing the signals from the plurality of sensors; and causing, by the processor of the computing device, the multimodal Al meeting assistive engine to determine, based on analyzing the signals, that at least one physical characteristic of the plurality of physical characteristics of the user is abnormal. The method may further include, in response to determining that the at least one physical characteristic of the user is abnormal, causing, by the processor of the computing device, the multimodal Al meeting assistive engine to modify the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

[0008] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.

[0009] The foregoing, together with other features and examples, will become more apparent upon referring to the following specification, claims, and accompanying drawings.Brief Description of The Drawings

[0010] FIG. 1 is a pictorial diagram depicting an example of an operating environment including an Al system having a multimodal Al meeting assistive engine for facilitating and optimizing enterprise personnel communications.

[0011] FIG. 2 is a pictorial diagram depicting an Al video component of the Al system, for implementing certain aspects of the present disclosure.

[0012] FIG. 3 is a pictorial diagram depicting an Al audio component of the Al system, for implementing certain aspects of the present disclosure.

[0013] FIG. 4 is a pictorial diagram depicting an Al speech component of the Al system, for implementing certain aspects of the present disclosure.

[0014] FIG. 5 is a pictorial diagram depicting an Al knowledge component of the Al system, for implementing certain aspects of the present disclosure.

[0015] FIG. 6 is a pictorial diagram depicting an Al involvement component of the Al system, for implementing certain aspects of the present disclosure.

[0016] FIG. 7 is a pictorial diagram depicting an Al psychology component of the Al system, for implementing certain aspects of the present disclosure.

[0017] FIG. 8 is a pictorial diagram depicting an Al feedback component of the Al system, for implementing certain aspects of the present disclosure.

[0018] FIG. 9 is a block diagram depicting an example of a computing device suitable for implementing the functionality of the Al meeting assistive engine according to an example of the present disclosure.

[0019] FIG. 10 is a block diagram depicting a transformer architecture that can be used with the Al meeting assistive engine to facilitate and optimizeenterprise personnel communications according to an example of the present disclosure.

[0020] FIG. 11 is a flowchart representing a method of facilitating and optimizing enterprise personnel communications according to an example of the present disclosure.Detailed Description

[0021] Certain aspects of systems and methods according to the present disclosure can address one or more issues identified above. For example, examples of the present disclosure are directed to the use of a system incorporating multi-modal artificial intelligence (Al) to facilitate and optimize enterprise personnel communications and to foster enterprise personnel growth and development. For example, and broadly speaking, systems and methods according to examples of the present disclosure may also operate to facilitate personnel activities, automate workarounds to problems, and improve the technical, domain-specific, and personal growth and development of enterprise personnel.

[0022] According to other aspects of the present disclosure, a system may include a multimodal Al meeting assistive engine that can acquire, from multiple sources, personnel characteristics, task information, schedule information, knowledge requirements, environmental characteristics, enterprise information, etc., and can use the acquired characteristics and other information to improve the communication experience and performance of enterprise personnel. The multiple sources by which the Al engine can acquire personnel / environmental characteristics and other information may be, for example and without limitation, video, audio, image, textual, or other signal sources that includes hardware devices such as cameras, microphones, smart phones and other portable smart devices, virtual assistant devices, physiological characteristic (e.g., fitness) tracking devices, etc.

[0023] According to other aspects of the present disclosure, an Al meeting assistive engine may also be communicatively coupled to various sources of information within, and possibly outside of, the enterprise. For example, various applications and software used by personnel may be mined for informationrelative to individual personnel characteristics or preferences. Examples of the Al meeting assistive engine may also be communicatively coupled to a knowledge graph of all the information available in the enterprise, and may include a relational mapping of different components of the information to individual personnel, or groupings of personnel, such as but not limited to, personnel within given enterprise departments.

[0024] According to other aspects of the present disclosure, an Al meeting assistive engine may be provided with, and may operate according to, various rules, parameters, principles, etc., that govern the learning, reasoning, and problem solving functions of the Al meeting assistive engine. For example, the rules may include ethical rules that set boundaries regarding the information that may be collected about personnel, shared with personnel, or said or showed to personnel, during the course of the Al meeting assistive engine performing its intended functions.

[0025] According to other aspects of the present disclosure, system and method examples can be used to assist enterprise personnel in a myriad of different ways, including but not limited to, problem solving, communications, scheduling and time management, collaboration, socialization, motivation, learning, and personal growth and development. For example, system and method examples according to the present disclosure can assist personnel with respect to scheduling, running, or otherwise participating in virtual or in-person meetings, and assisting or guiding personnel during such meetings. System and method examples according to the present disclosure can also connect personnel to an available / growing knowledge base and assist with extraction of information that is predicted to be useful / helpful to solving a problem or completing a task, or can help find workarounds to problems, some of which may be automatable by Al and data science. System and method examples according to the present disclosure can help personnel recognize and understand their weaknesses and strengths, and motivate and train personnel to voice their ideas and opinions. System and method examples according to the present disclosure can create collaboration opportunities, encourage employees to participate in social and group activities, and help to remove discriminations and prevent falling apart. System and method examples according to the present disclosure canassist personnel with finding and utilizing learning materials and may suggest mentors. System and method examples according to the present disclosure can provide frequent and beneficial feedback to improve personnel soft skills. System and method examples according to the present disclosure can additionally improve upon enterprise knowledge integration and utilization.

[0026] For example, the multimodal Al engine can receive data from a number of sources and establish a baseline of what constitutes usual patterns of a user’s voice, speech, mannerisms, expressions, and the like. By combining modalities (e.g., data types from multiple sensors), the multimodal Al engine can be trained on data sets that can contain complementary information, leading to improved predictions and accuracy in identifying abnormalities in patterns (e.g., speech, voice, mannerisms, etc.) of the user. Further, the use of a multimodal Al engine enables disclosed systems to vary the weight given to data from each source, such that the multimodal Al engine can determine which data sources are indicative of abnormalities. Thus, the multimodal Al engine can leverage large quantities of sensor data to make qualitative decisions in both analyzing and interpreting the data.

[0027] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. The following sections describe various additional features and examples with reference to the drawings in which like numerals indicate like elements, and directional descriptions are used to describe the illustrative examples but, like the illustrative examples, should not be used to limit the present disclosure.

[0028] Referring now to the drawings, FIG. 1 is a block diagram depicting one example of an operating environment in which an artificial intelligence (Al) system 100 with an Al meeting assistive engine 102 can be utilized to, among other things, facilitate and optimize enterprise personnel communications and, as a result, foster personnel growth and development. More specifically, FIG. 1 depicts examples of various hardware and Al components of the Al system 100 according to aspects of the present disclosure. The Al system 100 may be a specialized computing system that can be used for processing large amounts of data and effectuating learning by the Al meeting assistive engine 102.

[0029] As shown in FIG. 1 , the Al meeting assistive engine 102 of the Al system 100 may obtain various types of information from a multitude of sources. For example, the Al meeting assistive engine 102 may receive or otherwise obtain information stored in a data repository relative to a given user 104 of the Al system 100. The stored information may include, for example and without limitation, a photograph 106 of or other similar identifying information about the user 104. The Al meeting assistive engine 102 may also receive or otherwise obtain user information such as, an avatar 108 or another representation of the likeness of the user 104.

[0030] In addition to receiving or otherwise obtaining user personal information, the Al meeting assistive engine 102 may also receive or otherwise obtain information regarding a given meeting 110 that the user 104 will host or attend, such as but not limited to, the date, time, length, and subject matter of the meeting 110, as well as the identities of other participants in the meeting 110. The Al meeting assistive engine 102 may also receive or otherwise obtain technical information about the meeting 110, such as whether the meeting will include the use of video or audio hardware, or the use of particular programs or applications, such as presentation applications.

[0031] The Al meeting assistive engine 102 may obtain meeting information from a number of different meeting or scheduling applications, such as but not limited to, Microsoft Outlook®, Microsoft Teams®, Cisco Webex®, or Zoom®. The Al meeting assistive engine 102 may also obtain scheduling information from project management or workflow management applications 112, such as for example, and without limitation, Jira or ServiceNow®.

[0032] The Al meeting assistive engine 102 can also be provided with a number of different settings 114. The settings 114 may include, for example, default user settings such as user images or avatars, privacy settings, default meeting settings, hardware settings such as but not limited to camera and microphone settings, and settings of any other type related to user activities with which the Al meeting assistive engine 102 may assist. The settings 114 may also include, without limitation, information that governs the operation of the Al meeting assistive engine 102. Such operational settings may include rules directed to, for example, privacy, security, or ethics.

[0033] As may also be observed in FIG. 1 , the Al meeting assistive engine 102 can acquire or learn real time characteristics about the user 104 or the environment in which the user is located, such as via signals from one or more sensors. The sensors can take the form of hardware devices 116 such as cameras, microphones, smart phones, smart watches, and other portable smart devices, virtual assistant devices, physiological characteristic measurement devices, virtual or augmented reality devices, or other hardware devices that may be present in different enterprise environments or might be otherwise owned and utilized by the user 104. Correspondingly, the user characteristic-identifying signals received by the Al meeting assistive engine 102 may correspondingly be, in at least some examples, image, video, audio, other types of signals associated with such common hardware devices. In examples, where user characteristicidentifying signals received by the Al meeting assistive engine 102 may also emanate from user fitness monitors or other types of health monitors, the signals may be associated with physiological user characteristics such as, for example, body temperature or heart rate. The user 104 or user environment characteristics acquired or learned by the Al meeting assistive engine 102 may be used by the Al meeting assistive engine 102 to help predict how best to assist the user 104 with respect to effectively and efficiently completing a given task, to predictively provide the user 104 with advice or information relative to the user’s participation in a meeting, to advise the user 104 how to generally perform better in front of others, such as suggesting changes to the way the user 104 speaks, moves, or otherwise behaves in the real-world or virtual presence of others, etc.

[0034] According to examples of the present disclosure, the Al meeting assistive engine 102 can also acquire information and knowledge that may be useful to the user 104 in completing tasks, problem solving, or communicating with others. This information and knowledge maybe acquired from a number of sources, which may be both unique to the enterprise or publicly available. For example, having learned or having been otherwise informed of a problem or task on which the user 104 is working, the Al meeting assistive engine 102 may acquire information or knowledge from a knowledge source associated with the user 104 and / or the problem or task to be completed. With respect to the Al system 100 example illustrated in FIG. 1 , the Al meeting assistive engine 102may acquire information or knowledge through an enterprise intranet or by using the Internet 118. Thus, the Al meeting assistive engine 102 can access and search essentially any knowledge or information repository that is communicatively coupled to the intranet. Likewise, the Al meeting assistive engine 102 can also acquire information and knowledge from any publicly accessible website. As one example, FIG. 1 indicates that the Al meeting assistive engine 102 may use an Internet search engine 120 to access an online encyclopedia. As indicated, in at least some examples, the Al meeting assistive engine 102 may also save information and knowledge to the knowledge or information repositories and applications or programs to which it is communicatively coupled.

[0035] As may be further observed according to the example presented in FIG. 1 , the Al meeting assistive engine 102 can be, and is preferably, multimodal in nature. That is, while most Al and machine learning models focus on a single modality of data (e.g., natural language processing models, image classification models, or speech recognition models), the multimodal Al meeting assistive engine 102 model can flexibly interact with and utilize information from multiple modalities simultaneously. The multiple modalities may serve as both model inputs and as model outputs. Since humans experience the world in a multimodal manner - i.e. , through seeing, hearing, feeling, smelling, tasting, etc. - the Al meeting assistive engine 102 can better understand and predict user behavior and needs by interpreting multimodal signals. Combining modalities can make complementary information and metadata available and can consequently improve the accuracy of the Al meeting assistive engine 102 on even singlemodality tasks. However, the multiple modalities can have different quantitative influences over the prediction output of the Al meeting assistive engine 102, and there can be varying levels of noise and conflicts between modalities. Consequently, according to examples of the present disclosure, each Al component of the overall system may focus on only one or two of a larger number of possible modalities associated therewith, and the importance (weight) of those modalities may be considered.

[0036] As described above, the Al meeting assistive engine 102 can include or otherwise be associated with a plurality of individual Al components. Theseindividual Al components may include for example, and without limitation, an Al video component 150, an Al audio component 200, an Al speech component 250, an Al knowledge component 300, an Al involvement component 350, an Al psychology component 400, and an Al feedback component 450. The specific functionality of each of the individual Al components 150-450 is described in more detail below. Generally speaking, however, the Al meeting assistive engine 102 can provide to the individual Al components 150-450 various pieces and types of information it has learned or received from other sources, such as for example, the sources described above. Likewise, the Al meeting assistive engine 102 may receive input from any or all of the individual Al components 150-450 relative to providing predictive assistance to the user 104.

[0037] The Al video component 150 represented in FIG. 1 , is depicted in more detail in FIG. 2. As shown, the Al video component 150 may be communicatively coupled to the data repository referred to with respect to Figure 1 , which has stored therein one or more photographs 106 and one or more avatars 108 associated with the user 104. When, for example, the user 104 is preparing to participate in a meeting, the Al video component 150 may determine, such as by way of a query, 120 whether the user 104 prefers to be identified by way of a photograph 106 or an avatar 108. The answer to the query 120 may be obtained from settings 152 that are associated with the Al video component 150 or may be determined by a real time selection that is presented to the user 104.

[0038] As also indicated in FIG. 2, a camera 154 such as but not limited to a web camera, may be used to capture a real-time still image of the user 104 or to present live video of the user 104 during the meeting. According to one example, the Al video component 150 can apply facial expressions, etc., from the real-time still image to the stored photograph(s) 106 or avatar(s) 108 associated with the user 104. According to another example, the Al video component 150 can receive or otherwise obtain a live video signal from the camera 154, and may analyze the live video signal to evaluate various visible physical characteristics of the user 104. For example, the Al video component 150 may evaluate the facial expressions, and head or other body movements of the user 104 during the course of a meeting. The Al video component 150 may be trained to recognize abnormal (e.g., unusual or undesirable) facial expressions or bodymovements of the user 104, or of users in general, such as facial expressions or body movements that are related to anxiety or confusion.

[0039] When unusual or undesirable facial expressions or body movements of the user 104 are detected, the Al video component 150 may operate to modify the video image of the user 104 using one or more image processing techniques, which may include the use of or reference to preexisting stored still images or video of the user. An Al-modified user image 156 that is based on a modified representation of the user 104 generated through image processing, may be substituted for the actual video image of the user 104 by the Al video component 150. The length of time for which the image substitution persists may depend on whether there is a continuation of the unusual or undesirable user facial expressions or body movements that triggered the substitution of the Al-modified user image 156. Alternatively, the substitution may persist for some predefined period of time at which point the live video image of the user 104 is returned, or the substitution may persist until the end of the meeting.

[0040] The ability and permission of the Al video component 150 to modify the appearance of the user 104 may be set within configuration parameters of the Al video component 150. For example, the Al video component 150 may include user selectable and settable modification, filtering, restriction and other parameters that can be used to govern the actions of the Al video component 150 relative to the appearance of the user 104 while monitoring signals from the camera 154, such as during a virtual meeting or otherwise. In an example, the configuration parameters of the Al video component 150 may be accessed through the settings 152, which may be contained in a file that is associated with the Al video component 150. In this manner, the user 104 has the real-time option of allowing the Al video component 150 to filter or alter facial expressions or body movements that might result from anxiety or confusion, as well as the option to present an unfiltered or unaltered user image.

[0041] The Al audio component 200 represented in FIG. 1 , is depicted in more detail in FIG. 3. As shown, operation of the Al audio component 200 may be controlled, at least in part, by a settings file 202. The settings file 202 may include various configuration parameters of the Al audio component 200. For example, the Al audio component 200 may include user selectable and settablemodification, filtering, restriction and other parameters that can be used to govern the actions of the Al audio component 200 relative to the sound of the user’s voice, such as parameters that dictate the permissions of the Al audio component 200 with respect to correcting various vocal anomalies of the user 104.

[0042] As also indicated in FIG. 3, the Al audio component 200 can receive or otherwise obtain audio signals from a microphone 204 that may be used to capture the voice of the user 104 as the user speaks during the course of a meeting. According to one example, the Al audio component 200 can analyze the audio signal from the microphone 204 to detect the presence of any abnormalities. For example, the Al audio component 200 may be trained to recognize abnormal (e.g., unusual or undesirable) vocal characteristics, such as but not limited to shakiness that might result from nervousness; stammering, convoluted, or nonsensical speech that might result from anxiety; mumbling, rambling, or disjointed speech that might result from confusion; the use of extraneous expressions; or non-language noises such as coughs, sneezes, or deep breathing.

[0043] When unusual or undesirable vocal characteristics of the user 104 are detected, the Al audio component 200 may operate to modify the voice of the user 104 using one or more audio processing techniques, which may include the use of filters or synthesized substitute speech built on preexisting stored speech of the user. Filters may be used, for example, to reduce or remove voice shakiness or non-language noises such as coughs, sneezes, or deep breathing. Frequency changes may also be effectuated to make the voice of the user 104 stronger or softer.

[0044] The Al audio component 200 may create an Al-modified user voice 206 that is based on a modified voice of the user’s voice generated through audio processing (e.g., filtering), and may substitute the Al-modified user voice 206 for the actual voice of the user 104. The length of time for which the voice substitution persists may depend on whether there is a continuation of the unusual or undesirable vocal characteristics of the user 104 that triggered the substitution of the Al-modified user voice 206. Alternatively, the substitution maypersist for some predefined period of time at which point the actual voice of the user 104 is returned, or the substitution may persist until the end of the meeting.

[0045] The Al speech component 250 represented in FIG. 1 , is depicted in more detail in FIG. 4. The Al speech component 250 can be used in both virtual and in-personal meetings. In the case of in-person meetings, assistance from the Al speech component 250 can be provided to the user 104 through, for example, a headset or a smart device (e.g., a phone, a watch, a VR or AR headset or glasses, etc.).

[0046] As shown in FIG. 4, operation of the Al speech component 250 may be controlled, at least in part, by a settings file 252. The settings file 252 may include various configuration parameters of the Al speech component 250. For example, the Al speech component 250 may include user selectable and settable modification, filtering, restriction and other parameters that can be used to govern the actions of the Al speech component 250 relative to user’s speech, such as parameters that dictate the permissions of the Al speech component 200 with respect to correcting various anomalies associated with grammar or word pronunciation by the user 104, or with respect to alteration of an accent of the user 104.

[0047] In a similar manner to the Al audio component 200 illustrated in FIG. 3, the Al speech component 250 can receive or otherwise obtain audio signals from the microphone 204 that may be used to capture the voice of the user 104 as the user speaks during the course of a meeting. According to one example, the Al speech component 250 can analyze the audio signal from the microphone 204 to detect the presence of any abnormalities associated with the user’s grammar, vocabulary, accent, or pronunciation. For example, the Al speech component 250 may be trained to recognize anomalies in the speech of the user 104, such as when the user 104 exhibits generally bad grammar, or mispronounces or uses incorrect words. The Al speech component 250 may also be trained to recognize anomalies in the speech patterns of the user 104, such as when the user 104 mumbles, unexpectedly stops talking, or repeatedly searches for words. The Al speech component 250 may also be trained to analyze the accent of the user 104 and may modify the user’s accent if the accent is excessive and modification is necessary to understanding the user’s speech.This may be particularly useful, for example, to minimize or eliminate possible miscommunication or misunderstanding in cases where not all meeting participants speak the same native language.

[0048] When anomalies associated with, for example, the grammar, vocabulary, accent, or pronunciation of the user 104 are detected, the Al speech component 250 may operate to modify the speech of the user 104 using one or more audio or frequency processing techniques, which may include the use of filters or synthesized substitute speech in the voice of the user but with proper grammar, vocabulary, accent, and pronunciation. Filters may also be used, for example, to reduce or remove repeatedly used extraneous expressions, such as but not limited to, the “uh” or “umm” utterances that are commonly used during pauses or transitions in speech.

[0049] The Al speech component 250 may create Al-modified user speech 256 that is based on a modified user voice generated through audio or frequency processing, and may substitute the Al-modified user speech 256 for the actual speech of the user 104. The length of time for which the speech substitution persists may depend on whether there is a continuation of the speech anomalies of the user 104 that triggered the substitution of the Al-modified user speech 256. Alternatively, the substitution may persist for some predefined period of time at which point the actual speech of the user 104 is returned, or the substitution may persist until the end of the meeting.

[0050] The Al knowledge component 300 represented in FIG. 1 , is depicted in more detail in FIG. 5. As shown, the Al knowledge component 300 can be communicatively coupled to a knowledge source 302, which may include, for example and without limitation, a user profile element 304, a user learning element 306, a user history element 308, a user projects element 310, and a user reviews element 312. In an example, the user profile element 304 may include information such as the user’s curriculum vitae or resume, a description of the user’s role in the enterprise, an email history, a chat history, etc. In an example, the user learning element 306 may include information such as user learnings, a user meeting history, images or videos, etc. In an example, the user history element 308 may include information such as the user’s work history, a conference attendance history, a history of past (internal or external) searchesfor information conducted by the user, a description of the user’s skills, etc. In an example, the user projects element 310 may include information such as a history of projects completed by the user, user works in progress, plans, etc. In an example, the user reviews element 312 may include information such as past performance reviews, management suggestions, etc. The Al knowledge component 300 can extract information from or save information to any of the individual components 304-312 of the knowledge source 302.

[0051] In an example, the Al knowledge component 300 can be communicatively coupled to one or more applications 322 such as but not limited to J ira or ServiceNow, and may receive or otherwise obtain problem information such as, without limitation, incident tickets or problem notifications therefrom. The Al knowledge component 300 can also be communicatively coupled to one or more other sources of information. For example, as represented in FIG. 5, the Al knowledge component 300 is communicatively coupled to a search engine 324 by which the Al knowledge component 300 can obtain information of interest, such as from an online encyclopedia or other information source.

[0052] The Al knowledge component 300 may also receive or otherwise obtain information 314 about an upcoming meeting at which the user 104 will present certain subject matter to other meeting participants. The meeting information 314 may include but is not limited to, the meeting topic, a further explanation of the subject matter to be presented by the user 104 such as may be obtained from descriptions in meeting invites, and a list of meeting invitees or participants.

[0053] As may also be observed in FIG. 5, the Al knowledge component 300 may include or otherwise be associated with several modules 316, 318, 320 that can assist the Al knowledge component 300 in providing the user 104 with information that is predicted to improve communications between the user 104 and other meeting participants with or to whom the user will be speaking. For example, a speaker support module 316 and an audience support module 318 may receive or otherwise obtain from the Al knowledge component 300, information such as the meeting topic, subject matter description, and list of meeting invitees or participants that the Al knowledge component 300 previously received or otherwise obtained. Using this information, the speaker supportmodule 316 can make predictions regarding materials that the user 104 may wish to review offline prior to the meeting to help the user 104 provide a more comprehensive and understandable description of the meeting topic for the other meeting participants.

[0054] The speaker support module 316 and the audience support module 318 may also receive or otherwise obtain from the Al knowledge component 300, information such as but not limited to profiles and backgrounds of various other meeting participants. In this manner, the Al knowledge component 300 can predict the likelihood that, and the degree to which, a knowledge gap may exist between the user 104 and the other meeting participants regarding the meeting topic to be presented by the user 104 during the meeting. In a case where it is predicted that a knowledge gap exists, the first speaker support module 316 may suggest to the user 104 that the user present real time expansions or explanations of the subject matter when, for example, it is determined that the user 104 has spoken too technically (e.g., too domain specific, too business specific, etc.) in a meeting with non-technical participants.

[0055] Using the previously described information received or otherwise obtained from the Al knowledge component 300, the second audience support module 318 may operate, for example, to help fill in any knowledge gaps regarding the subject matter presented by the user during the meeting. In this regard, the audience support module 318 may, for example, clarify for the other participants, abbreviations employed by the user 104, or may provide the other participants with definitions or explanations of domain or business specific terms being used. The audience support module 318 may also expand certain topics discussed by the user 104 that are predicted to be within the knowledge gap of the other participants. For example, non-technical participants might need complementary explanations when attending a technical meeting.

[0056] The enterprise personnel network support module 320 can receive or otherwise obtain from the Al knowledge component 300, information that the Al knowledge component 300 previously received or otherwise obtained from other information sources such as the applications or programs 322. The enterprise personnel network support module 320 can connect enterprise personnel that are predicted by the Al knowledge component 300 to have similar backgroundsor experiences. For example, the enterprise personnel network support module 320 may receive or otherwise obtain from the Al knowledge component 300, information regarding a particular problem or incident currently being worked on by one or more personnel. Additionally, the enterprise personnel network support module 320 may also receive or otherwise obtain from the Al knowledge component 300, an identification of one or more other enterprise personnel who have previously worked on and resolved a similar problem or incident. Using this information, the enterprise personnel network support module 320 can suggest a meeting between the one or more personnel currently working on the problem or incident and one or more enterprise personnel who have previously worked on and resolved a similar problem or incident.

[0057] Using information received or otherwise obtained from the Al knowledge component 300, such as incident ticket or problem notification information, the enterprise personnel network support module 320 may also suggest who to contact in regard to a particular problem or incident and how to resolve the problem or incident by providing a summary of the problem or incident description associated with the incident ticket or problem notification, as well as any workaround information associated with the incident ticket or problem notification. Workaround suggestions may include automation of a problematic task, by using Al or otherwise. In this regard, the enterprise personnel network support module 320 may have knowledge of similar tasks that have already been automated with Al as well as other proposed solutions, and may identify other tasks that can be similarly automated and the Al solutions that can help facilitate the automation.

[0058] While not shown in FIG. 5, the Al knowledge component 300 can also be connected to the one or more human resource (HR) systems of the enterprise, and may be used create job descriptions or specify the criteria and various capabilities that newly hired personnel should possess to help fill personnel gaps in various areas of the enterprise.

[0059] The Al knowledge component 300 may also have access to surveys and other information, and can learn to provide a list of suggestions by mining of the survey (e.g., through topic modeling, sentiment analysis, text mining, etc.) to improve experiences associated with using enterprise products and services.

[0060] The Al involvement component 350 represented in FIG. 1 , is depicted in more detail in FIG. 6. As shown, the Al involvement component 350 can be communicatively coupled to a knowledge source, such as the knowledge source 302 shown in FIG. 5. As previously described, the knowledge source 302 may include, for example and without limitation, a user profile element 304, a user learning element 306, a user history element 308, a user projects element 310, and a user reviews element 312. In an example, the user profile element 304 may include information such as the user’s curriculum vitae or resume, a description of the user’s role in the enterprise, an email history, a chat history, etc. In an example, the user learning element 306 may include information such as user learnings, a user meeting history, images or videos, etc. In an example, the user history element 308 may include information such as the user’s work history, a conference attendance history, a history of past (internal or external) searches for information conducted by the user, a description of the user’s skills, etc. In an example, the user projects element 310 may include information such as a history of projects completed by the user, user works in progress, plans, etc. In an example, the user reviews element 312 may include information such as past performance reviews, management suggestions, etc. The Al involvement component 350 can extract information from or save information to any of the individual components 304-312 of the knowledge source 302.

[0061] In an example, the Al involvement component 350 can be communicatively coupled to one or more applications, such as but not limited to the Jira or ServiceNow applications 322 shown in FIG. 5, and may receive or otherwise obtain information therefrom. The Al involvement component 350 can also be communicatively coupled to one or more other sources of information. For example, the Al knowledge involvement component 350 is communicatively coupled to a search engine, such as the search engine 324 shown in FIG. 5, by which the Al involvement component 350 can obtain information of interest, such as from an online encyclopedia or other information source.

[0062] As may also be observed in FIG. 6, the Al involvement component 350 may receive or otherwise obtain signals from a camera 352, which may be deployed by the user 104 or other personnel, to image non-presenting meeting participants 354 during an in-person meeting and at time when the user 104 oranother person is speaking or otherwise presenting information. Alternatively, the Al involvement component 350 may receive or otherwise obtain signals from the cameras of other meeting participants 354 during a virtual meeting and at time when the user 104 or another person is speaking or otherwise presenting information.

[0063] The Al involvement component 350 can analyze still and or video images from the camera 352 or from individual participant cameras, to evaluate, such as by way of image processing techniques, the facial expressions, eye movements, etc., of the other meeting participants 354. The Al involvement component 350 may use the observed facial expressions, eye movements, etc., of the other meeting participants 354 to predict the participation rate of the other meeting participants 354.

[0064] Based on the predicted participation rate, examples of the Al involvement component 350 can, for example, advise the user 104 or other speaker to change the topic, provide more elaborations, ask questions, or improve their speech, so as to cause the other meeting participants 354 to become more engaged in the meeting. Alternatively, or in addition thereto, examples of the Al involvement component 350 can notify one or more of the other meeting participants 354 that a particular part of the presentation or speech of the user 104 or other speaker may be beneficial to a project(s) with which they are involved, or might cover a knowledge gap they have relative to the meeting topic. Examples of the Al involvement component 350 can also suggest to the other meeting participants 354 that they ask relevant or proper questions of the user 104 or other presenter or speaker to encourage the other meeting participants 354 to become more involved in the meeting.

[0065] The Al psychology component 400 represented in FIG. 1 , is depicted in more detail in FIG. 7. As shown, the Al psychology component 400 may receive or otherwise obtain signals from a camera 402 of the user 104. The signals from the camera 402 may convey facial expressions, eye movements, or other characteristics of the user 104. The Al psychology component 400 may also receive or otherwise obtain signals from another device, such as but not limited to a fitness monitoring device or a smart watch that outputs signals indicative of one or more physiological characteristics of the user 104. The oneor more physiological characteristics may be, but are certainly not limited to, the heartrate, blood pressure, blood oxygen level, or body temperature of the user 104.

[0066] Examples of the Al psychology component 400 can use a psychological issues analyzer 406 to evaluate one or more of the facial expressions, eye movements, or other similar characteristics of the user 104 captured by the camera 402, either alone or in combination with one or more of the physiological characteristics of the user 104 captured by the fitness monitoring device, smartwatch, etc., to recognize anxiety or other detrimental emotions in the user 104. Once such emotions of the user 104 have been recognized, examples of the Al psychology component 400 may consult online or internally stored psychology materials 408, such as textbooks, journals, and other sources of information, for guidance on how the quell the user’s anxiety or other undesirable emotions. In one example, the Al psychology component 400 can employ a psychological mechanism 410 to alter the atmosphere of the meeting in a way that is calming to the user 104. For example, when the meeting is a virtual meeting, the Al psychology component 400 can change the theme of a profile associated with the user 104 or one or more of the other meeting participants to a more relaxing color or theme, or may suggest ice-breaking or other generally lighthearted conversations to both the user 104 and one or more of the other meeting participants to help lower the anxiety level of the user 104.

[0067] The Al feedback component 450 represented in FIG. 1 , is depicted in more detail in FIG. 8. As shown, the Al feedback component 450 may receive or otherwise obtain recorded video of a past meeting 452 as well as a feedback log or survey data 454 obtained from meeting participants other than the user 104 or another speaker. When the acquired information is a feedback log 454, for example, the Al feedback component 450 may employ processing techniques 456 such as but not limited to video processing, image processing, audio processing, or text processing, to evaluate and understand the feedback of the other meeting participants. The Al feedback component 450 may also summarize feedback and provide the same to the user 104 or the other meeting participants to effectuate improvement of the users presentation skills or the involvement and participation skills of the other meeting participants. The Alfeedback component 450 may also provide constructive suggestions along with the summary. A summary compiled by the Al feedback component 450 may also allow personnel who were unable to attend the associated meeting to efficiently and effectively understand the main points of the topic discussed.

[0068] FIG. 9 is a block diagram depicting an example of a computing device 500 suitable for implementing the functionality of the Al meeting assistive engine 102 and the associated Al components 150-450 shown in and described with respect to FIGS. 1 -8, according to an example of the present disclosure. The components depicted in FIG. 9 are provided for illustrative purposes only. Therefore, while FIG. 9 depicts the computing device 500 as including certain components, other examples of the computing device may involve more, fewer, or different components than those shown in FIG. 9.

[0069] The computing device 500 can include a network interface 502 for communicating with other devices in the Al system 100 over a network 504. Different types of networks 502 can be used for communication within the Al system 100 of FIG. 1 , For example, the network 502 may be a public data network, a private data network, or some combination thereof. A data network may include one or more of a variety of network types, including a wireless network, a wired network, or a combination of a wired and wireless network. Examples of suitable networks include, without limitation, the Internet, a personal area network, a local area network (“LAN”), a wide area network (“WAN”), or a wireless local area network (“WLAN”). A wireless network may include a wireless interface or a combination of wireless interfaces. A wired network may include a wired interface. The wired or wireless networks may be implemented using routers, access points, bridges, gateways, or the like, to connect devices or components in the data network.

[0070] The other devices may be sensors 506 of various types. For example, and without limitation, the sensors 506 may comprise hardware devices such as video (web) cameras 508, microphones 510, physiological monitoring devices 512 such as personal fitness monitors or smart watches, or other types of sensors 514.

[0071] As shown in FIG. 9, the computing device 500 can include a processor 516 that is communicatively coupled to a memory 518. The processor 516 canexecute computer-executable program code 520 stored in the memory 518, can access information stored in the memory 518, or both. The memory 518 can store program code in the form of instructions 522 that, when executed by the processor 516, causes the processor 516 to perform the operations described herein. The program code 520 may include machine-executable instructions 522 that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others.

[0072] Examples of a processor 516 can include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other suitable processing device. The processor 516 can include any suitable number of processing devices, including one. In addition to communicating with the memory 518, the processor 516 can include a memory.

[0073] The memory 518 can include any suitable non-transitory computer- readable medium. The computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable program code or other program code. Non-limiting examples of a computer-readable medium can include a magnetic disk, memory chip, optical storage, flash memory, storage class memory, ROM, RAM, an ASIC, magnetic storage, or any other medium from which a computer processor can read and execute program code. The program code may include processorspecific program code generated by a compiler or an interpreter from code written in any suitable computer-programming language. Examples of suitable programming language can include Hadoop, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.

[0074] The computing device 500 may also include a number of external or internal devices such as input or output devices. For example, the computing device 500 is illustrated with an input / output interface 524 that can receive inputfrom input devices or provide output to output devices. A bus 526 can also be included in the computing device 500. The bus 526 can communicatively couple one or more components of the computing device 500.

[0075] In some examples, the computing device 500 can include one or more output devices. One example of such an output device may be the network interface device 502. The network interface device 502 can include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks, such as but not limited to, the public data network 504 depicted in FIG. 9. Non-limiting examples of the network interface device 502 can include an Ethernet network adapter, a modem, etc.

[0076] Another example of an output device can include a presentation device (not shown in FIG. 9). A presentation device can include any device or group of devices suitable for providing visual, auditory, or other suitable sensory output. Non-limiting examples of a presentation device can include a touchscreen, a monitor, a speaker, a separate mobile computing device, etc.

[0077] As further represented in FIG. 9, the computing device 500 can execute the program code 520 that includes the instructions 522 to cause the processor 516 of the computing device 500 to perform the enterprise personnel communications facilitation and optimization operations described herein. For example, the instructions 522 in the program code 520 may cause the processor to: cause a multimodal artificial intelligence (Al) meeting assistive engine to receive or otherwise obtain, from a sensor, a signal representing a physical characteristic of a plurality of physical characteristics of a user, the signal transmitted by the sensor via one or more communication mediums such that the physical characteristic of the user is observable by at least one third party; evaluate, by the multimodal Al meeting assistive engine, the physical characteristic of the user by analyzing the signal from the at least one sensor; determine, by the multimodal Al meeting assistive engine, based on analyzing the signal, that the physical characteristic of the user is abnormal; and in response to determining that the physical characteristic of the user is abnormal, modify the signal, by the multimodal Al meeting assistive engine, such that the physical characteristic of the user appears normal to the at least one third party.

[0078] The program code 520 may be resident in any suitable computer- readable medium, such as a non-transitory computer readable medium, and may be executed on any suitable processing device. For example, as depicted in FIG. 9, the program code 520 for performing the various operations described herein can reside in the memory 518 of the computing device 500 along with program data 528 associated with the program code 520. Executing training or application of an Al model on the computing device 500 can configure the processor 516 to perform the operations described herein.

[0079] Data from the sensors 506, setting or parameter data associated with the Al meeting assistive engine 102 or any or all of the Al components 150-450 can be stored by the computing device 500. For example, such data, as well as other data, can be stored in one or more local data stores 530 connected to the bus 526, or in one or more locally or remotely located databases 532.

[0080] FIG. 10 is a block diagram depicting a transformer architecture 600 that can be used with the Al meeting assistive engine 102 shown in FIG. 1 and described above to facilitate and optimize enterprise personnel communications according to an example of the present disclosure. Most known Al and machine learning (ML) work has focused on models that deal with a single modality of data (e.g., language models, image classification models, or speech recognition models). While there has been progress with single modality Al and ML models, multimodal models have the advantage of being able to flexibly handle many different modalities simultaneously, both as model inputs and as model outputs.

[0081] In light of the above disclosure, an Al meeting assistive engine according to an example of the present disclosure preferably utilizes a selfsupervised multimodal architecture and strategy. More specifically, since Transformers are effective for learning semantic video, audio, image, and text representations, the Al meeting assistive engine model can project each modality into a feature vector and feed it into a Transformer encoder. For example, the Transformer architecture employed with an Al meeting assistive engine according to examples of the present disclosure may be a combination of two Transformer architectures, such as a combination of a Vision Transformer (ViT) natural language processing (NLP) model and a Bidirectional Encoder Representations from Transformers (BERT) language representation model.

[0082] Self-supervised learning also refers to a machine learning paradigm, and corresponding methods, for processing unlabeled data. The same strategy can be applied along with a combined Transformer architecture within the various multimodal Al components associated with the Al meeting assistive engine to capture and further process data such as images, videos, audio, and text, through various required tasks such as, for example, classification and decision making. FIG. 11 is a flowchart representing a method of facilitating and optimizing enterprise personnel communications according to an example of the present disclosure. The example method of FIG. 11 can include, as depicted at block 700, receiving or otherwise obtaining from a sensor, by a multimodal Al meeting assistive engine of a computing device, a signal representative of a physical characteristic of a plurality of physical characteristics of a user, where the signal is transmitted by the sensor via one or more communication mediums such that the physical characteristic of the user is observable by at least one third party. At block 702, the method can additionally include causing, by a processor of the computing device, the multimodal Al meeting assistive engine to evaluate the physical characteristic of the user by analyzing the signal from the at least one sensor. At block 704, the processor of the computing device causes the multimodal Al meeting assistive engine to determine, based on analyzing the signal, that the physical characteristic of the user is abnormal, and at block 706, in response to determining that the physical characteristic of the user is abnormal, the processor of the computing device causes the multimodal Al meeting assistive engine to modify the signal such that the physical characteristic of the user appears normal to the at least one third party.

[0083] Effective and efficient personnel communications within an enterprise setting can be difficult, especially in the case of a very large enterprise. Artificial intelligence and machine learning models can be effective mechanisms for facilitating and optimizing enterprise personnel communications. However, due to the various types of personnel communications that commonly occur, known Al models that learn from only a single modality of data may be incapable of effectively facilitating and optimizing multimodal personnel communications.

[0084] System and method examples according to the present disclosure can alleviate the problems mentioned above with respect to the use of Al that learnsfrom only a single modality of data. For example, by employing a multimodal Al meeting assistive engine, personnel communications can be facilitated and optimized by the Al meeting assistive engine regardless of the particular mode of communication involved.

[0100] The foregoing description of certain examples, including illustrated examples, has been presented only for purposes of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications, adaptations, and uses thereof will be apparent to those skilled in the art without departing from the scope of the disclosure.

Claims

What Is Claimed Is:1 . A system comprising: a plurality of sensors configured to detect a plurality of different physical characteristic of a user, and to transmit via one or more communication mediums, signals representative of the plurality of physical characteristics of the user such that the plurality of physical characteristics of the user are observable by at least one third party; and a computing device comprising: a multimodal artificial intelligence (Al) meeting assistive engine configured to receive or otherwise obtain the signals from the sensors; a processor; and a non-transitory computer-readable medium including instructions that are executable by the processor for causing the processor to perform operations comprising: evaluating, by the multimodal Al meeting assistive engine, the plurality of physical characteristics of the user by analyzing the signals from the sensors; determining, by the multimodal Al meeting assistive engine, based on analyzing the signals, that at least one physical characteristic of the plurality of physical characteristics of the user is abnormal; and in response to determining that the at least one physical characteristic of the user is abnormal, modifying, by the multimodal Al meeting assistive engine, the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

2. The system of claim 1 , wherein the sensors are devices selected from the group consisting of a camera, a microphone, a physiological characteristicmeasurement device, a smart phone, a smart watch, a virtual assistant, and a virtual or augmented reality device.

3. The system of claim 2, wherein the physical characteristics of the user are selected from the group consisting of facial expressions, body movements, voice shakiness, stammering speech, convoluted or nonsensical speech, mumbling, rambling speech, disjointed speech, extraneous expressions, poor grammar, poor vocabulary, mispronunciation, coughs, sneezes, and deep breathing.

4. The system of claim 1 , wherein the multimodal Al meeting assistive engine includes or is communicatively coupled to one or more Al components selected from the group consisting of an Al video component, an Al audio component, an Al speech component, an Al knowledge component, an Al involvement component, an Al psychology component, and an Al feedback component.

5. The system of claim 4, wherein the signals from the plurality of sensors are useable by the multimodal Al meeting assistive engine and one or more of the Al components to facilitate and optimize various different modes of user communications.

6. The system of claim 1 , wherein the multimodal Al meeting assistive engine is communicatively coupled to at least one settings file, the settings file including various settings information that governs operation of the Al meeting assistive engine and is elected from the group consisting of default user settings including user images and avatars, default meeting settings, privacy settings, security settings, ethics settings, and hardware settings including camera and microphone settings.

7. The system of claim 1 , wherein the multimodal Al meeting assistive engine is communicatively coupled to a source of meeting information that is useable by the multimodal Al meeting assistive engine to assist the user inscheduling a meeting, preparing for a meeting, participating in a meeting, or effecting an improvement in performance over a past meeting.

8. The system of claim 1 , wherein the multimodal Al meeting assistive engine is communicatively coupled to a knowledge source from which the multimodal Al meeting assistive engine can extract information that is predicted to be of assistance to the user relative to solving a problem.

9. The system of claim 1 , wherein the multimodal Al meeting assistive engine utilizes a self-supervised architecture and strategy that models relationships between different modalities of data by projecting each modality into a feature vector and feeding it into a Transformer encoder.

10. The system of claim 9, wherein the Transformer encoder is a combination of a Vision Transformer (ViT) natural language processing (NLP) model and a Bidirectional Encoder Representations from Transformers (BERT) language representation model.

11. A non-transitory computer-readable medium comprising instructions that are executable by a processor of a computing device for causing the processor to perform operations comprising: causing a multimodal artificial intelligence (Al) meeting assistive engine to receive or otherwise obtain, from a plurality of sensors, signals representative of a plurality of physical characteristics of a user, the signals transmitted by the plurality of sensors via one or more communication mediums such that the plurality of physical characteristics of the user are observable by at least one third party; evaluating, by the multimodal Al meeting assistive engine, the plurality of physical characteristics of the user by analyzing the signals from the sensors; determining, by the multimodal Al meeting assistive engine, based on analyzing the signals, that at least one physical characteristic of the plurality of physical characteristics of the user is abnormal; andin response to determining that the at least one physical characteristic of the user is abnormal, modifying, by the multimodal Al meeting assistive engine, the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

12. The non-transitory computer-readable medium of claim 11 , wherein: the sensors are a devices selected from the group consisting of a camera, a microphone, a physiological characteristic measurement device, a smart phone, a smart watch, a virtual assistant, and a virtual or augmented reality device; and the physical characteristic of the user are selected from the group consisting of facial expressions, body movements, voice shakiness, stammering speech, convoluted or nonsensical speech, mumbling, rambling speech, disjointed speech, extraneous expressions, poor grammar, poor vocabulary, mispronunciation, coughs, sneezes, and deep breathing.

13. The non-transitory computer-readable medium of claim 11 , wherein the multimodal Al meeting assistive engine includes or is communicatively coupled to one or more Al components selected from the group consisting of an Al video component, an Al audio component, an Al speech component, an Al knowledge component, an Al involvement component, an Al psychology component, and an Al feedback component.

14. The non-transitory computer-readable medium of claim 11 , wherein the multimodal Al meeting assistive engine is communicatively coupled to at least one settings file, the settings file including various settings information that governs operation of the Al meeting assistive engine and is elected from the group consisting of default user settings including user images and avatars, default meeting settings, privacy settings, security settings, ethics settings, and hardware settings including camera and microphone settings.

15. The non-transitory computer-readable medium of claim 11 , wherein the multimodal Al meeting assistive engine is communicatively coupled to a source of meeting information that is useable by the multimodal Al meeting assistive engine assist the user in scheduling a meeting, preparing for a meeting, participating in a meeting, or effecting an improvement in performance over a past meeting.

16. A method comprising: receiving or otherwise obtaining, by a multimodal artificial intelligence (Al) meeting assistive engine of a computing device, from a plurality of sensors, signals representative of a plurality of physical characteristics of a user, the signals transmitted by the plurality of sensors via one or more communication mediums such that the plurality of physical characteristics of the user are observable by at least one third party; causing, by a processor of the computing device, the multimodal Al meeting assistive engine to evaluate the plurality of physical characteristics of the user by analyzing the signals from the plurality of sensors; causing, by the processor of the computing device, the multimodal Al meeting assistive engine to determine, based on analyzing the signals, that at least one physical characteristic of the plurality of physical characteristic of the user is abnormal; and in response to determining that the at least one physical characteristic of the user is abnormal, causing, by the processor of the computing device, the multimodal Al meeting assistive engine to modify the signal corresponding to the at least one physical characteristic of the user such that the at least one physical characteristic of the user appears normal to the at least one third party.

17. The method of claim 16, wherein: the sensors are devices selected from the group consisting of a camera, a microphone, a physiological characteristic measurement device, a smart phone, a smart watch, a virtual assistant, and a virtual or augmented reality device; andthe physical characteristics of the user are selected from the group consisting of facial expressions, body movements, voice shakiness, stammering speech, convoluted or nonsensical speech, mumbling, rambling speech, disjointed speech, extraneous expressions, poor grammar, poor vocabulary, mispronunciation, coughs, sneezes, and deep breathing.

18. The method of claim 16, wherein the multimodal Al meeting assistive engine includes or communicates with one or more Al components selected from the group consisting of an Al video component, an Al audio component, an Al speech component, an Al knowledge component, an Al involvement component, an Al psychology component, and an Al feedback component.

19. The method of claim 16, further comprising: receiving or otherwise obtaining, by the multimodal Al meeting assistive engine, from a different sensor, a signal representing a physical characteristic of at least one participant attending a virtual meeting or an in-person meeting at which the user is speaking; evaluating, by the multimodal Al meeting assistive engine, the physical characteristic of the at least one participant present in the meeting, by analyzing the signal from the different sensor; and based on evaluating the physical characteristic of the at least one participant present in the meeting, advising the user, by the multimodal Al meeting assistive engine in real time, to modify at least one speech characteristic of the user.

20. The method of claim 16, further comprising: receiving or otherwise obtaining, by the multimodal Al meeting assistive engine from a meeting information source, information about an upcoming meeting at which the user will present certain subject matter to other meeting participants; and based on the information received or otherwise obtained from the meeting information source, providing to the user, by the multimodal Al meeting assistive engine, information that is predicted by the multimodal Al meeting assistiveengine to improve communications between the user and the other meeting participants while the user is presenting the certain subject matter to the other meeting participants; wherein the information received or otherwise obtained from the meeting information source is selected from the group consisting of a meeting topic, a further explanation of the subject matter to be presented by the user, and a list of meeting invitees or participants.

Citation Information

Patent Citations

  • System and method for prediction based preemptive generation of dialogue content

    US20190251956A1

  • System and method for content focused conversation

    US20220293122A1