Privacy aware multi-modal generative auto-reply
By collecting multimodal information and using generative large-scale language models to generate personalized responses, this approach solves the problem that existing automatic reply systems cannot meet users' privacy preferences and current situations, achieving more flexible and accurate automatic replies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-03-10
AI Technical Summary
Existing automated response systems cannot generate personalized and detailed responses based on the user's current situation and privacy preferences, resulting in responses that are not flexible or personalized enough to meet the user's needs in different situations.
By collecting multimodal information, including text-based and non-text-based information, a generative large language model (LLM) is used to generate personalized responses. The model controller controls the activation and deactivation of information sub-modules and generates appropriate automatic reply messages based on user privacy preferences.
It enables the generation of personalized and flexible automatic responses based on the user's current status and privacy preferences, improving the adaptability and accuracy of the response and meeting the user's needs in different situations.
Smart Images

Figure CN121646779A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of priority of U.S. Non-Provisional Application No. 18 / 454,456, filed August 23, 2023; the entirety of which is incorporated by reference herein. BACKGROUND
[0003] Auto-reply is a feature in computer systems or software applications that generates an automatic action or response when certain events or conditions are met, such as receiving an email or support request, etc. Auto-reply functionality is commonly used in text messaging applications, email systems, customer support tools, and social media platforms to acknowledge receipt or provide preliminary information when a human responder can not be available to reply. Such automated responses can include generic information, pre-written answers to common questions, instructions for further action, etc.
[0004] In recent years, auto-reply has become an important feature of a variety of different platforms and devices, including traditional mobile operating systems such as Android and iOS. Auto-reply functionality can help users set up automatic replies to various forms of communication, including phone calls, text messages, or emails. Auto-reply functionality can be particularly useful in situations where it is prohibited to respond to an incoming communication in a timely manner, such as during a meeting, during vehicle transit, or in situations where the user finds themselves in an area lacking wireless service.
[0005] Messages generated by auto-reply can be preset, customized, or a mix of both, ranging from simple notifications to more detailed explanations. Certain systems incorporate auto-reply functionality natively within their built-in messages or email settings, while others can depend on third-party applications provided in an app store or software repository. After enabling the auto-reply feature, a pre-defined response can be automatically transmitted in response to an incoming call or message. Some platforms offer sophisticated configurations, allowing for customization of auto-reply messages, including specifying particular contacts for auto-reply, specifying unique time periods for auto-reply to be active, etc. SUMMARY
[0006] Various aspects include a method, which can be implemented in a processing system of a computing device, for providing an automated reply response, which can include: collecting multi-modal information about a user of the computing device; determining a current user situation based on the collected multi-modal information; determining a user privacy preference for an automated reply response; generating a prompt based on the selected multi-modal information and the user privacy preference for an automated reply response and inputting the prompt to a generative large language model (LLM); receiving a list of personalized response suggestions from the generative LLM; in response to presenting the received personalized response suggestions on an electronic display of the computing device, receiving a user input selection of one of the received personalized response suggestions; and performing an automated reply action based on the received user input.
[0007] Some aspects can further include activating or deactivating one or more information submodules configured to receive data inputs and output text suitable for prompting the generative LLM based on the determined current user situation and the determined user privacy preference for an automated reply response. Some aspects can further include processing non-text based information by at least one of the active information submodules to generate text suitable for input to the generative LLM. In some aspects, generating the prompt can include generating the prompt by combining text based information with non-text based information. In some aspects, the non-text based information includes descriptions based on audio or video sensor data.
[0008] Some aspects can further include selecting a model size for processing non-text based information in an information submodule based on one or more of the determined current user situation, a context of an incoming communication, or the user privacy preference for an automated reply response.
[0009] In some aspects, performing the automated reply action based on the received user input can include generating a privacy-aware multi-modal generative automated reply message based on the received user input and transmitting the generated privacy-aware multi-modal generative automated reply message to a computing device that initiated an incoming call or message. In some aspects, determining the user privacy preference for an automated reply response can include determining the user privacy preference for an automated reply response based on information that identifies a class of information that the user allows to be included in an automated reply response based on at least one of a context of an incoming call or message, an initiator of the incoming call or message, a current location of the user, or a current activity of the user.
[0010] Some aspects can further include providing the user input selection of one of the received personalized response suggestions to a machine learning module to implement the generation of improved suggestions based on the selected multi-modal information and the user privacy preference for automatic reply responses.
[0011] Further aspects can include a computing device having a processor configured with processor-executable instructions to perform various operations corresponding to the methods outlined above. Further aspects can include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor to perform various operations corresponding to the method operations outlined above. Further aspects can include a computing device having various means for performing functions corresponding to the method operations outlined above. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate exemplary embodiments of the claims and, together with the general description and detailed description given below, serve to explain features of the present disclosure.
[0013] Figure 1 is a component block diagram illustrating components of an example computing system that can be configured to implement some embodiments.
[0014] Figure 2A and Figure 2B is a component block diagram illustrating components in a computing system configured to generate privacy-aware multi-modal generative automatic replies, in accordance with some embodiments.
[0015] Figure 3 is a process flow diagram illustrating a method of generating privacy-aware multi-modal generative automatic replies, in accordance with some embodiments.
[0016] Figure 4 is a component block diagram illustrating an example computing device suitable for use with various embodiments.
[0017] Figure 5 is a component block diagram illustrating an example wireless communication device suitable for use with various embodiments. DETAILED DESCRIPTION
[0018] Various embodiments will be described in detail with reference to the drawings, where like reference numerals can be used to refer to like elements throughout. References made to particular examples and implementations are for illustrative purposes and are not intended to limit the scope of the claims.
[0019] These embodiments include computing devices configured to automatically generate personalized responses based on privacy level settings. The computing devices can be configured to perform an analysis of a current condition of a user of the computing device (e.g., a current situation of the device user, etc.) based on a diverse privacy-aware multi-modal information collected from various information sources and in-device sensors to generate various candidate responses to an incoming communication. The computing device can determine a current context of the incoming communication and a current activity and condition of the user by analyzing the collected multi-modal information. The computing device can use this contextual information in conjunction with user privacy settings to generate appropriate prompts for input to a generative large language model (LLM) to generate a plurality of selectable responses.
[0020] To prepare the context and other user information for generating prompts for the generative LLM, the computing device can use a number of information processing submodules that receive as input sensors or textual information and output text in a format suitable for input to the prompts for the generative LLM. The term “submodule” is used herein to refer to subsystems of the computing device and / or to refer to specialized programs that execute in a processing system that is configured to analyze a particular type of data or information and generate an output that is in a format suitable for input to the generative LLM. Some information submodules can be configured to receive information from a text-based source, such as a memory, format such textual content into a format suitable for use in generating prompts for the generative LLM. Other information submodules can be configured to receive non-textual data, such as sensor data (e.g., camera, microphone, accelerometer, etc.), and interpret the data to generate text-based output suitable for input to the generative LLM. The components within the computing system interact with each other and / or with other components to provide or implement high-level functionality.
[0021] Some non-limiting examples of some of the modules that can be used in various embodiments include: a visual understanding submodule configured to recognize scenes or objects in camera images and generate descriptive text; a speech recognition submodule configured to receive sound from a microphone, recognize speech, and generate text (i.e., perform speech-to-text transcription); a sound processing submodule configured to receive sound from a microphone and generate text describing the background sound; an electrocardiogram (ECG) analysis submodule configured to receive ECG data from a sensor (e.g., a user’s smartwatch) and generate text describing a user’s state indicated by the ECG data; and a motion analysis submodule configured to process accelerometer and gyroscope information from an inertial measurement unit (IMU) within the computing device and generate text describing a user’s motion or activity.
[0022] In various embodiments, the computing device can include a model controller configured to determine a context of an incoming communication and an activity of a user, access user privacy settings or preferences, and then control which of various information submodules provide output to the generative LLM. The model controller can be configured to limit the types of information used by the generative LLM to generate a list of proposed responses based on the user’s privacy preferences and according to the context of the incoming communication and the user’s situation. By enabling or disabling access to information sources provided by the perception submodules through the model controller, the generative LLM can be controlled to generate a list of proposed personalized automated reply messages that are appropriate to the situation and consistent with the user’s privacy preference settings. Various embodiments can provide a comprehensive and flexible automated reply solution that easily accommodates a wide variety of situations and contexts that are consistent with the user’s privacy preferences / settings.
[0023] The term “computing device” can be used herein to refer to any or all of the following: a personal computer, a laptop computer, a tablet computer, a user equipment (UE), a smartphone, a personal or mobile multimedia player, a personal data assistant (PDA), a palm- top computer, a wireless electronic mail receiver, a multimedia Internet-enabled cellular telephone, a game system (e.g., PlayStation ™ , Xbox ™ , Nintendo Switch ™ , etc.), a wearable device (e.g., a smartwatch, a head-mounted display, a fitness tracker, etc.), a media player (e.g., a DVD player, ROKU ™ , AppleTV ™ , etc.), a digital video recorder (DVR), an automobile display, a portable projector, a 3D holographic display, and other similar devices that include a display and a programmable processor that can be configured to provide the functionality of various embodiments.
[0024] The term “processing system” is used herein to refer to a system that includes one or more processors (including multi-core processors) that are organized and configured to perform various computing functions. A processing system can include at least one memory, interface circuitry, and other components integrated into the system. In a processing system, one or more of the processors can be configured to perform one or more operations of the methods of various embodiments.
[0025] The term "system on a chip" (SoC) is used herein to refer to a single integrated circuit (IC) chip that contains multiple processors, at least one memory, and support resources, which can form a processing system integrated on a single substrate. A SoC can contain circuitry for digital, analog, mixed-signal, and radio-frequency functions. A single SoC processing system can also include any number of general- purpose or special-purpose processors (e.g., network processors, digital signal processors, modem processors, video processors, etc.), memory blocks (e.g., ROM, RAM, Flash, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.). For example, a SoC processing system can include an application processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc., operating as a master processor of the SoC. A SoC processing system can also include software for controlling the integrated resources and processors, as well as for controlling peripheral devices.
[0026] The term "system in a package" (SIP) can be used herein to refer to a single module or package that contains a processing system comprising two or more IC chips, multiple resources, computing units, cores, or processors on a substrate, or SoC. For example, a SIP processing system can include a single substrate on which multiple IC chips or semiconductor dies are vertically stacked. Similarly, a SIP processing system can include one or more multi-chip modules (MCM) on which multiple ICs or semiconductor dies are packaged into a unified substrate. A SIP processing system can also include multiple independent SOCs coupled together via high-speed communication circuitry and packaged in close proximity, such as on a single motherboard, in a single UE, or in a single CPU device. The proximity of the SoCs facilitates high-speed communication and sharing of memory and resources.
[0027] Automatic replies are a feature in modern computing devices to generate automatic responses or actions in a computer system or software application when certain events or conditions are met (e.g., an email or support request is received, etc.). Automatic reply solutions can improve the user experience by allowing actions to be performed (e.g., messages to be transmitted, etc.) without interrupting an existing conversation. Various embodiments improve computing devices by improving the responsiveness and flexibility of automatic reply solutions that are consistent with the context of an incoming communication, user activity, and user privacy preference settings. Various embodiments further improve computing devices by learning over time how to generate proposed automatic reply responses that meet user preferences.
[0028] Automated response solutions support both message-based and non-message-based responses and actions. Non-message-based responses / actions may include blocking calls from specific phone numbers or allowing users to press an exit button to refuse incoming calls. Message-based responses / actions allow users to select from a list of optional text messages to send to the caller while they are in another call. For example, an automated response system may select from template response messages such as “Can’t speak, please send me a text message” or “I will call you back immediately.” Some systems also allow users to pre-craft and register customized messages such as “I’m driving” or “I’m in a meeting.”
[0029] Standard automated response solutions are insufficient to allow for personalized or detailed actions or messages. For example, with standard solutions, suggested text can be pre-written by the user, software, or network vendor and stored in memory. Therefore, automated response messages are often generic, as they are crafted to fit a wide range of possible situations. For instance, during an urgent meeting, a template message such as "I will call you back immediately" may not adequately convey the severity of the situation. While customized messages offer some adaptability, they still may not cover every possible scenario.
[0030] Some software systems (e.g., Microsoft Outlook) ® Similar features (e.g., ...) could generate suggestions for a concise response immediately after a reply is sent. For example, suggestions could be generated by an artificial intelligence (AI) system analyzing the main subject of an incoming email and producing suggested responses based on the words in the email. Such systems may be limited by the content of the incoming email and may not be suitable for generating responses tailored to individual users, specific situations, or users' privacy preferences.
[0031] Various implementations include computing devices (e.g., smartphones, tablets, laptops, etc.) with processing systems that execute Advanced Interaction System (AIS) components, collecting information about the user from various device sensors and memory. Information from such AIS can be processed by a model controller to determine the type of information to be provided to the generative LLM, automatically generating a menu of personalized response suggestions for incoming communications (e.g., calls, text messages, emails, etc.) in response to the context of the communication, current user activity, and user privacy preferences. In some implementations, the model controller can be configured to selectively restrict the generative LLM's access to information sources and information submodules, such that the generative LLM generates suggested responses based on the context of the incoming communication, the user's current situation, multimodal information collected in the computing device, and the user's privacy preferences or settings for the automated response.
[0032] In some implementations, the model controller is configured to enable or disable generative LLM access to collected multimodal information, such as activating or deactivating various information submodules, based on or in response to user privacy preferences or settings (e.g., the user's selected privacy level). In some implementations, the model controller may determine user privacy settings for different categories of user information during incoming calls, text messages, or emails (optionally via a user menu). In some implementations, user privacy settings may also be pre-configured for each different category of user information. In some implementations, user privacy preferences or settings for automated response responses may include or identify the categories of information the user allows to be included in the automated response response, wherein the preference or setting is specific to or based on at least one of the context of the incoming call or message, the initiator of the incoming call or message, the user's current location, and / or the user's current activity. In this way, the user can specify in advance the types of information that may be included in the automated response response, depending on who is on the call / message, what the message relates to, what the user is currently experiencing, and other criteria the user may include in their privacy preferences.
[0033] In some implementations, the model controller may be configured to detect incoming communications, determine or retrieve user privacy settings, select an information source and / or information submodule output based on the user privacy settings, generate LLM query information (e.g., a prompt for a generative LLM) based on the selected information source, transmit the LLM query information to the generative LLM to generate a list or set of proposed custom responses conforming to the user privacy settings, present the proposed custom responses on an electronic display, and allow the user to select one of the presented proposed custom responses. The computing device may then use the selected proposed custom responses as automated reply messages and / or as feedback to a machine learning system or module to learn user preferences over time.
[0034] In some implementations, the AIS component may be configured to determine, characterize, represent, and / or store the user’s current situation as an information structure (e.g., user situation, etc.), which includes symbols or values representing combinations of multimodal information (e.g., sensor information, location information, calendar insights, etc.).
[0035] In some implementations, the AIS component can be configured to collect multimodal information from any or all of a variety of sensors, information submodules, and memories within the computing device. Examples of multimodal information that can be collected in the device include sensor information, identification information, operating mode information, location information, calendar insights, audiovisual data, motion data, health data, connectivity information, data network activity information, system resource usage information, status information, driver statistics, hardware component information, software application information, and transmitted information.
[0036] In some implementations, the AIS component and / or model controller may be configured to determine the user's current situation based on the collected multimodal information. For example, the AIS component and / or model controller may be configured to detect triggering events (e.g., incoming calls, text messages, emails, etc.), collect multimodal information, determine the user's current situation based on the collected multimodal information, and provide the collected or determined information to the model controller. The model controller may be configured to use the received information and the user's privacy preferences / settings to determine which information sources should be provided to the generative LLM for generating a list or set of personalized response suggestions. Customized responses to the proposed responses received from the LLM may be presented on an electronic display for the user to select. The processing system of the computing device may receive the user's selection (e.g., touching a proposed response on a touchscreen display) and perform actions based on the selected response, such as blocking the caller, sending the selected response as an automated reply message, using the selected response to formulate and generate an automated reply message, or similar actions.
[0037] In some implementations, the AIS component, model controller, or another module executing within the processing system of the computing device may also use a user-selected custom response as feedback to the machine learning system, enabling the model controller to learn user preferences for generating custom responses and, over time, learn how to better meet user needs with proposed automated response modalities. In some implementations, the AIS component may control and fine-tune the information input to the generative LLM based on user expectations and privacy controls.
[0038] In some implementations, the model controller may be configured to select and / or exclude outputs provided by various information submodules and / or multimodal information sources based on user privacy settings. For example, the model controller may be configured to use user privacy settings to determine information sources and / or information submodules suitable for generating customized responses in various contexts and situations in response to user privacy preferences, and to provide only the selected multimodal information to the generative LLM to generate proposed automated response.
[0039] In some implementations, the AIS component can be configured to provide users with the option to set privacy levels for different information categories and / or generate responses that conform to the user's selected privacy level or settings. Examples of privacy levels include anonymous, public, acquaintance, professional, friend, and family levels.
[0040] Various implementation schemes can be implemented in the processing system of a computing device, which may include many single-processor and multi-processor computer systems, SoC processing systems, or SIP processing systems. Figure 1An example SIP processing system 100 architecture is illustrated that can be used in mobile computing devices implementing various implementation schemes.
[0041] refer to Figure 1 The illustrated example SIP processing system 100 includes two SOC processing systems 102 and 104, a clock 106, a voltage regulator 108, and a wireless transceiver 166. The first SOC processing system 102 and the second SOC processing system 104 can communicate via an interconnect bus 150. Various processors 110, 112, 114, 116, 118, 121, and 122 can be interconnected with each other and to one or more memory elements 120, system components and resources 124, and a thermal management unit 132 via an interconnect bus 126, which may include advanced interconnects such as high-performance network-on-chip (NOC). Similarly, processor 152 can be interconnected to a power management unit 154, a millimeter-wave transceiver 156, at least one memory 158, and various additional processors 160 via an interconnect bus 164. These interconnect buses 126, 150, and 164 may include arrays of reconfigurable logic gates and / or implement bus architectures (e.g., CoreConnect, AMBA, etc.). Communication can be provided by advanced interconnect components such as NOC.
[0042] In some implementations, the first SOC processing system 102 may operate as a central processing unit (CPU) of a mobile computing device, executing instructions for software applications by performing arithmetic, logic, control, and input / output (I / O) operations specified by the instructions. In some implementations, the second SOC processing system 104 may operate as a dedicated processing unit. For example, the second SOC processing system 104 may operate as a dedicated 5G processing unit, responsible for managing high-capacity, high-speed (e.g., 5Gbps) and / or very high frequency short-wavelength (e.g., 28GHz millimeter-wave spectrum) communications.
[0043] The first SOC processing system 102 may include a digital signal processor (DSP) 110, a modem processor 112, a graphics processor 114, an application processor 116, one or more coprocessors 118 (e.g., vector coprocessors) connected to one or more of these processors, at least one memory 120, a deep processing unit (DPU) 121, an artificial intelligence processor 122, system components and resources 124, an interconnect bus 126, one or more temperature sensors 130, a thermal management unit 132, and a thermal power envelope (TPE) component 134. The second SOC processing system 104 may include a 5G modem processor 152, a power management unit 154, an interconnect bus 164, multiple millimeter-wave transceivers 156, at least one memory 158, and various additional processors 160, such as application processors, packet processors, etc.
[0044] Each processor 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160 in processing systems 100, 102, 104 may include one or more cores, and each processor / core may perform operations independently of the other processors / cores. For example, the first SOC processing system 102 may include a processor running a first type of operating system (e.g., FreeBSD, LINUX, OSX, etc.) and a processor running a second type of operating system (e.g., MICROSOFT WINDOWS 11). Additionally, any or all of processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160 may be included as part of a processor cluster architecture (e.g., a synchronous processor cluster architecture, an asynchronous or heterogeneous processor cluster architecture, etc.).
[0045] Processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, and 160, or any or all of them, can operate as the CPU of a mobile computing device. Additionally, processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, and 160, or any or all of them, can be included as one or more nodes in one or more CPU clusters. A CPU cluster can be a group of interconnected nodes (e.g., processing cores, processors, SOCs, SIPs, computing devices, etc.) configured to work in a coordinated manner to perform computational tasks. Each node can run its own operating system and contains its own CPU, memory, and storage devices. Tasks assigned to the CPU cluster can be divided into smaller tasks, which are distributed across the nodes for processing. Nodes can work together to complete a task, with each node handling a portion of the computation. The results of the computations from each node can be combined to produce a final result. CPU clusters are particularly useful for tasks that can be parallelized and executed concurrently. This allows CPU clusters to complete tasks much faster than a single high-performance computer. Furthermore, because CPU clusters consist of multiple nodes, they are generally more reliable and less prone to failure than a single high-performance component.
[0046] The first SOC processing system 102 and the second SOC processing system 104 may include various system components, resources, and custom circuitry for managing sensor data, analog-to-digital conversion, wireless data transmission, and performing other specialized operations such as decoding data packets and processing encoded audio and video signals for presentation in a web browser. For example, the system components and resources 124 of the first SOC processing system 102 may include power amplifiers, voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting processors and software clients running on mobile computing devices. The system components and resources 124 may also include circuitry for interfacing with peripheral devices such as cameras, electronic displays, wireless communication devices, external memory chips, etc.
[0047] The first SOC processing system 102 and / or the second SOC processing system 104 may further include input / output modules (not illustrated) for communicating with external SOC resources such as clock 106, voltage regulator 108, and wireless transceiver 166 (e.g., cellular wireless transceiver, Bluetooth transceiver, etc.). External SOC resources (e.g., clock 106, voltage regulator 108, wireless transceiver 166) may be shared by two or more internal SOC processors / cores.
[0048] In addition to the example SIP processing system 100 discussed above, various implementations can be implemented in a wide variety of computing systems, which may include a single processor, multiple processors, multi-core processors, or any combination thereof.
[0049] Figure 2A and Figure 2B Logical configurations of example components in computing systems (e.g., SIP processing system 100, SOC processing system 102, etc.) suitable for implementing various implementation schemes are illustrated. Reference Figures 1 to 2B The computing system may include a model controller 202, an LLM 204 component (e.g., a generative LLM), a visual understanding submodule 212, a speech recognition submodule 214, an electrocardiogram (ECG) analysis submodule 216, a motion analysis submodule 218, a display module 220, and a user input module 222. The functionality of the model controller 202 and the various information submodules 212 to 218 can be executed within the processing system, such as reference... Figure 1 The processing systems described are 100, 102, and 104.
[0050] Model controller 202 can operate as a gating manager that restricts the output of information submodules 212 to 218 and other data sources provided to LLM 204 for generating multiple suggested auto-response responses. Model controller 202 can be configured to analyze multimodal information to determine the context and user status of incoming communications, and control the generative LLM's access to information submodules 212 to 218 (e.g., by activating or deactivating selected submodules) and other information sources in response to user privacy preferences or settings regarding auto-response responses.
[0051] Model controller 202 can be configured to receive user IDs and privacy preferences / settings as input, indicating data categories available to the user. Model controller 202 can also collect or receive multimodal data, which may include text-based and non-textual information. In some embodiments, model controller 202 can determine that submodules 212 to 218 should be activated to filter the collected information and use gating to activate or deactivate the determined submodules, such that the output of model controller 202 flows through the activated submodules 212 to 218 to LLM 204. LLM 204 can generate response recommendations based on the received outputs of submodules 212 to 218.
[0052] Model controller 202 may also collect or receive text-based and non-text information stored in the memory of the computing device and / or received from various sensors on or coupled to the computing device. In some embodiments, model controller 202 may receive information via submodules 212 to 218. Examples of text-based information include user location (from GPS), contact data (name, relationship to the user), calendar data (upcoming events and meetings), connectivity data (Bluetooth, Wi-Fi status), and the user's current activity (from operating system information). Examples of non-text information include environmental data (e.g., sound data from a microphone, visual data from a camera system, etc.) and the user's health data (e.g., from a smartwatch, etc.). Generally, text-based information can be sent to LLM 204 without further analysis. On the other hand, non-text information (e.g., received from various sensors) may need to be analyzed and converted into text before being input into LLM 204. Various submodules 212 to 218 may perform such additional multimodal analysis operations to convert non-text information into a format suitable for use as input to LLM 204.
[0053] In some implementations, the various information submodules 212 to 218 may include more than one model, with different models having different sizes or computational capabilities. Models of different sizes may be able to analyze different amounts of information. In such implementations, the model controller 202 may be configured to determine the appropriate model (e.g., appropriate model size) within each activated submodule based on the determined circumstances, the nature of the incoming message or telephone call, the source of the incoming message or telephone call, etc. For example, activating a small model in the visual understanding 212 submodule may cause the system to analyze only the types of objects present in the image, which may be suitable for responding to some incoming messages. Activating a larger model may cause the system to use image annotation techniques to generate extended text, explaining the relationships between objects in the image, which may be suitable for providing a detailed response to the incoming messages. For example, in response to determining that the information that can be generated by the visual understanding 212 submodule is crucial for prompting the LLM 204 in the current context, the model controller 202 may activate a large visual processing model in the visual understanding submodule 212 to perform more robust visual analysis on the data and pass additional information to the LLM 204.
[0054] In some implementations, model controller 202 can be configured to operate in a rule-based mode (without learning) or as a small LLM operation. In rule-based mode, model controller 202 can open specific model gates based on predefined conditions. Different individuals may expect different responses under the same conditions, and the privacy level of the expected response for the same individual may vary depending on the specific conditions. For example, user privacy preferences may specify the categories of information that a user allows to be included in an automated response, depending on factors such as the context of the incoming call or message, the initiator of the incoming call or message, and the user's location and / or current activity. Therefore, model controller 202 can use machine learning techniques to learn from feedback related to the text generated from LLM 204 and the user's selected responses, thereby improving the LLM's selective access to different information sources and information processing submodules 212 to 218 based on incoming communications, user conditions, and individual preferences.
[0055] The visual understanding submodule 212 can be configured to analyze visual information and generate text describing aspects of an imaging scene. For example, the visual understanding submodule 212 can perform image recognition, which may include classifying objects, people, animals, etc., in an image. The visual understanding submodule 212 can perform semantic segmentation, which may include analyzing detailed pixels of an image to identify the location and boundaries of different object categories. The visual understanding submodule 212 can also use image annotation, where it considers objects within the image, the background, and their relationships to generate more meaningful text descriptions suitable for input into the LLM 204.
[0056] The speech recognition submodule 214 can be configured to decode and interpret auditory signals, translate spoken words into text, facilitate voice-controlled interactions, and translate such messages into text format. For example, the speech recognition submodule 214 can be configured to recognize and analyze audio information from the device's microphone, analyze the input audio signal, and convert the audio information into a text description based on the analysis. The speech recognition submodule 214 can transcribe audio or determine the current situation based on the input sound (e.g., identify the speaker, analyze the speaker's emotion or mood, etc.). Similarly, the sound recognition submodule (not shown separately) can be configured to analyze the sound received by the microphone to determine the nature of background sounds and generate text descriptions of ambient sounds picked up by the microphone.
[0057] ECG analysis submodule 216 can be configured to interpret and evaluate electrical signals corresponding to a user's cardiac activity, providing insights into the user's health or emotional state, and translating sensor data into a form (e.g., text) that can be received and processed by LLM 204. Such ECG analysis can be performed via a smartwatch or other dedicated electronic device, enabling the analysis of electrical signals. ECG analysis submodule 216 can also be configured to allow the understanding of cardiac electrical signals, thereby enabling the calculation of heart rate and the detection of abnormalities associated with different cardiac conditions.
[0058] The motion analysis submodule 218 can be configured to evaluate motion sensor data (e.g., physical movement captured by accelerometers, gyroscopes, etc.) and translate the sensor data into a form (e.g., text) that can be received and processed by the LLM 204. For example, the motion analysis submodule 218 can analyze a user's motion based on motion information received via a smartwatch or other electronic device. The motion analysis submodule 218 can determine whether the user is moving or stationary based on the motion data. The motion analysis submodule 218 can analyze specific actions or activities of the user based on the user's motion patterns.
[0059] The LLM 204 component can be configured to receive prompt input from the model controller 202 and / or selected information submodules 212 to 218, and use the prompts to generate a list of proposed personalized responses in response to a comparison of multimodal information, the user's current context or activity, and the user's privacy preferences or settings for automated response. The LLM 204 component can receive text input generated by the model controller 202 and / or selected information submodules 212 to 218 based on the results of text-based and non-text information collected, analyzed, transformed, and filtered from the model controller 202. The LLM 204 can analyze the received information to generate several subtle and personalized responses that effectively take into account the user's current situation and privacy settings.
[0060] Display module 220 can be configured to present the proposed personalized response and accept user input for selecting a preferred automatic response option. For example, display module 220 can be a touchscreen display such as that on a smartphone.
[0061] User input module 222 can be configured to receive and process user selections or inputs, which can be fed back into the system for continuous improvement and learning.
[0062] As an example, the processing system may receive specific user privacy preferences / settings as input, which may be stored in memory. The processing system can use the user privacy preference / setting input to make decisions regarding selectively enabling or disabling LLM access to various information submodules 212-218 (e.g., by selective activation or deactivation) based on the user privacy settings, the context of incoming communications, and current user activity. For example, user input may include user privacy settings that set permissions for location, calendar, camera, and microphone settings to "on," and permissions for Bluetooth, health, ECG, and motion settings to "off." In response, model controller 202 may cause data transmitted to LLM 204 to include outputs from submodules 212-218, such outputs providing text related to the user's location, upcoming calendar events, visual information from the camera, and auditory signals from the microphone, while excluding information from Bluetooth connectivity, health monitors, ECG sensors, and motion detectors.
[0063] As another example, model controller 202 can be configured to determine data relevant to incoming communications. Model controller 202 can process multimodal data without any feedback, additional training, or updates. Model controller 202 can determine the data for multimodal analysis based on user privacy preferences or settings for automated response. Text-based information (e.g., location information indicating a user is located in “Building A203,” calendar information indicating an “academic seminar from 10:00 AM to 11:00 AM,” etc.) can be passed to the LLM 204 component without further processing or analysis. Model controller 202 and / or submodules 212 to 218 can perform multimodal analysis on non-textual information (e.g., visual and audio data from microphones and cameras) to convert the non-textual information into a suitable text format for input into LLM 204. LLM 204 can combine text-based input (such as location and calendar events) with visual and auditory data to generate context-relevant responses that respect the user's privacy settings.
[0064] In some implementations, model controller 202 can collect and filter multimodal information from different submodules based on user privacy preferences / settings to generate a robust context-aware understanding of the user's context. By selecting different filtering mechanisms based on user privacy settings, model controller 202 can control the information used to generate the proposed response. For example, visual understanding submodule 212 can determine that there is a lack of visual input or a completely black screen, which may be due to the computing device being placed in the user's pocket. In response, visual understanding submodule 212 can generate the text "No visual information detected" to be sent to LLM component 204. On the other hand, speech recognition submodule 214 can detect and analyze the speaker's speech to understand that the speaker is currently giving a report. Based on this auditory data, LLM 204 can generate a proposed automatic response as "multiple individuals, including the owner of the phone, are having a conversation, and the conversation topic is relevant to deep learning."
[0065] In some implementations, the system can combine text-based and non-textual information to generate a comprehensive understanding that enables the LLM 204 component to generate diverse response options relevant to the incoming communication, responsive to the user's current situation, and respecting the user's privacy preferences. The range of generated automated response options can range from general responses to highly detailed responses reflecting all relevant available information. For example, text-based information (such as location details (“Building A203”) and calendar events (“Academic seminar from 10:00 AM to 11:00 AM”) can be combined with non-textual information such as visual understanding (“No visual message detected.”) and speech recognition (“Multiple individuals, including the phone's owner, are having a conversation, and the conversation topic is related to deep learning.”). In response to such input, the LLM 204 component can generate nuanced response options, ranging from general responses that do not utilize any specific information to highly detailed responses that incorporate all available data. Examples of such options could include: “I am currently unable to answer the phone,” “I am currently in Building A203, I will contact you later,” “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. I will contact you later,” “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. My phone is in my pocket and I cannot access it right now. I will contact you later,” and “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. My phone is in my pocket, and I am currently discussing deep learning with several people. I will contact you later.”
[0066] In some implementations, the model controller can be retrained based on specific responses selected by the user to generate better response options for subsequent incoming communications. This retraining can be performed repeatedly or continuously to fine-tune the operation and better suit user preferences and / or improve the performance of automatic response functionality. For example, some users may prefer to share a lot of information in certain situations, while others may prefer to limit the sharing of personal information. Therefore, in response to determining that a device user is attending a meeting based on calendar and location data, the model controller can activate both the visual understanding and speech recognition submodules to analyze information collected by the camera and microphone. The model controller can then transmit the collected information to the LLM, present options to the user, and evaluate the selected options to determine whether to activate more submodules 212 to 218 or use a larger model to provide more detailed information, or deactivate modules and use a smaller model to minimize resource usage.
[0067] Figure 3 An example of method 300 for performing automatic replies in a computing device according to some implementation schemes is illustrated. Reference Figures 1 to 3 The operation of method 300 can be performed in a computing device by a processing system (e.g., 100, 102, 104), which may include one or more processors as described (e.g., processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160, etc.), components, or subsystems (e.g., model controller 202, submodules 212 to 218, LLM 204, etc.). Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations of the method. To encompass any of the processors, hardware elements, and software elements that may be involved in performing method 300, the element performing the method operations is generally referred to as the "processing system". Components for performing the functions of method 300 may include a processing system (e.g., 100, 102, 104) including one or more processors (e.g., processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160, etc.) as described herein, components, or subsystems (e.g., model controller 202, submodules 212 to 218, LLM 204, etc.).
[0068] In box 302, the processing system can detect triggering events such as incoming calls, text messages, and emails. For example, the processing system can detect incoming calls by monitoring signaling information on the telecommunications network, or detect incoming text messages or emails by inspecting new messages on a specific application or subsystem. The processing system can also be configured to identify other triggering events, such as calendar alerts or reminders, by interacting with scheduling and time management applications.
[0069] In box 304, the processing system can collect user-related multimodal information from multiple information sources on or accessible by the computing device. For example, the processing system can collect sensor information from embedded sensors that measure attributes such as location, movement, temperature, and pressure. The processor system can collect location information from a Global Navigation Satellite System (GNSS) module (e.g., a Global Positioning System (GPS) module) that geolocates a precisely positioned device. The processing system can collect calendar insights from scheduling applications, audiovisual data from microphones and cameras, motion data from accelerometers, health data from health monitoring components such as heart rate sensors, and connectivity information, data network activity information, system resource usage information, etc., by monitoring and controlling the corresponding hardware and software components of such subsystems in the computing device. The processing system can also collect user information stored in memory, such as the user's age, gender, marital status, education level, employment status, and employer. By aggregating various multimodal information, the processing system can build a nuanced understanding of the user's current situation and allow for more personalized and context-aware interactions.
[0070] In box 306, the processing system can determine the current user situation based on the collected multimodal information. For example, the processing system can combine sensor readings, location data, calendar insights, motion information, and connectivity details to determine the user's context. As a more detailed example, the processing system can use location information combined with calendar data to determine if the user is currently attending a scheduled meeting. The processing system can use motion data and sensor information to determine if the user is currently driving, running, or walking. The processing system can use connectivity information to determine if the device is connected to a home or office network, thereby determining whether the user is currently at home or in the office.
[0071] In box 308, the processing system may determine user privacy preferences / settings. In some embodiments, the processing system may determine the user privacy level based on a user identifier that identifies data categories available to the user and / or user input. In some embodiments, user privacy preferences or settings for automated response can be stored in memory, and the processing system may determine user privacy preferences / settings by accessing that memory. In some embodiments, user privacy preferences for automated response include information identifying information categories that the user is allowed to include in the automated response based on or in response to the context of an incoming call or message, the initiator of the incoming call or message, the user's current location, or the user's current activity. In some embodiments, the processing system may determine the user privacy level based on any or all of user preferences, pre-configured settings, and contextual analysis. For example, a user may define privacy levels across different information categories, selecting preferences for various conditions or characteristics, such as location sharing, access to personal data, or communication permissions. Users may set such preferences for different individuals, contexts, or relationships, such as spouse, family, career, or public privacy preferences. Furthermore, the processing system may allow users to pre-configure privacy settings for different information categories, thereby providing fine-grained control over what to share with whom. The processing system can also dynamically adjust privacy settings based on the user's current situation and / or based on predefined rules or user behavior patterns.
[0072] In box 310, the processing system that performs the functionality of the model controller can enable or disable the generative LLM's access to information sources, such as selectively activating or deactivating information submodules based on the current user context and the user's privacy preferences or settings for auto-response responses. For example, if user privacy settings allow location, calendar, camera, and microphone data to be included in auto-response responses but restrict the exposure of Bluetooth, health, ECG, and motion data in auto-response responses, the processing system can selectively activate information submodules corresponding to or related to the allowed information categories and deactivate information submodules related to the restricted information categories. In some embodiments, the complexity of the analysis and / or the size of the language model within each information submodule can be selected based on the current user context and user privacy settings. In some embodiments, activating or deactivating information submodules may include using selective gating mechanisms in the computing device.
[0073] In box 312, the processing system can generate cues suitable for input into a generative LLM by applying a relevant subset of the collected multimodal information to an active information submodule. For example, the processing system can apply a relevant subset of the collected multimodal information (including both text-based and non-text-based data) to the active information submodule. As part of this operation, the processing system can filter and / or guide the collected information through an activated information submodule. The processing system can process non-text-based information, such as audio or video data, by designing dedicated information submodules for auditory or visual analysis. For example, an audio understanding submodule can interpret spoken words or background sounds, and a visual understanding submodule can analyze objects and relationships within a video or image and generate text suitable for any LLM cues.
[0074] In box 314, the processing system can input the generated prompt into the generative LLM. In various implementations, the processing system can transmit the LLM prompt to a locally hosted generative LLM within the device or add an LLM service host remotely on a server. For example, after generating an LLM prompt by applying text-based and non-text-based information to an active submodule within the computing device, the processing system can encapsulate the resulting output as an LLM prompt, which is then sent to the generative LLM. This transmission can be facilitated through a secure communication channel that ensures data integrity and confidentiality. In response to receiving the prompt, the generative LLM can generate a personalized list of response suggestions based on the input.
[0075] In box 316, the processing system can receive a list of personalized response suggestions from the generative LLM. The list of personalized response suggestions can include a variety of nuanced response options, ranging from generic responses that do not utilize any specific information to highly detailed responses that incorporate all available data. Examples of such options could include: “I cannot answer the phone right now,” “I am currently in Building A203, I will contact you later,” “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. I will contact you later,” “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. My phone is in my pocket and I cannot access it right now. I will contact you later,” and “I am currently in Building A203, attending an academic seminar from 10:00 AM to 11:00 AM. My phone is in my pocket, and I am currently discussing deep learning with multiple people. I will contact you later.”
[0076] In box 318, the processing system can present the proposed personalized responses on an electronic display for the user to choose from. Presentation operations may include converting the received suggestions into a visually appealing format and organizing them within a user-friendly interface. The processing system can display the responses using lists, grids, or other visually attractive layouts, accompanied by interactive elements that allow the user to easily select the desired response.
[0077] In box 320, the processing system may receive input selecting a personalized response from one of the presented proposals. For example, a user may select an option by tapping the desired response on a touchscreen interface, or by clicking with a mouse when using a desktop or laptop computer. User input may also be a voice command or a gesture.
[0078] In box 322, the processing system can perform automated response actions based on the selected response (e.g., blocking the caller, sending the selected response as an automated response message, using the selected response to formulate and generate an automated response message, etc.). For example, if the selected response is to block the caller, the processing system can update the block list or enable specific call control features to block further communication from that number. If the selected option is to send an automated response message, the processing system can automatically formulate a text message or multimedia message based on the selected response and send it to the original caller / sender via an appropriate channel (e.g., SMS, email). The processing system can also use the selected response as a basis to generate more complex automated response messages that incorporate additional information or customize the content of the selected option. The processing system can also perform automated response actions using templates, scripts, or further processing via generative LLM and / or AI components. In some implementations, the processing system can perform automated response actions based on predefined rules, user preferences, and / or real-time context.
[0079] In some implementations, in block 322, the processing system may use user privacy preferences for auto-response responses, taking into account the context of the incoming message or call and the initiator, to generate a privacy-aware multimodal generative auto-response message based on received user input, and transmit the generated privacy-aware multimodal generative auto-response message to the computing device that initiated the incoming call or message.
[0080] In option 324, the processing system can feed the selected response as feedback to the machine learning module to support the training of the model controller, thereby learning user preferences over time. This feedback or retraining loop can be part of an adaptive system that continuously refines and personalizes the system's behavior and recommendations. For example, the processing system can analyze specific responses selected by the user over time and identify patterns, preferences, and / or tendencies that reflect the user's behavior and decision-making. The processing system can use any or all of this information to adjust or retrain the model within the system. For example, if the user consistently selects certain types of responses for a particular situation, the machine learning system can learn to prioritize or recommend similar responses in future interactions.
[0081] Various implementation plans (including but not limited to the above references) Figures 1 to 3 The described implementation scheme can be used in a wide variety of wireless devices and computing systems (including laptop computers 400, examples of which are shown in...). Figure 4 Implemented in the example (see below). Reference Figures 1 to 4 The laptop computer 400 may include a processor 402 coupled to at least one memory (such as volatile memory 404) and a disk drive 406 containing mass non-volatile memory (such as flash memory). The laptop computer 400 may include a touchpad touch surface 408 serving as a pointing device for the computer and thus capable of receiving drag, scroll, and tap gestures. Additionally, the laptop computer 400 may have one or more antennas 410 for transmitting and receiving electromagnetic radiation, connectable to a wireless data link, and / or a cellular transceiver 412 coupled to the processor 402. The computer 400 may also include a BT transceiver 414 coupled to the processor 402, a compact disc (CD) drive 416, a keyboard 418, and a display 420. Other configurations of the processing system may include a computer mouse or trackball, well-known to be coupled to the processor (e.g., via a Universal Serial Bus (USB) input), which may also be used in various implementations.
[0082] Figure 5 This is a component block diagram of a computing device 500 suitable for use with various implementation schemes. (Reference) Figures 1 to 5 Various implementation schemes can be implemented on a variety of computing devices (examples of which are available in 500). Figure 5 This is exemplified in the form of a smartphone. The computing device 500 may include a first SoC processing system 102 coupled to a second SoC processing system 104. The first SoC processing system 102 and the second SoC processing system 104 may be coupled to at least one internal memory 516, a display 512, and a speaker 514. The first SoC processing system 102 and the second SoC processing system 104 may also be coupled to at least one subscriber identity module (SIM) 540 and / or a SIM interface, which may store information supporting a first 5G NR subscription and a second 5G NR subscription, which support services on a 5G non-standalone (NSA) network.
[0083] The computing device 500 may include an antenna 504 for transmitting and receiving electromagnetic radiation, which may be connected to a wireless transceiver 166 coupled to one or more processors in the first SoC processing system 102 and the second SoC processing system 104. The computing device 500 may also include a menu selection button or rocker switch 520 for receiving user input.
[0084] The computing device 500 also includes a sound codec (CODEC) circuit 510 that digitizes sound received from a microphone into data packets suitable for wireless transmission and decodes the received sound data packets to generate an analog signal for use with a speaker to produce sound. Similarly, one or more processors in the first processing system 102, the second processing system 104, the wireless transceiver 166, and the CODEC 510 may include digital signal processor (DSP) circuitry (not shown separately).
[0085] The processing system and its included processors or processing units can be any programmable microprocessor, microcomputer, or multiprocessor chip, which can be configured by software instructions (applications) to perform a variety of functions, including those described in the various embodiments. In some computing devices, multiple processors may be provided, such as one processor within a first circuit dedicated to wireless communication functions and another processor within a second circuit dedicated to running other applications. Software applications may be stored in memory before being accessed and loaded into the processor of the processing system. The processor may include internal memory sufficient to store application software instructions.
[0086] Specific implementation examples are described in the following paragraphs. While some of the specific implementation examples below are described as example methods, further example implementations may include: example methods discussed in the following paragraphs implemented by a computing device including at least one memory coupled to at least one processor configured (e.g., configured with processor-executable instructions) to perform the operations of the methods of the following specific implementation examples; example methods discussed in the following paragraphs implemented by a computing device including components for performing the methods of the following specific implementation examples; and example methods discussed in the following paragraphs may be implemented as a non-transitory processor-readable storage medium storing processor-executable instructions configured to cause a processor of a computing device to perform the operations of the methods of the following specific implementation examples.
[0087] Example 1. A method for performing an automatic response by a processing system of a computing device, the method comprising: collecting multimodal information about a user of the computing device; determining a current user situation based on the collected multimodal information; determining user privacy preferences for the automatic response; generating a prompt based on the selected multimodal information and the user privacy preferences for the automatic response and inputting the prompt into a generative large language model (LLM); receiving a list of personalized response suggestions from the generative LLM; receiving a user input selection for one of the received personalized response suggestions in response to presenting the received personalized response suggestions on an electronic display of the computing device; and performing an automatic response action based on the received user input.
[0088] Example 2. According to the method of Example 1, the method further includes: activating or deactivating one or more information submodules based on the determined current user situation and the determined user privacy preferences for automatic response, wherein the one or more information submodules are configured to receive data input and output text suitable for prompting the generative LLM.
[0089] Example 3. According to the method of Example 2, the method further includes: processing non-text-based information by at least one of the activity information submodules to generate text suitable for input into the generative LLM.
[0090] Example 4. According to the method of Example 3, generating the prompt includes: generating the prompt by combining text-based information and non-text-based information.
[0091] Example 5. The method according to Example 3, wherein the non-text-based information includes descriptions based on audio or video sensor data.
[0092] Example 6. According to the method of Example 1, the method further includes: selecting a model size for processing non-text-based information in the information submodule based on one or more of the determined current user situation, the context of the incoming communication, or the user privacy preferences for the automatic reply response.
[0093] Example 7. According to the method of Example 1, performing the automatic reply action based on the received user input includes: generating a privacy-aware multimodal generative automatic reply message based on the received user input; and transmitting the generated privacy-aware multimodal generative automatic reply message to the computing device that initiated the incoming call or message.
[0094] Example 8. According to the method described in Example 1, determining the user privacy preference for an automatic reply response includes: determining the user privacy preference for an automatic reply response based on information identifying information categories that the user is allowed to include in the automatic reply response based on at least one of the context of an incoming call or message, the initiator of the incoming call or message, the user's current location, or the user's current activity.
[0095] Example 9. According to the method of Example 1, the method further includes: providing the user input selection of one of the received personalized response suggestions to the machine learning module to realize the generation of the improvement prompt based on the selected multimodal information and the user privacy preference for the automatic reply response.
[0096] As used in this application, the terms "component," "module," "system," etc., are intended to include computer-related entities such as, but not limited to, hardware, firmware, combinations of hardware and software, software, or software being executed, configured to perform specific operations or functions. For example, a component can be, but is not limited to, a process, processor, object, executable, execution thread, program, and / or computer running on a processor. By way of illustration, both an application running on a computing device and the computing device itself can be referred to as a component. One or more components may reside within a process and / or execution thread, and components may reside on a processor or core and / or be distributed across two or more processors or cores. Furthermore, these components may execute on various non-transitory computer-readable media on which various instructions and / or data structures are stored. Components may communicate via local and / or remote processes, function or procedure calls, electronic signals, data packets, memory read / write, and other known network, computer, processor, and / or process-related communication methods.
[0097] A variety of different memory types and memory technologies are available or conceivable in the future, and any or all of these different memory types and memory technologies can be included and used in systems and computing devices implementing various implementation schemes. Such memory technologies / types may include non-volatile random access memory (NVRAM), such as magnetoresistive RAM (M-RAM), resistive random access memory (ReRAM or RRAM), phase-change random access memory (PC-RAM, PRAM, or PCM), ferroelectric RAM (F-RAM), spin-transfer torque magnetoresistive random access memory (STT-MRAM), and 3D-XPOINT memory. Such memory technologies / types may also include non-volatile or read-only memory (ROM) technologies, such as programmable read-only memory (PROM), field-programmable read-only memory (FPROM), and one-time programmable non-volatile memory (OTP NVM). Such memory technologies / types may also include volatile random access memory (RAM) technologies, such as dynamic random access memory (DRAM), double data rate (DDR) synchronous dynamic random access memory (DDR SDRAM), static random access memory (SRAM), and pseudo static random access memory (PSRAM). Systems and computing devices implementing various embodiments may also include or use electronic (solid-state) non-volatile computer storage media, such as flash memory. Each of the memory technologies mentioned above includes, for example, elements suitable for storing instructions, programs, control signals, and / or data for use in or for use in: advanced driver assistance systems (ADAS) of vehicles, system-on-a-chip (SOC), or other electronic components. Any references to terms and / or technical details relating to individual memory types, interfaces, standards, or memory technologies are for illustrative purposes only and are not intended to limit the scope of the claims to a particular memory system or technology, unless expressly stated in the language of the claims.
[0098] The various embodiments illustrated and described are provided merely as examples illustrating the various features of the claims. However, the features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and may be used or combined with other embodiments shown and described. Furthermore, the claims are not intended to be limited to any one of the exemplary embodiments. For example, one or more operations of the method may substitute for or combine with one or more operations of the method.
[0099] The foregoing method descriptions and process flowcharts are provided as illustrative examples only and are not intended to require or imply that the operations of the various embodiments must be performed in the given order. As those skilled in the art will appreciate, the operations in the foregoing embodiments can be performed in any order. Words such as “afterward,” “then,” “next,” etc., are not intended to restrict the order of operations; these words are only used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements (e.g., references using the articles “a,” “an,” or “described”) should not be construed as limiting that element to the singular.
[0100] The various exemplary logic blocks, modules, circuits, and algorithmic operations described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and operations have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. While those skilled in the art may implement the described functionality in different ways for each specific application, such implementation decisions should not be construed as departing from the scope of the claims.
[0101] Hardware for implementing the various exemplary logics, logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein can be implemented or executed using a processing system, which may include a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof, designed to perform the functions described herein. While a general-purpose processor may be a microprocessor, in alternative embodiments, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0102] In one or more embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, such functionality may be stored as one or more instructions or code on a non-transitory computer-readable medium or a non-transitory processor-readable medium. The operation of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store object program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above may also be included within the scope of non-transitory computer-readable and processor-readable media. In addition, the operation of a method or algorithm may reside as a single piece of code and / or instruction, or any combination or set of code and / or instructions, on a non-transitory processor-readable medium and / or computer-readable medium that may be incorporated into a computer program product.
[0103] The above description of the disclosed embodiments is provided to enable any person skilled in the art to implement or use the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the claims. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be granted the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
1. A method of performing an automated reply response by a processing system of a computing device, the method comprising: collecting multi-modal information about a user of the computing device; determining a current user situation based on the collected multi-modal information; determining a user privacy preference for an automated reply response; generating a prompt and inputting the prompt to a generative large language model (LLM) based on the selected multi-modal information and the user privacy preference for an automated reply response; receiving a list of personalized response suggestions from the generative LLM; receiving a user input selection of one of the received personalized response suggestions in response to presenting the received personalized response suggestions on an electronic display of the computing device; and performing an automated reply action based on the received user input. activating or deactivating one or more information sub-modules based on the determined current user situation and the determined user privacy preference for an automated reply response, the one or more information sub-modules configured to receive data input and output text suitable for prompting the generative LLM.
2. The method of claim 1, further comprising: processing non-text based information by at least one of the active information sub-modules to generate text suitable for input to the generative LLM.
3. The method of claim 2, further comprising: generating the prompt by combining the text based information with the non-text based information.
4. The method of claim 3, wherein generating the prompt comprises:
5. The method of claim 3, wherein the non-text based information comprises a description based on audio or video sensor data. selecting a model size for processing non-text based information in an information sub-module based on one or more of the determined current user situation, a context of an incoming communication, or the user privacy preference for an automated reply response.
6. The method of claim 1, further comprising:
7. The method of claim 1, wherein performing the automated reply action based on the received user input comprises: generating a privacy-aware multi-modal generative automated reply message based on the received user input; and transmitting the generated privacy-aware multi-modal generative automated reply message to a computing device that initiated an incoming call or message. determining the user privacy preference for an automated reply response based on information that identifies a category of information that the user allows to be included in an automated reply response based on at least one of a context of an incoming call or message, an initiator of the incoming call or message, a current location of the user, or a current activity of the user. providing the user input selection of one of the received personalized response suggestions to a machine learning module to implement the generation of an improved prompt based on the selected multi-modal information and the user privacy preference for an automated reply response.
8. The method of claim 1, wherein determining the user privacy preference for automatic reply responses comprises:
10. A computing device, the computing device comprising:
9. The method of claim 1, further comprising: at least one memory; a display; and a processing system coupled to the at least one memory and the display and comprising one or more processors, one or more of the one or more processors configured to: collect multi-modal information about a user of the computing device; determine a current user situation based on the collected multi-modal information; determine a user privacy preference for an automated reply response; generate a prompt based on the selected multi-modal information and the user privacy preference for an automated reply response and input the prompt to a generative large language model (LLM); receive a list of personalized response suggestions from the generative LLM; receive a user input selection of one of the received personalized response suggestions in response to presenting the received personalized response suggestions on the display of the computing device; and perform an automated reply action based on the received user input.
11. The computing device of claim 10, wherein one or more of the processors of the processing system are further configured to activate or deactivate one or more information sub-modules based on the determined current user context and the determined user privacy preference for an automated reply response, the one or more information sub-modules configured to receive data input and output text suitable for prompting the generative LLM.
12. The computing device of claim 11, wherein one or more of the processors of the processing system are further configured to process non-text based information by at least one of the active information sub-modules to generate text suitable for input to the generative LLM.
13. The computing device of claim 12, wherein one or more of the processors of the processing system are further configured to generate the prompt by combining text based information with non-text based information.
14. The computing device of claim 12, wherein the non-text based information comprises a description based on audio or video sensor data.
15. The computing device of claim 10, wherein one or more of the processors of the processing system are further configured to select a model size for processing non-text based information in an information sub-module based on one or more of the determined current user context, a context of an incoming communication, or the user privacy preference for an automated reply response.
16. The computing device of claim 10, wherein one or more of the processors of the processing system are further configured to perform the automated reply action based on the received user input by: generating a privacy-aware multi-modal generative automated reply message based on the received user input; and transmitting the generated privacy-aware multi-modal generative automated reply message to a computing device that initiated an incoming call or message.
17. The computing device of claim 10, wherein one or more of the processors of the processing system are further configured to determine the user privacy preference for an automated reply response based on information that identifies a class of information that the user allows to be included in an automated reply response based on at least one of a context of an incoming call or message, an initiator of the incoming call or message, a current location of the user, or a current activity of the user. 18. The computing device of claim 10, wherein one or more of the processors of the processing system are further configured to provide the user input selection of one of the received personalized response suggestions to a machine learning module to implement the generation of an improved prompt based on the selected multi-modal information and the user privacy preference for an automated reply response.
19. A computing device comprising: means for collecting multi-modal information about a user of the computing device; means for determining a current user situation based on the collected multi-modal information; means for determining a user privacy preference for an automated reply response; means for generating a prompt and inputting the prompt to a generative large language model (LLM) based on the selected multi-modal information and the user privacy preference for an automated reply response; means for receiving a list of personalized response suggestions from the generative LLM; means for receiving a user input selection of one of the received personalized response suggestions in response to presenting the received personalized response suggestions on an electronic display of the computing device; and means for performing an automated reply action based on the received user input. means for activating or deactivating one or more information sub-modules configured to receive data inputs and output text suitable for prompting the generative LLM based on the determined current user situation and the determined user privacy preference for an automated reply response.
20. The computing device of claim 19, further comprising: means for processing non-text based information by at least one of the active information sub-modules to generate text suitable for input to the generative LLM.
21. The computing device of claim 20, further comprising: means for generating the prompt by combining text based information with non-text based information.
22. The computing device of claim 21, wherein the means for generating the prompt comprises:
23. The computing device of claim 21, wherein the non-text based information comprises a description based on audio or video sensor data. means for selecting a model size for processing non-text based information in an information sub-module based on one or more of the determined current user situation, a context of an incoming communication, or the user privacy preference for an automated reply response.
24. The computing device of claim 19, further comprising:
25. The computing device of claim 19, wherein the means for performing the automated reply action based on the received user input comprises: means for generating a privacy-aware multi-modal generative automated reply message based on the received user input; and means for transmitting the generated privacy-aware multi-modal generative automated reply message to a computing device that initiated an incoming call or message. means for determining the user privacy preference for an automated reply response based on information that identifies a class of information that the user allows to be included in an automated reply response based on at least one of a context of an incoming call or message, an initiator of the incoming call or message, a current location of the user, or a current activity of the user. 26. The computing device of claim 19, wherein means for determining the user privacy preference for an automatic reply response comprises: 27. The computing device of claim 19, further comprising: a component for providing the user input selection of one of the received personalized response suggestions to a machine learning module to enable the generation of improved prompts based on the selected multi-modal information and the user privacy preference for automated reply responses.
28. A non-transitory processor-readable medium having stored thereon processor- executable instructions configured to cause one or more processors of a processing system of a computing device to perform operations comprising: collecting multi-modal information about a user of the computing device; determining a current user situation based on the collected multi-modal information; determining a user privacy preference for automated reply responses; generating a prompt based on the selected multi-modal information and the user privacy preference for automated reply responses and inputting the prompt to a generative large language model (LLM); receiving a list of personalized response suggestions from the generative LLM; receiving a user input selection of one of the received personalized response suggestions in response to presenting the received personalized response suggestions on an electronic display of the computing device; and performing an automated reply action based on the received user input.
29. The non-transitory processor-readable medium of claim 28, wherein the stored processor-executable instructions are configured to cause one or more processors of the processing system of the computing device to perform operations further comprising activating or deactivating one or more information sub-modules configured to receive data input and output text suitable for prompting the generative LLM based on the determined current user situation and the determined user privacy preference for automated reply responses.
30. The non-transitory processor-readable medium of claim 28, wherein the stored processor-executable instructions are configured to cause one or more processors of the processing system of the computing device to perform operations further comprising providing the user input selection of one of the received personalized response suggestions to a machine learning module to enable the generation of improved prompts based on the selected multi-modal information and the user privacy preference for automated reply responses.