Electronic messaging method using image-based noisy content - Patent Application 20070122997

The method generates image-based noisy content to address the challenges of ambiguous and noisy information in electronic messaging, enhancing training and communication efficiency by introducing controlled ambiguity and incomplete information.

JP7798454B2Active Publication Date: 2026-01-14INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023559723
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-30
Filing Date
2022-03-08
Publication Date
2026-01-14
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

Existing electronic messaging systems struggle with managing ambiguous, noisy, or overly informative information, which complicates user training and communication efficiency.

Method used

A method that generates image-based noisy content related to the message intent, which can be used to enhance or replace the original message, improving training and communication by introducing ambiguity or incomplete information.

Benefits of technology

Enhances user training by increasing difficulty and improving engagement through image-based noisy content, while ensuring secure and contextually relevant communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798454000001
    Figure 0007798454000001
  • Figure 0007798454000002
    Figure 0007798454000002
  • Figure 0007798454000003
    Figure 0007798454000003
Patent Text Reader

Abstract

The present disclosure relates to a method including receiving an electronic message in an electronic communication system, where content having an image-based noise may be generated, the content having the image-based noise being distinct from the content of the received electronic message and related to a message intent of the received message, the electronic communication system may be configured to provide the image-based noise to a recipient in place of the received electronic message or to provide the received electronic message to the recipient in addition to the image-based noise.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of digital computer systems, and more particularly to electronic messaging methods. [Background technology]

[0002] Currently, user or automated system training or testing is done using a set of predefined questions and answers. However, in the real world, information is not always provided in a simple and structured way. It can be ambiguous, noisy, contain only part of the necessary information, or conversely, contain too much information. Summary of the Invention

[0003] In various embodiments, there are provided methods, computer systems and computer program products as described by the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. The embodiments of the invention may be freely combined with one another if they are not mutually exclusive.

[0004] In one aspect, the invention relates to a computer-implemented method for electronic messaging, the method including receiving an electronic message in an electronic communication system, the message being determined for a recipient; determining a message intent for the received electronic message; generating image-based noisy content that is different from content of the received electronic message and related to the message intent; and controlling the electronic communication system to provide the image-based noisy content to the recipient in place of the received electronic message or to provide the received electronic message to the recipient in addition to the image-based noisy content.

[0005] In another aspect, the present invention relates to a computer program product including a computer readable storage medium having computer readable program code embodied thereon, the computer readable program code being configured to perform all of the steps of the method according to the aforementioned embodiments.

[0006] In another aspect, the invention provides a method for receiving an electronic message, determining a message intent of the received electronic message, generating image-based noisy content that is different from content of the received electronic message and related to the message intent, and controlling the electronic communication system to provide the image-based noisy content in place of the received electronic message or to provide the image-based noisy content in addition to the image-based noisy content. Reception and a computer system configured to perform the method of providing an electronic message.

[0007] Embodiments of the invention will now be described in more detail, by way of example, with reference to the following drawings, in which: [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1 illustrates a computer system according to an example of the present subject matter. [Figure 1B] FIG. 10 illustrates a window showing a record of a chat session. [Figure 2] 1 is a flowchart of a method according to an example of the present subject matter. [Figure 3] 1 is a flowchart of a method according to an example of the present subject matter. [Figure 4] 1 is a flowchart of a method according to an example of the present subject matter. [Figure 5] FIG. 1 illustrates a method for linking text and image embedding spaces in accordance with the present subject matter. [Figure 6] FIG. 1 illustrates a method for testing a learned embedding space in accordance with the present subject matter. [Figure 7]FIG. 1 illustrates a method for generating an image in accordance with the present subject matter. [Figure 8] FIG. 1 is a diagram illustrating a general computerized system suitable for implementing at least some of the method steps associated with the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0009] Various embodiments of the present disclosure are described by way of example, but are not intended to be exhaustive or limited to the disclosed embodiments. As will be apparent to those skilled in the art, many modifications and variations are possible without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements to commercially recognized technologies of the embodiments, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0010] Messaging can be written communication sent through various digital channels, such as email, SMS, or in-app chat. Messaging can be beneficial because it can provide relevant information at the right time. In particular, messaging performed with the right information and frequency can improve the performance of electronic communication systems. However, electronic messages can be difficult to manage due to the proliferation of platforms, devices, and systems used to create these records. The present subject matter can be beneficial because it provides a systematic way to control and manage electronic message communications. This can be particularly beneficial in some institutions, where the use of text and chat / instant messaging may be essential to accomplishing the institution's mission.

[0011] The electronic message may be received, for example, by reading a conversation log to generate a message for the messaging session. In another example, the electronic message may be received, for example, by intercepting an ongoing messaging session. The intercepted electronic message may be provided, for example, during the messaging session. The messaging session may include exchanging electronic messages including text, multimedia, or audio, or a combination thereof, in a real-time or non-real-time format. The real-time format may include instant messaging or chat, while the non-real-time format may include email, posting to a dynamic forum or feed, etc. The messaging session may be associated with a context depending on the use case. For example, a messaging session may be conducted between a chatbot agent and a user for training or testing the user with the chatbot agent. In this case, the electronic message may be a question and answer. In another example, the messaging session may be conducted between multiple users via a mobile messaging application provided on each of the users' mobile client devices. In another example, the messaging session may include a group chat in which the users share and discuss various topics, including video or other types of multimedia. The terms "chatbot" (or "chat bot") or "chat agent" refer to a computer program designed to simulate a conversation, either auditory or textual, with one or more human users, primarily for small talk or training purposes. The goal of such a simulation is to trick the end user into thinking that the program's output was generated by a human.

[0012] The intercepted / received electronic message is processed to generate image-based noisy content. The image-based noisy content may be a single piece of content. The content is referred to as image-based noisy content because it involves images and has intents that are merely related to the message intent. For example, the content may include one or more images, be generated using one or more images, or both. The image-based noisy content may be generated such that its intent constitutes a subset of the message intents. Here, the subset is obtained by excluding important intents from the message intents or selecting unimportant intents from the message intents. The message intents may be ranked based on their importance. For example, the message intents may be ranked based on centrality indices of a knowledge graph, such that the subset consists of K (K≧1) less important message intents. In another example, the message intents may be ranked based on user input. For example, a user may be prompted to provide a ranking of the message intents, and the ranking may be received.

[0013] For example, the image-based noisy content may include one or more images that can be used to enhance an existing communication flow, such as a training session between a chat agent and a user. For training use cases, the image-based noisy content may be generated to increase the difficulty of a given question by providing incomplete information, adding confusing intents to increase ambiguity, or adding misleading intents to intercepted / received electronic messages, or a combination thereof. In particular, the present subject matter can be employed, for example, to test / coach trainees who will eventually provide after-sales support or product customer service, or senior medical students who will apply medical knowledge to solve virtual patient cases. In the latter case, the image-based noisy content may be in the form of error-prone medical diagnoses, erroneous test results, patient images / photos, etc. Since the final interaction may be a conversation between two human agents (such as a consumer and a customer service agent or a doctor and a patient), the chat agent may be a conversational chatbot capable of mimicking such human interactions. In one example, the image-based noisy content may include multimodal conversational noisy utterances based on context identified from a conversational situation of a chatbot, where multimodal may refer to introducing inconsistencies across both text and images.

[0014] In addition, in a communication system where confidential information should not be provided according to data access rules, the intercepted / received electronic message may be processed to identify and mask or remove intents that do not satisfy the data access rules. The generated image-based noise content may include unmasked information from the received electronic message, which may enable secure data communication.

[0015] According to one embodiment, the method further includes determining a domain of the electronic messages and determining a set of one or more text messages and a set of one or more images for the determined domain, wherein generating the image-based noisy content includes any one of the following: Generating one or more noisy images from a set of text messages (wherein the image-based noisy content is the noisy image). Extracting noisy text from a set of images (where the image-based noisy content is the noisy text). Generating noisy multimodal content from a set of images and a set of text messages, where the image-based noisy content is the noisy multimodal content.

[0016] A domain may represent concepts or categories belonging to a part of the world, such as biology or politics. A domain typically models the definition of domain-specific terms. For example, a domain may refer to healthcare, advertising, commerce, medicine, chemistry, physics, computer science, oil and gas, or transportation. The domain of an electronic message may be determined, for example, using natural language processing techniques. The set of text messages and the set of images may be obtained, for example, from a history of conversational utterances on text and images. Obtaining more text and images from the same domain as the received electronic message may enable better understanding and processing of differences between the message intents of the electronic message. Generating images from the set of text messages may be performed using an embedding space, as described in the following embodiments.

[0017] According to one embodiment, generating image-based noisy content includes determining a joint embedding space for representing an electronic message and images, representing the text message and a predetermined set of images in the embedding space, ranking the set of images based on a similarity between the representation of the set of images and the representation of the text message in the embedding space, and selecting a subset of the set of images based on the ranking. The subset of images may be noisy images. That is, the subset of images may have a message intent that is not similar to the intent of the text message but is related to the message intent of the text message. The text message represented in the embedding space may be a received electronic message. In another example, the text message may be a message similar to the electronic message. For example, the text message may have a number of message intents equal to or less than the message intent of the electronic message. In another example, the text message may be one of the sets of text messages in the aforementioned embodiments. The set of images may be, for example, screenshots. The screenshots may be of a user interface taken at successive stages of software development or device monitoring, for example.

[0018] According to one embodiment, the subset of images is the second k ranked images, where k is a configurable parameter. For example, if the set of images is 10 images and k=3, the subset of images may include the images ranked 7th, 6th, and 5th.

[0019] According to one embodiment, the method further includes setting the value of the parameter k according to a desired noise level of the noisy content. The noise may be, for example, user-based noise or problem-based noise. User-based noise may be determined based on user characteristics, such as age. For example, increasing the value of the parameter k may increase the noise level for an experienced user compared to a normal user. The problem-based noise level may be changed by, for example, adding more confusing intents or removing more important intents. The noise level may be controlled, for example, to ensure that the generated multimodal noise is relevant to the case context. For example, excessively noisy data may confuse a trainee, while irrelevant noise may not be useful for training. Therefore, according to this embodiment, noise may be introduced so that the multimodal conversational utterance expresses inconsistencies between text content and image content. Also, according to this embodiment, noise may be introduced so that the multimodal conversational utterance expresses incomplete information in the form of images, text, or both. For example, a chatbot may dynamically introduce multimodal noise depending on automatically generated textual and / or visual utterances and how the trainee deals with previous noisy information, potentially improving user engagement.

[0020] According to one embodiment, the joint embedding is trained using hinge loss to align images and text according to a Siamese neural network architecture, which can generate embeddings such that distances represent the semantic similarity of the represented objects.

[0021] According to one embodiment, generating the image-based noisy content includes determining related intents for a message intent using a knowledge graph, the knowledge graph including entities that represent the intents, and extracting from a predetermined set of images a subset of images that are related to the related intents and not related to the message intent, the image-based noisy content being the subset of images.

[0022] A set of images may be represented by a node in a knowledge graph. A knowledge graph may represent one or more domain ontologies. For example, a knowledge graph may represent the domain of software and / or hardware bug fixing. In this case, a node in the knowledge graph may represent, for example, an intent, an image, and a resolution. An intent may represent, for example, a specific software problem. For example, intents might be "#heating" and "#virus," which indicate problems caused by a computer overheating and the presence of a computer virus, respectively. A solution is associated with the problem defined by the intent in the knowledge graph. For example, a solution associated with the intent "#virus" might be "#install_antivirus." Images may represent stack traces, error logs, function calls, or command output. An image in a graph may be, for example, a screenshot of a log describing a problem related to an intent. For example, the image "#fan_noise" might be associated with the intent "#heating." The present subject matter can be beneficial because it can utilize knowledge graphs to accurately generate images.

[0023] However, the domain of a knowledge graph may be broad enough to cover multiple topics. Continuing with the bug problem example, multiple topics may be covered in the knowledge graph. For example, operating system issues may cover one topic, and displays may cover another. This can be resolved by clustering the knowledge graph. The knowledge graph may be clustered into multiple clusters. A cluster may be represented, for example, by a subgraph of the knowledge graph, where the data in that cluster represents a specific topic. A cluster may include multiple subclusters, where a subcluster may represent commonly occurring intents, commonly co-occurring intents, commonly suggested solutions, discouraged or incorrect solutions, etc. Related intents may be determined, for example, to belong to the same cluster.

[0024] A knowledge graph may be a graph. A graph may refer to a property graph in which data values ​​are stored as properties on nodes and edges. Property graphs may be managed and processed by a graph database management system or other database systems. These systems provide a wrapper layer that converts the property graph into, for example, relational tables for storage, and converts the relational tables back to a property graph upon retrieval or query. A graph may be, for example, a directed graph. A graph may be a collection of nodes (also called vertices) and edges. An edge of a graph connects any two nodes in the graph. An edge can be represented by an ordered pair of nodes (v1, v2) and can be traversed from node v1 to node v2. A node in a graph may represent an entity. An entity may refer to a problem, a solution, etc. An entity (and corresponding node) may have one or more entity attributes or characteristics to which values ​​can be assigned. For example, the entity attribute of a solution may include an attribute indicating whether the solution is a generally recommended solution or a not recommended solution. The attribute value representing a node is the value of the entity attribute of the entity represented by the node. An edge may be assigned one or more edge attribute values ​​that at least indicate the relationship between the two nodes connected to the edge. The attribute value representing an edge is the value of the edge attribute. The relationship may include, for example, an inheritance relationship (e.g., parent-child) or an associative relationship according to a specific hierarchy. For example, an inheritance relationship between nodes v1 and v2 may be called an "is-a relationship" between v1 and v2. For example, "v2 is-a parent of v1."An associative relationship between nodes v1 and v2 is sometimes called a "has-a relationship" between v1 and v2, such as "v2 has a has-a relationship with v1." This means that v1 is part of, a component of, or related to v2.

[0025] According to one embodiment, the method further includes using the electronic message to determine a context of a messaging session, the context of the messaging session being defined by at least a subgraph of a knowledge graph, and using the subgraph to determine related intents as intents that belong to the determined context and are different from the message intent. Intents, referred to as "related intents," can be used to add ambiguity or insert misleading information. In other words, the content of the generated image may provide incomplete, ambiguous, or misleading information.

[0026] The context of the messaging session may be the topic of the messaging session. The topic of the messaging session may be determined by analyzing the content of the electronic message. The analysis may be performed, for example, using data mining techniques. A subgraph of the knowledge graph may include intents that are expected to share the topic of the electronic message. The intents of the subgraph may include message intents of the intercepted electronic message. The subgraph may be used advantageously because the intents of the subgraph may be ranked based on importance, for example, using centrality.

[0027] According to one embodiment, electronic messages are intercepted from a chat application of an electronic communication system, the chat application being configured to simulate a conversation with a user during a messaging session, the method including intercepting electronic messages of the chat application at a particular point in time during the messaging session.

[0028] A chat application (e.g., a chatbot) may be used to conduct online chat conversations via text or text-to-speech. For example, the chat application may be used to test or train users by asking them questions. The users may provide answers to the questions. The electronic message may consist, for example, of the text of the question. This embodiment may be beneficial because it allows for control over when modifications or adaptations of the electronic message (e.g., question) need to be made in accordance with the present subject matter.

[0029] The point in time for generating new / modified questions may be, for example, predefined or dynamically determined. For example, the point in time may be dynamically defined based on user input. This allows, for example, different modes of operation for user conversation / training. For example, an easy training mode and a difficult training mode may be used. An easy training mode may only consider modifying a small portion of the questions (e.g., at the beginning of the conversation), while a difficult mode of operation may modify more questions (e.g., at different stages of the conversation).

[0030] The image may be determined or generated by modifying the intent of the intercepted electronic message, which may be performed by removing an intent from the intercepted electronic message to provide incomplete information, by adding a confusing intent to add ambiguity, or by adding a misleading intent to the intercepted electronic message, or any combination thereof.

[0031] According to one embodiment, the electronic communication system is a chat server configured to distribute messages between chat clients, e.g., electronic messages are intercepted and processed in accordance with the present subject matter before being distributed.

[0032] According to one embodiment, an electronic message is received from a first chat client and addressed to a second chat client, and the method further includes detecting confidential information in the received electronic message, wherein the determined image's noisy content includes non-confidential information, and the determined image is provided in place of the received electronic message. The confidential information may include, for example, personal information such as a full name.

[0033] According to one embodiment, the method further includes providing a knowledge graph representing a domain of computer-related bug fixes. The knowledge graph includes entities representing problems and solutions. The message intent represents a computer-related technical problem and is associated with a set of solutions in the knowledge graph. Determining the images includes using the knowledge graph to determine related intents for the message intent that are associated with solutions different from the set of solutions. The method further includes extracting, from the predetermined set of images, a subset of images associated with the related intents and not associated with the message intent.

[0034] According to one embodiment, the method further includes creating a knowledge graph using communication transcripts and / or logs of past data communications, and clustering the intents in the knowledge graph according to one or more graph characteristics of the knowledge graph, wherein the graph characteristics include any one of a centrality index of each node in the graph and a distance from each node to other nodes in the graph.

[0035] FIG. 1A illustrates a computer system 100 according to an example of the present subject matter. The computer system 100 includes an agent computer 105. The agent computer 105 may be an electronic communication system. The computer system 100 includes an electronic communication controller 101. The electronic communication controller 101 is connected to a chat link 103 and receives each successive message transmitted between the user 102 and the agent computer 105 during a conversation or chat session. The link 103 may connect the agent computer 105 to a remote user computer of the user 102 or may be a link to a display device of the agent computer 105. The link 103 may be established over the Internet or other data channel. The link 103 may enable a conversation or chat that includes a stream of text messages exchanged between the user 102 and the agent computer 105.

[0036] The successive messages of the user 102 are received by the analysis unit 106 of the agent computer 105. The analysis unit 106 performs the function of analyzing the messages to determine the problem or inquiry of the user 102 that is the subject of the chat with the agent computer 105. If the messages are in text format, the analysis unit 106 includes a text analysis function to perform this function. The function of the analysis unit 106 may be part of a process to identify the specific response that the user 102 provided to a question from the agent computer 105.

[0037] The agent computer 105 includes a question builder 116 that receives input from the analysis unit 106. The question builder 116 uses these inputs (e.g., a chatbot) to construct or compose requests or questions. A request is a statement of a goal related to a problem, to which the user 102 must provide a solution. The question builder 116 may also generate requests without receiving input from the analysis unit, for example, to initiate a conversation with the user 102. Messages generated by the question builder 116 and / or received from the user 102 may be intercepted or provided as input to the electronic communication controller 101 via the link 103. The electronic communication controller 101 may use the knowledge graph 118 and / or the set of images 120 as information sources for modifying the intercepted messages in accordance with the present subject matter.

[0038] Messages exchanged between user 102 and agent computer 105 may be displayed on window 130, as shown in Figure 1B. Window 130 may be displayed on the user interface of agent computer 105 if user 102 is physically in contact with agent computer 105, or may be displayed on user 102's remote user computer.

[0039] FIG. 1B shows an example of a timeline view of chat messages between user 102 and agent computer 105. A window 130 showing a record of a chat session includes a first display area 131 for displaying messages and a second display area 133 for displaying timestamps of the chat messages. The messages are aligned with their respective timestamps. In the example shown in FIG. 1B, agent computer 105 may provide messages 135.1 through 135.n. Each message may be a question for user 102. In one example, the question may be supplemented by an image related to the item in the question. In another example, the question may be provided as an image. User 102 may provide corresponding response messages 136.1 through 136.n.

[0040] Although shown as a separate component, in another example, the electronic communications controller 101 may be part of the agent computer 105 .

[0041] 2 is a flowchart of a method according to an example of the present subject matter. For convenience of explanation, the method of FIG. 2 may be implemented in the system illustrated in FIG. 1A, but is not limited to this implementation. The method of FIG. 2 may be performed by, for example, electronic communication controller 101.

[0042] In step 201, an electronic message to be provided by the electronic communication system 105, for example, to a user or recipient, may be received or identified. The electronic message may be retrieved from a conversation log file or may be intercepted by the electronic communication controller 101. The electronic message may be, for example, a message sent from a sending computer to a receiving computer, which is the electronic communication system 105. The electronic message may be, for example, an outgoing message from the electronic communication system 105. In this case, the electronic communication controller 101 may be configured to intercept the electronic message before it is sent from the electronic communication system 105. If the electronic message is an incoming message to the electronic communication system 105, the electronic communication controller 101 may be configured to intercept the incoming electronic message before it is provided to a receiving application of the electronic communication system 105. In another example, the electronic message may be a message generated by an application of the electronic communication system 105 and displayed on an interface of the electronic communication system 105. In this case, the electronic communication controller 101 may be configured to intercept the electronic message before it is displayed. The electronic communication controller 101 may or may not be part of the electronic communication system 105.

[0043] The electronic message may be any type of electronic communication data structure. The electronic message may be, for example, an email, an instant message, a voice message, or a text message. The electronic message may be a message in a conversation. The electronic message may be one of the chat messages in the conversation. The conversation may be a series of messages sent between a chat agent and one or more users. The electronic message may be, for example, the first chat message in the conversation or a randomly selected electronic message in the conversation. In another example, the electronic message may be a selected chat message in the conversation. The selection may be performed based on selection criteria. The selection criteria may, for example, require that the message to be modified be received after a correct answer by the user.

[0044] In step 203, the intent of the received or intercepted electronic message may be determined. The determined intent may be referred to as a message intent. A message intent may refer to, for example, a goal a chatbot agent has in mind when providing a question or comment. Classifying an intent may be automatically associating text with a particular purpose or goal. The message intent may be determined, for example, by a classifier. The classifier may analyze fragments of text and classify them into intents such as "computer virus," "fever," etc. The intent classifier may use, for example, a machine learning algorithm that can associate words or expressions with specific intents.

[0045] At step 205, image-based noisy content may be generated. The image-based noisy content may include noisy content that is distinct from the content of the received electronic message. Although distinct from the content of the received electronic message, the image-based noisy content is associated with the message intent. The image-based noisy content may be, for example, one or more images. In another example, the image-based noisy content may include text derived from one or more images.

[0046] For example, image-based noisy content may be generated using a variational autoencoder (VAE) to identify text messages that have the same domain as the received electronic message but a different message intent than the received electronic message, using a database of previously used text messages. These identified text messages may be used as input to the VAE to generate one or more images.

[0047] In step 207, the content having the generated image-based noise may be provided. As one example, the content having the generated image-based noise may be provided in place of the intercepted electronic message. For example, if the intercepted electronic message is displayed on a computer interface, the content having the generated image-based noise may be displayed instead. As another example, the content having the generated image-based noise may be provided in addition to the intercepted electronic message. For example, if the intercepted electronic message is displayed on a computer interface, the content having the generated image-based noise may be displayed before or after the display of the intercepted electronic message (in the chat flow). This allows the text of the electronic message to be presented along with any of the generated ambiguous images in the next utterance.

[0048] After providing the generated image-based noisy content, it may be determined whether the user identified the generated content as noise. If the user does not identify the content as noise, the electronic communications controller 101 may stop providing further noisy content according to an easy mode of operation. In another example, the electronic communications controller may remain in this state for a short period of time according to a difficult mode of operation to see if the user can finally identify the noise in the conversation.

[0049] 3 is a flowchart of a method for generating an image according to an example of the present subject matter. For illustrative purposes, the method of FIG. 3 may be implemented in the system illustrated in FIG. 1A, but is not limited to this implementation. The method of FIG. 3 may be performed by, for example, electronic communication controller 101.

[0050] Related intents for a message intent may be determined in step 301. For example, a knowledge graph may be used to discover a set of intents related to the message intent.

[0051] In step 303, images associated with related intents but not associated with the message intent may be extracted from the knowledge base. The extracted images may be, for example, image-based noisy content determined in step 205. For example, images tagged with intents that are present in association set Y but not in image set X may be extracted. These images correspond to intents that are not present in the intercepted electronic message but are associated with them. These images may introduce ambiguity.

[0052] FIG. 4 is a flowchart of a method according to an example of the present subject matter. For convenience of explanation, the method of FIG. 4 may be implemented in the system illustrated in FIG. 1A, but is not limited to this implementation. The method of FIG. 4 may be performed, for example, by the electronic communication controller 101. The method of FIG. 4 provides further details of step 205. That is, the method of FIG. 4 may be used to determine noisy content based on an image in step 205.

[0053] A simultaneous embedding space may be determined in step 401. The simultaneous embedding space may represent an electronic message and an image.

[0054] At step 403, the electronic message (received at step 201) and a set of predetermined images may be represented in an embedding space.

[0055] At step 405, the set of images may be ranked based on the similarity between the representation of the set of images in the embedding space and the representation of the electronic message.

[0056] At step 407, a subset of the set of images may be selected based on the ranking. The method of FIG. 4 can rank all images by their similarity to the text, for example, using joint embedding. The top k images in this ranking are directly related to the text. To introduce ambiguity, the next k images are selected. These images are somewhat relevant but not completely useful in the context of the text t, resulting in ambiguity. The value of k can be fine-tuned.

[0057] FIG. 5 illustrates a method for linking text and image embedding spaces according to the present subject matter. A Siamese network 500 may be trained to identify image-text relationships. Images are analyzed by a CNN-based encoder, and text descriptions are analyzed by a BERT-based transformer. Joint embeddings are learned using hinge loss to align images and text according to the Siamese network neural network architecture. The Siamese network 500 is trained using a labeled dataset. The labeled dataset consists of entries. Each entry consists of a triplet consisting of an image (e.g., a screenshot), associated or random question text, and a label. If the question text is non-random, the label may be set to 1; otherwise, the label may be set to 0.

[0058] Figure 6 illustrates a method for testing a learned embedding space according to the present subject matter. One or more past messages 601 in a conversation may be represented in the embedding space (606). Additionally, one current message 604 and an associated generated image 605 may be represented in the joint embedding space (607). The similarity 608 between the two representations may be determined, and a loss function may be evaluated accordingly (609). Based on the determined loss, it may be determined whether the generated image improved training (610).

[0059] Figure 7 illustrates a method for generating image-based noisy content in accordance with the present subject matter, which provides an alternative method for generating image-based noisy content.

[0060] An input 700 may be provided. The input 700 is intended to be initially provided to a user so that the user can provide a response or feedback to the input. The input 700 may, for example, represent a virtual customer complaint created by an SME. The input may, for example, include an electronic message T0 and an image I0 related to the complaint. For example, the image I0 may be a screenshot of a terminal log, and the electronic message T0 may be a question related to the issue described in the log. The method may allow the input to be modified (e.g., to make the question / complaint more difficult) to better test the user.

[0061] At step 701, text analysis of the electronic message T0 may be performed. The text analysis may be performed to determine the intent and relationships of the electronic message T0. The text analysis may be performed, for example, using the history 720 of conversational utterances on the text. At step 702, image analysis of the image I0 may be performed. The image analysis may be performed to determine the intent and relationships of the image I0. The image analysis may be performed, for example, using the history 720 of conversational utterances on the image. Steps 701 and 702 may provide an input intent for the input and may further provide a ranking of the input intents, for example, based on their importance. At step 704, possible solutions to the complaint may be identified using the determined intent and the knowledge graph.

[0062] Steps 705-706 provide a first method for generating content with image-based noise based on the results of steps 701-704. Steps 705-708 provide a second method for generating content with image-based noise based on the results of steps 701-704. Step 711 provides a third method for generating content with image-based noise based on the results of steps 701-704. Step 712 provides a fourth method for generating content with image-based noise based on the results of steps 701-704. Step 713 provides a fifth method for generating content with image-based noise based on the results of steps 701-704.

[0063] A first method may enable modifying the input to introduce incompleteness into the input. To that end, in step 705, input intents that can be dropped from the input may be identified. To that end, in step 706, image analysis may be performed on the conversational speech on the image based on the domain and context of the input and further based on the image metadata. For example, a knowledge graph may be used to discover intents to drop from the input. For example, the conversational speech on the image may be used to determine which intents in the input are more important than other intents in the input. These important intents may be the intents to drop. For example, if the input has 10 intents, content representing one or two of the 10 intents may be dropped from the input, and the resulting / remaining input may be provided to the user as an incomplete input. These identified intents to drop may be utilized in various scenarios. In a first scenario, a new text message may be generated using the content of image I0 and the results of the image analysis. For example, the new text message may include the content of image I0, excluding the content representing the intents to drop. In a second scenario, a new image may be generated using the content of the input text T0 and the results of text analysis (step 708) performed on the conversational utterances on the text. The text analysis (708) may enable identification of noisy intents (e.g., unimportant intents) among the intents of the input text T0 using a knowledge graph. The new image may include the content of the input text T0 that represent these noisy intents. In a third scenario, a new text message and a new image may be generated using the image I0 and the content of the text T0. The new text and / or the new image may include content that represents only unimportant intents. This new image and / or the new text message may be provided to the user as an incomplete input in place of the input.

[0064] A second method may enable the use of a VAE to generate a noisy image. To that end, an input image I0 and the identified intents to be removed may be provided as inputs to the VAE to generate the noisy image in step 707. For example, the VAE takes image I0 and contextual information extracted by identifying the intents that need to be removed from image I0. The noisy image may be provided in addition to or instead of the user input.

[0065] A third method may allow misleading information to be provided to a user by using a VAE to inject noise into an input image based on text metadata associated with a textual conversational utterance in step 711. The VAE takes an image I0 and contextual information extracted by identifying an intent from a text message T0 that needs to be added as noise.

[0066] A fourth method may allow misleading information to be provided to a user by inserting misleading snippets into the input text message T0 at step 712 based on image metadata associated with conversational speech on the image.

[0067] A fifth method may allow for introducing mismatches into the input using text and image metadata of conversational utterances on the text and images in step 713. For example, from an input image I0 and text message T0, a new image and a new text message may be generated.

[0068] Thus, any of the first to fifth methods may enable the user to be provided with conversational speech 715 with noise in step 714.

[0069] FIG. 8 is a diagram illustrating a general computerized system 900 suitable for implementing at least some of the method steps associated with the present disclosure.

[0070] It should be understood that the methods described herein are at least partially non-interactive and automated by a computerized system such as a server or embedded system. However, in exemplary embodiments, the methods described herein can be implemented in a (partially) interactive system. These methods can also be implemented in software 912, 922 (including firmware 922), hardware (processor) 905, or a combination thereof. In exemplary embodiments, the methods described herein are implemented in software as executable programs and executed by a special-purpose or general-purpose digital computer such as a personal computer, workstation, minicomputer, or mainframe computer. Thus, the most common system 900 includes a general-purpose computer 901.

[0071] In an exemplary embodiment, from a hardware architecture perspective, as shown in FIG. 8 , a computer 901 includes a processor 905, a memory (main memory) 910 coupled to a memory controller 915, and one or more input and / or output (I / O) devices (or peripherals) 10, 945 communicatively coupled via a local input / output controller 935. The input / output controller 935 can be, but is not limited to, one or more buses or other wired or wireless connections as known in the art. While omitted for simplicity, the input / output controller 935 may include additional elements to enable communication, such as controllers, buffers (caches), drivers, repeaters, and receivers. Furthermore, the local interface may include address, control, or data connections, or a combination thereof, to enable appropriate communication between the above components. As described herein, the I / O devices 10, 945 may generally include any general-purpose cryptographic or smart card known in the art.

[0072] Processor 905 is a hardware device for executing software, particularly software stored in memory 910. Processor 905 can be any custom or commercially available processor, a central processing unit (CPU), a coprocessor among multiple processors associated with computer 901, a semiconductor-based (microchip or chipset type) microprocessor, or generally any device for executing software instructions.

[0073] The memory 910 can include any one or any combination of volatile memory elements (e.g., random access memory (RAM, such as DRAM, SRAM, SDRAM)) and non-volatile memory elements (e.g., ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), programmable ROM (PROM)). Note that the memory 910 can have a distributed architecture where various components are located remotely from one another yet accessible by the processor 905.

[0074] The software in memory 910 may include one or more separate programs, each of which includes an ordered list of executable instructions for implementing logical functions, particularly functions related to embodiments of the present invention. In the example of Figure 8, the software in memory 910 includes instructions 912, such as instructions for managing a database, such as a database management system.

[0075] The software in memory 910 should also typically include a suitable operating system (OS) 911. The OS 911 essentially controls the execution of other computer programs, such as software 912, possibly for implementing the methods described herein.

[0076] The methods described herein may be in the form of a source program 912, an executable program 912 (object code), a script, or any other entity that includes a set of instructions 912 to be executed. In the case of a source program, the program must be translated by a compiler, assembler, interpreter, etc., to operate properly in conjunction with the OS 911. These may or may not be contained within the memory 910. The methods may also be written as an object-oriented programming language with classes of data and methods, or as a procedural programming language with routines, subroutines, or functions, or a combination thereof.

[0077] In an exemplary embodiment, a conventional keyboard 950 and mouse 955 may be coupled to the input / output controller 935. Other output devices, such as the I / O device(s) 945, may include input devices, such as, but not limited to, a printer, scanner, microphone, etc. Finally, the I / O device(s) 10, 945 may further include devices that communicate both input and output, such as, but not limited to, a network interface card (NIC) or modulator / demodulator (for accessing other files, devices, systems, or networks), a radio frequency (RF) or other transceiver, a telephone interface, a bridge, a router, etc. The I / O device(s) 10, 945 may also be any general-purpose cryptographic card or smart card known in the art. The system 900 may further include a display controller 925 coupled to the display 930. In an exemplary embodiment, the system 900 may further include a network interface for coupling to a network 965, which may be an IP-based network for communication between the computer 901 and any external servers, clients, etc. via a broadband connection. The network 965 transmits and receives data between the computer 901 and the external system 30. The external system 30 may be involved in performing some or all of the steps of the methods described herein. In an exemplary embodiment, the network 965 may be a managed IP network administered by a service provider. The network 965 may be implemented wirelessly using wireless protocols and technologies such as WiFi, WiMax, etc. The network 965 may also be a packet-switched network such as a local area network, a wide area network, a metropolitan area network, the Internet network, or other similar network environment.The network 965 may be a fixed wireless network, a wireless local area network (WLAN), a wireless wide area network (WWAN), a personal area network (PAN), a virtual private network (VPN), an intranet, or other suitable network system and includes devices for transmitting and receiving signals.

[0078] If computer 901 is a PC, workstation, intelligent device, etc., the software in memory 910 may further include a basic input / output system (BIOS) 922. The BIOS is a set of basic software routines that initializes and tests hardware at startup, starts the OS 111, and supports data transfer between hardware devices. The BIOS is stored in ROM so that it can be executed when computer 901 is started.

[0079] During operation of the computer 901, the processor 905 is configured to execute software 912 stored in the memory 910, to communicate data to and from the memory 910, and to generally control the operation of the computer 901 in accordance with the software. The methods and OS 911 described herein are read, in whole or in part, typically the latter, by the processor 905, and possibly buffered within the processor 905, before being executed.

[0080] 8, the methods can be stored on any computer-readable medium, such as storage 920, for use by or in connection with any computer-related system or method. Storage 920 may include disk storage, such as HDD storage.

[0081] This subject matter may include the following clauses:

[0082] Item 1: A computer-implemented method for electronic messaging, comprising: receiving an electronic message in an electronic communication system, the message being determined for a recipient; determining a message intent for the received electronic message; generating image-based noise content that is different from the content of the received electronic message and that is related to the message intent; controlling the electronic communication system to provide the content with the image-based noise to the recipient in place of the received electronic message, or to provide the received electronic message to the recipient in addition to the content with the image-based noise; 11. A computer-implemented method comprising:

[0083] Section 2: determining the domain of said electronic message; determining a set of one or more text messages and a set of one or more images for the determined domain; generating the image-based noisy content includes: generating one or more noisy images from the set of text messages, wherein the image-based noisy content is the noisy image; and extracting noisy text from the set of images, wherein the image-based noisy content is the noisy text; generating noisy multimodal content from the set of images and the set of text messages, wherein the image-based noisy content is the noisy multimodal content; 2. The method of claim 1, comprising any one of:

[0084] Clause 3: The method of clause 2, wherein the set of images and the set of text messages are determined such that their intents are related to a portion of the message intent.

[0085] Clause 4: Generating the image-based noise content determining a joint embedding space for representing the electronic message and the image; representing a text message and a set of predetermined images in the embedding space, the text message being the electronic message or another message having an intent that is at least a portion of the message intent; ranking the set of images based on a similarity between a representation of the set of images in the embedding space and a representation of the text message; selecting a subset of one or more noisy images from the set of images based on the ranking, wherein the image-based noisy content includes the noisy images; and The method according to any one of items 1 to 3 above, comprising:

[0086] Clause 5: The method of clause 4, wherein the subset of noisy images is the second k ranked images, where k is a configurable parameter value.

[0087] Clause 6: The method of clause 5, further comprising setting the value of the parameter k according to a desired noise level of the noisy content.

[0088] Clause 7: The method of any of clauses 4 to 6 above, wherein the joint embedding is trained using hinge loss to align images and text according to a Siamese neural network architecture.

[0089] Clause 8: Generating the image-based noise content includes: determining related intents of the message intent using a knowledge graph, the knowledge graph including entities representing intents; and extracting a subset of images from a predetermined set of images that are associated with the related intent and that are not associated with the message intent, wherein the image-based noisy content includes the subset of images; The method according to any one of items 1 to 3 above, comprising:

[0090] Clause 9: Using the electronic message to determine a context of a messaging session, the context of the messaging session being defined by at least a subgraph of the knowledge graph; determining the related intents using the subgraph as intents that belong to the determined context and are associated with the message intent; 9. The method of claim 8, further comprising:

[0091] Clause 10: Generating the image-based noise content includes: determining related intents of the message intent using a knowledge graph, the knowledge graph including entities representing intents; and The above-mentioned Intent generating an image using a variational autoencoder for The method according to any one of items 1 to 3 above, comprising:

[0092] Clause 11: A method according to any one of clauses 1 to 10 above, wherein the electronic message is intercepted from a chat application of the electronic communication system, the chat application being configured to simulate a conversation with a user during a messaging session, and the receiving includes intercepting a message from the chat application at a predetermined point in the messaging session.

[0093] Clause 12: A method according to any one of clauses 1 to 11 above, wherein the electronic communication system is a chat server configured to distribute messages between chat clients.

[0094] Clause 13: The electronic message is received from a first chat client and is destined for a second chat client, and the method further includes detecting confidential information in the received electronic message; Image-based The noisy content includes non-sensitive information, Image-based noise content 13. The method of claim 12, wherein the electronic message is provided in lieu of the received electronic message.

[0095] Clause 14: Providing a knowledge graph representing a domain of computer-related bug fixing, the knowledge graph including entities representing problems and solutions, wherein the message intent represents a computer-related technical problem and is associated with a set of solutions in the knowledge graph; The aforementioned Generating image-based noisy content teeth, using the knowledge graph to determine related intents for the message intent, the related intents being associated with solutions different from the set of solutions; and extracting a subset of images from a predetermined set of images that are associated with the related intent and that are not associated with the message intent; 14. The method according to any one of items 1 to 13 above, comprising:

[0096] Clause 15: Creating the knowledge graph using communication records and / or logs of past data communications; Clustering intents of the knowledge graph according to one or more graph characteristics of the knowledge graph, the graph characteristics including any one of a centrality index of each node in the graph and a distance from each node to other nodes in the graph; 15. The method of claim 14, further comprising:

[0097] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium having stored thereon computer-readable program instructions for causing a processor to carry out aspects of the present invention.

[0098] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), static random access memory (SRAM), CD-ROMs, DVDs, memory sticks, floppy disks, punch cards, or mechanically encoded devices having instructions recorded on ridge-in-groove structures, etc., and suitable combinations thereof. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over a wire.

[0099] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing device / processing device. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, a LAN, a WAN, or a wireless network, or a combination thereof). The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium in the respective computing device / processing device for storage.

[0100] The computer-readable program instructions for carrying out the operations of the present invention can be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the "C" programming language and similar programming languages. The computer-readable program instructions can execute entirely on the user's computer as a stand-alone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to customize the electronic circuitry for carrying out aspects of the present invention.

[0101] Aspects of the present invention are described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of such computer or other programmable data processing apparatus, create means for performing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner. The computer-readable storage medium having instructions stored thereon thereby constitutes an article of manufacture including instructions for performing aspects of the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.

[0103] Computer-readable program instructions may also be loaded into a computer, other programmable apparatus, or other device and a series of operational steps executed on the computer, other programmable apparatus, or other device to create a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device perform the functions / operations identified in one or more blocks in the flowcharts and / or block diagrams.

[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for performing specific logical functions. In some implementations, the functions shown in the blocks may be performed in an order different from that shown in the figures. For example, depending on the functionality involved, two blocks shown in succession may actually be accomplished as a single step, may be executed simultaneously or substantially simultaneously, may be executed in a partially or fully overlapping manner, or the blocks may even be executed in reverse order. Note that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs specific functions or operations or executes a combination of dedicated hardware and computer instructions.

Claims

1. 1. A computer-implemented method for electronic messaging, comprising: receiving an electronic message in an electronic communication system, the message being determined for a recipient; determining a message intent for the received electronic message; generating image-based noise content that is different from the content of the received electronic message and that is related to the message intent; controlling the electronic communication system to provide the content with the image-based noise to the recipient in place of the received electronic message, or to provide the received electronic message to the recipient in addition to the content with the image-based noise; Including, generating the image-based noisy content includes: determining a joint embedding space for representing the electronic message and the image; representing a text message and a set of predetermined images in the embedding space, the text message being the electronic message or another message having an intent that is at least a portion of the message intent; ranking the set of images based on a similarity between a representation of the set of images in the embedding space and a representation of the text message; selecting a subset of one or more noisy images from the set of images based on the ranking, wherein the image-based noisy content includes the noisy images; and 11. A computer-implemented method comprising:

2. determining a domain of the electronic message; determining a set of one or more text messages and a set of one or more images for the determined domain; generating the image-based noisy content includes: generating one or more noisy images from the set of text messages, wherein the image-based noisy content is the noisy image; and extracting noisy text from the set of images, wherein the image-based noisy content is the noisy text; generating noisy multimodal content from the set of images and the set of text messages, wherein the image-based noisy content is the noisy multimodal content; The method of claim 1 , comprising any one of:

3. The method of claim 2 , wherein the set of images and the set of text messages are determined such that their intents are related to a portion of the message intent.

4. The method of claim 1 , wherein the subset of noisy images is the second k ranked images, where k is a configurable parameter value.

5. The method of claim 4 , further comprising setting the value of the parameter k according to a desired noise level of the noisy content.

6. The method of claim 1 , wherein the joint embedding is trained using hinge loss to align images and text according to a Siamese neural network architecture.

7. generating the image-based noisy content includes: determining related intents of the message intent using a knowledge graph, the knowledge graph including entities representing intents; and extracting a subset of images from a predetermined set of images that are associated with the related intent and that are not associated with the message intent, wherein the image-based noisy content includes the subset of images; The method of claim 1 , comprising:

8. using the electronic message to determine a context of a messaging session, the context of the messaging session being defined by at least a subgraph of the knowledge graph; determining the related intents using the subgraph as intents that belong to the determined context and are associated with the message intent; The method of claim 7 further comprising:

9. generating the image-based noisy content includes: determining related intents of the message intent using a knowledge graph, the knowledge graph including entities representing intents; and generating an image using a variational autoencoder for the associated intent, wherein the image-based noisy content includes the generated image; and The method of claim 1 , comprising:

10. 10. The method of claim 1, wherein the electronic message is intercepted from a chat application of the electronic communication system, the chat application configured to simulate a conversation with a user during a messaging session, and wherein receiving includes intercepting a message from the chat application at a predetermined time during the messaging session.

11. The method of claim 1 , wherein the electronic communication system is a chat server configured to distribute messages between chat clients.

12. 12. The method of claim 11, wherein the electronic message is received from a first chat client and is intended for a second chat client, the method further comprising detecting confidential information in the received electronic message, wherein the image-based noisy content includes non-confidential information, and the image-based noisy content is provided in place of the received electronic message.

13. providing a knowledge graph representing a domain of computer-related bug fixing, the knowledge graph including entities representing problems and solutions, the message intent representing a computer-related technical problem and associated with a set of solutions in the knowledge graph; generating the image-based noisy content includes: using the knowledge graph to determine related intents for the message intent, the related intents being associated with solutions different from the set of solutions; and extracting a subset of images from a predetermined set of images that are associated with the related intent and that are not associated with the message intent; The method of claim 1 , comprising:

14. Creating the knowledge graph using communication records and / or logs of past data communications; Clustering intents of the knowledge graph according to one or more graph characteristics of the knowledge graph, the graph characteristics including any one of a centrality index of each node in the graph and a distance from each node to other nodes in the graph; 14. The method of claim 13, further comprising:

15. 1. A computer program product comprising one or more computer-readable storage media embodied with computer-readable program instructions for execution by one or more processors of one or more computers, the computer-readable program instructions comprising: receiving, by the one or more processors, an electronic message from an electronic communication system, the message being determined for a recipient; determining, by the one or more processors, a message intent for the received electronic message; generating, by the one or more processors, content having image-based noise that is different from content of the received electronic message and that is associated with the message intent; controlling, by the one or more processors, the electronic communication system to provide the content with the image-based noise to the recipient in place of the received electronic message, or to provide the received electronic message to the recipient in addition to the content with the image-based noise; and instructions for executing the generating the image-based noisy content includes: determining a joint embedding space for representing the electronic message and the image; representing a text message and a set of predetermined images in the embedding space, the text message being the electronic message or another message having an intent that is at least a portion of the message intent; ranking the set of images based on a similarity between a representation of the set of images in the embedding space and a representation of the text message; selecting a subset of one or more noisy images from the set of images based on the ranking, wherein the image-based noisy content includes the noisy images; and a computer program product,

16. The program instructions include: determining, by the one or more processors, a domain of the electronic message; determining, by the one or more processors, a set of one or more text messages and a set of one or more images for the determined domain; generating the image-based noisy content includes: generating one or more noisy images from the set of text messages, wherein the image-based noisy content is the noisy image; and extracting noisy text from the set of images, wherein the image-based noisy content is the noisy text; generating noisy multimodal content from the set of images and the set of text messages, wherein the image-based noisy content is the noisy multimodal content; 16. The computer program product of claim 15, comprising any one of:

17. The computer program product of claim 16 , wherein the set of images and the set of text messages are determined such that their intents are related to a portion of the message intent.

18. 1. A computer system comprising one or more processors and one or more tangible storage media having stored thereon program instructions for execution by the one or more processors, the program instructions comprising: receiving an electronic message; determining a message intent for the received electronic message; generating image-based noise content that is different from the content of the received electronic message and that is related to the message intent; controlling an electronic communication system to provide the content with the image-based noise in place of the received electronic message or to provide the received electronic message in addition to the content with the image-based noise; and instructions for executing the generating the image-based noisy content includes: determining a joint embedding space for representing the electronic message and the image; representing a text message and a set of predetermined images in the embedding space, the text message being the electronic message or another message having an intent that is at least a portion of the message intent; ranking the set of images based on a similarity between a representation of the set of images in the embedding space and a representation of the text message; selecting a subset of one or more noisy images from the set of images based on the ranking, wherein the image-based noisy content includes the noisy images; and 2. A computer system comprising:

19. The program instructions include: determining, by the one or more processors, a domain of the electronic message; determining, by the one or more processors, a set of one or more text messages and a set of one or more images for the determined domain; generating the image-based noisy content includes: generating one or more noisy images from the set of text messages, wherein the image-based noisy content is the noisy image; and extracting noisy text from the set of images, wherein the image-based noisy content is the noisy text; generating noisy multimodal content from the set of images and the set of text messages, wherein the image-based noisy content is the noisy multimodal content; 20. The computer system of claim 18, comprising one of:

Citation Information

Patent Citations

  • File transmission apparatus and file network system

    JP2006074526A

  • Automated masking of confidential information in unstructured computer text using artificial intelligence

    US20200226288A1