Electronic message filtering
By clustering message headers and using machine learning classifiers to filter purchase-related messages, the method addresses the challenge of aggregating purchase transaction data across multiple retailers, enhancing data extraction efficiency and accuracy.
Patent Information
- Application Number
- JP2024021128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-05-17
- Filing Date
- 2024-02-15
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2038-05-01
AI Technical Summary
The complexity of purchase transactions across multiple retailers, coupled with the diversity of customer profiles and the fragmentation of purchase history data, makes it difficult for merchants to efficiently share and aggregate customer purchase information.
A method for processing electronic messages that involves clustering message headers based on similarity, using a density-based clustering process, and generating machine learning classifiers to filter purchase-related messages, thereby enabling the efficient extraction and aggregation of purchase transaction data.
This approach significantly reduces the time and resources required to identify and filter purchase-related electronic messages, improving the accuracy and efficiency of purchase data aggregation and analysis across multiple retailers.
Smart Images

Figure 0007681140000002 
Figure 0007681140000003 
Figure 0007681140000004
Abstract
Description
[Background technology]
[0001] People buy goods from many merchants with a variety of payment options. Such purchase transactions are usually accompanied by a physical receipt from a store or a message from the buyer's messaging account. The delivery will be confirmed by an electronic confirmation message sent to your email account (e.g. The large number and variety of confirmation messages allows people to confirm their purchases and It is difficult to grasp the overall history. In addition, The high diversity of customers means that merchants have a lot of options to create accurate customer profiles. It is becoming more difficult to obtain purchase history data. Even if a common identifier (e.g., a loyalty card or credit card) is used, Their purchases are typically tracked only by the merchant who issued the identifier to the customer. This lack of customer information makes it difficult to efficiently share customer purchase transaction information across multiple retailers. There is a limit to what we can find.
[0002] In order to improve this issue, we will implement the following measures to improve the quality of messages, such as purchase confirmation messages and delivery confirmation messages. Extracting purchase-related information from data sources issued directly to consumers by merchants A reporting system has been developed. Summary of the Invention
[0003] The present invention provides a method for transmitting messages between network nodes and managed by one or more message servers. The one or more network data storage systems may include a network of storage systems that may store user accounts. A method for processing a set of associated and stored electronic messages, the method being performed by a computer device. Each electronic message is associated with a sender, a header, and a body. According to this method, one or more of the network data storage systems The headers in the stored set are sent to multiple user accounts from one or more of the message servers. For each of the one or more senders, a count is fetched by the network node. For each, we associate a cluster with its respective dense region in the clustering data space. The fetched packets associated with the sender are based on a density-based clustering process. The headers that are retrieved are grouped into clusters. Within the clustering data space, The fetched headers are spaced apart based on the similarity between each pair of fetched headers. For each of one or more of the clusters, a message is sent by the network node. associated with the fetched headers in the cluster from one or more of the message servers. electronic media in a collection stored in one or more of the network data storage systems; A sample of each electronic message is obtained. To generate a classification dataset for each cluster, we use one or more purchase-related labels. and each label in a given label set, including its associated confidence, is used to generate a machine learning classifier. The clusters are classified by the cluster labels. Based on at least one cluster classification rule that maps a given label set to a Each cluster is assigned a label selected from the list. One of the purchase-related labels is assigned. Filtering purchase-related electronic messages for each of the one or more assigned clusters A filter that matches the current setting is automatically generated.
[0004] The present invention also provides a method for transmitting messages between network nodes and managed by one or more message servers. Each user account in one or more network data storage systems managed by the a computer device for processing a set of stored electronic messages associated with a Each electronic message is characterized in that it is a method performed by In this way, the headers in the set are associated with each of one or more senders. and fetched from one or more of the network data storage systems. For each of the one or more senders, the fetched headers are grouped into clusters. The process of grouping the fetched headers is independent of the content of the message body. Based on the similarity between headers in a cluster, the fetched headers are sorted into multiple clusters. For each of the plurality of clusters, Fetched from one or more of the network data storage systems and stored in the cluster One or more samples of the electronic messages associated with the assigned header are taken. A cluster is obtained by dividing the header and Based on the content of the text, the machine may be classified as receipt-related or non-receipt-related. Each electronic message filter is assigned a set of rules that are related to the receipt of the message. For each of one or more of the clusters specified as Each electronic message filter checks the subject field string in the header of the electronic message. Define different rules to match different patterns.
[0005] In some examples, one or more of the filters may be processed by one or more message servers. Each user account in one or more network data storage systems managed by Select a purchase-related electronic message from the set of stored electronic messages associated with the purchase A processor is provided for at least one network communication channel to select can be done.
[0006] The invention also relates to a computer apparatus operable to carry out the above-described method and to a method for implementing the above-described method. A computer that stores computer readable instructions that cause a computer device to perform the method described above. The readable medium is characterized. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 is an explanatory diagram illustrating an example of a network communication environment. [Diagram 2] FIG. 2 is a general illustration of the electronic message processing stages performed by an exemplary purchase transaction data retrieval system. [Diagram 3] FIG. 10 is an explanatory diagram illustrating an example of an electronic message. [Figure 4] FIG. 2 is a flow diagram illustrating an example process for creating an electronic message filter. [Diagram 5] FIG. 5 is an illustration of data related to the various stages of the electronic message filter generation process of FIG. 4. [Figure 6] FIG. 2 is a flow diagram illustrating an example process for creating an electronic message filter. [Figure 7] FIG. 1 illustrates an example system for generating electronic message filters. [Figure 8] FIG. 13 is an explanatory diagram illustrating an example of a cluster of headers in a clustering data space. [Figure 9] FIG. 2 is a flow diagram illustrating an example process for grouping headers of electronic messages into clusters. [Figure 10] FIG. 1 illustrates an example system for generating electronic message filters. [Figure 11] FIG. 2 is a flow diagram illustrating an example process for creating an electronic message filter. [Figure 12] FIG. 1 is a block diagram illustrating an example of a computer device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] In the following description, like reference numbers are used to identify like elements. The drawings are intended to illustrate major features of the exemplary embodiments. It is not intended to show every feature, but rather to indicate the relative dimensions of the depicted elements. It is not intended as a legal document, nor is it drawn to scale.
[0009] [I. Definitions of Terms] "Product" means any tangible or intangible item or good that can be purchased or used. is a service.
[0010] An "electronic message" is sent between physical network nodes from sender to receiver. , a persistent text-based record of information stored in non-transitory computer-readable memory Electronic messages may be structured (e.g., hypertext containing structure tag elements). Text markup language (HTML) messages) or unstructured (e.g. For example, a plain text message.
[0011] A "purchase-related electronic message" is an electronic message related to the purchase of a product. Examples of related electronic messages include order confirmations, delivery confirmations, refunds, cancellations, back orders, and customer support. This includes promotions, offers, and sales promotions.
[0012] The "sender domain" of an electronic message is the domain from which the electronic message is sent. For example, if an electronic message address is "local-mail If the "local-part" has the format "rt@domain", then the "local-part" is the "domain" specifies the management scope of the message address. Multiple message addresses may share the same sender domain.
[0013] A "receipt" is an electronic message confirming the status of a purchase of one or more items. Examples of receipts include an order confirmation electronic message and a delivery confirmation electronic message. Included.
[0014] "Purchase Transaction Information" (also referred to as "Purchase Transaction Data") is information relating to the purchase of a product. Purchase transaction data may include, for example, invoice data, purchase confirmation data, and product order information. (e.g., seller name, order number, order date, product description, product name, product quantity, product price, consumption Taxes, shipping charges, and order amount) and product shipping information (e.g. billing address, shipping company, shipping address, location, estimated shipping date, estimated delivery date, and tracking number.
[0015] "Computer" means any computer program stored temporarily or permanently on a computer-readable medium. Any machine, device, or apparatus that processes data in accordance with computer-readable instructions. "Computer Device" means one or more independent computers. An operating system is a system that manages the execution of tasks and the use of computing resources and hardware. A software component of a computer that manages and coordinates the sharing of hardware resources. "Software application" (software, application, computer Software, computer applications, programs, and computer programs A program (also called a programming language) is something that a computer can interpret and execute to perform one or more specific tasks. A "data file" is a set of instructions that can be stored in a software application. It is a block of information that persistently stores data used by
[0016] The term "computer-readable medium" (also referred to as "memory") refers to a medium that is Stores information (e.g., instructions and data) that can be read by a computer This refers to any tangible, non-transitory device capable of making such information tangible. Suitable storage devices for implementation include, for example, random access memory (RAM). Semiconductor memory devices such as EEPROMs, EPROMs, EEPROMs, and flash memory devices , magnetic disks such as internal hard disks and removable hard disks, optical magnetic disks Any physical form of data, including hard disks, DVD-ROM / RAM, and CD-ROM / RAM. This includes, but is not limited to, non-transitory computer readable memory. .
[0017] A "network node" is a physical junction or connection point within a communications network. Examples of network nodes include terminals, computers, and network switches. "Server System" includes, but is not limited to, one or more network nodes. A "client node" is a node that provides information or services to a is a network node that requests a service from a server system.
[0018] In this specification, the term "comprises" includes its subject matter, but means, but is not limited to, and the term "including" means "Based on" means including (having) the subject matter, but is not limited to it. means based at least in part on that subject.
[0019] II. Filtering Purchase-Related Electronic Messages A. Introduction Every day, people around the world send and receive nearly 200 billion emails. Very few of these emails are related to purchases. sift through your messaging accounts to get a regular dose of actionable information Identifying and retrieving purchase-related e-mail requires a significant amount of time and resources.
[0020] According to the examples described herein, to communicate purchase-related information to designated recipients. Discover and filter purchase-related electronic messages sent between physical network nodes Improved systems and methods for tagging a video signal are provided. These systems and methods include: There are many different electronic message formats used by different merchants and they vary between merchants. This will solve the practical problems that have arisen as a result of the rapid increase in These examples show how machine-generated electronic message headers can be automatically learned in terms of structure and semantics. This allows new message sources, new markets, and different languages to be supported. These examples span a wide variety of electronic message formats. A purchase-related electronic message that can be identified and filtered with high accuracy. Provides electronic message detection and filtering services.
[0021] The examples described herein are specific to machine-generated purchase-related electronic messages. Uses knowledge of structural features of purchases to automatically discover and filter purchase-related electronic messages This process is carried out by a computer device. It improves the processing resources, data storage resources, and network resources compared to conventional methods. This significantly reduces the number of queries and filter generation times. In some instances, this improvement can be achieved by: A unique sequence of specific electronic message processing rules in a network communications environment is executed. In some instances, multiple sales may be performed. Automatically identify the individual characteristics of multiple machine-generated purchase-related electronic messages sent by a merchant Dynamically learn and configure computing devices so that processing automatically adapts to their characteristics. This provides further advantages over conventional approaches. Some examples are, e.g., Generate purchase-related electronic message templates for each set of multiple sellers To accommodate the various levels of variation in the templates used by different vendors By automatically adjusting the processing of the computer device, the accuracy of the message discovery process is improved. This substantially improves the quality and efficiency of the system.
[0022] In certain instances, these systems and methods may include: Each message template defines the structural elements of the body of the electronic message. An electronic message filter that matches the headers of purchase-related electronic messages generated by Improved special-purpose computer device that can be programmed to automatically learn These systems and methods also include one or more message servers. Each user account in one or more network data storage systems A purchase-related electronic message can be selected from a set of associated stored electronic messages. At least one network communication channel must have a learned electronic message filter installed. A special-purpose computer that is programmed to install a This includes computer devices.
[0023] These improved systems and methods allow for the transfer of merchandise purchase information via a wide variety of electronic messages. Visualize and organize individual purchase history by identifying, extracting, and aggregating page types. Providing individuals with enhanced tools and targeted, less intrusive advertising Improved sales across various consumer segments enabling strategies and other marketing initiatives These improved systems can provide inter-party purchasing information to distributors and other organizations. Systems and methods are developed to monitor consumer purchases over time and to analyze consumer purchases for individual consumers or obtains updated purchase history information that can be aggregated across many consumers and It can provide actionable information to guide consumer behavior and organizational marketing strategies. For example, these improved systems and methods include the ability to extract distinct The new product purchase information allows consumers to organize their previous purchases and better understand their purchasing behavior. It can be used to provide marketing and other organisations with information about their own marketing strategies. to actionable data that can be used to improve campaign accuracy and return on investment It can be organized as follows.
[0024] B. Operating environment examples FIG. 1 illustrates an example of a network communication environment 10 having a network 11. The system 11 includes a purchase transaction data search system 12 and one or more product sellers 1 4, one or more product delivery companies 16 that deliver purchased products to the purchaser, and a message processing service one or more message providers 18 providing product and market information and One or more purchase transaction information users 2 purchasing services from a purchase transaction data retrieval system 12 0 and connect them together.
[0025] Network 11 is a local area network (LAN), a metropolitan area Networks (MAN), and wide area networks (WAN) (e.g., Internet The network 11 may be any of the following networks: A data retrieval system 12, one or more product sellers 14, product delivery companies 16, and a message A variety of transactions are carried out between the service provider 18 and the purchase transaction information user 20. Transmission of various different media types (e.g., text, voice, audio, and video) It has multiple computing platforms and distribution facilities that support A transaction data search system 12, a product seller 14, a product delivery company 16, and a message Each of the providers 18 and purchase transaction information consumers 20 typically comprises a network node (e.g. For example, a client computer or a server system) connected to the network 11 The network node comprises a tangible computer readable memory, a processor, and and input / output (I / O) hardware (which may include a display).
[0026] One or more of the product distributors 14 typically use web browsers and other network devices to Direct your products to the internet using network-enabled software applications At least one of the 14 product sellers allows individuals and companies to purchase products In either case, the goods may be purchased at a retail store. After the purchase transaction is completed, the product seller 14 sends a message associated with the product purchaser. An electronic message may be sent to your email address to confirm your purchase. The message may include, for example, the name of the seller, order number, order date, expected delivery date, product description, product name, This may include product order information such as product quantity, product price, sales tax, shipping fee, and order amount. The product seller 14 may have the product delivered by one of the product delivery companies 16. Depending on the type of product purchased, the product delivery company 16 may The goods may be delivered to the buyer either physically or electronically. In either case, the goods may be delivered by a delivery company16 Alternatively, the product seller 14 may send a shipping notification email to the message address associated with the purchaser. This delivery notification electronic message can be sent, for example, to Document information, billing address, shipping company, shipping address, expected shipping date, expected delivery date, tracking number, etc. This may include product shipping information.
[0027] Generally, the buyer's message address is any address to which electronic messages can be sent. Such a message address can be a network address of the following type: Examples include electronic mail (email) addresses, text message addresses (e.g., telephone sender identifiers, such as numbers or text messaging service user identifiers), social It includes a user identifier for the networking service and a facsimile telephone number. Electronic messages related to are usually associated with the buyer's messaging address. The message is routed to the buyer via each of the message providers 18. Provider 18 typically comprises one or more networks managed by one or more message servers. associated with the buyer's message address in the work data storage system Each message folder stores the buyer's electronic messages.
[0028] The purchase transaction data retrieval system 12 retrieves purchase transaction information from electronic messages of product purchasers. In some examples, the purchase transaction data retrieval system may include a message provider 1 8. Permission to access each message folder of the purchaser of the product managed by the purchaser In another example, the purchaser may obtain the purchase transaction data from the purchase transaction data retrieval system 12. The information is recorded on the user's local communication device (e.g., personal computer or mobile phone). This allows access to stored electronic messages.
[0029] As shown in FIG. 2, the purchase transaction data retrieval system 12 retrieves the purchaser's electronic messages 22. 22 through multiple stages after obtaining permission to access the Generates processed data 24 that is provided to purchase transaction information consumers 20. These stages include a message discovery stage 26, a field extraction stage 28, and a data and a data processing stage 30.
[0030] In the message discovery stage 26, the purchase transaction data retrieval system 12 retrieves the message related to the purchase of the product. Identifying the relevant electronic messages 22. In some examples, rule-based filters and A machine learning classifier is used to identify purchase-related electronic messages.
[0031] In the field extraction stage 28, the purchase transaction data retrieval system 12 extracts the fields from the electronic message. The product purchase information is extracted from the identified items among the items 22. Examples include vendor name, order number, order date, product description, product name, product quantity, product price, consumption Tax, shipping fee, order amount, billing address, shipping company, shipping address, estimated shipping date, estimated delivery date, and and tracking number.
[0032] In the data processing stage 30, the purchase transaction data retrieval system 12 processes various types of purchase The extracted product purchase information is processed according to the purchase transaction information user 20. For example, In the case of a user, the extracted product purchase information may, for example, indicate information about the user's purchases. This information is used to track your order during delivery and to collect your purchase details. and aggregated purchase summary information. The collected product purchase information may be used, for example, to target advertisements to consumers based on their purchasing history. In the case of market analysts, extracted product purchase information For example, anonymous item-level purchase details across retailers, categories, and devices. It is processed so that it can be provided.
[0033] C. Identifying and Filtering Purchase-Related Electronic Messages In an example described in detail below, the purchase transaction information data retrieval system 12 includes a filter learning The filter learning system includes a filter learning system for detecting structural elements of purchase-related electronic messages. Each message template defines a machine-generated electronic message. An electronic message matching system that matches the headers of each pair of similar purchase-related electronic messages, such as a pair of purchase-related messages The filter is automatically learned.
[0034] An example of a confirmation electronic message 32 for a merchandise order is shown in FIG. , a header 34 and a body 35. The header 34 contains the following standard structural elements: Includes "From:", "To:", "Date:" and "Subject:". The header also contains the following structural elements not shown in Figure 3: "Cc:" and "Content-Type", "Precedence:", and "Message- "ID:", "In-Reply-To:", "References:", and "Re ply-To:, Archived-At:, Received:, Return-Path:" The machine-generated structural elements of each message, i.e. the opening “Dear”36 and the standard information-carrying elements 37 (i.e. "Thank you for placing your order ... once your item " has been shipped."), "Order Number:" 38, "Order Summary" 40, and "Pro duct Subtotal: 42, Discounts:, Shipping Charges: 46, and Tax: 48, "Total:" 50, "Part No" 52, "Product Price" 54, and "Discount " 56 , "Part No" 58 , "Product Price" 60 , and "Discount" 62 . The structural elements 34 to 50 are fixed elements, and the set of structural elements 52 to 56 and 58 to 62 are The set contains the same fixed element repeated in each repeating element. The non-structural elements of the product purchase information provider 12 (e.g., price, order number, and part number) are These are the data fields that are extracted and classified by the parser portion of the
[0035] FIG. 4 illustrates, by way of example, a method 66 for automatically creating one or more electronic message filters. In this manner, a computer device may transmit a message between network nodes and one or more media In one or more network data storage systems managed by a message server, and process the set of stored electronic messages associated with each user account. Each electronic message in this collection is associated with a sender, a header, and a body. is.
[0036] In the illustrated example, the computing device generates a graph for each of one or more electronic message senders. The sender is programmed to execute method 4 (Figure 4, block 68). A single electronic messaging address (e.g., sales@store.com) or multiple Sender domains that can be associated with an electronic message address (e.g., * @sto re.com).
[0037] The computer device receives data from one or more of the network data storage systems. The fetched headers are then used to determine whether the particular sender Sometimes associated with the domain, other times fetched independently of the sender domain The computing device groups the fetched headers into clusters (see Figure 1). 4, block 72). In this process, for each sender, each fetched header is Classify messages based on the similarity of headers within a cluster, regardless of the content of the message body. For each of the one or more clusters (FIG. 4, blocks 74, 80), The computer device fetches data from one or more network data storage systems. One of the electronic messages associated with a header that has been assigned to a cluster Each of the above samples is acquired (FIG. 4, block 76). The learning classifier is used to identify the headers and sequences of one or more of the electronic messages found in the sample. Based on the content of the sentence, the clusters are classified as receipt-related or receipt-unrelated. (FIG. 4, block 78). The computer device is Automatically generate an electronic message filter for each of one or more specified clusters. Each electronic message filter checks the subject field string in the header of an electronic message. Each pattern is matched by a rule, and one or more network data sets are In one or more data storage systems managed by the data storage system Crawl the electronic messages stored in association with each user account (Figure 4, Block 82).
[0038] The approach shown in Figure 4 has three main stages: (i) a header is split into a a header structure learning stage for grouping electronic messages having objective elements into clusters; (ii) which header clusters correspond to one or more purchase-related electronic message types; (iii) a filter generation stage. The processing of a complete electronic message (e.g., including headers and body) includes This consumes substantially more resources than processing only the search and classification stages. The number of complete electronic messages processed can be significantly reduced through sampling. Thus, the method shown in Fig. 4 saves processing resources and data storage compared to conventional methods. The message resources, network resources, and fields required to create an electronic message filter. This can significantly reduce the number of filter generation times. In addition, this system uses the header and A complete electronic message (usually corresponding to a machine-generated electronic message such as a receipt) that corresponds to a cluster Users' personal electronic messages are not collected in order to obtain only a small sample of messages. The likelihood of inadvertent capture is low, so user privacy is inherently protected .
[0039] This method is based only on a sample of complete electronic messages associated with each cluster. Although the headers are classified according to the filter, machine-generated electronics are used to generate high-precision filters. It exploits the inherent structural nature of messages. In particular, the header structure learning stage Now, when applied to a machine-generated electronic message, the same message template It is possible to generate dense clusters of possible electronic message headers. As a result, only a few Or just one sample complete electronic message.
[0040] FIG. 5 shows an example of data that may be processed in the various stages of the filter construction method of FIG. In this example, the fetching stage 84 retrieves electronic mail corresponding to a particular sender domain. This involves fetching all 10 million headers in an example collection of messages. The clustering stage 86 splits the 10 million headers into 200 header clusters. The cluster classification stage 88 includes sorting 10 copies of each of the 200 clusters. 2000 complete emails corresponding to a fixed-size sample of electronic messages and detecting the presence of a plurality of electronic messages using a machine learning classifier. The filter generation stage 90 includes classifying the samples according to the purchase-related This involves constructing a filter for each cluster that is classified as a In the example, the number of complete electronic messages that are processed to generate an electronic message filter (i.e., 2000) represents only 0.02 of the total number of electronic messages in this set. As a result, it requires less processing resources, data storage resources, and This significantly reduces the amount of data, network resources, and filter generation times.
[0041] In some instances, these substantial advantages are due, at least in part, to the computing device A specific computer-readable method for identifying headers of purchase-related electronic messages. It results from programming instructions into a computer device. The computer device's ability to identify purchase-related headers is such that the computer device is The data is then classified into dense clusters, and a machine learning classifier is then used to classify the data for each header cluster. Purchase-related header clusters based on a small sample of associated complete electronic messages This is achieved by setting specific instructions in a computer device that identify the
[0042] FIG. 6 is a flow diagram of an example electronic message filter construction process 98 of FIG. According to the method, a computer device transmits one or more messages between network nodes. one or more network data storage systems managed by a message server. and processing a set of stored electronic messages associated with each user account in the Each electronic message in this collection is associated with a sender, a header, and a body. It is being done.
[0043] In this example, the computing device receives electronic messages from one or more electronic message senders. 6. The method of claim 5, wherein the method is programmed to perform one or more elements of the method of FIG. (FIG. 6, block 100). As previously mentioned, the sender may use a single electronic message address. (e.g., sales@store.com) or associated with multiple electronic messaging addresses The sender domain that can be attached (for example, * @store.com) It can be said that:
[0044] According to the method of FIG. 6, a computing device (e.g., a client network node) can be sent from one or more message servers, associated with a sender, and to multiple user accounts. A collection stored in one or more of the network data storage systems across In one example, the computer device may fetch a header in the , fetch all electronic message headers in a collection of electronic messages. The computer device fetches one or more samples of the electronic message headers in the collection. To play.
[0045] Before fetching the header, the computing device typically receives a Indirectly accessing your messaging account through third-party services such as access authorization services The computing device then obtains permission to access the user's Fetches sender-related headers from the message account. In an example, the computing device may receive (e.g., by invoking an electronic message API) ) crawling users' messaging accounts and analyzing the contents of electronic message headers Implement an electronic message crawling engine that scans and evaluates electronic messages. The message crawling engine uses the "From:" and "Subject:" fields. Parse one or both of the fields and apply one or more filters (e.g., regular expression filters) to the parsed results to identify the headers that correspond to the target sender.
[0046] The computer device classifies the clusters into respective dense regions in the clustering data space. Based on a density-based clustering process that associates the fetched headers with In the clustering data space, The fetched headers are then compared against each other based on the similarity between each pair of fetched headers. In general, any density-based clustering process can be used. In some instances, an iterative clustering process, as described below in conjunction with Figures 8 and 9, can be used. In another example, the computer device Density-Based Spatial Clustering (DBSCAN) We used a clustering process called "Actual Clustering of Applications with Noise" to Then, the fetched headers are divided into clusters.
[0047] In some examples, the computer device may further include a processor that performs a step of dividing the headers of the electronic message into clusters before dividing the headers into clusters. In some examples, the computer device performs preprocessing of the subject field in the message header. A sequence of characters (e.g., alphanumeric characters) separated by a blank space Tokenize the text-based content of the subject field in the header by extracting columns A sequence of symbols usually corresponds to a word and a number. The computer device can convert uppercase letters to lowercase letters, remove punctuation marks, and Tokens that match integer and real number patterns in message headers are treated as wildcard tokens. Normalize the content of the subject field by replacing Integers are replaced with the wildcard token "INT" and real numbers with "FLOAT". The subject field is normalized to This increases the ability of the computer device to find purchase-related electronic messages.
[0048] In some examples, the similarity between each pair of fetched headers is calculated based on the headers of the electronic message. Based on a content similarity criterion that compares the degree of similarity and dissimilarity of the content of pairs of strings in the header. In some of these examples, the subject field of each header is In some of these examples, The similarity measure corresponds to the Jaccard similarity coefficient. The similarity coefficient between two headers is calculated as the size of the intersection of the bigrams of both headers. It is a measure based on the size of the union of all the nodes in a set.
[0049] After the headers are divided into clusters, the computer device For each, the following process is performed (FIG. 6, block 108).
[0050] A computing device (e.g., a client network node) may include a message server. One or more of the following are associated with the fetched header in the cluster: electronic messages in a collection stored in one or more of the data storage systems. In some examples, a computer The computer device then sends a predetermined number (e.g., 10, 5, or 1) of electronic copies to each cluster. In another example, the computer device may retrieve a message by, for example, A variable number of electronic messages for each cluster according to statistical criteria that characterize Get the.
[0051] The computing device identifies, via a machine learning classifier, one or more purchase-associated labels and associated signals. Each label in the sample is searched for using each label of a predetermined set of labels including the reliability and the We classify the messages and generate a classification dataset for each cluster (Figure 6, Block Q110).
[0052] In some examples, the machine learning classifier may be a supervised machine learning model (e.g., logistic regression). A bag-of-words model (a log-based regression model or a naive Bayes model) is used to analyze purchase-related electronic messages. Some examples of these are In this paper, a bag-of-words representation is a set of descriptive features that describe a particular purchase-related electronic message. In some examples, each feature may include a string (e.g., a word) and whether the string matches a predefined dictionary. In some cases, the dictionary represents the number of times a word or The message, the sender address (e.g., the text before the "@" sign), and the Includes the words in the body and the number of images in the message body.
[0053] In some examples, a given set of labels may be associated with receiving an electronic message, An example of this type is a label that indicates whether the item is related to the receipt or not. The label set is {"received", "unknown"}. Categorize child messages into multiple purchase-related electronic message categories. An example of this type is The label set as follows: "Cancel", "Back Order", "Coupon", "Promotion", "Unknown" Some or all of these may be included.
[0054] In some examples, the machine learning classifier may use a sampled electronic mail for each cluster. For each message, we provide a respective predicted label selected from a predefined label set, and We then assign a confidence level to each predicted label. The class dataset includes predicted labels and their corresponding samples of electronic messages. The confidence level for each of these associated labels is then calculated.
[0055] The computer system maps each classification data set to a respective cluster label. Based on at least one cluster classification rule that is pinged, each cluster is assigned a predefined label. assign each cluster label selected from the rule set (Figure 6, block 112 In some examples, the cluster classification rules may include determining whether a computer device is Instructs the system to label clusters with a particular label based on a confidence factor. The factors are the number of electronic messages in the corresponding samples that are assigned the same label, and The confidence associated with the assigned label is also used. According to an example of the rule, a particular label is generated for each item that satisfies a confidence threshold (e.g., 98% or more). If a certain confidence level is assigned to all the messages in the sample, the cluster is created. In some cases, the confidence factor of a particular cluster is If a child does not meet the confidence threshold, the electronic message in the cluster is flagged for manual classification. A tag can be added.
[0056] In some instances, the predicted label of a particular electronic message may be below a confidence threshold. If it is determined that the electronic message is a Flagging, in some instances with manually labeled electronic messages ,expanding the training set of machine learning classifiers.
[0057] The computer device generates a purchase-related electricity Automatically generate filters for each child message (Figure 6, Block In some cases, the process of generating filters involves and identifying a common substring in the headers of the and generating a respective filter (e.g., a regular expression filter) based on the A filter typically matches a set of match patterns for each of the subject field strings in an electronic message. In some cases, sequence mining is used to extract n-grams from the headers. (i.e., a sequence of n consecutive items for a given sequence of text. ) to generate filters based on parsing the subject fields of the headers in each cluster. The number of n-grams that appear in the field is calculated, and each filter is assigned a associated with a significant number of occurrences (e.g., the n-gram appears in this header at a high rate) In some of these examples, Sequence mining involves analyzing bigrams in the subject field of the header. An example of how to automatically generate filters from each cluster is shown in Figure 11. This will be explained below.
[0058] After generating the respective filters for each purchase-related cluster, the processor one or more network data storage systems managed by a message server. from a set of stored electronic messages associated with each user account in At least one network communication channel for selecting purchase-related electronic messages One or more of the filters can be installed in the In an example, the computing device crawls a user's messaging account and retrieves electronic messages. Implement an electronic message crawling engine that analyzes and evaluates the content of message headers. In some examples, the electronic message crawling engine may crawl the electronic messages of users. The "From:" and "Subject:" header fields of It parses the field and resolves one or more of the generated filters (e.g., regular expression filters). The results of the analysis are then applied to identify purchase-related headers that correspond to the sender of interest. The electronic message crawling engine extracts the complete electronic message that corresponds to the identified purchase-related headers. Search for child messages. In some examples, each filter searches for one or more electronic messages. Each set of body extraction parsers is associated with a For each filter that matches a respective one of the electronic messages, the computer device uses one or more electronic message body extraction parsers associated with the matched filters An example of a message body extraction parser is the US Patent No. 8,844,010, U.S. Patent No. 9,563,915 and U.S. Patent No. 9,56 This is described in No. 3,904.
[0059] FIG. 7 is an illustration of an example system 118 for constructing message filters. 118 is transmitted between network nodes and communicates with the respective message servers (e.g., Message Provider 1, Message Provider 2, ..., Message Provider M) One or more network data storage systems 122, 124, 1 26, each user account (e.g., alice, bob, clark, dan, eric, peter, and The method processes a set of stored electronic messages associated with each of the plurality of electronic messages (e.g., rob).
[0060] The system 118 may, for each of the one or more senders, generate a plurality of user IDs associated with the sender. 122-126 of network data storage systems across user accounts a sample of each of the headers in a set of electronic messages stored in one or more of the The header sampler 120 fetches all headers associated with the sender. By fetching only a sample of available headers instead of fetching The header sampler 120 utilizes processing resources, data storage resources, and network resources. and reducing the number of generations required to create electronic message filters. In another example, the header sampler 120 may include a header sampler for detecting a sender domain. Fetch a sample of headers spanning
[0061] The preprocessor 128 performs fading of electronic message headers before dividing the headers into clusters. Preprocessing of subject fields in samples retrieved. In some cases, preprocessing The processor 128 performs the pre-processing steps described above with respect to the fetch process described above for the method of FIG. In some examples, the pre-processor 122 also performs one or more of the same Treat all headers with the same subject field content as a single instance. In this way, the preprocessor 128 removes duplication of header data. management resources, data storage resources, network resources, and electronic message facilitation resources. This further reduces the number of iterations required to build a filter.
[0062] The cluster engine 130 then parses the preprocessed headers in the sample by sender domain. In some examples, the grouping is done by dividing the clusters into clusters based on the clustering data. It is based on a density-based clustering process that associates each dense region in the data space with In the clustering data space, the preprocessed header is The distances between the pairs are based on the similarity between them. The similarity between each pair of headers obtained is the similarity of the content in the headers of the electronic messages and The content similarity is calculated based on the criteria of comparing the degree of difference between the two. The header content is the text in the subject field and in the sender message address (for example, the In some of these examples, the headers are The similarity of subject fields is measured using the Jaccard similarity coefficient. The similarity coefficient of both headers is calculated by the size of the intersection of bigrams in the subject field. It is a measure based on the size of the union of all the nodes in a set.
[0063] As shown in Figure 8, the similarity score calculated between each pair of preprocessed headers is defines how close the data are to each other in the clustering data space 132 In some instances, the clustering process identifies elements that are connected in the graph. If their connection similarity scores are greater than a similarity threshold level, There is a relationship between the circular nodes (representing headers). Figure 8 shows the preprocessed 20 headers. The example shows the sample divided into 12 clusters (shown in dashed lines). .
[0064] Figure 9 shows an alternative clustering process with a variable similarity threshold. The authors report that the variability of machine-generated electronic messages generated by multiple senders is The purpose of this method is to iteratively determine an optimal set of relationships between headers that essentially represent
[0065] The clustering process sets the current similarity threshold level to the initial similarity threshold T 0 Set to The method starts by determining whether the similarity threshold is set to 1 (FIG. 9, block 140). In some examples, the initial value of the similarity threshold T 0 From 0 In some of these examples, the similarity threshold is set to an initial level for the similarity measure of 1. Initial value T 0 0.6≦T 0 The range is ≦0.8.
[0066] Next, the header in the sample is updated to the current threshold level T 0 Based on set C 0 Classes in (FIG. 9, block 142), set C 0 Within a set of separate clusters The number of clusters in N 0is stored (FIG. 9, block 144). In some examples, the header The process of classifying messages is based on the text in the subject field of each message header (e.g. For example, it compares the fetched messages based on a comparison of strings, n-grams, and / or words. calculating a similarity score between each pair of dihedrases; and Dividing the fetched headers into clusters based on a comparison of the current threshold level to the Includes:
[0067] The second iteration of the process is performed using another threshold T 1 (Figure 9, Block 1 46, 148, 142, 144). In some cases, for each iteration, the current threshold is The set of thresholds {T i} is a (e.g., mathematical formula or algorithm) It can be dynamically determined (based on In the example, the previous similarity is reduced by a given value (e.g., 0.1 on a similarity scale of 0 to 1). Each successive threshold is determined by decrementing the threshold.
[0068] In the second iteration of the clustering process, the headers in the sample are Value Level T 1 A set of clusters based on C 1 (Figure 9, block 142) , divided cluster C 2 The number of clusters in the set, N 1 is stored (Figure 9, block 1 44). C consists of a header containing a unique subject line. 1 Clusters in the C consists of a header with a unique subject line. 1 The number of all identified headers in M 1 but In some instances, the subject line may be different from other headers in the cluster (FIG. 9, block 150). A header is said to have a unique subject if it does not contain any words in common with the subject of the In another example, a unique subject field line is used to determine the header subject of a cluster. Subject field content, such as comparisons between strings or n-grams within a name field line Identified based on other text-based comparisons.
[0069] The number of headers in a cluster with unique subject lines, M j is the threshold M TH twist If the number of unique subjects is too large (Figure 9, block 152), the number of unique subjects is deemed to be too large and multiple The clusters in the leading set of clusters (i.e., C i-1 ) in the header The output cluster set 160 for the current sample is returned by the cluster engine 130. (FIG. 9, block 156). Also, the number of clusters in the current iteration, N i and the preceding anti- Number of clusters in the loop N i-1 If satisfies the similarity criterion (FIG. 9, block 152), The number of clusters is considered to have converged, and the number of clusters in the previous set is (i.e., C i-1 ) as output cluster set 160 for the current sample in the header , are returned by cluster engine 130 (FIG. 9, block 156).
[0070] In some examples, the similarity criterion is the number of clusters between the current cluster and the preceding cluster. The ratio of the differences is compared to the number of clusters in the preceding iteration. The similarity criterion corresponds to the following formula:
number
[0071] If neither the tests in blocks 152 nor 154 are satisfied, Another iteration of the clustering process is repeated using the next clustering threshold (Figure 9 , block 148).
[0072] Referring again to FIG. 7, the cluster engine 130 may assign the header samples to cluster 16. After sorting into 0, the electronic message sampler 162 sorts the headers in each header cluster 160 into Select each sample and from the message provider, select the The result is that each block of the header 160 is , i, respectively, of the electronic message 164 for raster i.
[0073] For each sample i of the electronic message 164, the electronic message classifier 166 In some instances, the electronic message classification may include: The classifier 166 is a machine learning classifier of the type described above with respect to Figures 4 and 6. The classifier generates a set of labels and associated confidences 168 for each sample. In some instances, a prediction of a particular electronic message may be assigned to the electronic message 164. In response to determining that the assigned label is below the confidence threshold, the computing device may Certain electronic messages are flagged for classification (FIG. 7, block 170). In one example, manually labeled electronic messages were used to generate an electronic message classifier. Let 166 learn.
[0074] In some examples, the cluster classification rules may be used to assign the same label to computer devices. The number of electronic messages in the corresponding sample that were assigned and the labels assigned to them Clustering with a particular label based on one or more confidence factors, such as an associated confidence level In some examples, the confidence factor may be one or more confidence thresholds. If not (FIG. 7, block 172), the electronic messages in the cluster are manually classified. The file is then flagged as such (FIG. 7, block 170).
[0075] For each cluster that is assigned a respective purchase-related label (Figure 7, block 17 4) a computer device for filtering purchase-related electronic messages; The filter 175 is automatically generated (FIG. 7, block 176).
[0076] In the illustrated example, if a cluster is not assigned each purchase-related label (Figure 7, block 174), the computer device sends an electronic message regarding the next cluster of the header 160. 7. Proceed directly to processing the next sample i=i+1 of step 164 (FIG. 7, block 177). In this process, the computer device calculates the electronic structure of the components in the next sample i=i+1. Repeat the cluster labeling process based on the classification of messages (Figure 7, block 1 66~172).
[0077] In an alternative example, the next sample i=i+1 of the electronic message 164 may be processed. 7, block 177), the computing device receives the purchase-related label. For each unassigned cluster, filter non-purchase related electronic messages. The filter 179 for each of these filters is automatically generated (FIG. 7, block 178). In some of the examples, the non-purchase-related electronic message filter 179 As a component of the prefilter 120 (shown in FIG. 7) or as part of a separate prefilter. It is installed at the front end of the message filter construction system 118. The non-purchase related electronic message filter 179 is used to filter the header sampler 120 The headers of non-purchase related electronic messages fetched by the and then sending an electronic message having a header corresponding to the non-purchase-related electronic message identified thus far. Removes child messages and reduces processing resources, data storage resources, and network resources. Further reducing the number of generation times required to build a service and purchase-related electronic message filter It is possible.
[0078] FIG. 10 incorporates elements of the message filter construction system 118 of FIG. Implement an iterative process to build a filter from a sample of the fetched header data. 1 shows an example of a message filter construction system 180.
[0079] In this example, the header sampler 120 generates, for each of one or more senders, a header sampler 120 associated with the sender. 1 22 to 126, the headers in a set of electronic messages stored in one or more of In another example, the header sampler 120 may fetch each sample of The preprocessor 128 fetches a sample of the header across the before the subject field in the fetched sample of electronic message headers. The cluster engine 130 processes the clusters by dividing them into their sub-clusters in the clustering data space. The preprocessing is based on a density-based clustering process that associates each dense region with The preprocessed headers are divided into clusters 160. In the clustering data space, The headers are separated from each other based on the similarity between each pair of preprocessed headers. For each set i of header clusters 160 that are assigned purchase-related labels, The computer device filters the purchase-related electronic messages according to the method described above with respect to FIG. For each set of filters i to be filtered, we automatically generate (See Q176.)
[0080] In the first iteration of the filter construction process, the message filter construction system 1 80 calculates the number of samples associated with each sender from the first sample of each of the headers associated with the sender. Construct the first set of filters (i.e., {filter set i}) for each filter .
[0081] This process includes a second iteration of each of the headers in the collection of electronic messages for each sender. In this second iteration of the filter construction process, the message The filter construction system 180 generates a second substring for each of the headers associated with the sender. From the sample, the second set of filters for each sender (i.e., {filters Then, construct a set {i+1} of
[0082] The filter results are compared on a sender-by-sender basis (Figure 10, block 182). This process In the method, for each sender, the computer device 10, block 183. The first and second sets of filters (i.e., {filter set i} and {Filter result i} and {Filter result i+1}) are then 1} for each set of all headers in the set corresponding to the sender. It is used.
[0083] If the filtering results are similar (FIG. 10, block 184), the filter construction process The process ends (FIG. 10, block 186). In some examples, the first set of filters and does not match any filter in the second set, The first filtering result and the second filtering result are based on a comparison of the number of headers that are filtered. The similarity between the header and the first and second sets of filters is calculated. If the numbers of are similar, the set of filters is considered to be similar enough, and the filter construction The process then ends (Figure 10, block 186).
[0084] If the filter results are not similar (FIG. 10, block 184), the filter construction process In some cases, the number of filters shared between the compared filter sets may be Filters are stored in memory 188 for use in filtering electronic messages ( 10, block 190). Increase the previous header sample size (Figure 10, block 19 2) Construct a filter using a larger respective sample of headers for each sender. Another iteration of the process is performed (Figure 10, block 194).
[0085] FIG. 11 shows the process of generating filters for header clusters 198. In one example, the computing device may identify each purchase-related cluster (e.g., by a purchase-related label). This process is performed for each labeled cluster (Figure 11, block 200 The computer system determines whether each bigram occurs within the subject line of all headers in cluster 198. The computer device counts how many times the cluster 1 appears (FIG. 11, block 202). 98 (FIG. 11, block 204). For each bigram in the field, the computer device determines whether the bigram corresponds to a subject field in the header. Then, a measure of the frequency of occurrence of each of the fields is determined (FIG. 11, block 206). An example of such a frequency criterion is the number of subject fields that contain bigrams, The ratio of the number of subject fields that contain bigrams divided by the number of subject fields that do not contain bigrams is The percentage of subject fields that contain bigrams divided by the total number of subject fields. The computer device detects whether a threshold (e.g., 80%) is satisfied and the ratio of the frequency of the occurrence of the threshold in the field. Combine each of the bigrams in the selected header associated with each frequency criterion. Then, the filter for the cluster is constructed by adding the The computer system calculates the set of bigrams that are incorporated into the cluster filter. In some examples, the bigram is A set is considered to have converged if it did not change in the last iteration. If the set of bigrams has converged (FIG. 11, block 210), the bigrams in the set are In some cases, the bigram filter is converted to a bigram filter (FIG. 11, block 212). The bigrams are converted into one or more regular expressions that define the filters. If not (FIG. 11, block 210), the process selects a The process is then repeated for the other headers that were retrieved (FIG. 11, blocks 204-210).
[0086] Another example of the filter construction process in Figure 11 uses the header subject instead of bigrams. Field analysis includes strings, n-grams, and words that appear within the subject field, among others. This is done on the text features.
[0087] [III. Examples of Computer Devices] The computer apparatus includes an improved method for carrying out the functions of the processes described herein. In some examples, the electronic message processing system is programmed to The process of constructing a message filter and filtering electronic messages with one or more electronic message filters. The process of filtering messages is performed by a separate and distinct computing device. In other instances, the same computing device executes these processes.
[0088] FIG. 12 illustrates an example of a computer device implemented by computer system 320. The computer system 320 includes a processing unit 322 and a system memory 324. A memory 324 and a processor 325 connect the processing unit 322 to the various elements of the computer system 320. The processing unit 322 includes one or more data processors. Each of these data processors may be implemented using a variety of commercially available computer The system memory 324 may be in the form of any one of a number of processors. A software application that typically defines the addresses available to the software application. The present invention includes one or more computer readable media associated with an application addressing space. The system memory 324 includes basic input / output (I / O) memory, including the start-up routines for the computer system 320. The read-only memory (ROM) stores the system (BIOS) and the random access memory (RAM). The system bus 326 may include a memory bus, a peripheral bus, and or local bus, PCI, VESA, Microchannel ( Any of a variety of bus protocols including Microchannel, ISA, and EISA The computer system 320 may be compatible with a persistent storage memory 3 28 (e.g., hard drives, floppy drives, CD-ROM drives, magnetic tapes It also includes a hard disk drive, flash memory device, and digital video disk. The persistent storage memory is connected to the system bus 326 and stores data, data structures, and computer One or more computer-implemented devices that provide non-volatile or persistent storage of computer-executable instructions. This includes computer readable medium disks.
[0089] The user may input data using one or more input devices 330 (e.g., one or more keyboards, Mouse, microphone, camera, joystick, physical motion sensors, and touch pad) to interact with the computer system 320 (e.g., to send commands or data). The information is displayed on a display controlled by a display controller 334. Through a graphical user interface (GUI) presented to the user on monitor 332 The computer system 320 can be presented in a variety of ways, including other input / output hardware ( Peripheral output devices, such as speakers and a printer, may also be included. The computer system 320 includes a network adapter 336 ("Network Interface Card"). It connects to other network nodes through a single interface (also called a "interface card" or NIC).
[0090] A number of program modules may be stored in the system memory 324. These modules provide an application programming interface (API), Operating System (OS) 340 (e.g. Microsoft Corporation Windows operating system available from Microsoft Corporation, Redmond, New York. a process for constructing an electronic message filter and a method for constructing an electronic message filter; to execute one or more of the processes for filtering electronic messages in One or more software applications for programming the computer system 320 Software applications 341, drivers 342 (e.g., GUI drivers ), network transport protocol 344, and data 346 (e.g., input data data, output data, program data, registry, and configuration settings).
[0091] The disclosed subject matter, including systems, methods, processes, functional operations, and logical flows, is Examples of subject matter described in the above include devices that perform functions by manipulating input and generating output. A data processing device (e.g. computer hardware and digital Examples of the subject matter described herein may also be implemented in a data One or more tangible, non-transitory carrier media (e.g., machine readable) for execution by a data processor. A read-only storage device, substrate, or sequential access memory device ( Software or firmware as one or more sets of computer instructions It can be embodied as something tangible.
[0092] The details of the specific implementations described herein may be adapted to specific embodiments of particular inventions. may be specific and shall not be considered as limiting the scope of any claimed invention. For example, features described in the context of separate embodiments should not be construed as limiting the scope of the present invention. and features described in the context of a single embodiment may be combined into multiple separate implementations. Further, the invention may be implemented in a variety of forms, including within a particular order of steps, tasks, or Disclosure of an operation or process does not necessarily mean that a step, task, operation, or process is It is not required that the steps be performed in a particular order; rather, in some cases, One or more of the steps, tasks, actions, and processes illustrated may be performed in a different order or multiple times. The processes can be executed according to multiple task schedules or in parallel.
[0093] [IV. Conclusion] According to the embodiments described herein, a purchase-related electronic message filter is configured. SYSTEMS, METHODS, AND COMPUTER SYSTEMS FOR BUILDING AND FILTERING PURCHASE-RELATED ELECTRONIC MESSAGES - Patent application A computer readable medium is provided.
[0094] Other embodiments are within the scope of the claims.
Claims
1. An interface circuit; machine-readable instructions; one or more processor circuits; Equipped with The one or more processor circuits execute the machine-readable instructions to grouping a plurality of first electronic message headers into a plurality of clusters based on similarity, the plurality of clusters being generated without regard to content of body text corresponding to the plurality of first electronic message headers; obtaining, for each of a plurality of said clusters, a sample of an entire electronic message associated with each of said plurality of first electronic message headers, said entire electronic message having a header and a body; classifying each of the samples for each of a plurality of the clusters to identify a first set of the plurality of the clusters, the first set of the plurality of clusters being classified by a machine learning classifier as being associated with a purchase transaction; reducing computational resource consumption by generating a respective filter for each cluster in the first set of clusters, each of the filters for identifying electronic messages associated with a purchase transaction; To carry out Device.
2. The apparatus of claim 1 , wherein each of the plurality of first electronic message headers is associated with a source, the source being at least one of an email address and a source domain.
3. 2. The apparatus of claim 1, wherein one or more of the one or more processor circuits execute the machine-readable instructions to group each of the plurality of first electronic message headers into a plurality of the clusters based on a Jaccard similarity coefficient.
4. one or more of the one or more processor circuits execute the machine-readable instructions to further classify the first set of the clusters into a plurality of subcategories; the plurality of subcategories are based on a type of electronic message associated with a purchase transaction; the plurality of subcategories include one or more of: order notification, shipping notification, refund, cancellation, back order, coupon, promotion, and unknown; 2. The apparatus of claim 1.
5. One or more of the one or more processor circuits execute the machine-readable instructions to Parsing the plurality of second electronic message headers; applying each of the filters for each of the clusters to the plurality of second electronic message headers to identify those of the plurality of second electronic message headers that are purchase-related electronic messages; obtaining a plurality of entire electronic messages corresponding to each of the plurality of second electronic message headers; 2. The apparatus of claim 1.
6. 6. The apparatus of claim 5, wherein one or more of the one or more processor circuits execute the machine-readable instructions to extract purchasing information from an entire electronic message corresponding to each of a plurality of the second electronic message headers.
7. 7. The device of claim 6, wherein the purchase information includes one or more of a merchant name, an order number, an order date, a product description, a product name, a quantity of the product, a price of the product, a sales tax, a shipping charge, an order amount, a billing address, a shipping company, a shipping address, an expected shipping date, an expected delivery date, and a tracking number.
8. at least, grouping a plurality of first electronic message headers into a plurality of clusters based on similarity, the plurality of clusters being generated without regard to content of body text corresponding to the plurality of first electronic message headers; obtaining, for each of a plurality of said clusters, a sample of an entire electronic message associated with each of said plurality of first electronic message headers, said entire electronic message having a header and a body; classifying each of the samples for each of a plurality of the clusters to identify a first set of the plurality of the clusters, the first set of the plurality of clusters being classified by a machine learning classifier as being associated with a purchase transaction; reducing computational resource consumption by generating a respective filter for each cluster in the first set of clusters, each of the filters for identifying electronic messages associated with a purchase transaction; At least one machine-readable medium having machine-readable instructions for causing at least one processor circuit to execute the method.
9. 9. The at least one machine-readable medium of claim 8, wherein each of the plurality of first electronic message headers is associated with a source, the source being at least one of an email address and a source domain.
10. 10. At least one machine-readable medium as recited in claim 8, wherein the machine-readable instructions cause one or more of the at least one processor circuit to perform the step of grouping each of the plurality of first electronic message headers into a plurality of the clusters based on a Jaccard similarity coefficient.
11. The machine-readable instructions cause one or more of the at least one processor circuit to further classify the first set of clusters into a plurality of subcategories; the plurality of subcategories are based on a type of electronic message associated with a purchase transaction; the plurality of subcategories include one or more of: order notification, shipping notification, refund, cancellation, back order, coupon, promotion, and unknown; 9. At least one machine-readable medium as recited in claim 8.
12. The machine-readable instructions may be for one or more of the at least one processor circuit to: parsing a plurality of second electronic message headers; applying each of the filters for each of the clusters to the plurality of second electronic message headers to identify those of the plurality of second electronic message headers that are purchase-related electronic messages; obtaining a plurality of entire electronic messages corresponding to each of the plurality of second electronic message headers; Execute the 9. At least one machine-readable medium as recited in claim 8.
13. 13. At least one machine-readable medium according to claim 12, wherein the machine-readable instructions cause one or more of the at least one processor circuits to perform a step of extracting purchasing information from an entire electronic message corresponding to each of a plurality of the second electronic message headers.
14. 14. The at least one machine-readable medium of claim 13, wherein the purchase information includes one or more of a merchant name, an order number, an order date, a product description, a product name, a quantity of the product, a price of the product, sales tax, a shipping charge, an order amount, a billing address, a shipping carrier, a shipping address, an expected shipping date, an expected delivery date, and a tracking number.
15. at least one processor circuit grouping a plurality of first electronic message headers into a plurality of clusters based on similarity, the plurality of clusters being generated independent of body content corresponding to the plurality of first electronic message headers; at least one processor circuit obtaining, for each of a plurality of said clusters, a sample of an entire electronic message associated with each of said plurality of first electronic message headers, said entire electronic message having a header and a body; classifying, by at least one processor circuit, each of the samples for each of a plurality of the clusters to identify a first set of the plurality of the clusters, the first set of the plurality of clusters being classified by a machine learning classifier as being associated with a purchase transaction; reducing computational resource consumption by generating, by at least one processor circuit, a respective filter for each cluster in the first set of clusters, each of the filters for identifying electronic messages associated with a purchase transaction; The method includes:
16. 16. The method of claim 15, wherein each of the plurality of first electronic message headers is associated with a source, the source being at least one of an email address and a source domain.
17. 16. The method of claim 15, wherein the grouping of each of the plurality of first electronic message headers into a plurality of the clusters is based on a Jaccard similarity coefficient.
18. further classifying the first set of clusters into a number of subcategories; the plurality of subcategories are based on a type of electronic message associated with a purchase transaction; the plurality of subcategories include one or more of: order notification, shipping notification, refund, cancellation, back order, coupon, promotion, and unknown; The method of claim 15.
19. parsing a plurality of second electronic message headers; applying each of the filters for each of the clusters to the plurality of second electronic message headers to identify those of the plurality of second electronic message headers that are purchase-related electronic messages; obtaining a plurality of entire electronic messages corresponding to each of the plurality of second electronic message headers; 16. The method of claim 15 further comprising:
20. 20. The method of claim 19, further comprising extracting purchase information from an entire electronic message corresponding to each of a plurality of the second electronic message headers.
21. 21. The method of claim 20, wherein the purchase information includes one or more of a merchant name, an order number, an order date, a product description, a product name, a quantity of the product, a price of the product, a sales tax, a shipping charge, an order amount, a billing address, a shipping carrier, a shipping address, an expected shipping date, an expected delivery date, and a tracking number.
Citation Information
Patent Citations
Free text and attribute searching of electronic program guide (EPG) data
JP2004289848A
Device, method and program for determining junk mail
JP2011034417A
Illegal mail determination device, illegal mail determination method, and program
JP2014102708A
Method, device and system for determining mail class
US20090019171A1
Extracting product purchase information from electronic messages
US20160110763A1