Data mining system and method for mail data
By using an email data mining system and pre-trained language models and the BM25 algorithm, the problem of unified processing of email data from multiple channels was solved, achieving efficient semantic analysis and accurate retrieval, and improving the accuracy and efficiency of email data processing.
Patent Information
- Application Number
- CN202511232222.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies lack multi-source integration capabilities, making it difficult to simultaneously support the acquisition of email data from multiple channels such as mail servers, local clients, and real-time streams. The data cleaning, integration, and transformation processes lack standardized procedures, resulting in low processing efficiency and unstable quality. Keyword-based retrieval methods cannot understand semantics, which can easily lead to missed detection of relevant emails or false detection of irrelevant emails.
This paper provides an email data mining system, which includes a data storage management module, a semantic retrieval and analysis module, a comprehensive mining and processing module, and an email data acquisition module. It generates semantic vectors through a pre-trained language model and combines the BM25 algorithm and cosine similarity calculation to achieve unified processing and accurate retrieval of multi-source data.
It improves data relevance and retrieval accuracy. By unifying multi-source data formats, constructing a computable semantic space, and dynamically adjusting keyword weights and similarity thresholds, it achieves adaptive optimization of the system, thereby improving the accuracy and efficiency of email data processing.
Smart Images

Figure CN121070985A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data retrieval processing, in particular to a mail data data mining system and method. BACKGROUND
[0002] With the deepening of enterprise digital transformation, mail has become the core carrier of daily office and business communication, and has accumulated a large amount of mail information containing structured and unstructured data.
[0003] The traditional mail management system usually only supports simple keyword-based retrieval, which is difficult to meet the needs of enterprises for deep mining and intelligent analysis of mail content. At the same time, in the prior art, the processing of mail data has the following limitations:
[0004] Lack of multi-source integration capability, difficult to support mail data acquisition from mail servers, local clients and real-time streams at the same time, lack of standardized process in data cleaning, integration and conversion process, resulting in low processing efficiency and unstable quality, keyword-based retrieval method cannot understand semantics, easy to cause related mail missed detection or irrelevant mail false detection. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a mail data data mining system and method, which solves the problem of inaccurate retrieval matching caused by single semantic analysis retrieval.
[0006] To achieve the above purpose, the present application realizes the following technical scheme: a mail data data mining system, comprising:
[0007] A data storage management module is used for cleaning, integrating and converting the mail data transmitted by the mail data acquisition module to obtain preprocessed data, and storing the structured data and unstructured data in the preprocessed data to generate storage management information and transmit it to the semantic retrieval analysis module;
[0008] A semantic retrieval analysis module is used for obtaining input information of a user, generating a semantic vector through a pre-trained language model, and matching a database to obtain a preliminary screening result, segmenting the input information, segmenting according to common words to obtain a plurality of segmented words, matching the segmented words with the preliminary screening result to determine the keywords corresponding to the input information, and comprehensively matching the keywords and the semantic vector with the preliminary screening result to obtain a preselected result, and transmitting it to the comprehensive mining processing module;
[0009] The comprehensive mining processing module is configured to optimize the preselected result, calculate the accuracy and recall rate of the preselected result, analyze the accuracy and recall rate, and display the preselected result to the corresponding operator if the accuracy and recall rate meet the requirements, or generate a specific analysis signal if any of the accuracy and recall rate is incorrect, process the specific analysis signal, take different optimization measures according to different situations, generate corresponding optimization processing information, and transmit the optimization processing information to the data management information output module.
[0010] As a further scheme of the present application, the mail data acquisition module and the data management information output module are further included.
[0011] The mail data acquisition module is configured to acquire mail data, wherein the mail data includes structured data and unstructured data, and the acquisition mode includes grabbing from a mail server, extracting from a local client, and real-time stream acquisition.
[0012] The data management information output module is configured to display the preselected result and the optimization processing information to the corresponding operator.
[0013] As a further scheme of the present application, the specific manner in which the data storage management module generates storage management information is as follows:
[0014] For the structured data in the preprocessed data, a relational database is used to store mailbox addresses, timestamps, and department structured fields, and supports multi-condition queries; for the unstructured data, a document database is used to store original texts, and corresponding storage management information is generated and transmitted to the semantic retrieval analysis module.
[0015] As a further scheme of the present application, the specific manner in which the semantic retrieval analysis module obtains multiple segmented words is as follows:
[0016] The input information of the user is acquired, and retrieval analysis is performed in combination with the corresponding storage management information; a pre-trained language model is used to identify the input information, and a corresponding semantic vector is obtained; the obtained semantic information is matched with the established database, and a preliminary screening result is obtained; the input information is segmented, and multiple segmented words are obtained according to common words.
[0017] As a further scheme of the present application, the specific manner in which the semantic retrieval analysis module determines the keywords corresponding to the input information is as follows:
[0018] The segmented words are acquired and screened to obtain pre-selected segmented words, and are marked as i, and i=1, 2, …, j, wherein j represents the type of the pre-selected segmented words, and a preliminary screening result is acquired and is marked as n, and n=1, 2, …, m, wherein m represents the number of the preliminary screening results, then the pre-selected segmented words i are sequentially matched with the preliminary screening results n in the order of the pre-selected segmented words i, and the corresponding matching times are acquired, and the matching proportion values corresponding to the pre-selected segmented words i are calculated, and the pre-selected segmented words with the matching proportion values greater than a preset proportion value are screened and are marked as keywords.
[0019] As a further scheme of the present application, the semantic retrieval analysis module performs secondary matching screening on the preliminary screening results according to the keywords and the semantic vectors, and the specific manner of obtaining the pre-selected results is as follows:
[0020] The scores of the keywords and the preliminary screening results are calculated, the keywords are acquired and are marked as o, and o=1, 2, …, p, wherein p represents the type of the keywords, and the scores of the keywords o and the preliminary screening results are calculated according to the formula The BM25 score corresponding to the keyword o is calculated, wherein n is the preliminary screening result, f(q o , n) is the frequency of the keyword o in the preliminary screening result n, IDF(q o ) is the inverse document frequency of the keyword o, |n| is the length of the preliminary screening result, avgdl is the average length of all the preliminary screening results, and k1 and b are adjustment parameters.
[0021] Then the similarity value of the semantic vector and the preliminary screening result n is calculated by using the cosine similarity, and the BM25 score and the similarity value are weighted and summed, and the comprehensive index Z of the keyword o and the preliminary screening result is calculated according to the formula Z=score BM25 ×α+score 相似度值 ×(1-α), wherein α is the weight coefficient of the BM25 score, score BM25 is the BM25 score, score 相似度值 is the similarity value, and the preliminary screening results with the comprehensive index greater than an index threshold value are screened and are marked as pre-selected results.
[0022] As a further scheme of the present application, the specific manner in which the comprehensive mining processing module optimizes the obtained pre-selected results is as follows:
[0023] All the pre-selected results are acquired, and the obtained pre-selected results are classified, and are classified into true positive results TP and false positive results FP, and the number of the TP results and the FP results is acquired, and then the formula The corresponding precision Precision is calculated, and all preliminary screening results are obtained, and the corresponding false negative results FN are obtained, which specifically represent the relevant documents that are not retrieved, and the formula The corresponding recall Recal l is calculated, and the precision and the recall are analyzed.
[0024] As a further scheme of the present application, the specific way of analyzing the comprehensive precision and recall is:
[0025] The obtained precision Precision and recall Recal l are respectively compared with the corresponding comparison threshold, if the precision Precision is greater than the precision comparison threshold, it indicates that the corresponding preselected result is accurate, otherwise it indicates that the preselected result is inaccurate;
[0026] If the recall Recal l is greater than the recall comparison threshold, it indicates that the preselected result is accurate, otherwise it indicates that the preselected result is inaccurate;
[0027] If the precision Precision and the recall Recal l are both correct, it indicates that the preselected result is correct, and the preselected result is displayed to the corresponding operator, otherwise if any one of the precision Precision and the recall Recal l is incorrect, it indicates that the preselected result is abnormal, and a specific analysis signal is generated and analyzed.
[0028] As a further scheme of the present application, the specific way of generating a specific analysis signal and analyzing it is:
[0029] The obtained specific analysis signal is processed, if the precision is low and the recall is high, the keyword filtering is strengthened and the vector similarity threshold is increased, if the precision is high and the recall is low, the vector retrieval threshold is reduced and the keyword synonym is expanded, if both are low, the vector model is retrained, the mixed weight is optimized, and the corresponding optimization processing information is generated and transmitted to the data management information output module.
[0030] A data mining method of mail data, which specifically comprises the following steps:
[0031] Step 1, collect mail data, and perform data cleaning, integration and conversion to obtain preprocessed data, and store structured data and unstructured data to generate storage management information;
[0032] Step 2, identify the input information and match it with the database to obtain the preliminary screening result, and segment the input information into common words to obtain preselected segmented words, and calculate the matching proportion to screen the keywords;
[0033] Step three, calculate the score of the keyword, semantic vector and preliminary screening result, and weighted sum to get the comprehensive index, and get the pre-selection result based on the comprehensive index;
[0034] Step four, classify the TP and FP results of the pre-selection result, and calculate the corresponding accuracy and recall rate, and judge whether the pre-selection result is accurate or not, if not, adjust based on the accuracy and recall rate to generate optimization processing information.
[0035] The present application provides a kind of data mining system and method of mail data.Compared with prior art, it has the following beneficial effects:
[0036] The present application unifies multi-source data format, combines enterprise address book supplementary metadata, improves data correlation, standardizes structured data and vectorizes unstructured text, constructs a computable semantic space, filters keywords by matching proportion, calculates word frequency weight by combining BM25 algorithm, improves keyword matching accuracy, generates semantic vector using pre-trained language model, calculates semantic relevance by cosine similarity, dynamically fuses keyword matching score and semantic similarity score, balances accuracy and comprehensiveness, calculates accuracy and recall rate in real time, automatically identifies search bias type, adjusts keyword weight, similarity threshold or expands synonym table, and realizes adaptive optimization of system. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 The present application is a system principle block diagram;
[0038] Figure 2 The present application is a step method diagram. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0040] Embodiment one
[0041] Please refer to Figure 1 The present application provides a kind of data mining system of mail data, including: mail data acquisition module, data storage management module, semantic retrieval analysis module, comprehensive mining processing module and data management information output module, and combined with the drawings Figure 1 It can be known that the information between the above functional modules is one-way transmission.
[0042] The mail data collection module is configured to collect mail data, which includes structured data and unstructured data, and the specific collection methods include mail server crawling, local client extraction, and real-time stream collection, and the collected mail data is transmitted to the data storage management module.
[0043] Specifically, the mail server crawling is performed by connecting enterprise mail servers (such as Outlook and Gmail servers) through IMAP, POP3 protocols or vendor APIs (such as Exchange Web Services) to batch obtain data.
[0044] The local client extraction is performed by parsing the storage files (such as PST and MBOX formats) of local mail clients (such as Thunderbird and Foxmail).
[0045] The real-time stream collection is performed by capturing real-time generated mails (such as real-time sent and received mails) through port listening or message queues (such as Kafka).
[0046] The data storage management module is configured to preprocess the obtained mail data, and the preprocessing operations include data cleaning, data integration and data conversion to obtain corresponding preprocessed data, and the specific processing methods are as follows:
[0047] The data cleaning includes deduplication (deleting duplicate mails based on mail ID and content hash value), noise removal (filtering garbled codes by encoding conversion such as UTF-8 correction, invalid fields such as empty subject and mail without content), and preliminary filtering of spam mails (obviously spam mails are removed by rules (such as containing “unsubscribe” and “promotion” high-frequency words) or simple models (such as based on a sender blacklist)), the data integration includes merging multi-source data (standardizing data from different mail servers and clients (such as unifying mailbox format and timestamp time zone)), and associating and supplementing (combining enterprise address book and department information to match real name, department and other metadata for mailbox address), the data conversion includes structured data standardization (such as unifying timestamp to “YYYY-MM-DDHH:MM:SS”, extracting domain name from mailbox address (judging internal and external mailbox)), and unstructured text conversion (segmenting mail subject and body (Chinese using Jieba, English using NLTK), removing stop words (such as “of” and “the”), word vector conversion (such as Word2Vec and BERT embedding), and converting text to computable numerical features);
[0048] According to the obtained preprocessed data, corresponding data storage is performed, specifically, for the structured data in the preprocessed data, a relational database (such as MySQL, PostgreSQL) is used to store the structured fields such as mailbox address, timestamp, department, and support multi-condition query (such as “filtering mails from external mailbox in 2023”), for unstructured data, a document database (such as MongoDB) is used to store the original text and segmentation results, or a search engine (such as Elasticsearch) is used to implement full-text retrieval (such as quickly locating mails containing the keyword “contract”), and corresponding storage management information is generated, which is transmitted to the semantic retrieval analysis module.
[0049] The semantic retrieval analysis module is used to obtain the input information of the user, and performs retrieval analysis in combination with the corresponding storage management information, identifies the input information by using a pre-trained language model, and obtains a corresponding semantic vector, matches the obtained semantic information with the established database, and the matching is performed by a semantic analysis model, which is not described in detail herein, and a preliminary screening result is obtained by screening;
[0050] The input information is segmented and processed, specifically segmented according to commonly used words, to obtain a plurality of segmented words, and the segmented words are matched with the preliminary screening result to determine the corresponding keywords of the input information, and the specific matching method is as follows:
[0051] The segmented words are obtained and screened to obtain preselected segmented words, and the stop words (such as “de” and “in” without meaning) and low-frequency noise words (such as rare words with an occurrence frequency of ≤1) are removed, and are marked as i, and i = 1, 2, …, j, where j represents the type of preselected segmented words, and the preliminary screening result is obtained and marked as n, and n = 1, 2, …, m, where m represents the number of preliminary screening results, then according to the order of the label of the preselected segmented words i, the preliminary screening results n are matched and analyzed in turn, and the corresponding matching times are obtained, and the matching times represent the number of preliminary screening results matched with the corresponding preselected segmented words, for example, the number of preliminary screening results is 20 groups, and one group of preselected segmented words matches with 12 groups of preliminary screening results, then the corresponding matching times are 12, and the matching proportion value corresponding to the preselected segmented words i is calculated, and the preselected segmented words with a matching proportion value greater than a preset proportion value are screened and marked as keywords according to the order from large to small;
[0052] Then, the preliminary screening results are twice matched and screened by combining the keywords and the semantic vector to obtain preselected results, and the specific matching and screening method is as follows:
[0053] The score of the keyword and the preliminary screening result is calculated, the keyword is obtained and labeled as o, and o = 1, 2, …, p, wherein p represents the type of the keyword, and the BM25 score corresponding to the keyword o is calculated according to the formula , wherein n is the preliminary screening result, f(q o , n) is the frequency of the keyword o in the preliminary screening result n, IDF(q o ) is the inverse document frequency of the keyword o, |n| is the length of the preliminary screening result, avgdl is the average length of all preliminary screening results, and k1 and b are adjustment parameters.
[0054] The specific calculation method of IDF(q o ) is, The inverse document frequency is calculated according to the above formula, wherein h(q o ) is the number of preliminary screening results containing the keyword o.
[0055] Then the similarity value of the semantic vector and the preliminary screening result n is calculated by the cosine similarity, and the BM25 score and the similarity value are weighted and summed, and the comprehensive index Z of the keyword o and the preliminary screening result is calculated according to the formula Z = score BM25 × α + score 相似度值 × (1-α), wherein α is the weight coefficient of the BM25 score, score BM25 is the BM25 score, score 相似度值 is the similarity value, and the comprehensive index is sorted from large to small, then the preliminary screening result whose comprehensive index is greater than the index threshold is screened and labeled as a preselected result, and the preselected result is transmitted to the comprehensive mining processing module.
[0056] The comprehensive mining processing module is used for optimizing the obtained preselected result, and the accuracy and recall rate of the preliminary screening result with respect to the preselected result are calculated respectively, and the two are analyzed comprehensively, and the specific processing method is as follows:
[0057] All preselected results are obtained, and the obtained preselected results are classified, which are classified into TP (true positive) results and FP (false positive) results, wherein the TP (true positive) result represents a document that is retrieved and relevant (correct hit), the FP (false positive) result represents a document that is retrieved but not relevant (misjudged as relevant), the number of TP (true positive) results and FP (false positive) results is obtained, and the accuracy Precision is calculated according to the formula , and all preliminary screening results are obtained, and the FN (false negative) results corresponding thereto are obtained, wherein the FN (false negative) result represents a document that is not retrieved but relevant (misses relevant documents), and the recall rate Recall is calculated according to the formula The corresponding recall rate Recal l is calculated;
[0058] The obtained accuracy rate Precision and the recall rate Recal l are compared with the corresponding comparison thresholds, respectively. If the accuracy rate Precision is greater than the accuracy rate comparison threshold, it indicates that the corresponding preselected result is accurate, otherwise it indicates that the preselected result is inaccurate;
[0059] If the recall rate Recal l is greater than the recall rate comparison threshold, it indicates that the preselected result is accurate, otherwise it indicates that the preselected result is inaccurate;
[0060] If the accuracy rate Precision and the recall rate Recal l are both correct, it indicates that the preselected result is correct, and the preselected result is displayed to the corresponding operator, otherwise if any one of the accuracy rate Precision and the recall rate Recal l is incorrect, it indicates that the preselected result is abnormal, and a specific analysis signal is generated;
[0061] The obtained specific analysis signal is processed. If the accuracy rate is low and the recall rate is high, the keyword filtering is strengthened (such as increasing the entity matching weight), the vector similarity threshold is increased, if the accuracy rate is high and the recall rate is low, the vector retrieval threshold is reduced, the keyword synonyms are expanded, if both are low, the vector model is retrained (such as domain fine-tuning), the mixed weight is optimized, and the corresponding optimization processing information is generated, and is transmitted to the data management information output module.
[0062] The data management information output module is used to display the obtained preselected result and optimization processing information to the corresponding operator.
[0063] Embodiment two
[0064] Please refer to Figure 2 The application provides a data mining method of mail data, which specifically comprises the following steps:
[0065] Step one, collect mail data, and perform data cleaning, integration and conversion to obtain preprocessed data, and store structured data and unstructured data, generate storage management information, and the specific processing method is the same as the processing process of the data storage management module.
[0066] Step two, identify the input information and match it with the database to obtain a preliminary screening result, and segment the input information into common words to obtain preselected segmented words, and calculate the matching proportion to screen out keywords, and the specific processing method is the same as the processing process of the semantic retrieval analysis module;
[0067] Step three, calculate the score of the keyword, semantic vector and the preliminary screening result, and weighted sum to get the comprehensive index, based on the comprehensive index screening to get the pre-selected results, the specific processing mode is the same as the processing process of the semantic retrieval analysis module;
[0068] Step four, classify the TP and FP results of the pre-selected results, and calculate the corresponding accuracy and recall rate, and judge whether the pre-selected results are accurate or not, if not, adjust based on the accuracy and recall rate to generate optimization processing information, and the specific processing mode is the same as the processing process of the comprehensive mining processing module.
[0069] Part of the data in the above formula is calculated by taking its value, not by substituting the parameter unit, and the contents not described in detail in the specification are all prior art known to those skilled in the art.
[0070] The above examples are only used to illustrate the technical method of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical method of the present application.
Claims
1. A data mining system of mail data, characterized by, The application comprises: a data storage management module for cleaning, integrating and converting the email data transmitted by the email data collection module to obtain preprocessed data, storing the structured data and unstructured data in the preprocessed data, generating storage management information, and transmitting the storage management information to the semantic retrieval analysis module; a semantic retrieval analysis module for obtaining input information of a user, generating a semantic vector through a pre-trained language model, and matching a database to obtain a preliminary screening result, segmenting the input information, segmenting the input information according to common words to obtain a plurality of segmented words, matching the segmented words with the preliminary screening result to determine the keywords corresponding to the input information, and performing secondary matching and screening on the preliminary screening result based on the keywords and the semantic vector to obtain a preselected result and transmitting the preselected result to the comprehensive mining processing module; a comprehensive mining processing module for optimizing the obtained preselected result, calculating the accuracy and recall rate of the preselected result on the preliminary screening result respectively, analyzing the accuracy and recall rate, and if the accuracy and recall rate meet the requirements, indicating that the preselected result is correct, and displaying the preselected result to the corresponding operator, otherwise, if any group is incorrect, indicating that the preselected result is abnormal, and generating a specific analysis signal, processing the specific analysis signal, taking different optimization measures according to different situations, and generating corresponding optimization processing information, and transmitting the optimization processing information to the data management information output module.
2. The data mining system for mail data according to claim 1, wherein, The application further comprises an email data collection module and a data management information output module; the email data collection module is used for collecting email data, and the email data comprises structured data and unstructured data, and the collection methods comprise email server grabbing, local client extraction and real-time stream collection; the data management information output module is used for displaying the obtained preselected result and optimization processing information to the corresponding operator.
3. The data mining system for mail data according to claim 1, wherein, The specific method of the data storage management module for generating storage management information is: for the structured data in the preprocessed data, a relational database is used to store mailbox addresses, timestamps and department structured fields, and supports multi-condition queries, and for the unstructured data, a document database is used to store original texts, and corresponding storage management information is generated and transmitted to the semantic retrieval analysis module.
4. The data mining system for mail data according to claim 1, wherein, The specific method of the semantic retrieval analysis module for obtaining a plurality of segmented words is: obtaining input information of a user, and combining corresponding storage management information for retrieval analysis, recognizing the input information by using a pre-trained language model to obtain a corresponding semantic vector, matching the obtained semantic information with an established database to obtain a preliminary screening result, and segmenting the input information, and segmenting the input information according to common words to obtain a plurality of segmented words.
5. The data mining system for mail data according to claim 1, wherein, The method of the semantic retrieval analysis module for determining the keywords corresponding to the input information is: The segmented words are obtained and the segmented words are filtered to obtain pre-selected segmented words, and the pre-selected segmented words are marked as i, and i=1, 2, …, j, wherein j represents the type of the pre-selected segmented words, and a preliminary filtering result is obtained and marked as n, and n=1, 2, …, m, wherein m represents the number of the preliminary filtering results, then the pre-selected segmented words i are sequentially matched with the preliminary filtering results n in the order of the pre-selected segmented words i, and the corresponding matching times are obtained, and the matching proportion of the pre-selected segmented words i is calculated, and the pre-selected segmented words are sorted in descending order of the matching proportion, and the pre-selected segmented words with a matching proportion greater than a preset proportion are selected as keywords.
6. The data mining system for mail data according to claim 1, wherein, The semantic retrieval analysis module performs secondary matching filtering on the preliminary filtering results based on the keywords and semantic vectors, and the specific manner of obtaining the pre-selected results is as follows: The score of the keyword and the preliminary screening result is calculated, the keyword is obtained and labeled as o, and o=1, 2, …, p, wherein p represents the type of the keyword, and the BM25 score corresponding to the keyword o is calculated according to the formula , wherein n is the preliminary screening result, f(q o , n) is the frequency of the keyword o in the preliminary screening result n, IDF(q o ) is the inverse document frequency of the keyword o, |n| is the length of the preliminary screening result, avgdl is the average length of all preliminary screening results, and k1 and b are adjustment parameters. Then the similarity value between the semantic vector and the preliminary screening result n is calculated by cosine similarity, and the BM25 score and the similarity value are weighted and summed, according to the formula Z = score BM25 × α + score 相似度值 × (1- α) to calculate the comprehensive index Z of the keyword o and the preliminary screening result, wherein α is the weight coefficient of the BM25 score, score BM25 is the BM25 score, score 相似度值 is the similarity value, and the comprehensive index is sorted from large to small, and then the preliminary screening result whose comprehensive index is greater than the index threshold is screened out, which is recorded as the preselected result.
7. The data mining system for mail data according to claim 1, wherein, The specific manner in which the comprehensive mining processing module optimizes the obtained pre-selected results is as follows: All pre-selected results are obtained, and the obtained pre-selected results are classified, specifically classified into true positive results TP and false positive results FP, and the number of TP results and FP results is obtained, and then the formula The corresponding accuracy rate Precision is calculated, all preliminary screening results are obtained, and the corresponding false negative results FN are obtained, specifically indicating that the relevant documents are not retrieved, and the formula The corresponding recall rate Recal l is calculated, and the accuracy rate and the recall rate are analyzed.
8. The data mining system for mail data according to claim 7, wherein, The specific manner in which the comprehensive accuracy and recall rate are analyzed is as follows: The obtained accuracy Precision and recall rate Recal l are compared with the corresponding comparison thresholds, and if the accuracy Precision is greater than the accuracy comparison threshold, it indicates that the corresponding pre-selected result is accurate, otherwise it indicates that the pre-selected result is not accurate; If the recall rate Recal l is greater than the recall rate comparison threshold, it indicates that the pre-selected result is accurate, otherwise it indicates that the pre-selected result is not accurate; If the accuracy Precision and the recall rate Recal l are both correct, it indicates that the pre-selected result is correct, and the pre-selected result is displayed to the corresponding operator, otherwise if any one of the accuracy Precision and the recall rate Recal l is incorrect, it indicates that the pre-selected result is abnormal, and a specific analysis signal is generated and analyzed.
9. The data mining system for mail data according to claim 8, wherein, The specific manner in which the specific analysis signal is generated and analyzed is as follows: The obtained specific analysis signal is processed, if the accuracy is low and the recall rate is high, the keyword filtering is strengthened and the vector similarity threshold is increased, if the accuracy is high and the recall rate is low, the vector retrieval threshold is reduced and the keyword synonyms are expanded, if both are low, the vector model is retrained, the mixed weight is optimized, and the corresponding optimization processing information is generated, which is transmitted to the data management information output module.
10. A method of data mining of mail data for implementing the data mining system of any one of claims 1 to 9, characterized by, The method specifically comprises the following steps: Step one, collecting mail data, and performing data cleaning, integration and conversion to obtain pre-processed data, and storing structured data and unstructured data to generate storage management information; Step two, identifying the input information and matching it with the database to obtain a preliminary filtering result, segmenting the input information to obtain pre-selected segmented words, and calculating the matching proportion to obtain keywords; Step three, calculating the scores of the keywords, semantic vectors and preliminary filtering results, and weighting and summing to obtain a comprehensive index, and selecting the pre-selected results based on the comprehensive index; Step four, classifying the pre-selected results into TP and FP results, calculating the corresponding accuracy and recall rate, and determining whether the pre-selected result is accurate based on the accuracy and recall rate, if not, adjusting the accuracy and recall rate to generate optimization processing information.
Citation Information
Patent Citations
Synonym mining method and device for question and answer retrieval system
CN110442760A
Mail analysis method based on text mining
CN115599909A
Phishing mail deep detection method based on word frequency and context semantic multi-feature fusion
CN119814391A
Hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm
CN119961376A
Intelligent document duplicate checking system and method based on vector database and large language model
CN120373283A
Cited By
Big data quick retrieval method and system
CN121502059A