Method for classifying mailbox addresses, related apparatus and computer program product

By using name adjustment and word segmentation methods, the risks of misclassification and computational complexity in email address classification are resolved, achieving efficient and accurate automated classification of email addresses.

CN122490150APending Publication Date: 2026-07-31SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HODE INFORMATION TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for email address classification suffer from high risks of misclassification and high computational complexity, especially when dealing with a large number of email addresses, making it difficult to achieve efficient and high-quality automated classification.

Method used

The email address is restored using name adjustment rules, an initial email address is generated, and then word segmentation is performed. The results are then categorized based on the average length of the segmentation result set and other conditions.

Benefits of technology

This effectively reduces the risk of misclassification due to name distortion and improves the accuracy and quality of email address classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490150A_ABST
    Figure CN122490150A_ABST
Patent Text Reader

Abstract

This application provides a method, related apparatus, and computer program product for classifying email addresses. Based on name adjustment rules, the application restores the original email address of the email address to be classified, obtaining an initial email address corresponding to the original email address. The name adjustment rules instruct the generation of an email address by adding target characters to the target position in the initial email address. The initial email address is then segmented into words to obtain a set of segmentation results. In response to the average word length of the segmentation result set being less than or equal to a first length threshold, the email address to be classified is determined as the target category. Therefore, not only can automated email address classification be achieved, but the risk of misclassification due to name distortion can also be effectively reduced, thereby improving the accuracy and quality of classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, computer-readable medium, and computer program product for classifying email addresses. Background Technology

[0002] In modern information and communication technologies, email has become one of the most important means for users to interact with information systems. Email addresses serve as the fundamental identifiers of email systems and support the reliable delivery of information. With the widespread use of the internet and email, email addresses have become an important carrier for user communication and business activities.

[0003] Against this backdrop, due to security and risk control requirements, the accurate and efficient automated classification of email addresses has gradually become a key technological need for improving data management efficiency and business intelligence. Therefore, how to classify email addresses more efficiently and with higher quality is a matter of concern and urgent need. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for classifying email addresses, which can not only automate the classification of email addresses but also effectively reduce the risk of misclassification due to name distortion, thereby improving the accuracy and quality of classification.

[0005] One aspect of this application provides a method for classifying email addresses, comprising: restoring the email names of the email addresses to be classified based on name adjustment rules to obtain the initial email names corresponding to the email addresses to be classified, wherein the name adjustment rules are used to indicate that the email names are generated by adding target characters to the target positions of the initial email names; performing word segmentation on the initial email names to obtain a word segmentation result set; and determining the email addresses to be classified as the target category in response to the average word segmentation length of the word segmentation result set being less than or equal to a first length threshold.

[0006] Another aspect of this application provides an apparatus for classifying email addresses, comprising: an email address restoration module configured to restore the email address to be classified based on a name adjustment rule to obtain an initial email address corresponding to the email address to be classified, wherein the name adjustment rule is used to indicate that an email address is generated by adding a target character at a target position in the initial email address; an email address segmentation module configured to segment the initial email address to obtain a segmentation result set; and a first address classification module configured to determine the email address to be classified as the target category in response to the average segmentation length of the segmentation result set being less than or equal to a first length threshold.

[0007] In another aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for classifying email addresses as provided above.

[0008] Another aspect of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the method for classifying email addresses as provided above.

[0009] Another aspect of this application is a computer program product that includes a computer program having computer program instructions stored thereon, which, when executed by a processor, can implement the method for classifying email addresses as provided above.

[0010] The solution provided in this application firstly restores the email address to be classified based on name adjustment rules, obtaining the initial email address corresponding to the email address to be classified. The name adjustment rules instruct the generation of the email address by adding target characters to the target position in the initial email address. Then, the initial email address is segmented into words to obtain a set of segmentation results. Finally, in response to the average segmentation length of the set of segmentation results being less than or equal to a first length threshold, the email address to be classified is determined as the target category. Therefore, not only can automated email address classification be achieved, but the risk of misclassification due to name distortion can also be effectively reduced, thereby improving the accuracy and quality of classification. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating a process for classifying email addresses, as provided in one embodiment of this application; Figure 2 A flowchart illustrating another process for classifying email addresses as provided in an embodiment of this application; Figure 3 A flowchart illustrating a process for classifying email addresses in a specific application scenario, as provided in an embodiment of this application; Figure 4 A schematic diagram of a device for classifying email addresses provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing the solutions in the embodiments of this application.

[0013] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] In a typical configuration of this application, the terminal and the service network devices each include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0016] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0017] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer program instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0018] As discussed above, how to classify email addresses more efficiently and with higher quality is a matter of concern and urgent need.

[0019] In some solutions, email addresses can be categorized based on domain name or email address similarity. For example, after collecting email addresses, those that meet the similarity requirements can be clustered, and the email addresses included in the clustering results can be classified as similar addresses.

[0020] However, this method suffers from high computational complexity, making it difficult to effectively classify email addresses as the number of addresses to be categorized and processed increases. Furthermore, it lacks the ability to recognize and classify email addresses generated using random characters, hindering the efficient and high-quality automation of email address classification.

[0021] To address this issue, this application provides a method for classifying email addresses. The method first restores the original email address based on a name adjustment rule, obtaining an initial email address corresponding to the original address. The name adjustment rule instructs the generation of the email address by adding target characters to the target position in the initial email address. Then, the initial email address is segmented into words, resulting in a segmentation result set. Finally, in response to the average segmentation length of the result set being less than or equal to a first length threshold, the email address to be classified is determined as the target category. This not only enables automated email address classification but also effectively reduces the risk of misclassification due to name distortion, thereby improving the accuracy and quality of classification.

[0022] In practical scenarios, the execution entity of this method can be a user device, a device composed of a user device and a network device integrated through a network, or an application running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.

[0023] When the executing entity is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules, or as a single software program or software module, without specific limitations.

[0024] Figure 1 The present application illustrates a process 100 for classifying email addresses, which includes at least the following processing steps: (Step) S101, Based on the name adjustment rules, restore the email names of the email addresses to be classified to obtain the initial email names corresponding to the email addresses to be classified. In the embodiments of this application, the executing entity may first obtain the email addresses to be categorized. The email address to be categorized (the email address to be categorized) may consist of two parts: the email address and the domain name. For example, these two parts may be connected by, or separated by, an "@" symbol. For example, for an email address to be categorized, "XX@yy.com", "XX" can be understood as the "email address", and correspondingly, "@yy.com" can be understood as the "domain name".

[0025] In this step, after obtaining the email addresses to be categorized, the executing entity can extract the email names included in them and obtain the name adjustment rules corresponding to the email addresses to be categorized.

[0026] Name adjustment rules can be used to instruct rules for generating email addresses by adding target characters (e.g., special characters, characters, numbers, etc.) to the target position in the initial email address name.

[0027] For example, some providers, in order to facilitate users' use of their email services, allow users to create multiple associated or derived email addresses under the same username by adding special symbols such as "+" or ".", or specific letters, to the end or middle of the email address. In such cases, adding special symbols such as "+" or "." can be understood as the aforementioned name adjustment rules.

[0028] For example, for the "XX@yy.com" example above, it can be used to generate new email names such as "XX.@yy.com" or "XX+@yy.com" that are associated with and indicate the same user, based on name adjustment rules that include special symbols such as "+" or ".".

[0029] Accordingly, after obtaining the name adjustment rules, the email names of the email addresses to be categorized can be restored based on these rules. For example, the added "." and "+" can be removed to restore and identify the original email name from which the email name to be categorized originated.

[0030] It should be understood that in practice, the implementing entity may attempt to perform such a restoration operation for each email address. However, for those unclassified email addresses (or more specifically, email addresses of unclassified email addresses) that have not been adjusted, their status may not change after the restoration attempt (because they are not obtained by adjusting email addresses).

[0031] In some embodiments, name adjustment rules can typically be pre-maintained and configured, allowing the executing entity to obtain these rules locally. In other cases, after obtaining the email addresses to be categorized, the executing entity can retrieve the name adjustment rules from the provider offering the service for those addresses, based on their corresponding domain names.

[0032] In some embodiments, to improve restoration efficiency, a set of domain names can be maintained in advance based on standards that allow such name adjustments. For example, the providers or suppliers corresponding to the domain names in the set of domain names are allowed to make such "name adjustments".

[0033] In this scenario, the executing entity can optionally or alternatively choose to first read the domain name of the email address to be categorized. If the domain name is included in a pre-determined set of domain names (i.e., this set of domain names can be pre-set based on the aforementioned standard that allows adjusting email names according to name adjustment rules to form a group of email addresses belonging to the same user, and that for the same user, there are at least two providers or providers of email addresses with different email names), then the executing entity responds by restoring the email name of the email address to be categorized based on the name adjustment rules, obtaining the initial email name corresponding to the email address to be categorized. This avoids identification errors due to email addresses provided by providers or providers that lack or refuse to use name adjustment rules to adjust and generate email address groups, thus improving the overall identification quality.

[0034] It should be understood that the acquisition, storage, use, processing, transportation, provision and disclosure of any type of information involved in the technical solutions of this application, such as user personal information, business-related information (e.g., email addresses to be categorized), and information provided by third parties (e.g., name adjustment rules), all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0035] For example, when an implementing entity needs to categorize email addresses and obtain the email addresses to be categorized and their corresponding name adjustment rules, it will notify relevant parties of the acquisition of this information, as well as the processing method and purpose of this information, through methods such as pop-ups or sending requests, in order to request authorization from the holder of this information. Accordingly, after authorization, the implementing entity will obtain the corresponding information according to the pre-announced scope and usage method, and process it accordingly. This ensures that the acquisition and use of information are both known and permitted by the users and information providers, guaranteeing that the processes of acquisition, storage, use, processing, transportation, provision, and disclosure comply with relevant laws and regulations and do not violate public order and good morals.

[0036] S102, perform word segmentation on the initial email address to obtain a set of word segmentation results; In the embodiments of this application, after obtaining the initial email address based on the above S101, the executing entity can perform word segmentation (or word cutting) processing on the restored initial email address to obtain one or at least two word segmentation results, and form a corresponding word segmentation result set based on these word segmentation results.

[0037] For example, the executing entity can use a tokenizer to segment the initial email address into smaller, independent units, such as tokens, to discretize and convert the initial email address into discrete units.

[0038] S103, in response to the average word segmentation length of the word segmentation result set being less than or equal to the first length threshold, the email address to be classified is determined as the target category.

[0039] In the embodiments of this application, after obtaining the word segmentation result set based on S102 above, the executing entity can count the number of word segmentation results included in the word segmentation result set in this step, and determine the average word segmentation length of the word segmentation result set (i.e., the average length of the word segmentation results) based on the ratio of the length of the initial email name to the number of such results. For example, after the executing entity segments the initial email name "XXYYZZ" into three word segmentation results of "XX", "YY", and "ZZ", the average word segmentation length can be exemplarily set to "2" (6 characters minus 3 word segmentation results = "2").

[0040] If the average word length is less than or equal to the first length threshold (which is usually set based on the assumption that the initial email address is generated regularly according to certain rules and habits), it means that the initial email address may not have been generated by the user based on real information and real usage habits, but rather by a random method (or, such an initial email address may be a string of characters lacking effective information and regularity generated by a random or machine-generated method). In such a case, the executing entity can respond to this and classify it into the target category.

[0041] In some embodiments, the target classification can be understood as a classification that deviates from common, conventional practices, may require special attention, or may pose a risk. For example, the aforementioned classification might be of email addresses that are mass-generated and registered by machine. This tagging and classification method allows for the differentiation of these email addresses, enabling subsequent processing (e.g., managing access permissions, identifying associated email addresses of the same user for convenient batch management of email addresses, etc.) to meet specific needs such as authorization management and security control scenarios.

[0042] The method for classifying email addresses provided in this application firstly restores the original email address based on name adjustment rules to obtain the initial email address corresponding to the original email address. The name adjustment rules instruct the generation of the email address by adding target characters to the target position in the initial email address. Then, the initial email address is segmented into words to obtain a set of segmentation results. Finally, in response to the average word length of the segmentation result set being less than or equal to a first length threshold, the email address to be classified is determined as the target category. Therefore, this method not only enables automated classification of email addresses but also effectively reduces the risk of misclassification due to name distortion, thereby improving the accuracy and quality of classification.

[0043] In some embodiments, during the above process, considering that numbers may often be difficult to segment effectively, resulting in each segmentation result being an independent number, the executing entity can first check whether the initial email address actually contains letters during the segmentation process to obtain the segmentation result set. If letters are included, the executing entity can respond by further segmenting the initial email address as discussed above to obtain the segmentation result set. This avoids misjudgments in such "purely numerical" cases and improves the classification quality of email addresses.

[0044] In some embodiments, as discussed above, the management and classification process of email addresses to be classified may be for at least two email addresses to be classified. In such cases, the executing entity may also choose to refer to the specific circumstances of the email addresses to be classified (e.g., whether there is "co-occurrence information" in the email address part) to complete the classification process and achieve the classification purpose.

[0045] For easier understanding, you can also refer to the following: Figure 2 Let's have a discussion. Figure 2 A flowchart of another process 200 for classifying email addresses according to an embodiment of this application is shown.

[0046] Process 200 may specifically include the following steps: S201, Based on the name adjustment rules, restore the email names of the email addresses to be classified to obtain the initial email names corresponding to the email addresses to be classified; S202, perform word segmentation on the initial email address to obtain a set of word segmentation results; The above S201-S202 and as follows Figure 1 The S101-S102 shown are the same. For the same parts, please refer to the corresponding parts of the previous embodiment. They will not be repeated here.

[0047] S203, in response to the fact that the number of word segmentation result sets including the target word segmentation result is greater than or equal to the number threshold, read the target word segmentation result set including the target word segmentation result; Specifically, compared to the above process 100, in addition to completing the classification of email addresses to be classified through S103, the executing entity can also choose to use "co-occurrence information" for classification as discussed above (it should be understood that in different embodiments, S103 and S203 can actually coexist and are not "exclusively" related).

[0048] For example, the executing entity can detect the specific word segmentation results included in the word segmentation result set corresponding to each initial email address, and count the number of initial email addresses that include the target word segmentation results.

[0049] In practice, the executing entity can use word segmentation results in the form of "Chinese character spelling" as target word segmentation results to count the number of times the same target word segmentation result appears. For example, for "heng", the executing entity can use it as the "target word segmentation result" and count the number of initial email addresses that have "target word segmentation results" in the form of "heng" (or, the number of word segmentation result sets that include the target word segmentation result).

[0050] If the number of word segmentation result sets including the target word segmentation result is greater than or equal to the number threshold (for ease of understanding, the word segmentation result set including the target word segmentation result can also be described as the target word segmentation result set), then the execution entity can choose to read the target word segmentation result set including the target word segmentation result.

[0051] In some embodiments, a second length threshold can be set to select and mark target segmentation results for reference and searching from the segmentation results. For example, classification results with a length greater than or equal to the second length threshold can be selected as target classification results to control and reduce the amount of computation while using richer information that is less likely to be duplicated due to randomness to determine whether such "co-occurrence" is caused by the same user registering with the same or similar information, thereby improving classification quality.

[0052] Accordingly, the specific value of the second length threshold can be adaptively set according to this objective and different scenarios, which will not be repeated here.

[0053] S204, Based on the type and arrangement structure of the word segmentation results included in the target word segmentation result set, perform clustering operation based on each target word segmentation result set; Specifically, after determining the target word segmentation result set based on the above S203, the executing entity can choose to perform clustering operations based on the type and arrangement structure of the word segmentation results included therein in order to attempt to cluster these target word segmentation result sets.

[0054] In some embodiments, the types of word segmentation results can be divided into letters, numbers, special characters, etc., and the arrangement structure can correspondingly represent the order and arrangement relationship of the word segmentation results including the types (for example, corresponding to a specific target word segmentation result set, its arrangement structure can be exemplarily "letters", "numbers", "letters", "special characters").

[0055] Accordingly, in this step, the executing entity can perform "clustering" on each target word segmentation result set based on the included "types" and the arrangement structure of each "type" to cluster the target word segmentation results that meet the similarity requirements between the included "types" and the arrangement structure of each "type" to a cluster center.

[0056] Then, if a cluster center can cluster at least two related target word segmentation result sets, the executing entity can recognize it as a "cluster result".

[0057] It should be understood that after the clustering operation performed in this step, the clustering results of the executing entity can be zero (for example, each target word segmentation result set is a cluster center, and there are no two target word segmentation result sets that can be clustered) to multiple.

[0058] S205, in response to the ability to cluster at least one clustering result, the email addresses to be classified corresponding to the target word segmentation result set that can be clustered into the clustering result are determined as the target category.

[0059] Specifically, if at least one clustering result can be obtained in the above S204, that is, there is at least one clustering result that can cluster and associate the two target word segmentation result sets, then the executing entity can respond to this and determine that they are related and similar (for example, belonging to the same user), and determine the email address to be classified corresponding to the target word segmentation result set that can be clustered to the clustering result as the target category.

[0060] It should be understood that the "standards for clustering results" can be set differently depending on the specific classification and recognition strength. For example, in the process of classifying and processing a large number of email addresses to be classified, the criteria for determining and recognizing "clustering results" can be increased accordingly (for example, increasing the number of target word segmentation result sets that can be identified as a clustering result). This will not be elaborated on here.

[0061] In this way, we can use some co-occurrence information in the initial email addresses to efficiently and with low computational cost query and mine email addresses that may belong to the same user, and then classify and label them.

[0062] In some embodiments, if there are at least two email addresses to be classified, after the executing entity restores the initial email address (e.g., restores the initial email address based on S101), it may also choose to detect whether there is a target initial email address in the initial email address that appears more or less than or equal to a threshold number of times (i.e., multiple email addresses are restored to the target initial email address after being restored).

[0063] Accordingly, if such target initial email names exist, it indicates that there may be a situation where a large number of email names are adjusted using name adjustment rules. In this case, the executing entity can directly determine the email addresses to be classified corresponding to these target initial email names as target categories for labeling and representation. This allows subsequent processing based on the classification results of the target categories, combined with actual business and scenario strategies. Consequently, in this case, the executing entity does not need to perform word segmentation on the email addresses to be classified corresponding to the target initial email names that have already been classified into target categories, thus saving computational resources. That is, in such an embodiment, the executing entity will only perform actions such as S102 on those initial email names that do not belong to the target initial email names to perform word segmentation and obtain a set of word segmentation results.

[0064] To enhance understanding, this application also provides a specific implementation scheme based on a particular application scenario. Please refer to it. Figure 3 , Figure 3This is a flowchart of a process 300 for classifying email addresses in a specific application scenario, provided as an embodiment of this application.

[0065] In process 300, for example, the aforementioned “executing entity” can classify the (to be classified) email addresses 310, 320, 330 and 340, or determine whether each of them is or belongs to “target category 370”.

[0066] It should be understood that the number of “email addresses” mentioned above is merely an illustrative choice for ease of understanding and is not intended to limit the number of (to be categorized) email addresses that need to be processed and classified.

[0067] In process 300, the executing entity can first (simultaneously or sequentially) execute S301a on email address 310, S301b on email address 320, S301c on email address 330, and S301d on email address 340 respectively, in order to restore their email names respectively.

[0068] Accordingly, after S301a, email address 310 can be restored to its initial email address 311; after S301b, email address 320 can be restored to its initial email address 321; after S301c, email address 330 can be restored to its initial email address 331; and after S301d, email address 340 can be restored to its initial email address 341.

[0069] For example, since the number of times the initial email address "Xxyyzz" appears repeatedly (e.g., "3" corresponding to initial email address 311, initial email address 321 and initial email address 331) is greater than or equal to the number of occurrences threshold (the number of occurrences threshold in process 300 can be "2" for example), the executing entity can respond to this by executing S302 to classify email address 310, email address 320 and email address 330 into target category 370.

[0070] For the initial email address 341, the executing entity can continue to execute S303 to segment the initial email address 341 into words, and obtain segmentation results 351 to 356.

[0071] Then, the executing entity can continue to execute S304 to determine the average segmentation length 360 based on the number of segmentation results 351 to 356 (e.g., "6") (e.g., based on 8 / 6, the average segmentation length 360 is 1.3333).

[0072] For example, if the first length threshold is "2", then in such a case, since the average word segmentation length 360 (1.3333) is less than "2", the executing entity can choose to continue executing S305 to classify email address 340 into target category 370.

[0073] This application also provides an apparatus for classifying email addresses, the structure of which is as follows: Figure 4 The apparatus 400 shown includes: an email address restoration module 410, configured to restore the email address of the email address to be classified based on a name adjustment rule, to obtain the initial email address corresponding to the email address to be classified, wherein the name adjustment rule is used to indicate that the email address is generated by adding target characters to the target position of the initial email address; an email address word segmentation module 420, configured to segment the initial email address to obtain a word segmentation result set; and a first address classification module 430, configured to determine the email address to be classified as the target category in response to the average word segmentation length of the word segmentation result set being less than or equal to a first length threshold.

[0074] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for classifying email addresses provided in this embodiment can not only realize the automatic classification of email addresses, but also effectively reduce the risk of misclassification caused by name distortion, thereby improving the accuracy and quality of classification.

[0075] In some embodiments, in response to the existence of at least two email addresses to be classified, the apparatus 400 further includes: a second address classification module, configured to determine the email address to be classified corresponding to the target initial email name as the target category in response to the existence of a target initial email name, wherein the number of times the target initial email name appears repeatedly is greater than or equal to a frequency threshold.

[0076] In some embodiments, in response to the existence of at least two email addresses to be classified, the apparatus 400 further includes: a target set reading module, configured to read a target word segmentation result set including the target word segmentation result in response to the number of word segmentation result sets including the target word segmentation result being greater than or equal to a number threshold; a clustering execution module, configured to perform clustering operations based on the type and arrangement structure of the word segmentation results included in the target word segmentation result set; and a third address classification module, configured to determine the email addresses to be classified corresponding to the target word segmentation result set that can be clustered into the clustering result as the target classification in response to the ability to cluster at least one clustering result.

[0077] In some embodiments, the length of the target word segmentation result is greater than or equal to a second length threshold.

[0078] In some embodiments, the email name restoration module 410 includes: a domain name reading submodule, configured to read the domain name of the email address to be classified; and an email name restoration submodule, configured to restore the email name of the email address to be classified based on name adjustment rules in response to the domain name being included in a predetermined set of domain names, to obtain the initial email name corresponding to the email address to be classified.

[0079] In some embodiments, the email name segmentation module 420 is further configured to segment the initial email name in response to the inclusion of letters in the initial email name, thereby obtaining a set of segmentation results.

[0080] Based on the same concept, this application also provides an electronic device, a readable storage medium, and a computer program product. The method corresponding to the electronic device can be the method for classifying email addresses in the foregoing embodiments, and its problem-solving principle is similar to that method. The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the foregoing embodiments of this application.

[0081] Electronic devices can be user devices, or devices composed of user devices and network devices integrated through a network, or applications running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, and can be used to implement some processing functions when setting an alarm clock. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.

[0082] Figure 5The diagram illustrates the structure of an electronic device suitable for implementing the methods and / or technical solutions in the embodiments of this application. The electronic device 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on a program stored in a Read Only Memory (ROM) 502 or a program loaded from a storage portion 508 into a Random Access Memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.

[0083] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, touchscreen, microphone, infrared sensor, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, etc., and speakers, etc.; a storage section 508 including one or more computer-readable media such as hard disk, optical disk, magnetic disk, semiconductor memory, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet.

[0084] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, it performs the functions defined in the methods of this application.

[0085] Another embodiment of this application provides a computer-readable storage medium and a computer program product having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.

[0086] Specifically, this embodiment may employ any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, a system, apparatus, or device that is, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0087] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0088] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0089] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0090] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0092] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules and units is only a logical functional division, and in actual implementation, there may be other division methods. Taking units as examples, multiple units or page components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0093] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] Furthermore, the functional modules and units in the various embodiments of this application can be integrated into one processing module or unit, or each module or unit can exist physically separately, or two or more units can be integrated into one module or unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules and units.

[0095] The integrated modules and units implemented as software functional modules and units described above can be stored in a computer-readable storage medium. These software functional modules and units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

[0097] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

Claims

1. A method for classifying email addresses, characterized in that, include: Based on the name adjustment rules, the email names of the email addresses to be classified are restored to obtain the initial email names corresponding to the email addresses to be classified. The name adjustment rules are used to indicate how to generate the email names by adding target characters to the target positions of the initial email names. The initial email address is segmented into words to obtain a set of segmentation results; In response to the average word segmentation length of the word segmentation result set being less than or equal to a first length threshold, the email address to be classified is determined as the target category.

2. The method according to claim 1, characterized in that, In response to the existence of at least two of the email addresses to be classified, the method further includes: In response to the existence of a target initial email name, the email address to be classified corresponding to the target initial email name is determined as the target category, wherein the number of times the target initial email name appears repeatedly is greater than or equal to a frequency threshold.

3. The method according to claim 1, characterized in that, In response to the existence of at least two of the email addresses to be classified, the method further includes: In response to the fact that the number of the word segmentation result set including the target word segmentation result is greater than or equal to the number threshold, the target word segmentation result set including the target word segmentation result is read; Based on the type and arrangement structure of the word segmentation results included in the target word segmentation result set, clustering operations are performed on each target word segmentation result set; In response to the ability to cluster at least one clustering result, the email address to be classified corresponding to the target word segmentation result set that can be clustered into the clustering result is determined as the target category.

4. The method according to claim 3, characterized in that, The length of the target word segmentation result is greater than or equal to the second length threshold.

5. The method according to claim 1, characterized in that, The process of restoring the email addresses to be categorized based on the name adjustment rules to obtain the initial email addresses corresponding to the email addresses to be categorized includes: Read the domain name of the email address to be categorized; In response to the domain name being included in a pre-determined set of domain names, the email names of the email addresses to be classified are restored based on the name adjustment rules to obtain the initial email names corresponding to the email addresses to be classified.

6. The method according to any one of claims 1-5, characterized in that, The initial email address is segmented into words to obtain a set of segmentation results, including: In response to the initial email address containing letters, the initial email address is segmented into words to obtain a set of segmentation results.

7. A device for classifying email addresses, characterized in that, include: The email name restoration module is configured to restore the email name of the email address to be classified based on the name adjustment rule to obtain the initial email name corresponding to the email address to be classified. The name adjustment rule is used to indicate that the email name is generated by adding target characters to the target position of the initial email name. The email name segmentation module is configured to segment the initial email name to obtain a set of segmentation results; The first address classification module is configured to determine the email address to be classified as the target category in response to the average word length of the word segmentation result set being less than or equal to a first length threshold.

8. An electronic device, the electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.

9. A computer-readable medium having stored thereon computer program instructions that can be executed by a processor to implement the method as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.