A method and apparatus for contact recovery based on multi-source heterogeneous data collaborative processing

By integrating heterogeneous data from multiple sources through standardization and similarity calculation methods, the inconsistency and redundancy issues in contact data management in existing technologies are resolved, enabling high-precision contact data merging and recovery.

CN120892558BActive Publication Date: 2026-01-30深圳市乐数科技有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511440996.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-30
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing contact management methods cannot effectively integrate heterogeneous data from multiple sources, resulting in data loss, duplication, or inconsistency. In particular, the accuracy and reliability of the processing results are low when migrating across platforms and devices.

Method used

By standardizing the processing of multi-source heterogeneous data, a unified multi-source contact data set is formed. The same data under each data type is matched, and when the data is inconsistent, the contact data is merged by similarity calculation. The similarity between the different data groups and the same data groups is combined to determine whether to merge.

Benefits of technology

It achieves cross-platform and cross-device data integration, improves the accuracy of contact matching, reduces data loss or redundancy, enhances merging accuracy, and avoids the risk of erroneous or missed merging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892558B_ABST
    Figure CN120892558B_ABST
Patent Text Reader

Abstract

This invention relates to the technical field of data recognition and provides a contact recovery method and apparatus based on multi-source heterogeneous data collaborative processing. The contact recovery method includes: standardizing multi-source heterogeneous data to obtain a multi-source contact data set; matching identical data under each data type in the multi-source contact data set; obtaining two first contact data sets corresponding to the identical data; merging the two first contact data sets to obtain second contact data when data under multiple data types is identical in the two first contact data sets; and using the second contact data, third contact data, and remaining contact data as the final contact data and synchronizing it to the target device. This similarity-based intelligent merging method effectively handles situations where there are incomplete matches between different data sources, improves merging accuracy, and reduces the risk of incorrect or missed merging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of data recognition, and in particular relates to a contact recovery method and device based on multi-source heterogeneous data collaborative processing. Background Technology

[0002] With the widespread use of smartphones and other mobile devices, users' contact data is no longer limited to a single data source but is widely stored across different devices and applications. This contact data comes from various heterogeneous data sources, including but not limited to phone address books, email systems, social media platforms, instant messaging applications, and enterprise management systems. Due to the differences in data formats, structures, and update frequencies among these data sources, inconsistencies and redundancies in contact data have arisen, causing considerable difficulties for users in managing their contact information.

[0003] Currently, most traditional contact management methods rely on a single data source or static data synchronization. These methods often fail to effectively integrate data from multiple heterogeneous sources, especially when facing cross-platform and cross-device data migration, easily leading to data loss, duplication, or inconsistency. To eliminate these problems, many research and technological developments have proposed contact data merging schemes based on data fusion and matching algorithms. However, these schemes typically cannot effectively handle the fusion of multi-source heterogeneous data, and the accuracy and reliability of the processing results are low when dealing with complex data types and poor data quality. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a contact recovery method and apparatus based on multi-source heterogeneous data collaborative processing, in order to solve the technical problem that the accuracy and reliability of the processing results are low when the data types are complex and the data quality is poor.

[0005] A first aspect of this invention provides a contact recovery method based on multi-source heterogeneous data collaborative processing, the contact recovery method based on multi-source heterogeneous data collaborative processing including:

[0006] Acquire multi-source heterogeneous data and standardize the multi-source heterogeneous data to obtain a multi-source contact data set; wherein, the data type of each contact data in the multi-source contact data set includes telephone number, name, email address, company name, date of birth and / or home address;

[0007] In the multi-source contact data set, match the same data under each data type;

[0008] Obtain the data of the two first contacts corresponding to the same data;

[0009] merge the two first contact data to obtain second contact data when the data of the multiple data types in the two first contact data are all the same;

[0010] merge the two first contact data to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in the multiple data types when the data of the multiple data types in the two first contact data are not all the same, wherein the difference data group refers to two data of the same data type that are different, and the same data group refers to a data group of the same data type that has the same data;

[0011] merge the two first contact data to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in the multiple data types when the data of the multiple data types in the two first contact data are not all the same, wherein the difference data group refers to two data of the same data type that are different, and the same data group refers to a data group of the same data type that has the same data;

[0012] Further, the step of merging the two first contact data to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in the multiple data types when the data of the multiple data types in the two first contact data are not all the same includes:

[0013] extracting a difference data group and a same data group in the multiple data types corresponding to the two first contact data;

[0014] calculating a first similarity between two current data in the difference data group according to a data type corresponding to the difference data group;

[0015] taking a first preset value as the second similarity of the same data group, wherein the first preset value includes 1;

[0016] weighting and summing the first similarity or the second similarity corresponding to each of the multiple data types according to a weight coefficient corresponding to the multiple data types to obtain a final similarity;

[0017] if the final similarity is greater than a first threshold value, obtaining a channel priority of the two first contact data;

[0018] taking the first contact data corresponding to the maximum channel priority in the two first contact data as the third contact data, and eliminating the other first contact data.

[0019] Further, the step of calculating a first similarity between two current data in the difference data group according to a data type corresponding to the difference data group includes:

[0020] when the data type corresponding to the difference data set is a telephone number, a birthday, or an email address, a first similarity is calculated according to a number of continuous same digits in the difference data set;

[0021] when the data type corresponding to the difference data set is a name, a first similarity is calculated according to pinyin corresponding to the name;

[0022] when the data type corresponding to the difference data set is a home address, a first similarity is calculated according to a word distribution feature corresponding to the home address;

[0023] when the data type corresponding to the difference data set is a company name, a first similarity is calculated according to a character distribution feature of the company name.

[0024] Further, the step of calculating the first similarity according to the number of continuous same digits in the difference data set when the data type corresponding to the difference data set is a telephone number, a birthday, or an email address comprises:

[0025] when the data type corresponding to the difference data set is a telephone number or a birthday, a first number of continuous same digits in the two current data is extracted; wherein the first number of continuous same digits refers to a number of continuous same digits in the two current data;

[0026] a second number of digits in the two current data is respectively counted;

[0027] the first number is divided by a maximum second number to obtain the first similarity;

[0028] if the data type corresponding to the difference data set is an email address, a prefix and a suffix in the email address are extracted;

[0029] if the two suffixes are same, a third number of continuous same digits in the two prefixes is extracted; wherein the third number of continuous same digits refers to a number of continuous same digits in the two prefixes;

[0030] a fourth number of digits in the two prefixes is respectively counted;

[0031] the third number is divided by a maximum fourth number to obtain the first similarity;

[0032] if the two suffixes are different, the first similarity is determined as 0.

[0033] Further, the step of calculating the first similarity according to the pinyin corresponding to the name when the data type corresponding to the difference data set is a name comprises:

[0034] if the data type corresponding to the difference data set is a name, obtaining a pinyin corresponding to the name;

[0035] if the two Chinese names corresponding to the difference data set are the same, a second preset value is taken as the first similarity;

[0036] if the two Chinese names corresponding to the difference data set are different, and the two pinyins corresponding to the difference data set are the same, a third preset value is taken as the first similarity;

[0037] if the two Chinese names corresponding to the difference data set are different, and the two pinyins corresponding to the difference data set are different, 0 is taken as the first similarity.

[0038] Further, when the data type corresponding to the difference data set is a home address, the step of calculating the first similarity according to the word distribution characteristics corresponding to the home address comprises:

[0039] when the data type corresponding to the difference data set is a home address, the home address is processed by word segmentation to obtain a plurality of geographical words;

[0040] matching the number of same words between the plurality of geographical words corresponding to the two current data in the difference data set respectively;

[0041] respectively counting the number of current words in the two current data;

[0042] dividing the number of same words by the maximum number of current words to obtain the first similarity.

[0043] Further, when the data type corresponding to the difference data set is a home address, the step of processing the home address by word segmentation to obtain a plurality of geographical words comprises:

[0044] when the data type corresponding to the difference data set is a home address, geographical division words in the address are extracted; the geographical division words include country, province, city, district, street, road, community, building, unit and room;

[0045] based on the positions of the geographical division words in the home address, the home address is segmented to obtain a plurality of strings;

[0046] the string and the adjacent geographical division word located on the right side of the string are taken as initial words;

[0047] According to the original order of the plurality of initial words in the address, the left initial word and the right initial word of each current initial word are obtained; wherein the left initial word refers to the adjacent initial word located on the left side of the current initial word, and the right initial word refers to the adjacent initial word located on the right side of the current initial word;

[0048] Each current initial word, the left initial word corresponding to each current initial word, and the right initial word corresponding to each current initial word are combined into a to-be-recognized word group according to the original order in the address;

[0049] The to-be-recognized word group is input into a word marking model to obtain a word marking corresponding to each character output by the word marking model; wherein the word marking includes a word head, a word middle, a word tail, and a single word;

[0050] According to the distribution position of each character in the current initial word, it is judged whether each character conforms to the word marking;

[0051] If the characters in the current initial word all conform to the word marking, the current initial word is taken as the geographical word;

[0052] If the characters in the current initial word do not all conform to the word marking, the current initial word is divided based on the word marking to obtain the geographical word.

[0053] Further, when the data type corresponding to the difference data set is a company name, the step of calculating the first similarity according to the character distribution characteristics of the company name comprises:

[0054] If the data type corresponding to the difference data set is a company name, the character strings of the two current data corresponding to the difference data set are input into a similarity model to obtain a first similarity output by the similarity model;

[0055] The similarity model is:

[0056]

[0057]

[0058]

[0059]

[0060] wherein, the first similarity, the nonlinear prefix weighting function, the number of characters of the first current data in the difference data set, represents the number of characters of the second current data in the difference data set, represents taking the maximum value, represents the number of characters of the same prefix in the two strings, represents a dynamic penalty factor, represents the character similarity, represents the position difference of the same string in the two current data of the high correlation character, represents the number of high correlation characters, which refers to the same string existing in the two current data and the position difference is not more than / 2-1.

[0061] The second aspect of the embodiment of the application provides a contact recovery device based on multi-source heterogeneous data collaborative processing, which comprises:

[0062] a first acquisition unit, configured to acquire multi-source heterogeneous data, and perform standardization processing on the multi-source heterogeneous data to obtain a multi-source contact data set; wherein the data type of each contact data in the multi-source contact data set comprises a telephone number, a name, an email address, a company name, a birthday and / or a home address;

[0063] a matching unit, configured to match the same data under each data type in the multi-source contact data set;

[0064] a second acquisition unit, configured to acquire two first contact data corresponding to the same data;

[0065] a first merging unit, configured to merge the two first contact data to obtain second contact data when the data under multiple data types in the two first contact data are all the same;

[0066] a second merging unit, configured to merge the two first contact data to obtain third contact data according to the first similarity corresponding to the difference data set and the second similarity corresponding to the same data set in multiple data types when the data under multiple data types in the two first contact data are not all the same; wherein the difference data set refers to two data having differences under the same data type, and the same data set refers to a data set having the same data under the same data type;

[0067] a synchronization unit, configured to take the second contact data, the third contact data and remaining contact data as final contact data, and synchronize the final contact data to a target device; wherein the remaining contact data refers to contact data in the multi-source contact data set which has not been processed by merging.

[0068] The third aspect of the embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the contact recovery method based on multi-source heterogeneous data collaborative processing in the first aspect.

[0069] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps in the contact recovery method based on multi-source heterogeneous data collaborative processing in the first aspect.

[0070] Compared with the prior art, the embodiment of the present application has the beneficial effects that: since the current contact data is stored in multiple heterogeneous data sources, including mobile phone address book, social media, email system, etc., each data source uses different formats, structures and storage methods, which brings great challenges to the management and recovery of contact data. The present application converts the multi-source heterogeneous data into a unified standard format through standardized processing, forms a multi-source contact data set, and realizes cross-platform and cross-device data integration, providing a standardized and unified data basis for subsequent contact recovery. In the multi-source contact data set, the same contact in different data sources can be effectively identified by matching the same data under each data type (such as phone number, name, email address, etc.), reducing the possibility of data loss or redundancy. Especially in the case of multiple similar contact records in different data sources, the accuracy of contact matching is improved by accurately matching the same data of each data type, avoiding repeated and erroneous data merging. The present application proposes a similarity evaluation method based on multiple data types. In the merging process of two contact data, if the data under multiple data types are consistent, they are directly merged into new contact data; if there are differences, similarity calculation is performed to determine whether merging is needed and how to merge, combining the similarity of the difference data group and the same data group. This intelligent merging method based on similarity can effectively handle the case of incomplete matching in different data sources, improve the merging accuracy, and reduce the risk of false merging or missing merging. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or related technical description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0072] Figure 1A schematic flow chart of a contact recovery method based on multi-source heterogeneous data collaborative processing is shown.

[0073] Figure 2 A schematic diagram of a contact recovery device based on multi-source heterogeneous data collaborative processing is shown.

[0074] Figure 3 A schematic diagram of a terminal device is shown. DETAILED DESCRIPTION

[0075] In the following description, for the purposes of explanation and not limitation, specific details are set forth, such as particular sequences of steps, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0076] The embodiments of the present application provide a contact recovery method and device based on multi-source heterogeneous data collaborative processing, to solve the technical problem of low accuracy and reliability of processing results in the case of complex data types and poor data quality in the traditional scheme.

[0077] Firstly, the present application provides a contact recovery method based on multi-source heterogeneous data collaborative processing. Please refer to Figure 1 , Figure 1 A schematic flow chart of a contact recovery method based on multi-source heterogeneous data collaborative processing is shown. As shown in Figure 1 The contact recovery method based on multi-source heterogeneous data collaborative processing can include the following steps:

[0078] Step 101: Obtain multi-source heterogeneous data, and standardize the multi-source heterogeneous data to obtain a multi-source contact data set; wherein the data type of each contact data in the multi-source contact data set includes a phone number, a name, an email address, a company name, a birthday, and / or a home address;

[0079] Collect multi-source heterogeneous data from multiple different sources (such as mobile phone address books, local backup data, third-party application data, and social media, etc.), which may have different formats and contain different types of contact information.

[0080] The data type of each contact data in the multi-source contact data set includes but is not limited to a phone number, a name, an email address, a company name, a birthday, and / or a home address.

[0081] Due to the inconsistent data formats from different sources, the standardization process aims to convert all data into a unified format. This facilitates subsequent processing and comparison, ensuring that data from different sources can be seamlessly integrated into a unified data set.

[0082] Name normalization: remove non-alphabet characters (such as removing spaces, punctuation, special characters, etc.), convert Chinese to pinyin, and convert to lowercase uniformly.

[0083] Address normalization: remove non-key information (such as removing house numbers, floors, and other easily changing information), standardize regional names, convert Chinese to pinyin, and convert to lowercase uniformly.

[0084] Birthday normalization: format standardization (such as converting to YYYY-MM-DD format), convert to timestamp (convert date to Unix timestamp for numerical comparison).

[0085] Step 102: In the multi-source contact data set, match the same data under each data type;

[0086] In multiple data sources, the system checks whether there is the same contact information (such as phone number, name, email address, company name, birthday, or home address) under each data type.

[0087] Step 103: Obtain two first contact data corresponding to the same data;

[0088] After matching the same data, the system extracts two contact records matching the data type from multiple data sources, which may come from different data sources.

[0089] Step 104: When the data under multiple data types in the two first contact data are the same, merge the two first contact data to obtain second contact data;

[0090] If the data under multiple data types of the two contact data are completely consistent (for example, the phone number, name, email address, company name, birthday, and home address in the two data sources are completely the same), the two data can be directly merged into one contact data (that is, randomly remove any one contact data), that is, form the second contact data.

[0091] Step 105: When the data in multiple data types in the two first contact data are not uniform, according to the first similarity corresponding to the difference data group and the second similarity corresponding to the same data group in the multiple data types, the two first contacts are merged to obtain third contact data; wherein the difference data group refers to two data with differences in the same data type, and the same data group refers to a data group with the same data in the same data type.

[0092] If two contact data have differences in some data types, these data are classified as "difference data groups".

[0093] If some data types are the same in two contact data (such as the same phone number), these data belong to "same data groups".

[0094] By calculating the similarity corresponding to the difference data group and the same data group, it is determined whether the two contact data belong to the same person. If the overall difference is small and the similarity is high, the two contact data can be merged to form third contact data.

[0095] This processing mode considers the case of incomplete matching of data, and determines whether to merge by calculating the similarity.

[0096] Specifically, step 105 specifically includes steps 1051 to 1056:

[0097] Step 1051: Extracting the difference data group and the same data group in the multiple data types corresponding to the two first contact data;

[0098] First, it is necessary to identify which data types are the same (for example, the phone numbers of the two data sources are the same) and which data types are different (for example, the names of the two data sources are different) from the two contact data. These data will be classified as "same data group" and "difference data group" respectively.

[0099] Same data group: the content in these data types is completely consistent and can be considered the same.

[0100] Difference data group: the content in these data types has differences and needs to be further compared and processed.

[0101] Step 1052: According to the data type corresponding to the difference data group, calculating the first similarity between the two current data in the difference data group;

[0102] For each pair of data in the difference data group, the system calculates their similarity. The calculation of the similarity can use different algorithms. This similarity value will be used to measure the degree of similarity of two data items in different data types.

[0103] Specifically, step 1052 specifically comprises steps A1 to A4:

[0104] Step A1: when the data type corresponding to the difference data set is a phone number, a birthday, or an email address, a first similarity is calculated according to the number of consecutive identical digits in the difference data set;

[0105] All digits in the two data are extracted and compared one by one. The number of identical digits is calculated. The number of identical digits is divided by the total number of digits to obtain a normalized similarity value (0-1).

[0106] Specifically, step A1 specifically comprises steps A11 to A18:

[0107] Step A11: when the data type corresponding to the difference data set is a phone number or a birthday, a first number of consecutive identical digits in the two current data is extracted; wherein the first number of consecutive identical digits refers to the number of consecutive identical digits in the two current data;

[0108] For phone number or birthday data, first, the consecutive identical digits in the two data are found. This matching of consecutive digits can reflect the similarity of the two.

[0109] Step A12: a second number of digits in the two current data is counted respectively;

[0110] The digits in each data are counted, that is, the digits in each data are counted.

[0111] Step A13: the first number is divided by the maximum second number to obtain the first similarity;

[0112] The matching of consecutive identical digits can effectively handle the partial similarity problem of the phone number or the birthday, especially in the case where the two data are similar but not completely consistent, the continuity of the digits is used to improve the accuracy.

[0113] Step A14: if the data type corresponding to the difference data set is an email address, a prefix and a suffix in the email address are extracted;

[0114] When processing an email, first, the suffix part (i.e., the content after “@”) of the two email addresses is compared. If the suffixes of the two emails are different, the similarity is directly set to 0 (no further comparison is performed).

[0115] Step A15: If the two suffixes are the same, extract a third number of consecutive identical digits in the two prefixes; wherein the third number of consecutive identical digits refers to the number of consecutive identical digits in the two prefixes;

[0116] The comparison of digits is performed in the prefix part of the mailbox address, the consecutive identical digits (e.g., the number part of the mailbox has consecutive identical characters) are extracted, and the length of the consecutive digits is counted.

[0117] Step A16: Count the fourth number of digits in the two prefixes respectively;

[0118] Step A17: Divide the third number by the maximum fourth number to obtain the first similarity;

[0119] Step A18: If the two suffixes are not the same, determine the first similarity as 0.

[0120] The suffixes are directly determined as 0, because if the domain name parts of the two mailbox addresses are different, it means that they are likely to belong to completely different service providers or users, and should not be considered as the same user.

[0121] For the consecutive digit matching of the prefix part, this way can better handle the case where the mailbox has similar number parts but different other parts. For example, some mailbox addresses may only differ in the number suffix, but in fact they are the mailbox of the same person.

[0122] In the embodiments corresponding to steps A11 to A18, different similarity calculation strategies are adopted for the three types of data: phone number, birthday, and email address. The accuracy of consecutive identical digit matching is improved, and corresponding rules are designed according to the different characteristics of the data types. By comparing the maximum number, it avoids false judgments caused by partial digit repetition or different positions. Overall, this method can effectively improve the accuracy of contact data recovery, especially when dealing with similar format or content of digital data, it has higher fault tolerance and precision.

[0123] Step A2: When the data type corresponding to the difference data set is a name, calculate a first similarity according to the pinyin corresponding to the name;

[0124] Convert the Chinese character name to pinyin. Use the string similarity algorithm between pinyin strings to compare. The distance value obtained can be normalized into a similarity (1 indicates complete identity, 0 indicates complete difference).

[0125] Chinese character names may have different shapes but similar pronunciations. Using pinyin can avoid misjudgment of "different shapes with the same pronunciation" and is more in line with the habit of name comparison in the Chinese context.

[0126] Specifically, step A2 specifically includes steps A21 to A24:

[0127] Step A21: when the data type corresponding to the difference data set is a name, obtaining the pinyin corresponding to the name;

[0128] Chinese to Pinyin processing is performed on the two names respectively (such as using the open source library pypinyin), and Chinese characters are converted into corresponding pinyin strings.

[0129] After using pinyin to represent, the situation of homophonic different characters can be handled, and the semantic tolerance is improved. A unified format is laid for subsequent comparison.

[0130] Step A22: if the two Chinese names corresponding to the difference data set are the same, a second preset value is taken as the first similarity;

[0131] If the two names are completely consistent (literally consistent), a high preset value, for example 1.0, is directly assigned, indicating that they are completely the same.

[0132] Accurate matching has the highest priority, and its similarity is directly confirmed as the highest value. It can improve matching efficiency without additional calculation.

[0133] Step A23: if the two Chinese names corresponding to the difference data set are different, and the two pinyins corresponding to the difference data set are the same, a third preset value is taken as the first similarity;

[0134] If the pinyins of the two names are the same, but the literal is different (may be homophonic miswriting, simplified and traditional Chinese characters, variant characters, etc.), a higher but lower than the complete match "third preset value" (such as 0.6) is assigned.

[0135] The same pinyin indicates the same pronunciation, which may be:

[0136] Homophonic different characters (commonly seen in input errors or nickname writing), simplified and traditional Chinese character difference or common alternative characters. It is not directly equivalent to complete match, but should be determined as "strong correlation".

[0137] Step A24: if the two Chinese names corresponding to the difference data set are different, and the two pinyins corresponding to the difference data set are different, 0 is taken as the first similarity.

[0138] If the pinyins of the two names are also different (i.e. the shapes and pronunciations are different), it is considered that the similarity between them is 0, that is, there is no semantic association.

[0139] The pinyin and the character are different, and there is no similarity. It can be directly skipped for further matching, saving resources.

[0140] In the embodiments corresponding to steps A21 to A24, the regularization branch is used to avoid unnecessary complex similarity calculation; the pinyin comparison is introduced to tolerate errors caused by input methods; the fuzzy recognition processing of Chinese name data in the contact recovery process is applicable; and through the setting of the "second preset value" and the "third preset value", the matching sensitivity can be flexibly adjusted.

[0141] Step A3: When the data type corresponding to the difference data set is a home address, a first similarity is calculated according to a word distribution feature corresponding to the home address.

[0142] Chinese word segmentation is performed on the two address texts to extract key place words. The similarity or coincidence rate of the words in the semantic space is analyzed.

[0143] Chinese addresses have different expression methods (redundancy, omission, and sequence change). Word segmentation and word distribution models can effectively identify addresses that have the same semantics but different expressions.

[0144] Specifically, step A3 specifically includes steps A31 to A34:

[0145] Step A31: When the data type corresponding to the difference data set is a home address, the home address is subjected to word segmentation processing to obtain a plurality of geographical words.

[0146] Since a home address usually contains multiple geographical information (such as province, city, district, street, etc.), it is appropriate to use word segmentation to subdivide and process each geographical word.

[0147] The home address is subjected to word segmentation processing to cut the address string into multiple individual geographical words. For example:

[0148] Address: "AA City BB District CC Road xx No."

[0149] Word segmentation result: ["AA City", "BB District", "CC Road", "xx No."]

[0150] Through word segmentation, the geographical information (such as city, area, street name, etc.) in the address is separated out, which facilitates subsequent word matching. This step is equivalent to semantic segmentation of the address, which improves the accuracy of comparison.

[0151] Specifically, step A31 specifically includes steps A311 to A319:

[0152] Step A311: When the data type corresponding to the difference data set is a home address, geographical division words in the address are extracted; the geographical division words include country, province, city, district, street, road, community, building, unit, and room.

[0153] Geographical division words refer to keywords in a household address that can indicate different geographical levels, such as country, province, city, etc. Specifically, they include country, province, city, district, street, road, community, building, unit, room, and number.

[0154] These words help clarify different areas and levels in an address. For example, in the address "AA City BB District CC Road XX Number", the extracted geographical division words include city, district, road, and number.

[0155] By extracting these specific geographical division words, the household address can be structured and processed hierarchically, making subsequent segmentation and matching more accurate.

[0156] Step A312: Based on the position of the geographical division words in the household address, the household address is segmented into multiple strings;

[0157] The goal of segmentation is to obtain the intermediate part connecting the geographical division words, further preparing for segmentation.

[0158] Step A313: Take the string and the adjacent geographical division word on the right side of the string as the initial word;

[0159] After extracting the string between adjacent geographical division words, the string and the adjacent geographical division word on the right side are taken as the initial word.

[0160] This method helps accurately extract core words in the address, avoiding missing key geographical information in the address during simple segmentation.

[0161] Step A314: According to the original order of multiple initial words in the address, obtain the left initial word and the right initial word of each current initial word; wherein the left initial word refers to the adjacent initial word on the left side of the current initial word, and the right initial word refers to the adjacent initial word on the right side of the current initial word;

[0162] For each initial word, its left and right initial words are obtained, which is to maintain the integrity of address information during segmentation. For example:

[0163] For the initial word "AA District", the left word is "BB City" and the right word is "CC Road".

[0164] The left and right adjacent initial words help maintain the context information of segmentation, ensuring the reasonableness of segmentation and providing more accurate context for subsequent judgment.

[0165] Step A315: Combining each current initial word, the left initial word corresponding to each current initial word, and the right initial word corresponding to each current initial word into a to-be-recognized word group according to the original order in the address;

[0166] Combining each initial word and its left and right adjacent words into a to-be-recognized word group according to the original address order.

[0167] Combining words according to the original order of the address helps to ensure that the words can maintain their actual meaning and positional relationship in subsequent processing, avoiding misinterpretation of the address.

[0168] Step A316: Inputting the to-be-recognized word group into a word marking model to obtain a word marking corresponding to each character output by the word marking model; wherein the word marking includes a word beginning, a word middle, a word end, and a single word;

[0169] Inputting the to-be-recognized word group into a word marking model, the word marking model will mark each character to determine the role of the character in the word. The word marking model includes but is not limited to traditional processing methods such as recurrent neural networks.

[0170] Word marking includes:

[0171] Word beginning: the beginning part of the word.

[0172] Word middle: the middle part of the word.

[0173] Word end: the end part of the word.

[0174] Single word: a character itself is an independent word.

[0175] Using a word marking model can determine the specific grammatical role of each character according to the context, ensuring that the characters in the original address are correctly classified as geographical words.

[0176] Step A317: According to the distribution position of each character in the current initial word, judging whether each character meets the word marking;

[0177] According to the word marking, judging whether each character meets the position requirement in the initial word. If the character meets the word marking, the word can be considered as a geographical word.

[0178] For example, if "AB Road" meets the word beginning, word middle, and word end marking rules, "AB Road" is considered as a valid geographical word.

[0179] Since the words obtained by geographical division may have some interference (original data has errors or sequence errors, etc.), it is necessary to verify the correctness through the word tagging model, and to ensure the correctness of the word segmentation by verifying whether each character meets the word tagging, avoiding information loss or mismatch caused by incorrect word segmentation.

[0180] Step A318: If the characters in the current initial word meet the word tagging, the current initial word is taken as the geographical word.

[0181] Step A319: If the characters in the current initial word do not meet the word tagging, the current initial word is divided based on the word tagging to obtain the geographical word.

[0182] If the characters in the current initial word do not completely meet the word tagging requirements, the initial word is re-divided according to the word tagging until a geographical word that meets the tagging rules is obtained.

[0183] This step ensures the accuracy of word segmentation. Even if the characters of the initial word do not completely meet the tagging requirements, a reasonable geographical word can be obtained through re-division, avoiding the appearance of non-standard word segmentation results.

[0184] In the embodiments corresponding to steps A311 to A319, the accuracy of word segmentation is ensured by carefully extracting the positions of geographical division words and division characters. Through precise control of adjacent words and positions, the original information of the address can be maintained as much as possible during word segmentation. Through the word tagging model and the subsequent division mechanism, it is ensured that inappropriate words are corrected, improving the flexibility and robustness of word segmentation.

[0185] Step A32: Match the number of identical words between the multiple geographical words corresponding to the two current data in the difference data set.

[0186] The geographical words of the two family addresses are compared one by one, and the number of their identical geographical words is calculated. For example:

[0187] Data 1: "AA City BB District CC Road XX No." → Word segmentation result: ["AA City", "BB District", "CC Road", "XX No."]

[0188] Data 2: "AA City BB District CC Road YY No." → Word segmentation result: ["AA City", "BB District", "CC Road", "YY No."]

[0189] In this case, the number of identical words is 3 ("AA City", "BB District", "CC Road").

[0190] The home address often contains multiple similar geographical terms (e.g. the same city, the same area), so by comparing the number of the same terms, the similarity of the two addresses can be accurately reflected.

[0191] This step focuses on the commonality of geographical terms to calculate the similarity of the two addresses.

[0192] Step A33: Count the number of current terms in the two current data respectively.

[0193] For the segmentation results of the two home addresses, the total number of terms in each address is calculated respectively.

[0194] Data 1: "AA City BB District CC Road 88" → Segmentation result: ["AA City", "BB District", "CC Road", "88"] → Current term number is 4

[0195] Data 2: "AA City BB District CC Road 99" → Segmentation result: ["AA City", "BB District", "CC Road", "99"] → Current term number is 4

[0196] In this example, the number of terms of the two addresses is the same, both 4.

[0197] Counting the number of terms in each address helps to evaluate the baseline when assessing similarity. Because the address with fewer terms, if most of the terms are the same, can also show a high degree of similarity.

[0198] This step can reflect the influence of address length and complexity, avoiding errors between long and short addresses.

[0199] Step A34: Divide the number of the same terms by the maximum number of current terms to get the first similarity.

[0200] The number of the same terms reflects the degree of overlap of the two addresses in geographical information, and the maximum term number as a standard can effectively eliminate the influence of length difference between addresses.

[0201] Using this way, the similarity between home addresses can be clearly quantified, avoiding errors that may be introduced by simple string comparison methods.

[0202] In the embodiments corresponding to steps A31 to A34, by segmentation processing and comparison of the number of the same terms, errors that may be introduced by direct string comparison are avoided, ensuring that only geographical terms are compared. It is suitable for various formats of home address data, even if there are different writing methods or orders, it can better evaluate the similarity. The maximum term number as a standard effectively reduces the influence of length difference on comparison.

[0203] Step A4: When the data type corresponding to the difference data set is a company name, a first similarity is calculated according to the character distribution characteristics of the company name.

[0204] It is worth noting that, due to the strong geographical division of the home address and the hierarchical dependence, and the characteristics of the company name such as industry classification, regional information or company type, there is a big difference in the similarity calculation of the home address and the company name, and a more suitable way needs to be combined with the characteristics of each other to calculate, so in this embodiment, two ways are used to calculate the first similarity of the home address and the company name.

[0205] The company name often contains industry, regional or registered suffix, and the character distribution is more easy to capture the similarities and differences of the name structure. It can adapt to the comparison needs of various variant naming forms.

[0206] Specifically, step A4 specifically includes: if the data type corresponding to the difference data set is a company name, input the string of the two current data corresponding to the difference data set into the similarity model to obtain the first similarity output by the similarity model.

[0207] The calculation of the first similarity is directly related to whether the company names in different data sources can be accurately classified into the same entity next. This is one of the key steps in the company name recovery process, because the similarity model provides a quantitative basis for subsequent data matching, merging or correction.

[0208] The similarity model is:

[0209]

[0210]

[0211]

[0212]

[0213] wherein, the first similarity, the nonlinear prefix weighting function, the number of characters of the first current data in the difference data set, the number of characters of the second current data in the difference data set, the maximum value, the number of characters with the same prefix in the two strings, the dynamic penalty factor, the character similarity, the position difference of the same string in the two current data of the high correlation character, represents the number of high correlation characters, which means the same string existing in both current data and the difference of position is no more than / 2-1.

[0214] is the final similarity value between two company names, ranging from 0 to 1, the value closer to 1 means more similar. represents the preliminary similarity of two strings, which is obtained by calculating the number of character matches and the difference in the order of characters.

[0215] The preliminary similarity calculated is adjusted by a nonlinear prefix weighting function and a dynamic penalty factor. The preliminary similarity calculated is adjusted by a nonlinear prefix weighting function and a dynamic penalty factor.

[0216] In the similarity model The weight of prefix matching is nonlinearly increased by the ratio of prefix length and its square. Longer prefix matching will have a greater impact on the overall similarity. Company names usually have longer prefixes, so prefix matching is particularly important for judging similarity. By square weighting, longer prefix matching will have a greater impact on the overall similarity.

[0217] The penalty factor is used to adjust the penalty degree of the unmatched part in the similarity calculation. The penalty factor is dynamically calculated according to the length of the two company names. The mismatch difference between long names will give greater penalty. The purpose of this penalty factor is to adjust for the length difference of company names. Longer names (especially when there are large differences in some parts) will increase the penalty factor, thus increasing the penalty in similarity calculation. The longer the company name, the greater the impact of the penalty factor, so that large differences (such as longer names compared to shorter names) can be properly amplified.

[0218] In modern company names, there are often industry classification, regional information, company type (such as "Limited Company", "Group") and so on. By introducing weighted matching distance and character weight, the matching accuracy can be adjusted according to the industry, region, etc. in the company name. For specific fields or industries, the improved algorithm can be flexibly adjusted. For example, technology companies, financial companies, manufacturing companies, etc. usually have fixed industry vocabulary (such as "technology", "shares"), and the improved algorithm can identify and weight these industry vocabulary through character weighting, thereby enhancing industry adaptability.

[0219] In the embodiments corresponding to steps A1 to A4, by setting "customized similarity calculation rules" for different data types, the semantic judgment ability and comparison accuracy in the contact recovery process are improved, especially in the face of ambiguity, variants, typos or format differences, the intelligent processing advantage can be better reflected. This approach also lays the foundation for subsequent machine learning model access and accuracy improvement.

[0220] Step 1053: The first preset value is used as the second similarity of the same data group; wherein the first preset value includes 1;

[0221] For the same data group (i.e. the part of the data is completely consistent), there is no need to calculate the similarity, and a fixed value is directly given as the similarity. This fixed value is set to 1, indicating that the similarity of completely consistent data is full marks.

[0222] Step 1054: According to the weight coefficients corresponding to the plurality of data types, the first similarity or the second similarity corresponding to each of the plurality of data types is weighted and summed to obtain a final similarity;

[0223] Different data types may have different effects on the final judgment result. Therefore, each data type (such as name, phone number, email, etc.) may be assigned different weight coefficients. Some data types (such as phone numbers) are more important, so they may be assigned higher weights.

[0224] According to the similarity of each data type and the corresponding weight coefficient, the final weighted similarity is calculated. In this way, the final similarity takes into account the importance of different data types.

[0225] Step 1055: If the final similarity is greater than a first threshold, the channel priority of the two first contact data is obtained;

[0226] If the final similarity after weighting exceeds the first preset threshold, it is considered that the two contact data are similar enough and can be merged.

[0227] A threshold is set to ensure that only when the similarity is high enough, the two contact data are considered to be the same person, avoiding false merging due to low similarity.

[0228] Step 1056: The first contact data corresponding to the maximum channel priority of the two first contact data is used as the third contact data, and the other first contact data is excluded.

[0229] In some cases, contact data from different sources can have different reliability. The concept of "channel priority" is introduced here to determine which of the two contact data to choose as the final contact data. For example, some data sources may be more reliable, or a data source may provide more complete information.

[0230] In this step, according to the value of channel priority, it is determined which contact data to keep. The data with higher channel priority will be selected as the final contact data (i.e. the third contact data), while the data with lower channel priority will be discarded.

[0231] In the corresponding embodiments of steps 1051 to 1056, it is refined how to consider similarity calculation and channel priority in the case of incomplete matching between two contact data, to make the decision of whether to merge and finally select which contact data as the recovered contact data. By introducing the mechanism of similarity calculation, weighted summation and channel priority, it can ensure that the contact recovery process is more intelligent and accurate, thereby improving the quality and reliability of the final recovery result.

[0232] Step 106: The second contact data, the third contact data and the remaining contact data are taken as the final contact data and synchronized to the target device; wherein the remaining contact data refers to the contact data in the multi-source contact data set that has not been processed by merging.

[0233] The final contact data includes the second contact data after complete matching and merging, the third contact data after similarity high merging, and those contact data that have not been processed by any merging (i.e. the remaining contact data).

[0234] Once these data are processed and merged, the system will synchronize the final contact data to the target device (such as mobile phone, computer, etc.). This ensures the synchronization and consistency of contact data between multiple devices.

[0235] The remaining contact data refers to those contact data that have not been merged in the entire process. Usually these data do not have enough information or similarity to be merged with other data, so they remain as independent contact data.

[0236] In the embodiments corresponding to steps 101 to 106, since the current contact data is stored in multiple heterogeneous data sources including mobile phone address book, social media, email system, etc., each data source uses different formats, structures and storage methods, which brings great challenges to the management and recovery of contact data. The present application converts multiple source heterogeneous data into a unified standard format through standardization processing, forms a multi-source contact data set, and realizes cross-platform and cross-device data integration, providing a standardized and unified data basis for subsequent contact recovery. In the multi-source contact data set, the same contact in different data sources can be effectively identified by matching the same data under each data type (such as phone number, name, email address, etc.), reducing the possibility of data loss or redundancy. Especially in the case of multiple similar contact records in different data sources, the accuracy of contact matching is improved by accurately matching the same data of each data type, avoiding repeated and erroneous data merging. The present application proposes a similarity evaluation method based on multiple data types. In the merging process of two contact data, if the data under multiple data types are consistent, they are directly merged into new contact data; if there are differences, similarity calculation is performed to determine whether merging is needed and how to merge by combining the similarity of the difference data group and the same data group. This intelligent merging method based on similarity can effectively handle the case of incomplete matching in different data sources, improve the merging accuracy and reduce the risk of false merging or missing merging.

[0237] As Figure 2 The present application provides a contact recovery device based on multi-source heterogeneous data collaborative processing, please see Figure 2 , Figure 2 A contact recovery device based on multi-source heterogeneous data collaborative processing provided by the present application is shown in the schematic diagram, as Figure 2 shown, a contact recovery device based on multi-source heterogeneous data collaborative processing includes:

[0238] The first acquisition unit 21 is used for acquiring multi-source heterogeneous data and standardizing the multi-source heterogeneous data to obtain a multi-source contact data set; wherein the data type of each contact data in the multi-source contact data set includes phone number, name, email address, company name, birthday and / or home address;

[0239] The matching unit is used for matching the same data under each data type in the multi-source contact data set;

[0240] The second acquisition unit 22 is used for acquiring two first contact data corresponding to the same data;

[0241] The first merging unit 23 is configured to merge the two first contact data into second contact data when the data of the multiple data types in the two first contact data are all the same.

[0242] The second merging unit 24 is configured to merge the two first contact data into third contact data according to the first similarity corresponding to the difference data group and the second similarity corresponding to the same data group in the multiple data types when the data of the multiple data types in the two first contact data are not all the same; wherein the difference data group refers to two data of the same data type that are different, and the same data group refers to a data group of the same data type that has the same data.

[0243] The synchronization unit 25 is configured to synchronize the second contact data, the third contact data and the remaining contact data as the final contact data to a target device; wherein the remaining contact data refers to the contact data in the multi-source contact data set that has not been processed by merging.

[0244] The contact recovery device based on multi-source heterogeneous data collaborative processing provided by the application can effectively identify the same contact in different data sources by matching the same data under each data type (such as a telephone number, a name, an email address, etc.), thereby reducing the possibility of data loss or redundancy. Especially in the case where there are multiple similar contact records in different data sources, the accuracy of contact matching is improved by accurately matching the same data of each data type, and repeated and erroneous data merging is avoided. The application provides a similarity evaluation method based on multiple data types. In the merging process of two contact data, if the data of multiple data types are consistent, the two contact data are directly merged into new contact data; if the data are different, similarity calculation is performed, and the similarity of the difference data group and the same data group is combined to determine whether merging is needed and how to merge. This intelligent merging method based on similarity can effectively handle the case where there is incomplete matching in different data sources, improve the merging accuracy, and reduce the risk of false merging or missed merging.

[0245] Figure 3 is a schematic diagram of a terminal device provided by an embodiment of the application. As shown in FIG. 1, the terminal device includes a processor 10, a memory 20 and a communication interface 30. Figure 3As shown, the terminal device 3 of this embodiment includes a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a contact recovery program based on collaborative processing of multi-source heterogeneous data. The processor 30 implements the steps in each of the above contact recovery methods based on collaborative processing of multi-source heterogeneous data when executing the computer program 32, such as Figure 1 Steps 101 to 106 shown above. Alternatively, the processor 30 implements the functions of each unit in each of the above apparatus embodiments when executing the computer program 32, such as Figure 2 The functions of the units shown above.

[0246] For example, the computer program 32 can be divided into one or more units stored in the memory 31 and executed by the processor 30 to complete the present application. The one or more units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 32 in the terminal device 3. For example, the computer program 32 can be divided into units with specific functions as follows:

[0247] A first acquisition unit configured to acquire multi-source heterogeneous data and standardize the multi-source heterogeneous data to obtain a multi-source contact data set; wherein the data types of each contact data in the multi-source contact data set include a phone number, a name, an email address, a company name, a birthday, and / or a home address;

[0248] A matching unit configured to match the same data under each data type in the multi-source contact data set;

[0249] A second acquisition unit configured to acquire two first contact data corresponding to the same data;

[0250] A first merging unit configured to merge the two first contact data to obtain second contact data when the data under multiple data types in the two first contact data are all the same;

[0251] A second merging unit configured to merge the two first contact data to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in multiple data types when the data under multiple data types in the two first contact data are not all the same; wherein the difference data group refers to two data that differ under the same data type, and the same data group refers to a data group having the same data under the same data type;

[0252] The synchronization unit is configured to synchronize the second contact data, the third contact data and the remaining contact data as final contact data to the target device, wherein the remaining contact data refers to contact data in the multi-source contact data set which has not been processed by the merging.

[0253] The terminal device includes, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that, Figure 3 The terminal device 3 is only an example and does not constitute a limitation on the terminal device 3, and can include more or fewer components than shown, or combine certain components, or different components, for example, the terminal device can also include an input / output device, a network access device, a bus, etc.

[0254] The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0255] The memory 31 can be an internal storage unit of the terminal device 3, for example, a hard disk or a memory of the terminal device 3. The memory 31 can also be an external storage device of the terminal device 3, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 31 can include both the internal storage unit and the external storage device of the terminal device 3. The memory 31 is used to store the computer program and other programs and data required by the terminal device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0256] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0257] It should be noted that the information interaction, execution process and the like between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and the specific functions and brought technical effects can be referred to the method embodiments part, and will not be repeated here.

[0258] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs. The internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific name of each functional unit and module is only for convenient distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0259] The embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps in each method embodiment.

[0260] The embodiment of the present application provides a computer program product, when the computer program product runs on a mobile terminal, so that the mobile terminal executes to realize the steps in each method embodiment.

[0261] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods, which can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer readable storage medium, and the computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can at least include any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, such as U disk, mobile hard disk, magnetic disk or optical disk.

[0262] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0263] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0264] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / network device and method can be implemented in other ways. For example, the above-described apparatus / network device embodiments are merely schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0265] The units described as separate components can or can not be physically separate, and the components displayed as separate components can or can not be physical separate, and can be located in one position or distributed on a plurality of network units.

[0266] It should be understood that the term "comprises" or "comprising," when used in this specification and the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0267] It should also be understood that the term "and / or" when used in this specification and the following claims, refers to one or more of the associated listed items and can be used interchangeably with "or / or".

[0268] As used in this specification and the appended claims, the term "if' can be interpreted as meaning "when," or "once," or "in response to determining," or "in response to ascertaining," depending on the context. Similarly, the phrase "if it is determined" or "if it is ascertained" can be interpreted to mean "once it is determined," or "in response to determining," or "once it is ascertained," or "in response to ascertaining," depending on the context.

[0269] In addition, the terms "first," "second," "third," etc. as used in the description of the application and the following claims are not used to denote or imply relative importance and are merely identify the names of common elements.

[0270] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in some embodiments" or "in other embodiments" or "in still other embodiments," or the like in various places throughout this specification are not necessarily all referring to the same embodiment, unless otherwise indicated. Furthermore, the term "comprising" or "containing" or "including" or "having" or the like as used herein is specifically intended to mean "including, but not limited to."

[0271] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A contact recovery method based on multi-source heterogeneous data collaborative processing, characterized in that, The contact recovery method based on multi-source heterogeneous data collaborative processing comprises: Obtaining multi-source heterogeneous data, and standardizing the multi-source heterogeneous data to obtain a multi-source contact data set; wherein the data type of each contact data in the multi-source contact data set comprises a phone number, a name, an email address, a company name, a birthday, and / or a home address; Matching the same data under each data type in the multi-source contact data set; Obtaining two first contact data corresponding to the same data; When the data under multiple data types in the two first contact data are not uniform, merging the two first contacts to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in the multiple data types; wherein the difference data group refers to two data having differences under the same data type, and the same data group refers to a data group having the same data under the same data type; The second contact data, the third contact data, and the remaining contact data are regarded as final contact data, and are synchronized to a target device; wherein the remaining contact data refers to contact data in the multi-source contact data set that has not been processed by merging; The method comprises: Extracting a difference data group and a same data group in the multiple data types corresponding to the two first contact data; Calculating a first similarity between two current data in the difference data group according to the data type corresponding to the difference data group; A first preset value is used as the second similarity of the same data group; wherein the first preset value comprises 1; According to the weight coefficient corresponding to the multiple data types, the first similarity or the second similarity corresponding to each of the multiple data types is weighted and summed to obtain a final similarity; If the final similarity is greater than a first threshold value, obtaining a channel priority of the two first contact data; The first contact data corresponding to the maximum channel priority in the two first contact data is used as the third contact data, and the other first contact data is excluded; The method comprises: When the data type corresponding to the difference data group is a phone number, a birthday, or an email address, calculating the first similarity according to the number of consecutive identical digits in the difference data group; When the data type corresponding to the difference data group is a name, calculating the first similarity according to the pinyin corresponding to the name. ​ When the data type corresponding to the difference data set is a family address, a first similarity is calculated according to a word distribution feature corresponding to the family address; When the data type corresponding to the difference data set is a company name, a first similarity is calculated according to a character distribution feature of the company name.

2. The contact recovery method based on multi-source heterogeneous data collaborative processing according to claim 1, characterized in that, The step of calculating the first similarity according to the number of continuous same digits in the difference data set when the data type corresponding to the difference data set is a telephone number, a birthday or an email address comprises: When the data type corresponding to the difference data set is a telephone number or a birthday, a first number of continuous same digits in the two current data is extracted; wherein the first number of continuous same digits refers to the number of continuous same digits in the two current data; A second number of digits in the two current data is counted respectively; The first number is divided by the maximum second number to obtain the first similarity; When the data type corresponding to the difference data set is an email address, a prefix and a suffix in the email address are extracted; When the two suffixes are the same, a third number of continuous same digits in the two prefixes is extracted; wherein the third number of continuous same digits refers to the number of continuous same digits in the two prefixes; A fourth number of digits in the two prefixes is counted respectively; The third number is divided by the maximum fourth number to obtain the first similarity; When the two suffixes are different, the first similarity is determined as 0.

3. The contact recovery method based on multi-source heterogeneous data collaborative processing according to claim 1, characterized in that, The step of calculating the first similarity according to the pinyin corresponding to the name when the data type corresponding to the difference data set is a name comprises: When the data type corresponding to the difference data set is a name, the pinyin corresponding to the name is obtained; When the two Chinese names corresponding to the difference data set are the same, a second preset value is taken as the first similarity; When the two Chinese names corresponding to the difference data set are different and the two pinyins corresponding to the difference data set are the same, a third preset value is taken as the first similarity; When the two Chinese names corresponding to the difference data set are different and the two pinyins corresponding to the difference data set are different, 0 is taken as the first similarity.

4. The contact recovery method based on multi-source heterogeneous data collaborative processing according to claim 1, characterized in that, The step of calculating the first similarity according to the word distribution feature corresponding to the family address when the data type corresponding to the difference data set is a family address comprises: When the data type corresponding to the difference data set is a family address, the family address is subjected to word segmentation processing to obtain a plurality of geographical words; The number of same words between the plurality of geographical words corresponding to the two current data in the difference data set is matched; The number of current words in the two current data is counted respectively; The number of same words is divided by the maximum number of current words to obtain the first similarity.

5. The contact recovery method based on multi-source heterogeneous data collaborative processing according to claim 4, characterized in that, The step of subjecting the family address to word segmentation processing to obtain a plurality of geographical words when the data type corresponding to the difference data set is a family address comprises: When the data type corresponding to the difference data set is a family address, geographical division words in the address are extracted; the geographical division words comprise country, province, city, district, street, road, community, building, unit and room. Split the family address into multiple strings based on the location of the geographical division words in the family address; Take the string and the adjacent geographical division word located to the right of the string as initial words; According to the original order of the multiple initial words in the address, obtain the left initial word and the right initial word of each current initial word; wherein the left initial word refers to the adjacent initial word located to the left of the current initial word, and the right initial word refers to the adjacent initial word located to the right of the current initial word; Combine each current initial word, the left initial word corresponding to each current initial word, and the right initial word corresponding to each current initial word into a to-be-recognized word group according to the original order in the address; Input the to-be-recognized word group into a word marking model to obtain the word marking corresponding to each character output by the word marking model; wherein the word marking includes the beginning of a word, the middle of a word, the end of a word, and a single word; Determine whether each character in the current initial word conforms to the word marking according to the distribution position of each character in the current initial word; If all the characters in the current initial word conform to the word marking, take the current initial word as the geographical word; If the characters in the current initial word do not all conform to the word marking, divide the current initial word based on the word marking to obtain the geographical word.

6. The contact recovery method based on multi-source heterogeneous data collaborative processing according to claim 1, characterized in that, When the data type corresponding to the difference data group is a company name, the step of calculating the first similarity based on the character distribution characteristics of the company name includes: If the data type corresponding to the difference data group is a company name, input the string of the two current data corresponding to the difference data group into a similarity model to obtain the first similarity output by the similarity model; The similarity model is: wherein, denotes the first similarity, denotes a non-linear prefix weighting function, denotes the number of characters of the first current data in the difference data set, denotes the number of characters of the second current data in the difference data set, denotes taking the maximum value, denotes the number of characters with identical prefix in the two strings, denotes a dynamic penalty factor, denotes the character similarity, denotes the position difference of the identical string in the two current data of the high correlation character, denotes the number of high correlation characters, which are identical strings in the two current data and the position difference is no more than / 2-1.

7. A contact recovery device based on multi-source heterogeneous data collaborative processing, characterized in that, The contact person recovery device based on collaborative processing of multi-source heterogeneous data includes: A first acquisition unit is configured to acquire multi-source heterogeneous data, and standardize the multi-source heterogeneous data to obtain a multi-source contact person data set; wherein the data type of each contact person data in the multi-source contact person data set includes a phone number, a name, an email address, a company name, a birthday, and / or a family address; A matching unit is configured to match the same data under each data type in the multi-source contact person data set; A second acquisition unit is configured to acquire two first contact person data corresponding to the same data; A first merging unit is configured to merge the two first contact person data to obtain second contact person data when the data under multiple data types in the two first contact person data are the same; A second merging unit is configured to merge the two first contact person data to obtain third contact person data according to the first similarity corresponding to the difference data group and the second similarity corresponding to the same data group in the multiple data types when the data under multiple data types in the two first contact person data are not the same; wherein the difference data group refers to two data with differences under the same data type, and the same data group refers to a data group with the same data under the same data type. The synchronization unit is configured to synchronize the second contact data, the third contact data and remaining contact data as final contact data to a target device, wherein the remaining contact data refers to contact data in the multi-source contact data set that has not been processed by the merging. The step of merging the two first contacts to obtain third contact data according to a first similarity corresponding to a difference data group and a second similarity corresponding to a same data group in a plurality of data types when the data in the plurality of data types in the two first contact data is uneven includes the following steps. Extracting a difference data group and a same data group in a plurality of data types corresponding to the two first contact data; Calculating a first similarity between two current data in the difference data group according to a data type corresponding to the difference data group; Taking a first preset value as the second similarity of the same data group, wherein the first preset value includes 1; According to a weight coefficient corresponding to a plurality of data types, weighting and summing the first similarity or the second similarity corresponding to each of the plurality of data types to obtain a final similarity; If the final similarity is greater than a first threshold value, obtaining a channel priority of the two first contacts; Taking the first contact data corresponding to the maximum channel priority in the two first contact data as the third contact data and eliminating the other first contact data; The step of calculating the first similarity between the two current data in the difference data group according to the data type corresponding to the difference data group includes the following steps. When the data type corresponding to the difference data group is a phone number, a birthday or an email address, calculating the first similarity according to a number of consecutive identical digits in the difference data group; When the data type corresponding to the difference data group is a name, calculating the first similarity according to a pinyin corresponding to the name; When the data type corresponding to the difference data group is a home address, calculating the first similarity according to a word distribution feature corresponding to the home address; When the data type corresponding to the difference data group is a company name, calculating the first similarity according to a character distribution feature of the company name.

8. A terminal device, comprising: The terminal device includes a memory, a processor and a contact recovery program based on multi-source heterogeneous data collaborative processing stored on the memory and executable on the processor, and the contact recovery program based on multi-source heterogeneous data collaborative processing is configured to implement the steps in the contact recovery method based on multi-source heterogeneous data collaborative processing in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Name matching method and device based on multiple data sources, equipment and medium

    CN116010562A

  • Multi-source heterogeneous data fusion method and system based on edge calculation

    CN120541795A