Family relationship-based lost customer tracking method and device, medium and product
By extracting contact and family member numbers from the historical bills of lost customers, building a family relationship network, screening suspicious numbers that match behavioral patterns and language styles, and calculating similarity to determine the target number, the problem of difficulty in tracking lost loan customers is solved, and more efficient loan recovery is achieved.
Patent Information
- Application Number
- CN202510760906.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-26
AI Technical Summary
When existing technologies are used to track down lost loan customers, it is difficult for traditional methods to obtain effective clues after the customer changes their number or cuts off contact, making it difficult for financial institutions to recover the loans, resulting in economic losses.
By extracting contact and family member numbers from the lost customer's historical bills, calculating intimacy, building a family relationship network, screening suspicious numbers that match behavioral patterns and language styles, and calculating similarity to determine the target number.
It improves the accuracy and effectiveness of tracking lost customers, provides effective contact clues, increases the possibility of loan recovery, and reduces economic losses.
Smart Images

Figure CN120707320A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of financial recovery technology, and in particular to a method, system, electronic device, storage medium and computer program product for tracking lost customers based on family relationships. Background Art
[0002] In the financial recovery sector, tracking lost loan customers typically relies on direct contact via their mobile phone number or by contacting emergency contacts. However, when a customer changes their phone number or loses contact with their original contact, these traditional methods become ineffective due to a lack of effective leads. This makes it difficult for financial institutions to effectively contact the lost loan customer, hindering loan recovery and resulting in financial losses. Therefore, current loan recovery solutions struggle to effectively track lost customers.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, electronic device, storage medium and computer program product for tracking lost customers based on family relationships, aiming to solve the technical problem of difficulty in effectively tracking lost customers.
[0005] To achieve the above objectives, the present application proposes a method for tracking lost customers based on family relationships, which includes:
[0006] Extracting the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact, calculating the intimacy between the number before the loss of contact and the family member number, and screening the family member numbers whose intimacy exceeds a preset intimacy threshold as close numbers;
[0007] Extracting the associated number of the intimate number from the historical bills corresponding to the intimate number, and building a family relationship network based on the number before the loss of contact, the contact number, the intimate number and the associated number;
[0008] Screening a first suspicious number from the family relationship network, detecting a second suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identifying a third suspicious number whose language style matches the language style of the number before the loss of contact;
[0009] The similarity between the third suspicious number and the number before the loss of contact is calculated based on the historical bills of the third suspicious number, and the third suspicious number whose similarity exceeds a preset similarity threshold is used as the target number.
[0010] In one embodiment, the step of calculating the intimacy between the number before the loss of contact and the family member number includes:
[0011] The total call duration, call contact days, total number of text messages, and number of text message contact days between the number before the loss of contact and the family member number are counted from the historical bills corresponding to the number before the loss of contact, and positive weights are assigned to the total call duration, the number of call contact days, the total number of text messages, and the number of text message contact days;
[0012] Obtain the latest communication date between the number before the loss of contact and the family member number, calculate the number of days between the latest communication date and the date of loss of contact of the lost customer, and assign a reverse weight to the number of days;
[0013] The intimacy between the number before the loss of contact and the family member number is obtained by comprehensively considering the total call duration, the number of call contact days, the total number of text messages, the number of text message contact days and the number of interval days.
[0014] In one embodiment, the step of constructing a family relationship network based on the number before the loss of contact, the contact number, the close number, and the associated number includes:
[0015] The number before the loss of contact, the contact number, the intimate number, and the associated number are defined as network nodes;
[0016] According to the communication records of the numbers corresponding to the network nodes, network node pairs with communication behaviors among the network nodes are identified, and the network nodes in the network node pairs are associated to obtain a family relationship network.
[0017] In one embodiment, the step of screening the first suspicious number from the family relationship network includes:
[0018] Traversing each network node in the family relationship network, and determining a target network node among the network nodes, wherein the target network node corresponds to the close number;
[0019] Suspicious network nodes associated with at least two of the target network nodes are screened out from the network nodes, numbers corresponding to the suspicious network nodes are compared with the contact numbers, and the contact numbers are removed from the numbers corresponding to the suspicious network nodes to obtain a first suspicious number.
[0020] In one embodiment, the step of detecting a second suspicious number whose behavior pattern in the first suspicious number matches the historical behavior pattern of the number before the loss of contact includes:
[0021] Extracting time series data within a preset time period after the lost customer loses contact from historical bills corresponding to the first suspicious number, and identifying abnormal numbers with abnormal behavior patterns among the first suspicious numbers based on the time series data;
[0022] Determine the historical behavior pattern of the number before the lost contact with the lost customer before the lost contact, and screen out matching numbers from the abnormal numbers whose time period distribution matching degree and change trend matching degree with the historical behavior pattern reach corresponding preset matching degree thresholds, and determine the matching numbers as second suspicious numbers.
[0023] In one embodiment, the step of identifying a third suspicious number in the second suspicious number that matches the language style of the number before the loss of contact includes:
[0024] Extracting text from the second suspicious number's SMS logs, analyzing sentence structure, high-frequency vocabulary, and punctuation usage habits of the text from the SMS logs, and generating language style features;
[0025] Calculating language style similarity between the language style feature and the historical language style feature of the number before the loss of contact;
[0026] The second suspicious number whose language style similarity exceeds a preset style similarity threshold is determined to have a language style match and is marked as a third suspicious number.
[0027] In one embodiment, the step of calculating the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number includes:
[0028] establishing a first usage habit matrix of the third suspicious number based on historical bills of the third suspicious number, and performing similarity calculation between the first usage habit matrix and the second usage habit matrix of the number before the loss of contact, to obtain usage habit similarity;
[0029] Determining a common contact number between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact overlap degree based on the common contact number;
[0030] determining a common contact location between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact location overlap based on the common contact location;
[0031] A weighted sum is taken of the usage habit similarity, the contact overlap, and the contact location overlap to generate a similarity between the third suspicious number and the number before the loss of contact.
[0032] In addition, to achieve the above purpose, the present application also proposes a lost customer tracking system based on family relationships, the lost customer tracking system based on family relationships comprising:
[0033] A close number confirmation module is used to extract the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact, calculate the intimacy between the number before the loss of contact and the family member number, and select the family member number whose intimacy exceeds a preset intimacy threshold as a close number;
[0034] A family relationship network construction module is used to extract the associated number of the close number from the historical bills corresponding to the close number, and to construct a family relationship network based on the number before the loss of contact, the contact number, the close number and the associated number;
[0035] a suspicious number confirmation module, configured to screen a first suspicious number from the family relationship network, detect a second suspicious number in the first suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identify a third suspicious number in the second suspicious number whose language style matches the language style of the number before the loss of contact;
[0036] The target number confirmation module is used to calculate the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and use the third suspicious number whose similarity exceeds a preset similarity threshold as the target number.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the method for tracking lost customers based on family relationships as described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the lost customer tracking method based on family relationships as described above are implemented.
[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the lost customer tracking method based on family relationships as described above.
[0040] The present application provides a method for tracking a lost customer based on family relationships. The method for tracking a lost customer based on family relationships includes: extracting the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact of the lost customer, calculating the intimacy between the number before the loss of contact and the family member number, and screening the family member numbers whose intimacy exceeds a preset intimacy threshold as intimate numbers; extracting the associated numbers of the intimate number from the historical bills corresponding to the intimate number, and building a family relationship network based on the number before the loss of contact, the contact number, the intimate number and the associated numbers; screening a first suspicious number from the family relationship network, detecting a second suspicious number in the first suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identifying a third suspicious number in the second suspicious number whose language style matches the number before the loss of contact; calculating the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and taking the third suspicious number whose similarity exceeds the preset similarity threshold as the target number.
[0041] This application extracts contact numbers and family member numbers from historical bills corresponding to the lost customer's pre-missing number, calculates the degree of affinity between the pre-missing number and the family member numbers, and screens out intimate numbers. This allows for the initial identification of numbers closely related to the lost customer, providing a richer source of leads for subsequent tracking. By extracting associated numbers from historical bills corresponding to intimate numbers and constructing a family relationship network based on the pre-missing number, contact numbers, intimate numbers, and associated numbers, the application can more intuitively demonstrate the social relationship context of the lost customer, providing a basis for subsequent screening of suspicious numbers. By first screening a first suspicious number from the family relationship network, then detecting a second suspicious number within the first suspicious number whose behavior pattern matches the historical behavior pattern of the pre-missing number, and finally identifying a third suspicious number within the second suspicious number whose language style matches the pre-missing number, this multi-level screening can gradually narrow the scope of suspicious numbers and improve tracking accuracy. By calculating the similarity between the third suspicious number and the pre-missing number based on its historical bills, and selecting the third suspicious number whose similarity exceeds a preset similarity threshold as the target number, the application can accurately identify the number with the strongest association with the lost customer, providing effective contact leads for financial institutions. Compared with related solutions that rely on mobile phone numbers to directly reach and contact emergency contacts, when the customer changes his number or cuts off contact with the original contact after losing contact, the solution will fail due to the inability to obtain effective clues. This solution expands the source of clues by mining contact numbers and family member numbers from the historical bills of lost customers, and further constructing a family relationship network. Even if the customer changes his number or cuts off contact with the original contact, new clues can be found from the family relationship network. Through multi-level screening steps such as calculating intimacy, matching behavioral patterns and language styles, and calculating similarity, the target number can be determined, which improves the accuracy and effectiveness of tracking, can more effectively track lost customers, and provide financial institutions with effective contact clues, thereby increasing the possibility of recovering loans and reducing economic losses caused by customer loss of contact. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 A flowchart of the first embodiment of the method for tracking lost customers based on family relationships provided in this application;
[0045] Figure 2 A family relationship network diagram provided for the missing customer tracking method based on family relationships in this application;
[0046] Figure 3 A flowchart of the second embodiment of the method for tracking lost customers based on family relationships provided in this application;
[0047] Figure 4 This is a schematic diagram of the module structure of the lost customer tracking system based on family relationships according to an embodiment of the present application;
[0048] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the lost customer tracking method based on family relationships in an embodiment of the present application.
[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0051] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0052] The main solution of the embodiment of the present application is: extract the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact of the lost customer, calculate the intimacy between the number before the loss of contact and the family member number, and screen the family member numbers whose intimacy exceeds the preset intimacy threshold as intimate numbers; extract the associated numbers of the intimate number from the historical bills corresponding to the intimate number, and build a family relationship network based on the number before the loss of contact, the contact number, the intimate number and the associated number; screen a first suspicious number from the family relationship network, detect a second suspicious number in the first suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identify a third suspicious number in the second suspicious number whose language style matches the number before the loss of contact; calculate the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and take the third suspicious number whose similarity exceeds the preset similarity threshold as the target number.
[0053] In this embodiment, for ease of description, the following description is made using a lost customer tracking system based on family relationships as the execution subject.
[0054] Existing technologies for tracking lost loan customers usually rely on direct contact with mobile phone numbers, contact with emergency contacts, etc. However, when a customer changes his number or cuts off contact with the original contact after losing contact, traditional methods will fail due to the inability to obtain effective clues, making it difficult for financial institutions to effectively contact the lost loan customers, making it difficult to recover the loans, resulting in economic losses.
[0055] This application provides a solution that expands the source of clues by mining contact numbers and family member numbers from the historical bills of lost customers and further building a family relationship network. Even if the customer changes his number or cuts off contact with the original contact, new clues can be found from the family relationship network. Through multi-level screening steps such as calculating intimacy, matching behavioral patterns and language styles, and calculating similarity, the target number can be determined, which improves the accuracy and effectiveness of tracking, and can more effectively track lost customers, providing financial institutions with effective contact clues, thereby increasing the possibility of recovering loans and reducing economic losses caused by lost customers.
[0056] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, mobile phone, or other electronic device, or a system, program, or application capable of implementing the aforementioned functions. This embodiment and the following embodiments will be described below using a family relationship-based missing customer tracking system (hereinafter referred to as the system) as an example.
[0057] All actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0058] Based on this, the embodiment of the present application provides a method for tracking lost customers based on family relationships. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the method for tracing lost customers based on family relationships of this application.
[0059] In this embodiment, the method for tracking lost customers based on family relationships includes steps S01 to S04:
[0060] Step S01: extracting the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact, calculating the intimacy between the number before the loss of contact and the family member number, and screening the family member numbers whose intimacy exceeds a preset intimacy threshold as the intimacy number;
[0061] It should be noted that the system obtains the pre-disconnection number of the lost customer and retrieves the historical bills corresponding to the number. In the financial recovery scenario, the lost customer refers to a loan customer who cannot be contacted through conventional means (such as directly dialing a mobile phone number, contacting an emergency contact, etc.) for various reasons. The pre-disconnection number refers to the number used by the lost customer before the loss of contact. The historical bills record in detail the communication status of the number within a preset time period (for example, 3 months), including call records (such as call time, call duration, the other party's number, etc.), SMS records (such as SMS sending and receiving time, the other party's number, SMS content, etc.) and other information. Other numbers that have had communication contact with the pre-disconnection number, i.e., contact numbers, are extracted from the historical bills. At the same time, the numbers of the lost customer's family members are identified through the preset family relationship network of the lost customer. The preset family relationship network is pre-built based on the family relationships of the lost customer. The system calculates the intimacy between the number before the loss of contact and the numbers of each family member. The intimacy is used to measure the closeness of the communication connection between the number before the loss of contact and the family member numbers. It can be calculated comprehensively through multiple indicators. These indicators can be call frequency (the more calls in a certain period of time, the higher the intimacy may be), call duration (the longer the single call time, the higher the intimacy may be), the number of text message exchanges (the more frequent the text message exchanges, the higher the intimacy may be), etc. The calculated intimacy is compared with the preset intimacy threshold (for example, 0.7), and the family member numbers whose intimacy exceeds the threshold are screened out and marked as intimate numbers, indicating that they have close interactions with the lost customer.
[0062] It can be understood that in step S01, by extracting the contact number and family member number, and calculating the intimacy between the number before the loss of contact and the family member number, the close numbers are screened out. Even if the customer cuts off contact with the original contact, it is still possible to find the new number of the lost customer through family members, which increases the possibility of obtaining the new contact information of the lost customer, expands the scope of tracking clues, and increases the probability of finding the new number of the lost customer.
[0063] Step S02: extracting the associated numbers of the intimate number from the historical bills corresponding to the intimate number, and constructing a family relationship network based on the number before the loss of contact, the contact number, the intimate number, and the associated numbers;
[0064] It should be noted that after obtaining the close numbers, the system retrieves the historical bills corresponding to these close numbers and extracts the associated numbers, that is, other numbers that have had communication contact with the close numbers. Taking the number before the loss of contact as the core, combined with the contact number, close number and associated number, by analyzing the communication relationship between them (such as who has had calls, text messages, etc.), a family relationship network is constructed. This network can be represented by a graph structure, with nodes representing each number and edges representing the communication connection between numbers.
[0065] It can be understood that step S02, by constructing a family relationship network, integrates the lost customers and their related personnel into a network structure, clearly showing the social relationship between the lost customers and their related personnel (contact numbers, close numbers, associated numbers), so that financial institutions can analyze the social dynamics of lost customers from a more macro perspective, which helps to discover possible new contact methods or social circles of lost customers, and thus discover potential tracking clues, which is more comprehensive and in-depth than existing technologies.
[0066] Step S03, screening a first suspicious number from the family relationship network, detecting a second suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identifying a third suspicious number whose language style matches the language style of the number before the loss of contact;
[0067] It should be noted that the system screens out first suspicious numbers from the family relationship network. These numbers are initially screened from the family relationship network and are likely related to the lost customer. The screening criteria can be customized based on actual needs. For example, a number that has a certain frequency of communication with the number before the loss of contact (e.g., the number of calls exceeding a certain threshold within a certain time period) can be considered the first suspicious number. The system then analyzes the first suspicious number to determine whether its behavior pattern matches the historical behavior pattern of the number before the loss of contact. The behavior pattern refers to the regularities and habits exhibited by the first suspicious number during a preset time period, including communication time distribution (e.g., the number before the loss of contact had a habit of making calls between 8 and 10 p.m., and the first suspicious number also has a similar call time distribution), and frequently used contacts (e.g., the number before the loss of contact frequently contacted certain numbers, and the first suspicious number also had contact with these numbers). The historical behavior pattern refers to the behavior pattern of the number before the loss of contact during a preset time period. The behavior pattern of the first suspicious number is compared with the historical behavior pattern of the number before the loss of contact, and numbers that meet a preset matching threshold (e.g., 80%) are selected as second suspicious numbers. Further identify the number in the second suspicious number that matches the language style of the number before the loss of contact as the third suspicious number. The language style refers to the language characteristics used by the number during the communication process, such as word usage habits, expression methods, tone, etc., which can be judged by analyzing the word usage habits (such as the frequent use of certain specific words), expression methods (such as concise and clear or detailed descriptions), and tone (such as formal, casual, etc.) in the text message content and call recordings. For example, if the number before the loss of contact often used expressions such as "Hi, how are you lately" in text messages, and a number with similar wording and expression methods in the second suspicious number is also identified as the third suspicious number.
[0068] It can be understood that step S03, through dual matching screening of behavioral patterns and language styles, screens suspicious numbers from multiple dimensions, can more accurately locate numbers that may be related to lost customers, making tracking more precise, and can more effectively exclude irrelevant numbers, thereby improving the accuracy of tracking targets and saving financial institutions a lot of manpower and time costs.
[0069] Step S04: Calculate the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and use the third suspicious number whose similarity exceeds a preset similarity threshold as the target number.
[0070] It should be noted that after the system obtains the third suspicious number, it retrieves the historical bills of these numbers and calculates the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number. The similarity can be calculated by combining multiple indicators, such as usage habit similarity, contact overlap, and contact location overlap, etc. These indicators are assigned different weights, and the degree of similarity between the third suspicious number and the number before the loss of contact is obtained through weighted calculation. The calculated similarity is compared with a preset similarity threshold (for example, 0.8), and the third suspicious number whose similarity exceeds the threshold is used as the target number. The target number is considered to be the new number used by the lost customer.
[0071] It can be understood that step S04, through similarity calculation and threshold screening, can more accurately find the new number that the lost customer may use, providing strong support for financial institutions to recover loans. When the similarity exceeds the preset threshold, it can be highly suspected that the number is the new number of the lost customer, so that corresponding contact measures can be taken, thereby improving the success rate of loan recovery.
[0072] In a feasible implementation, in step S01, the step of calculating the intimacy between the number before the loss of contact and the family member number includes steps A01 to A03:
[0073] Step A01: Count the total call duration, call duration, total number of text messages, and number of days of text message contact between the number before the loss of contact and the family member's number from the historical bills corresponding to the number before the loss of contact, and assign positive weights to the total call duration, call duration, total number of text messages, and number of days of text message contact;
[0074] It should be noted that the system retrieves the historical bills corresponding to the number before it lost contact. The historical bills record in detail the communication between the number and each family member number, conducts data mining and analysis on the historical bills, and calculates the total call duration, call contact days, total number of text messages and text message contact days between the number before it lost contact and each family member number. Among them, the total call duration refers to the sum of the duration of all calls between the number before it lost contact and the family member numbers within the preset time period. For example, in the past year, the number before it lost contact had 10 calls with family member number A, and the duration of each call was 5 minutes, 8 minutes, 3 minutes, etc., then the total call duration is the accumulation of these call durations; the call contact days refer to the total number of calls between the number before it lost contact and the family member numbers within the preset time period. The number of days in which there was phone contact. For example, in the past 30 days, the number before the loss of contact and family member number A had phone calls on the 1st, 5th, and 10th days, then the number of phone contact days is 5 days; the total number of text messages refers to the total number of text messages sent and received between the number before the loss of contact and the family member number within the preset time period. For example, in the past month, the number before the loss of contact and family member number A sent 20 text messages to each other, then the total number of text messages is 20; the number of text message contact days refers to the number of days in which the number before the loss of contact and the family member number had text message exchanges within the preset time period. For example, in the past 15 days, the number before the loss of contact and family member number A had text message exchanges on the 2nd, 7th, and 12th days, then the number of text message contact days is 3 days.
[0075] Additionally, it should be noted that the system assigns positive weights to call duration, call days, total text message count, and text message days. Positive weights indicate that these metrics contribute positively to intimacy; that is, the larger the value of these metrics, the higher the intimacy. For example, the weight of call duration could be set to 0.4, the weight of call days to 0.2, the weight of total text message count to 0.25, and the weight of text message days to 0.15. These weights can be adjusted based on actual needs and experience.
[0076] Step A02: Obtain the latest communication date between the number before the loss of contact and the family member number, calculate the number of days between the latest communication date and the date of loss of contact of the lost customer, and assign a reverse weight to the number of days between the latest communication date and the date of loss of contact of the lost customer;
[0077] It should be noted that the system obtains the date of the most recent communication between the number before the loss of contact and each family member's number. By analyzing the communication records in the historical bills, the system finds the date of the last call or text message exchange as the most recent communication date, and obtains the lost contact date of the lost customer. The lost contact date is known and is the date when banks, financial institutions, and major financial platforms determine that the lost customer cannot be contacted through conventional means. The number of days between the most recent communication date and the lost contact date is calculated. For example, if the most recent communication date is May 10, 2024, and the lost contact date is May 20, 2024, the number of days between the most recent communication date and the lost contact date is 10 days.
[0078] Additionally, it's important to note that the system assigns a negative weight to the number of days between encounters. This weight indicates that the longer the interval, the lower the intimacy. For example, a negative weight of 0.1 would indicate that for every additional day between encounters, the intimacy would decrease proportionally.
[0079] Step A03, comprehensively consider the total call duration, call contact days, total number of text messages, text message contact days, and interval days to obtain the intimacy between the number before the loss of contact and the family member's number.
[0080] It should be noted that the system calculates the intimacy between the number before the loss of contact and the family member's number based on the following formula:
[0081]
[0082] Among them, a represents the number before the loss of contact, y n Indicates the nth family member number, V(a, y n 0 means the number a before the loss of contact and the family member number y n The intimacy Indicates the total duration of the call, w A represents the total call duration weight, q represents the number of call days, w q Indicates the weight of call contact days, Indicates the total number of text messages, w N represents the weight of the total number of SMS messages, p represents the number of SMS contact days, w p Indicates the weight of SMS contact days, d diff Indicates the number of days between d Indicates the weight of the interval days.
[0083] In addition, it should be noted that, without considering the weight, the system can also calculate the intimacy between the number before the loss of contact and the family member's number using the intimacy formula:
[0084]
[0085] Among them, d lostIndicates the date of loss of contact for the lost customer. Represents a and family member number y n All dates with call logs, Represents a and family member number y n All dates with text message records, The number of days between diff , among which, the total duration of calls, the number of call contact days, the total number of text messages and the number of text message contact days are directly proportional to the intimacy, and the number of days between the date of loss of contact and the date of the most recent communication is inversely proportional to the intimacy. Because banks, financial institutions and major financial platforms have discovered that the date of loss of contact for lost customers may be the same as or later than the actual date of loss of contact, that is, Therefore, in order for the intimacy formula to make sense, add 1 to the denominator in the formula.
[0086] In this embodiment, by counting the total call duration, call contact days, total number of text messages and text message contact days between the number before the loss of contact and the family member number, we can have a more comprehensive understanding of the closeness of the connection between the lost customer and the family member, and can dig out the actual communication activity between the lost customer and the family member, providing basic data for subsequent intimacy evaluation, solving the problem of lack of tracking clues caused by the inability to obtain such communication details in traditional methods. By obtaining the latest communication date between the number before the loss of contact and the family member number, and calculating the number of days between the number and the date of loss of contact with the lost customer, we can evaluate the timeliness of the family member's understanding of the current situation of the lost customer, and through By assigning reverse weights to the number of days between calls, when calculating intimacy, the intimacy scores of family member numbers with a long time of recent communication can be lowered, making the screened out intimate numbers more likely to provide effective clues, thereby improving the targeted nature of tracking. By comprehensively integrating multiple indicators such as total call duration, number of call contact days, total number of text messages, number of text message contact days, and number of days between calls, the intimacy between the number before the loss of contact and the family member number can be comprehensively and accurately assessed. The screened out intimate numbers are more likely to keep in touch with the lost customer or know the lost customer's new contact information, thereby increasing the possibility of financial institutions tracking the lost customer, helping to recover loans and reduce economic losses.
[0087] In a feasible implementation, in step S02, the step of constructing a family relationship network based on the number before the loss of contact, the contact number, the close number, and the associated number includes steps A11 to A12:
[0088] Step A11, defining the number before the loss of contact, the contact number, the intimate number, and the associated number as network nodes;
[0089] It should be noted that the system obtains the number before loss of contact, contact number, close number and associated number, and defines these numbers as network nodes in the family relationship network. Each network node represents a specific number and is the basic component unit of the family relationship network.
[0090] Step A12: identifying network node pairs with communication behaviors in each network node according to the communication records of the corresponding numbers of each network node, and associating the network nodes in the network node pairs to obtain a family relationship network.
[0091] It should be noted that the system retrieves the communication records for the numbers corresponding to each network node. The communication records detail the communication between the numbers, including call records (such as call time, call duration, and the other party's number) and SMS records (such as the time when SMS messages were sent and received, the other party's number, and the content of the SMS messages). Based on the communication records, network node pairs with active communication are identified within each network node. If there is a call or SMS exchange record between two network nodes, then these two network nodes constitute a network node pair with active communication. The network nodes in these pairs with active communication are associated, that is, the two network nodes are connected with edges. The edges between network nodes can represent the communication relationship between network nodes. In this way, all pairs of network nodes with active communication are associated, ultimately resulting in a family relationship network. This network can be a graph structure, with nodes representing numbers and edges representing the communication relationships between numbers.
[0092] For example, to help understand the technical concept or technical principle of this application, please refer to Figure 2 , Figure 2 Provides a family relationship network diagram, where a is the number before the loss of contact, C1 and C2 are contact numbers, B1 and B2 are close numbers, and D1 is the associated number. Figure 2 From the association relationships, we can see that a has a communication relationship with C1, C2, B2 and D1, C1 has a communication relationship with a, C2 and B1, C2 has a communication relationship with C2 and a, B1 has a communication relationship with C1 and D1, B2 has a communication relationship with a, and D1 has a communication relationship with B1 and a.
[0093] In this embodiment, by defining the number before loss of contact, contact number, close number and associated number as network nodes, an infrastructure is provided for the subsequent construction of a unified relationship network, solving the problem of scattered number information and difficulty in integration and utilization. This structured processing enables subsequent unified analysis and operation based on these nodes, laying the foundation for exploring the potential relationship between numbers, providing an initial data organization form for effectively tracking lost customers, and improving the integration and operability of information related to lost customers. By analyzing communication records, identifying and associating network node pairs with communication behavior, the problem of unclear relationship between numbers and difficulty in judging the degree of association is solved, and the relationship network around lost customers can be more accurately portrayed. Through this association, it can be clearly seen which numbers have frequent communications, thereby inferring the possible close relationship between them, allowing staff to more effectively track lost customers based on the relationship chain in the network and improve the success rate of loan recovery.
[0094] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 In step S03, the step of screening the first suspicious number from the family relationship network includes steps S11 to S12:
[0095] Step S11, traversing each network node in the family relationship network, determining a target network node among each network node, where the target network node corresponds to the intimate number;
[0096] It should be noted that the system checks each network node in the family relationship network one by one. The family relationship network is a structure containing multiple network nodes. Each network node represents a number (or an entity related to the number). According to the type of number corresponding to the network node, the network node corresponding to the close number is found in each network node. These network nodes are the target network nodes.
[0097] Step S12: Filter out suspicious network nodes associated with at least two target network nodes from each network node, compare the numbers corresponding to the suspicious network nodes with the contact numbers, remove the contact numbers from the numbers corresponding to the suspicious network nodes, and obtain a first suspicious number.
[0098] It should be noted that the association between each network node and the target network node in the family relationship network is checked. If a network node is associated with at least two target network nodes (for example, there is a direct communication record, joint participation in an event, etc.), the network node is marked as a suspicious network node, and the number corresponding to the suspicious network node is obtained. These numbers are compared one by one with the known contact numbers. If the number corresponding to a suspicious network node intersects with the contact number, the number is removed from the numbers corresponding to the suspicious network node. After the above comparison and removal operations, the number corresponding to the remaining suspicious network node is the first suspicious number.
[0099] Specifically, the first suspicious number can be determined by the following formula:
[0100] M=((KC1∩(KC2∪KC3∪…KC r ))∪(KC2∩(KC1∪KC3∪…KC r ))∪…
[0101] ∪(KC r ∩(KC1∪KC2∪…KC r-1 )))
[0102] -((KC1∩(KC2∪KC3∪…KC r ))
[0103] ∪(KC2∩(KC1∪KC3∪…KC r ))∪…
[0104] ∪(KC r ∩(KC1∪KC2∪…KC r-1 )))∩{c1,c2,…,c l}
[0105] Where M is the set of the first suspicious numbers, KC r Indicates intimate number C r Contact number in C r is the number with the highest intimacy ranking r with the number a before the loss of contact, ((KC1∩(KC2∪KC3∪…KC r ))∪KC2∩KC1∪KC3∪…KCr∪…∪KCr∩KC1∪KC2∪…KCr-1 can calculate the number set that has communication relations with more than two target network nodes in C1, C2,…, Cr. Since some of the numbers in this set may come from {c1, c2,…, c k}, that is, the contact number, so we must subtract ((KC1∩(KC2∪KC3∪…KC r ))∪(KC2∩(KC1∪KC3∪…KC rThe intersection of ))∪…∪KCr∩KC1∪KC2∪…KCr-1 and c1, c2,…, ck, that is, the intersection of the suspicious network node and the contact number, is used to obtain the first suspicious number M.
[0106] In this embodiment, by traversing each network node in the family relationship network and determining the target network node corresponding to the close number, accurate positioning is provided for subsequent screening operations, so that the subsequent screening process can be targeted and avoid blindly searching the entire network, greatly improving the search efficiency, reducing unnecessary computing overhead, and laying the foundation for the subsequent effective tracking of lost customers. Suspicious network nodes that are associated with at least two target network nodes are screened out from each network node, which increases the accuracy of screening out nodes that are truly related to the lost customer. The corresponding numbers of the suspicious network nodes are compared with the contact numbers and the contact numbers are eliminated, further purifying the list of suspicious numbers and improving the effectiveness of the suspicious numbers. The first suspicious number finally obtained is more likely to point to the lost customer or a person who has an important relationship with the lost customer, providing a more accurate target for subsequent loan recovery work, effectively solving the problem of difficult to effectively track lost customers in the current loan recovery plan, and improving the success rate of loan recovery.
[0107] In a feasible implementation, in step S03, the step of detecting a second suspicious number whose behavior pattern in the first suspicious number matches the historical behavior pattern of the number before the loss of contact includes steps B01 to B02:
[0108] Step B01: extracting time series data within a preset time period after the lost customer loses contact from historical bills corresponding to the first suspicious number, and identifying abnormal numbers with abnormal behavior patterns among the first suspicious numbers based on the time series data;
[0109] It should be noted that from the historical bills corresponding to the first suspicious number, time series data within a preset time period after the lost customer lost contact is extracted, including but not limited to call frequency, communication time period distribution, etc., and through a time series analysis model (such as a mutation point detection algorithm), abnormal behavior patterns such as sudden changes in call frequency (sudden increase / decrease) and communication time period offset (for example, the original active period was daytime, and it turned to late at night after the loss of contact) are identified. The abnormal behavior pattern indicates that the user behavior has changed significantly from the behavior in the previous preset time period, and the numbers with abnormal behavior patterns in the first suspicious number are marked as abnormal numbers.
[0110] Step B02: Determine the historical behavior pattern of the number before the lost customer loses contact, select matching numbers from the abnormal numbers whose time period distribution matching degree and change trend matching degree with the historical behavior pattern reach corresponding preset matching degree thresholds, and determine the matching numbers as second suspicious numbers.
[0111] It should be noted that the system will also obtain the historical behavior patterns of the number before the lost customer lost contact. The historical behavior pattern refers to the behavior pattern within the user's preset time period, including the lost customer's fixed active time period (such as the daily call peak from 20:00 to 22:00) and periodic behavior trends (such as a sudden increase in the proportion of Internet traffic on weekends). The overlapping ratio of the abnormal number and the number before the loss of contact in a fixed active time period is counted to obtain the time period distribution matching degree (for example, the original customer and the abnormal number both made calls between 22:00 and 24:00, with an overlap rate of 80%). Specifically, the time period distribution matching degree can be evaluated by calculating the cosine similarity or Euclidean distance of the time period distribution histograms corresponding to the abnormal number and the number before the loss of contact. The time period distribution matching degree refers to the overlapping ratio of the active time periods of both the abnormal number and the number before the loss of contact, reflecting the temporal consistency of behavior; and the consistency of the periodic behavior trends of the abnormal number and the number before the loss of contact is analyzed to obtain the change trend matching degree (for example, the original customer had a sudden increase in calls due to the approaching repayment date, and the abnormal number showed the same trend in a similar period). Specifically, the dynamic time warping (DTW) algorithm can be used to calculate the similarity of the periodic behavior trend curves corresponding to the abnormal number and the number before the loss of contact to obtain the change trend matching degree. The change trend matching degree refers to the similarity of the behavior fluctuation pattern (such as mutation, periodicity) between the abnormal number and the number before the loss of contact, reflecting the logical correlation of behavior. If the time period distribution matching degree and change trend matching degree of an abnormal number reach the corresponding preset matching degree thresholds (such as time period distribution matching degree ≥ 70%, change trend matching degree ≥ 60%), the abnormal number (ie, the matching number) will be determined as the second suspicious number.
[0112] In this embodiment, by identifying abnormal behavior patterns based on time series data, it is possible to screen out numbers with abnormal behavior patterns from massive number data. These numbers are more likely to be associated with lost customers, providing clues for subsequent accurate tracking. By determining the historical behavior patterns of the numbers before the loss of contact and matching the abnormal numbers with them, it is possible to dig out clues that are highly relevant to the lost customers, and more accurately screen out numbers that may be associated with the lost customers. By setting two indicators, time period distribution matching and change trend matching, and setting corresponding preset matching thresholds, the matching process can be quantified, allowing the machine to screen according to clear rules, thereby improving the accuracy and reliability of screening.
[0113] In a feasible implementation, in step S03, the step of identifying a third suspicious number in the second suspicious number that matches the language style of the number before the loss of contact includes steps B11 to B13:
[0114] Step B11: extracting the text of the SMS record of the second suspicious number, analyzing the sentence structure, high-frequency vocabulary, and punctuation usage habits of the SMS record text, and generating language style features;
[0115] It should be noted that all text message content (including sent and received text) is extracted from the historical bills of the second suspicious number to form a complete text message log. Natural language processing technology is used to perform syntactic analysis on the text message log to identify sentence features. These features can include sentence length (long / short), sentence type (declarative, interrogative, imperative), and paragraph segmentation. Sentence features are analyzed to obtain the sentence structure of the text message log. Sentence structure represents the grammatical structure of sentences in the text message log, such as simple sentences and complex sentences. Different sentence structures reflect the speaker's expression habits and thinking style. A word segmentation algorithm is used to break the text message log into lexical phrases. Lexical phrases that appear in the text message log at a preset frequency threshold are counted as high-frequency words in the text message log. High-frequency words are words that appear frequently in the text message log and can reflect the speaker's common expressions and topics of interest. The use of punctuation marks in text messages is recorded, such as the frequency of sentence-ending punctuation (periods, exclamation marks) and special symbols (ellipsis, dashes). This information is used to identify punctuation habits, which can reflect the speaker's tone and emotion. Sentence structure, high-frequency vocabulary, and punctuation habits are converted into structured feature vectors, or language style features, to characterize the lost customer's language style.
[0116] Step B12, calculating the language style similarity between the language style feature and the historical language style feature of the number before the loss of contact;
[0117] It should be noted that, similar to step B11, the historical language style features of the number before the loss of contact are extracted from the historical SMS records. The historical language style features refer to the language style features corresponding to the SMS record text of the number before the loss of contact. Based on a similarity algorithm such as the cosine similarity algorithm or the Euclidean distance algorithm, the language style similarity value between the second suspicious number and the number before the loss of contact is calculated. The language style similarity is a quantitative indicator used to measure the degree of similarity between the language styles of two numbers. The similarity value range is usually between 0 and 1. The closer the value is to 1, the more similar the language styles of the two numbers are.
[0118] Step B13: The second suspicious number whose language style similarity exceeds a preset style similarity threshold is determined to have a language style match and marked as a third suspicious number.
[0119] It should be noted that the language style similarity value calculated in step B12 is compared with the preset style similarity threshold. If there is a second suspicious number whose language style similarity value exceeds the preset style similarity threshold, the number is determined to have a language style match and is marked in the system as a third suspicious number.
[0120] In this embodiment, by extracting the language style characteristics of the second suspicious number, the language style of the number user is fully reflected, providing rich data support for subsequent similarity calculations. Through similarity calculations, second suspicious numbers that are relatively close to the language style of the number before the loss of contact can be quickly screened out, narrowing the tracking scope and improving tracking efficiency. Third suspicious numbers that match the language style of the number before the loss of contact are screened out from the second suspicious numbers. The users of these numbers are likely to be the lost customers themselves or people closely related to them. By marking the third suspicious numbers, resources can be concentrated to conduct in-depth investigations on these numbers, thereby improving the success rate of loan recovery.
[0121] In a feasible implementation, in step S04, the step of calculating the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number includes steps B21 to B24:
[0122] Step B21: establishing a first usage habit matrix of the third suspicious number based on the historical bills of the third suspicious number, and calculating similarity between the first usage habit matrix and the second usage habit matrix of the number before the loss of contact, to obtain usage habit similarity;
[0123] It should be noted that the usage habit matrix refers to a matrix composed of data on the calling habits, text messaging habits, and Internet usage habits of lost users in 24 time periods of a day, which is compiled based on the historical bills corresponding to the numbers. The row attributes of the usage habit matrix are 24 time periods, and the column attributes include the average number of calls, average call duration, average number of text messages, average Internet traffic, average Internet usage duration, and average number of network connections in time periods 0-23.
[0124] For example, the second usage habit matrix of the number before the loss of contact is expressed as follows:
[0125]
[0126] Among them, a represents the number before the loss of contact, Y a The second usage habit matrix of number a before the loss of contact, m = 0, 1, 2, ..., 23, It represents the average number of calls made by number a in the mth time period of 24 time periods in a day before the loss of contact. It represents the average call duration of number a in the mth time period before the loss of contact. It represents the average number of SMS messages sent by number a in the mth time period before the loss of contact. It represents the average Internet traffic of number a in the mth time period before the connection is lost. It represents the average online time of number a in the mth time period before the connection is lost. Indicates the average number of times number a connected to the Internet in the mth time period before losing contact.
[0127] Similarly, the first usage habit matrix of the third suspicious number is expressed as follows:
[0128]
[0129] Among them, M j Indicates the third suspicious number, Indicates the third suspicious number M j The third usage habit matrix, m = 0, 1, 2, ..., 23, Indicates the third suspicious number M j The average number of calls in the mth time period among the 24 time periods in a day, Indicates the third suspicious number M j The average call duration in the mth time period, Indicates the third suspicious number M j The average number of text messages in the mth time period, Indicates the third suspicious number M j The average Internet traffic in the mth time period, Indicates the third suspicious number M j The average online time duration in the mth time period, Indicates the third suspicious number M j The average number of network connections in the mth time period.
[0130] In addition, it should be noted that usage habit similarity refers to the similarity between the third suspicious number and the number before the loss of contact during the 24 time periods of the day. Due to a combination of factors such as a person's physiology, education, personality, social relationships, environment, economy, and equipment, each person's mobile phone usage habits are different. The same person's mobile phone usage habits are generally relatively stable. The Pearson correlation coefficient method is used to measure the usage habit similarity between the third suspicious number and the number before the loss of contact. The calculation formula for usage habit similarity is as follows:
[0131]
[0132] Among them, S habit (a,M j ) represents the third suspicious number M j Similarity in usage habits with the number a before the loss of contact, It's Y a The sample mean of yes The sample mean of is the matrix components, when S habit (a,M j )>0, M jThere is a certain similarity between M and A in terms of mobile phone usage habits. j It is unrelated or negatively correlated with a's mobile phone usage habits.
[0133] Step B22: determining a common contact number between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact overlap degree based on the common contact number;
[0134] It's important to note that everyone's social relationships vary, influenced by factors such as education, personality, environment, and finances. The more contact phone numbers two people share, the more similar their social relationships are, and the closer their relationship is, or they might be the same person. We extract the contact number lists from the third suspicious number and the previous contact number's historical bills, respectively. We perform an intersection operation on the two lists to obtain the common contact numbers. Based on these common contact numbers, we generate a contact overlap degree, which measures the degree of similarity between the corresponding contact numbers of the two numbers.
[0135] For example, the Jaccard similarity coefficient is used to measure the contact similarity between the number before the loss of contact and the third suspicious number. The calculation formula is as follows:
[0136]
[0137] Among them, S MobileNumber (a,M j ) represents the third suspicious number M j The overlap degree with the contact of number a before the loss of contact, {c1, c2, ..., c k} represents the contact number of a, Indicates m j Contact number, Indicates a and M j The number of common contact numbers, Indicates a and M j The ratio of the number of common contact numbers to the total number of contact numbers of the two numbers is calculated as the contact overlap degree.
[0138] Step B23, determining a common contact location between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact location overlap based on the common contact location;
[0139] It should be noted that the location information of the communication or transaction is extracted from the historical bills of the third suspicious number and the number before the loss of contact, and the location information is matched and intersected to obtain the common contact location. The contact location overlap is generated based on the common contact location. The contact location overlap refers to the similarity of the locations where two mobile phone numbers make calls, send and receive text messages or surf the Internet within a certain period of time.
[0140] For example, similar to step B22, the Jaccard similarity coefficient is used to measure the overlap of the contact locations of the number before the loss of contact and the third suspicious number. The calculation formula is as follows:
[0141]
[0142] Among them, S city (a,M j 0 represents the third suspicious number M j The overlap degree with the contact location of number a before the loss of contact, {e a1 , e a2 ,…,e ab} represents the contact location set of a, Indicates M j A collection of contact locations, Indicates a and M j The number of common contact locations, Indicates a and M j The ratio of the number of common contact locations to the total number of contact locations of the two numbers is calculated as the contact location overlap.
[0143] Step B24: performing a weighted summation of the usage habit similarity, the contact overlap, and the contact location overlap to generate a similarity between the third suspicious number and the number before the loss of contact.
[0144] It should be noted that since the importance of similarity in each dimension varies, the similarity weights assigned to each dimension are also different. Usage similarity is the basic condition for determining whether two numbers belong to the same customer, so usage similarity can be given a higher weight. Contact overlap is also the main dimension for measuring whether two mobile phone users are the same person, so contact similarity can also be given a higher weight. Contact location overlap measures the similarity between two mobile phone users to a certain extent, but since the range of people's activities changes dynamically, contact location similarity is given a lower weight. The weight calculation formula is as follows:
[0145] S(a,M j )=[w1×s habit (a,M j )+w2×S MobileNumber (a,M j )+w3
[0146] ×S city (a,M j )]×100%
[0147] Among them, w1 is the usage habit similarity S habit (a,M j) corresponds to the weight, w2 is the usage habit similarity S MobileNumber (a,M j ) corresponds to the weight, w3 is the usage habit similarity S city (a,M j ) corresponding weights. Multiply usage habit similarity, contact overlap, and contact location overlap by their corresponding weights, and add the results to obtain the similarity between the third suspicious number and the number before the loss of contact.
[0148] In this embodiment, by calculating the similarity between the first usage habit matrix and the second usage habit matrix of the number before the loss of contact, a usage habit similarity value is obtained. This value can intuitively reflect the similarity between the two numbers in usage habits, and provides an important basis for subsequent judgment of whether the two numbers are related. By extracting and comparing the contact information in the historical bills of the third suspicious number and the number before the loss of contact, the common contact number of the two can be accurately found. The contact overlap can intuitively reflect the similarity between the two numbers in terms of contacts. The higher the overlap, the closer the connection between the two numbers and the common contact, and the more likely they are the same user or have a close relationship. The extraction and integration of communication location information from the number and the historical bills of the number before it lost contact can accurately find the common contact location of the two numbers. The overlap of contact locations can intuitively reflect the similarity between the two numbers in terms of contact locations. The higher the overlap, the higher the overlap of the two numbers in contact locations, and the more likely they are the same user or have a close relationship. By taking a weighted summation of the similarity of usage habits, contact overlap and contact location overlap, it is possible to comprehensively consider information from multiple dimensions and generate a comprehensive and accurate similarity value between the third suspicious number and the number before it lost contact. This value can accurately reflect the degree of association between the two numbers and provide a more reliable basis for loan recovery plans.
[0149] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the method of tracking lost customers based on family relationships in the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0150] This application also provides a lost customer tracking system based on family relationships, please refer to Figure 4 , the lost customer tracking system based on family relationships includes:
[0151] The intimate number confirmation module 10 is used to extract the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact, calculate the intimacy between the number before the loss of contact and the family member number, and select the family member number whose intimacy exceeds a preset intimacy threshold as the intimate number;
[0152] A family relationship network construction module 20 is used to extract the associated numbers of the close numbers from the historical bills corresponding to the close numbers, and to construct a family relationship network based on the number before the loss of contact, the contact number, the close numbers, and the associated numbers;
[0153] Suspicious number confirmation module 30, configured to screen a first suspicious number from the family relationship network, detect a second suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identify a third suspicious number whose language style matches the language style of the number before the loss of contact;
[0154] The target number confirmation module 40 is configured to calculate the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and use the third suspicious number whose similarity exceeds a preset similarity threshold as the target number.
[0155] Optionally, the intimate number confirmation module 10 is further configured to:
[0156] The total call duration, call duration, total number of text messages, and number of days of text message contact between the number before the loss of contact and the family member's number are counted from the historical bills corresponding to the number before the loss of contact, and positive weights are assigned to the total call duration, call duration, total number of text messages, and number of days of text message contact;
[0157] Obtain the most recent communication date between the number before the loss of contact and the family member's number, calculate the number of days between the most recent communication date and the loss date of the lost customer, and assign a reverse weight to the number of days between them;
[0158] The intimacy between the number before the loss of contact and the family member's number is obtained by comprehensively considering the total call duration, number of call contact days, total number of text messages, number of text message contact days and number of interval days.
[0159] Optionally, the family relationship network construction module 20 is further configured to:
[0160] The number before loss of contact, contact number, intimate number and associated number are defined as each network node;
[0161] According to the communication records of the corresponding numbers of each network node, the network node pairs with communication behaviors in each network node are identified, and the network nodes in the network node pairs are associated to obtain a family relationship network.
[0162] Optionally, the suspicious number confirmation module 30 is further configured to:
[0163] Traverse each network node in the family relationship network, determine the target network node in each network node, and the target network node corresponds to the close number;
[0164] Suspicious network nodes associated with at least two target network nodes are screened out from each network node, numbers corresponding to the suspicious network nodes are compared with contact numbers, and the contact numbers are removed from the numbers corresponding to the suspicious network nodes to obtain a first suspicious number.
[0165] Optionally, the suspicious number confirmation module 30 is further configured to:
[0166] Extracting time series data within a preset time period after the lost customer lost contact from historical bills corresponding to the first suspicious number, and identifying abnormal numbers with abnormal behavior patterns among the first suspicious numbers based on the time series data;
[0167] Determine the historical behavior pattern of the number before the lost customer lost contact, screen out matching numbers from the abnormal numbers whose time period distribution matching degree and change trend matching degree with the historical behavior pattern reach the corresponding preset matching degree thresholds, and determine the matching numbers as second suspicious numbers.
[0168] Optionally, the suspicious number confirmation module 30 is further configured to:
[0169] Extracting the text of the SMS logs from the second suspicious number, analyzing the sentence structure, high-frequency vocabulary, and punctuation usage habits of the SMS logs to generate language style features;
[0170] Calculate the language style similarity between the language style features and the historical language style features of the number before the loss of contact;
[0171] The second suspicious number whose language style similarity exceeds a preset style similarity threshold is determined to have a language style match and is marked as a third suspicious number.
[0172] Optionally, the target number confirmation module 40 is further configured to:
[0173] establishing a first usage habit matrix of the third suspicious number based on historical bills of the third suspicious number, and performing similarity calculation between the first usage habit matrix and the second usage habit matrix of the number before the loss of contact, to obtain usage habit similarity;
[0174] Determine the common contact number between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number, and generate a contact overlap degree based on the common contact number;
[0175] Determining a common contact location between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact location overlap based on the common contact location;
[0176] The similarity between the third suspicious number and the number before the loss of contact is generated by performing a weighted summation of the usage habit similarity, the contact overlap, and the contact location overlap.
[0177] The family-based missing customer tracking system provided in this application utilizes the family-based missing customer tracking method described in the aforementioned embodiment, resolving the technical issue of difficulty effectively tracking missing customers. Compared to the prior art, the family-based missing customer tracking system provided in this application has the same beneficial effects as the family-based missing customer tracking method described in the aforementioned embodiment. Other technical features of the family-based missing customer tracking system are the same as those disclosed in the aforementioned embodiment, and are not further elaborated upon here.
[0178] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the lost customer tracking method based on family relationships in the above-mentioned embodiment one.
[0179] Reference below Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, PADs (Portable Application Description: tablet computers), etc., and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0180] like Figure 5As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or a hard disk; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wired to exchange data. Although the figures show electronic devices with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.
[0181] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0182] The electronic device provided in this application, using the family relationship-based missing customer tracking method described in the above embodiment, can resolve the technical problem of difficulty in effectively tracking missing customers. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the family relationship-based missing customer tracking method described in the above embodiment, and the other technical features of the electronic device are the same as those disclosed in the above embodiment, and are not further described here.
[0183] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0184] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0185] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the method for tracking lost customers based on family relationships in the above-mentioned embodiment.
[0186] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0187] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0188] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by an electronic device, the lost customer tracking device based on family relationships: extracts the lost customer's contact number and family member number from the historical bills corresponding to the lost customer's number before loss, calculates the intimacy between the number before loss and the family member number, and screens family member numbers whose intimacy exceeds a preset intimacy threshold as intimate numbers; extracts associated numbers of the intimate number from the historical bills corresponding to the intimate number, and builds a family relationship network based on the number before loss, the contact number, the intimate number and the associated number; screens a first suspicious number from the family relationship network, detects a second suspicious number in the first suspicious number whose behavior pattern matches the historical behavior pattern of the number before loss, and identifies a third suspicious number in the second suspicious number whose language style matches the number before loss; calculates the similarity between the third suspicious number and the number before loss based on the historical bills of the third suspicious number, and takes the third suspicious number whose similarity exceeds the preset similarity threshold as the target number.
[0189] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0190] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0191] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0192] The computer-readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for tracking lost customers based on family relationships. This computer-readable storage medium can address the technical issue of the difficulty in effectively tracking lost customers. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for tracking lost customers based on family relationships provided in the aforementioned embodiments, and are not further elaborated here.
[0193] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for tracking lost customers based on family relationships.
[0194] The computer program product provided in this application can solve the technical problem of the difficulty in effectively tracking lost customers. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the lost customer tracking method based on family relationships provided in the above embodiment, and will not be elaborated here.
[0195] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for tracking lost customers based on family relationships, characterized in that: The method for tracking lost customers based on family relationships includes: Extracting the contact number and family member number of the lost customer from the historical bills corresponding to the number before the loss of contact, calculating the intimacy between the number before the loss of contact and the family member number, and screening the family member numbers whose intimacy exceeds a preset intimacy threshold as close numbers; Extracting the associated number of the intimate number from the historical bills corresponding to the intimate number, and building a family relationship network based on the number before the loss of contact, the contact number, the intimate number and the associated number; Screening a first suspicious number from the family relationship network, detecting a second suspicious number whose behavior pattern matches the historical behavior pattern of the number before the loss of contact, and identifying a third suspicious number whose language style matches the language style of the number before the loss of contact; The similarity between the third suspicious number and the number before the loss of contact is calculated based on the historical bills of the third suspicious number, and the third suspicious number whose similarity exceeds a preset similarity threshold is used as the target number.
2. The method for tracking lost customers based on family relationships according to claim 1, characterized in that: The step of calculating the intimacy between the number before the loss of contact and the family member number includes: The total call duration, call contact days, total number of text messages, and number of text message contact days between the number before the loss of contact and the family member number are counted from the historical bills corresponding to the number before the loss of contact, and positive weights are assigned to the total call duration, the number of call contact days, the total number of text messages, and the number of text message contact days; Obtain the latest communication date between the number before the loss of contact and the family member number, calculate the number of days between the latest communication date and the date of loss of contact of the lost customer, and assign a reverse weight to the number of days; The intimacy between the number before the loss of contact and the family member number is obtained by comprehensively considering the total call duration, the number of call contact days, the total number of text messages, the number of text message contact days and the number of interval days.
3. The method for tracking lost customers based on family relationships according to claim 1, characterized in that: The step of constructing a family relationship network based on the number before the loss of contact, the contact number, the close number, and the associated number includes: The number before the loss of contact, the contact number, the intimate number, and the associated number are defined as network nodes; According to the communication records of the numbers corresponding to the network nodes, network node pairs with communication behaviors among the network nodes are identified, and the network nodes in the network node pairs are associated to obtain a family relationship network.
4. The method for tracking lost customers based on family relationships according to claim 1, wherein: The step of screening the first suspicious number from the family relationship network includes: Traversing each network node in the family relationship network, and determining a target network node among the network nodes, wherein the target network node corresponds to the close number; Suspicious network nodes associated with at least two of the target network nodes are screened out from the network nodes, numbers corresponding to the suspicious network nodes are compared with the contact numbers, and the contact numbers are removed from the numbers corresponding to the suspicious network nodes to obtain a first suspicious number.
5. The method for tracking lost customers based on family relationships according to claim 1, characterized in that: The step of detecting a second suspicious number whose behavior pattern in the first suspicious number matches the historical behavior pattern of the number before the loss of contact comprises: Extracting time series data within a preset time period after the lost customer loses contact from historical bills corresponding to the first suspicious number, and identifying abnormal numbers with abnormal behavior patterns among the first suspicious numbers based on the time series data; Determine the historical behavior pattern of the number before the lost contact with the lost customer before the lost contact, and screen out matching numbers from the abnormal numbers whose time period distribution matching degree and change trend matching degree with the historical behavior pattern reach corresponding preset matching degree thresholds, and determine the matching numbers as second suspicious numbers.
6. The method for tracking lost customers based on family relationships according to claim 1, characterized in that: The step of identifying a third suspicious number in the second suspicious number that matches the language style of the number before the loss of contact includes: Extracting text from the second suspicious number's SMS logs, analyzing sentence structure, high-frequency vocabulary, and punctuation usage habits of the text from the SMS logs, and generating language style features; Calculating language style similarity between the language style feature and the historical language style feature of the number before the loss of contact; The second suspicious number whose language style similarity exceeds a preset style similarity threshold is determined to have a language style match and is marked as a third suspicious number.
7. The method for tracking lost customers based on family relationships according to claim 1, characterized in that: The step of calculating the similarity between the third suspicious number and the number before the loss of contact based on the historical bills of the third suspicious number includes: establishing a first usage habit matrix of the third suspicious number based on historical bills of the third suspicious number, and performing similarity calculation between the first usage habit matrix and the second usage habit matrix of the number before the loss of contact, to obtain usage habit similarity; Determining a common contact number between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact overlap degree based on the common contact number; determining a common contact location between the third suspicious number and the number before the loss of contact based on historical bills of the third suspicious number, and generating a contact location overlap based on the common contact location; A weighted sum is taken of the usage habit similarity, the contact overlap, and the contact location overlap to generate a similarity between the third suspicious number and the number before the loss of contact.
8. An electronic device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for tracking a lost customer based on family relationships as described in any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the lost customer tracking method based on family relationships as described in any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method for tracking lost customers based on family relationships according to any one of claims 1 to 7 are implemented.