Method of validating anonymized user data
An automated method for validating anonymized user data enhances personalization and reduces costs by optimizing communication channels and timing using intelligent algorithms, addressing inefficiencies in existing platforms.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OBSHCHESTVO S OGRANICHENNOJ OTVETSTVENNOSTYU BOTTO
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Existing customer data platforms lack dynamic optimization and real-time personalization capabilities, leading to inefficient communication strategies and increased costs.
An automated method for validating anonymized user data using intelligent algorithms to analyze communication history, user preferences, and behavioral data, optimizing communication channels and timing for personalized interactions.
Improves communication accuracy and reduces costs by ensuring personalized and timely interactions based on comprehensive user data analysis.
Smart Images

Figure RU2024000330_07052026_PF_FP_ABST
Abstract
Description
[0001] METHOD OF VALIDATION OF ANIMALIZED USER DATA
[0002] AREA OF TECHNOLOGY
[0003] The claimed technical solution generally relates to the field of computing technology, and in particular to an automated method for validating anonymized user data.
[0004] LEVEL OF TECHNOLOGY
[0005] A well-known technology is the Mindbox cloud-based marketing automation platform, which helps collect and process customer data from online and offline, automate communications, manage them from a single window, and launch omnichannel communications, promotions, or advertising (https: / / mindbox.ru / ).
[0006] Another example of a customer data platform is the state-of-the-art Twilio Segment platform, which aggregates clean, consistent customer data to generate real-time analytics. Twilio Segment enables data teams to easily prepare, enrich, and activate existing data in the warehouse, enabling marketers to quickly act on personalized communications.
[0007] In addition, the state of the art includes the HLR or Home Location Register technology, which allows one to check the subscriber's telephone number for activity: whether it exists and is accessible.
[0008] The disadvantage of known solutions in this area of technology is that they have a number of limitations in personalization, dynamic optimization and real-time analytics compared to the proposed solution:
[0009] • First, Mindbox focuses primarily on automating existing communication channels, but lacks the functionality to dynamically optimize timing and communication channels based on real-time user portraits. This can lead to insufficiently personalized and less effective campaigns.
[0010] • Second, Twilio Segment provides tools for collecting and processing data, but data integration and omnichannel communications management require additional resources and technologies. This solution does not provide a built-in system for analysis and recommendations for optimizing message sending times depending on individual user preferences, which reduces responsiveness and personalization. • Finally, HLR technology is limited to checking phone number activity without providing data on user preferences and other digital portrait characteristics, making it highly specialized and insufficient for a full-fledged omnichannel strategy.
[0011] The proposed solution differs from existing technologies in its ability to deliver deep personalization based on a comprehensive analysis of the user's digital profile using anonymized user data, including preferences, behavioral data, and timing patterns. The model considers multiple communication channels (email, SMS, calls, etc.) and automatically selects the most optimal time for communication based on the collected data. This significantly improves the accuracy and effectiveness of interactions, reducing communication costs and increasing target audience response.
[0012] ESSENCE OF THE INVENTION
[0013] The proposed technical solution proposes a new approach to validating anonymized user data. It utilizes algorithms for intelligent validation, scoring, and analytics of anonymized user data based on phone numbers and email addresses.
[0014] The technical result consists in expanding the arsenal of technical means for solutions of this type.
[0015] An additional technical result achieved by solving this problem is ensuring the efficiency of communications and increasing the accuracy of scoring.
[0016] The claimed technical result is achieved by implementing a method for validating anonymized user data, which includes the following stages: a) in a database of records concerning user telephone numbers, data is collected regarding successful and unsuccessful communication attempts over a long period of time, factors influencing the success of contact are determined, such as: the frequency of successful calls, the time of day when users most often answer, preferences for communication channels, the data is analyzed, to which the number is linked, normalization and determination of the type of telephone numbers are carried out by converting the telephone number into a string containing only a sequence of digits, as well as by removing all characters except digits, after which the converted number is checked for compliance with the international standard using a given template, the country and type of telephone number are determined using a database of all international codes,and determine whether the telephone number is still active and belongs to the user or inactive; b) in the database for records concerning user email addresses: normalize the data by converting the email address record into a unified format, as well as by excluding extraneous characters from the email address record and cleaning the record format; based on the normalized data, check the syntax of the email address by checking the correctness of the spelling of the record, taking into account the pre-established rules for the formation of the record; determine the existence of the domain to which the email address belongs, and if the domain exists, this indicates the possibility of sending letters to this address; check the existence of an electronic mailbox on the specified mail server by sending a request to the specified electronic mail server,to determine the validity of the specified e-mail address; check the e-mail address for its registration on web resources; c) update each individual entry of the telephone number and e-mail address in the database, taking into account the validation operations performed in stages a) and b).
[0017] DESCRIPTION OF DRAWINGS
[0018] The invention will be further described in accordance with the accompanying drawings, which are provided to illustrate the invention and in no way limit its scope. The following drawings are attached to the application:
[0019] Fig. 1 illustrates a scheme for validating anonymized user data by telephone numbers.
[0020] Fig. 2 illustrates the validation scheme of anonymized user data by email addresses.
[0021] DETAILED DESCRIPTION OF THE INVENTION
[0022] In the following detailed description of the embodiment of the invention, numerous implementation details are provided to provide a clear understanding of the present invention. However, it will be obvious to those skilled in the art how the present invention can be used with or without these implementation details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the features of the present invention. Furthermore, it will be clear from the foregoing that the invention is not limited to the implementation shown. Numerous possible modifications, changes, variations, and substitutions, while maintaining the spirit and form of the present invention, will be apparent to those skilled in the art.
[0023] Below is a detailed example of the implementation of the method for validating anonymized user data.
[0024] The intelligent validation and scoring algorithm analyzes large volumes of data (Big Data) on user communications, including phone calls, instant messaging, and email addresses. The analysis is based on historical interaction data over a long period, taking into account:
[0025] • Successful and unsuccessful communication attempts;
[0026] • Data on user preferences obtained from a big data repository, including profiles in social networks, instant messengers and the history of email interactions;
[0027] • User segmentation based on behavioral data;
[0028] • Predicting the "liveness" of a number.
[0029] At the first stage (Fig. 1) in the claimed solution a) data is collected in a database of records concerning the telephone numbers of users regarding successful and unsuccessful communication attempts over a long period of time, factors influencing the success of the contact are determined, such as: the frequency of successful calls, the time of day when users most often answer, preferences for communication channels, the data is analyzed, what the number is linked to is normalized and the type of telephone numbers is determined by converting the telephone number into a string containing only a sequence of digits, as well as by removing all characters except digits, after which the converted number is checked for compliance with the international standard using a given template, the country and type of telephone number are determined using a database of all international codes, and it is determined whether the telephone number is still active and belongs to the user or inactive.
[0030] Step a) described above is implemented using an ensemble model, gradient boosting (XGBoost) for efficient classification of structured data, and recurrent neural networks (RNNs) for analyzing time series sequences. Gradient boosting classifies structured data (communication success / failure), while RNNs analyze temporal dependencies to identify behavioral patterns, allowing for the consideration of temporal aspects of user interactions and improving the accuracy of predictions regarding number activity and the effectiveness of the selected communication channel. This model analyzes large volumes of data (Big Data) on user communications, such as calls, instant messaging, and email. The analysis is based on historical interaction data over a long period. The model is trained on this data to estimate the probability of successful communication with the subscriber, taking into account:
[0031] • Successful and unsuccessful communication attempts;
[0032] • Segmentation of users based on behavioral data in communications.
[0033] Model training involves several key steps:
[0034] • Data collection: The first stage involves collecting data on communications with users over a long period. This data includes successful and unsuccessful communication attempts across various channels (calls, messages, email, etc.), interaction timestamps, and other behavioral metrics.
[0035] • Data pre-processing: the data then undergoes pre-processing, which includes: o Normalization of phone numbers to bring them to a uniform format; o Removal of duplicates and gaps in the data; o Categorization of interactions (successful / unsuccessful calls, message deliverability, email open rate, etc.); o Highlighting temporal patterns of interactions.
[0036] • Feature engineering: During the feature engineering stage, features are created for the model, such as the frequency of successful contacts, the most successful time periods (e.g., morning or evening), and the communication channel with the highest probability of completing the target action. Timestamps are also added to analyze changes over time.
[0037] • Model training: The data is split into training and testing sets. The model is trained on the training data using gradient boosting to classify successful and unsuccessful interactions and an RNN to analyze the sequence of temporal events. Precision, recall, Fl-score, and / or ROC-AUC metrics are used to evaluate the quality of the model.
[0038] • Model validation: After training, the model is tested on an independent dataset that has not previously been used in training. The model estimates the probability of successful contact and segments users by temporal patterns, providing the result in the form of predictions. The use of cross-validation minimizes possible overfitting errors and improves the stability of predictions.
[0039] User segmentation is based on their behavior across different communication channels. This process involves several steps:
[0040] • Time pattern analysis: Depending on the time of day when users most often respond to communications (e.g. morning, afternoon or evening), users are divided into segments by time preference.
[0041] • Analysis of preferred communication channels: based on successful and unsuccessful communication attempts, the channels through which users prefer to receive information are identified (for example, some users answer calls faster, others prefer instant messengers or email). In this way, users are segmented by preferred channels.
[0042] • Behavioral activity: frequency of interactions is taken into account. Users who actively respond to messages make up one segment, and those who interact less often make up another.
[0043] • Additional behavioral data: An important aspect is using data from past interactions on social media and other platforms to create more precise segments. For example, users who actively engage with advertising offers on social media can be segmented for targeted communication through these channels.
[0044] As a result of segmentation, users are grouped into clusters, for each of which an optimal communication strategy is selected based on channels, time, and message content, as well as the formation of user groups with similar characteristics using K-means methods.
[0045] By analyzing data on successful and unsuccessful calls, the model can estimate the likelihood that a phone number is still active and belongs to the user. This is critical for ensuring effective communications, reducing costs from invalid contacts, and improving scoring accuracy.
[0046] The model constantly enriches the data received from microservices to update digital user profiles and improve prediction accuracy. This is achieved through the integration of new data sources and the continuous self-training of the model based on new input data. The more data the model receives, the more accurate its predictions regarding the likelihood of successful contact become.
[0047] Below is an example of how the algorithm works for scoring phone numbers:
[0048] 1. Data collection: • Data collection and preparation:
[0049] - Historical data: All interactions with users are collected, including successful and unsuccessful communication attempts.
[0050] - User profiles: information from social media, instant messaging, and email is analyzed to create digital profiles. The model studies correlations between various factors (activity time, communication channel preferences) and creates a user profile, predicting their behavior based on this data.
[0051] - Pre-processing: Data are cleaned of gaps (missing values in the data that may arise due to errors in data collection or transmission) and outliers (abnormally high or low values that may distort the analysis), normalized for analysis by cleaning up the gaps by removing records with gaps or filling with median / mean values, statistical methods (interquartile range (IQR) method) are used to remove outliers.
[0052] • Modeling and feature extraction: a) Feature selection: identifying factors that influence the success of a contact, such as:
[0053] Frequency of successful calls.
[0054] Time of day when users most often respond. Communication channel preferences. b) Creation of new features using the following methods:
[0055] - One-hot encoding for categorical data by types of communication channels with their transformation into numerical features in machine learning algorithms - the algorithms classify the data, distributing it into categories (successful / unsuccessful, preferred channels, etc.): logistic regression for classifying successful / unsuccessful communication attempts; gradient boosting for identifying complex dependencies between communication channels and interaction results; decision trees for determining key factors of communication success).
[0056] - A moving average for time series to account for trends in activity, using a weighted moving average (WMA), which places more weight on recent data but also takes into account past values with a given weight, or an exponential moving average (EMA), which is more sensitive to recent data, allowing for better tracking of recent changes in behavior and activity. Both methods help quickly adapt the model to new data, which is especially important in user behavior analysis tasks, where activity can change quickly due to factors such as stocks, seasonality, or news. • Training the model using machine learning algorithms (after data collection, preprocessing and normalization of the data, splitting the data into training and test sets), such as:
[0057] Logistic regression for binary classification (success / failure).
[0058] - Regression tree as a decision support tool and combining classifier algorithms into an ensemble to improve predictive power.
[0059] - Deep learning using recurrent neural networks (RNNs) for more complex patterns that consider data sequences, such as user behavior over time. Complex patterns are regularities that are not always obvious and may include dependencies between multiple data parameters: temporal dependencies (RNNs process data sequences, preserving information about temporal dependencies to infer preferred times for contacting users); multivariate dependencies (dependencies between content type (informational messages or advertising), time of day, and communication channel, to identify user preferences for communication types and channels at specific times). Patterns may include responses to different types of content (advertising, informational messages, etc.).
[0060] 2. Data analysis:
[0061] • Using machine learning algorithms to process historical data: gradient boosting, logistic regression, RNN, decision trees to process time series data, such as frequency of interactions, response to certain communication channels, time of day preferences
[0062] • Application of classification methods with logistic regression to determine the probability of successful contact based on:
[0063] Frequencies of successful and unsuccessful communication attempts.
[0064] - User profiles, including information from social networks and instant messengers.
[0065] Segmentation of users based on communication behavior.
[0066] Based on this data, the model assigns a probability of successful contact, learning from historical data. For example, if a user answers calls more often in the morning, the model will predict a higher probability of success in the morning.
[0067] 3. Number scoring: each number is given a score based on a set of parameters:
[0068] • Success / failure of communication: +1 point for successful contact, -1 point for unsuccessful. • Source reliability: assessment on a scale (from 1 to 10 points depending on the reliability of the data source; the reliability of the source is also determined by the expert assessment method (Delphi Method) and the purity of the source data).
[0069] • Data age: More points are awarded for more recent data.
[0070] • Summing up the points to form the overall scoring of the room.
[0071] 4. Activity forecasting:
[0072] Using models to estimate the probability of number activity based on collected data and scoring:
[0073] • Scoring model: based on the trained data, a model for assigning points to records in Big Data is formed, which assigns activity points to a number, taking into account:
[0074] Historical frequency of successful interactions.
[0075] Data about user preferences and behavior.
[0076] - Digital portrait data.
[0077] • Activity probability assessment:
[0078] - Logistic regression to predict the probability that a number is active. The model analyzes the input data and predicts the probability based on:
[0079] Successes and failures for each issue.
[0080] - Reliability of the source data.
[0081] • Threshold: A threshold (e.g. 0.7) is defined above which the number is considered active (the threshold can change dynamically based on actual results).
[0082] • Testing and validation:
[0083] - Cross-validation: The model is cross-validated to check its accuracy and robustness.
[0084] - Metrics: Metrics such as precision, recall, and Fl-score are evaluated to determine the performance of the model.
[0085] Advantages :
[0086] • Optimization of communication costs by increasing the accuracy of predictions.
[0087] • Improve user experience through personalized offers.
[0088] • Continuous self-learning and adaptation of the model based on new data.
[0089] Technologies used:
[0090] • Big Data for analysis and prediction. • Continuous Learning to improve the accuracy of predictions based on incoming data.
[0091] • Microservices for flexibility and scalability of architecture.
[0092] The process of normalizing telephone numbers is carried out as follows.
[0093] 1. The resulting record containing the phone number is converted into a string containing only a sequence of digits. All characters except digits are removed. This is necessary to standardize data entry and further work with the number in its standard form.
[0094] 2. Check the received record for compliance with the international standard E.164 (this is an international standard that is used for telephone numbers all over the world, including public networks) using a regular expression A \+[l-9]\d{ l,14}$.
[0095] 2.1 After the first step of normalizing the number (removing unnecessary characters), we check it for compliance with the international E.164 standard, which defines the format of telephone numbers. The E.164 standard specifies that a telephone number must begin with the "+" symbol, followed by a country code (1 to 3 digits), and any valid telephone number can contain up to 15 digits.
[0096] 2.2. To perform the check, the regular expression " Л \+[1- 9]\d{ 1 ,14}$":
[0097] - Checks that the number starts with a symbol
[0098] - Checks that the first digit after "+" is in the range from 1 to 9 (country code cannot start with 0);
[0099] - Limits the length of the remaining part of the number after the "+" symbol to 15 digits.
[0100] 2.3. If the number does not match the pattern " A \+[l-9]\d{l,14}$" - it is considered invalid.
[0101] Example:
[0102] - The entered number: "+123456789012345" complies with the E.164 standard, as it starts with "+" and contains the permitted number of digits.
[0103] - The entered number: "012345678901234" does not match because the "+" symbol is missing and it begins with an invalid digit "0".
[0104] 3. If the number complies with the E.164 standard, the country and type of number (mobile, landline, etc.) are determined programmatically:
[0105] 3.1. After successful validation of the number according to the E.164 standard, the software determines the country to which the processed number belongs. This is achieved by analyzing the initial digits of the number, which correspond to the international country code. For example: io - code "1" - for countries of the North American Numbering Plan (NANP), such as the USA and Canada, etc.
[0106] . "44" — для Great Britain;
[0107] - "7" - for Russia, Kazakhstan, South Ossetia and Abkhazia;
[0108] - "90" - for Turkey;
[0109] - "998" - for Uzbekistan, etc.
[0110] 3.2. The country is determined using a database of all international codes.
[0111] 3.3. After identifying the country, the software verifies the number type—mobile, landline, or special (e.g., service numbers or premium-rate numbers)—using current databases containing information on the allocation of number ranges for each country. For example, in some countries, mobile numbers begin with specific prefixes (e.g., in the UK, mobile numbers begin with "7" followed by the UK international country code "44").
[0112] Example:
[0113] - The number "+12012031434" will be recognized as USA-FIXED_LINE_OR_MOBILE
[0114] - The number "+5491120028197" will be recognized as Argentina-MOBILE
[0115] - The number "+74955316650" will be recognized as Russian Federation-FIXED LINE, etc.).
[0116] The next step of the claimed method (Fig. 2) is b) in the database for records concerning user e-mail addresses: normalizing the data by converting the e-mail address record into a unified format, as well as by excluding extraneous characters from the e-mail address record and cleaning the record format; based on the normalized data, checking the syntax of the e-mail address by checking the correctness of the spelling of the record, taking into account the pre-established rules for forming the record; determining the existence of the domain to which the e-mail address belongs, and if the domain exists, this indicates the possibility of sending letters to this address; checking the existence of an e-mail box on the specified mail server by sending a request to the specified e-mail server to determine the validity of the specified e-mail address;check the email address to see if it is registered on web resources;
[0117] When processing email data, a separate validation and scoring function is used: li - Normalization: Converts email to a standard format by removing unnecessary characters and correcting the format. This eliminates errors caused by incorrect spelling and makes the email readable and acceptable.
[0118] - Syntax checking: After normalization, the email is checked for correct spelling, taking into account all email address formation rules. This step excludes addresses with invalid characters or incorrect formatting and is implemented using built-in methods and regular expressions.
[0119] - MX check: Determines whether the domain to which the address belongs exists. If the domain exists, this indicates that emails can be sent to this address.
[0120] - SMTP check: Checks the existence of a mailbox on the specified server. The system sends a request to the mail server to determine whether the mailbox exists. This allows you to exclude non-existent or blocked emails.
[0121] - Activity Status Check: For email services such as Gmail and Mail.ru, this feature allows you to check the activity status. For example, you can determine the last time a user's Gmail account was updated, as well as access the user's profile photo and comments. For Mail.ru, you can access the last login time, as well as additional user information, if available (age, name, country, city).
[0122] - Web Resource Registration Check: Determines whether a given email address is registered on various web resources. This feature is useful for marketing, security, and user activity analysis.
[0123] The next step is to c) update each individual phone number and email address record in the database, taking into account the validation operations performed in steps a) and b).
[0124] In the proposed solution, the model continuously enriches data received from microservices to update the repository and improve prediction accuracy. This is achieved by integrating new data sources and continuously training the model based on new input data. The more data the model receives, the more accurate its predictions regarding user preferences and the likelihood of successful contact become.
[0125] Below is an example of the general appearance of a computing system that ensures the implementation of the declared method of validating anonymized user data or is part of a computer system, for example, a server, a personal computer, part of a computing cluster, processing the necessary data to implement the declared technical solution.
[0126] A computing system that provides the data processing necessary for the implementation of the claimed solution generally contains the following components: one or more processors, at least one memory, a data storage means, input / output interfaces, an input means, and network interaction means.
[0127] When executing machine-readable commands contained in the RAM, the processor of the device is configured to perform the basic computing operations necessary for the operation of the device or the functionality of one or more of its components.
[0128] Memory is typically implemented as RAM, where the necessary software logic is loaded to provide the required functionality. When implementing the proposed solution, the memory capacity required for its implementation is allocated.
[0129] The data storage device can be an HDD, SSD, RAID array, network storage, flash memory, etc. The device allows for long-term storage of various types of information.
[0130] Interfaces are standard means for connecting and operating peripherals and other devices, such as USB, RS232, RJ45, COM, HDMI, PS / 2, Lightning, etc.
[0131] The choice of interfaces depends on the specific device design, which may be a personal computer, mainframe, server cluster, thin client, smartphone, laptop, etc.
[0132] A keyboard may be used as a data input device in any embodiment of the system implementing the described method. The keyboard hardware may be any known hardware implementation: it could be a built-in keyboard used on a laptop or netbook, or a separate device connected to a desktop computer, server, or other computing device. The connection may be either wired, in which the keyboard cable is connected to a PS / 2 or USB port located on the desktop computer's system unit, or wireless, in which the keyboard exchanges data via a wireless communication channel, such as a radio channel, with a base station, which, in turn, is directly connected to the system unit, for example, to one of the USB ports.In addition to the keyboard, data input devices may also include: a joystick, display (touch screen), projector, touchpad, mouse, trackball, light pen, speakers, microphone, etc.
[0133] Network communication tools are selected from devices that provide network data reception and transmission, such as an Ethernet card, WLAN / Wi-Fi module, Bluetooth module, BLE module, NFC module, IrDA, RFID module, GSM modem, etc. These tools enable data exchange via a wired or wireless data transmission channel, such as WAN, PAN, LAN, Intranet, Internet, WLAN, WMAN, or GSM.
[0134] The device components are connected via a common data bus.
[0135] These application materials present a preferred implementation of the claimed technical solution, which should not be used as limiting other particular embodiments of its implementation that do not go beyond the scope of the requested scope of legal protection and are obvious to specialists in the relevant field of technology.
Claims
PC17RU2024 / 000330 CLAUSES OF THE INVENTION L A method for validating anonymized user data, which includes the steps of: d) collecting data on successful and unsuccessful communication attempts over a long period of time from a database of records concerning user telephone numbers, determining factors that influence the success of the contact, such as: the frequency of successful calls, the time of day when users most often answer, preferences for communication channels, analyzing the data to which the number is linked, normalizing and determining the type of telephone numbers by converting the telephone number into a string containing only a sequence of digits, as well as by removing all characters except digits, after which checking the converted number for compliance with an international standard using a given template, determining the country and type of the telephone number using a database of all international codes, and determining whether the telephone number is still active and belongs to the user or inactive;e) in the database for records relating to user email addresses: normalise the data by converting the email address record into a uniform format, as well as by excluding extraneous characters from the email address record and cleaning the record format; on the basis of the normalised data, check the syntax of the email address by checking the correctness of the spelling of the record, taking into account the pre-established rules for forming the record; determine the existence of the domain to which the email address belongs, and if the domain exists, this indicates the possibility of sending letters to this address; check the existence of the email box on the specified mail server by sending a request to the specified email server to determine the validity of the specified email address;check the email address for registration on web resources; f) update each individual entry of the telephone number and email address in the database, taking into account the validation operations performed in stages a) and b).
Citation Information
Patent Citations
Monitoring user activity on smart mobile devices
EP2608041B1
Method of forming and structuring an electronic database
RU2696295C1
Method and system of depersonalized assessment of clients of organizations for carrying out operations between organizations
RU2795371C1
Reputation scoring and reporting system
US20100114744A1
System and method for analyzing user device information
US20150370814A1