Method and system for identifying natural persons

WO2026177630A1PCT designated stage Publication Date: 2026-08-27PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/RU2025/000046
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2025-02-21
Publication Date
2026-08-27

Smart Images

  • Figure RU2025000046_27082026_PF_FP_ABST
    Figure RU2025000046_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method for identifying natural persons (NPs) in a register of profiles of NPs, implemented by a computing device, comprises the steps of: extracting, from an external database, information about NPs, which comprises parameters of NPs; forming a preliminary list of NPs based on the extracted information; extracting a set of values for significant parameters of the NPs; determining, with the use of a pre-trained neural network, weight values which are to be assigned when the values for NP parameters match or do not match; comparing the values of all parameters of a profile of an NP with the NP parameters from the preliminary list in order to determine the degree of similarity of the retrieved NP profiles to NPs from the preliminary list, taking the aforementioned weight values into account; and determining, on the basis of the values for the degree of similarity of the retrieved NP profiles, an NP profile which is similar to an NP from the preliminary list. The technical result consists in increasing the accuracy of identifying NPs in a register of profiles of NPs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD AND SYSTEM FOR IDENTIFICATION OF INDIVIDUALS

[0002] AREA OF TECHNOLOGY

[0003]

[0001] The present invention relates generally to computing technology, and in particular to a method and system for identifying individuals (PI) by analyzing personal data of PI obtained from various sources, which may differ both in the composition of attributes and in the values ​​of these attributes, for the subsequent combination and analysis of data that relate to the same PI, for example, in CRM systems, social networks, geolocation services or for authorization of users in various services or the formation, on the basis of this data, of various notifications containing personal data, both to the user and to third-party organizations, for example, tax returns to the Federal Tax Service, etc.

[0004] LEVEL OF TECHNOLOGY

[0005]

[0002] The prior art provides solutions aimed at processing customer data to form a single profile with current customer data.

[0006]

[0003] For example, a method and device for collecting data for a single client profile are known, disclosed in Russian Patent No. 2781 767, published on October 17, 2022. This document describes a method for updating data of a single client profile (SCP), performed by at least one computing device, comprising the steps of: receiving a request to save personal data (PD) of a client in the SCP from a PD source; checking the client's PD for compliance with specified requirements and standardizing the PD; assigning a status to the client's PD, indicating whether the PD is verified, or characterizing the level of trust in the client's PD based on a validity marker contained in the request, or based on the level of trust in the data source; determining the type of client's PD; searching for the SCP in the SCP database; determining that the found SCP already stores PD of a given type of client; determining whether the PD of a given type is unique; determining the status of the client's PD stored in the found SCP;compare the status of the client's personal data received in the aforementioned request and the status of the client's personal data stored in the retrieved EPC; based on the results of the comparison of the statuses of the client's personal data, update the client's personal data in the EPC.

[0004] A disadvantage of the known solution is the lack of mechanisms to identify a client as an individual based on the information contained in the register of individuals, or to identify two sets of data as belonging to the same individual if some data in the aforementioned sets differ. Also, the known solution is not equipped with technical means, in particular a neural network, to determine the weighting coefficients used to determine the similarity of profiles.

[0007] DISCLOSURE OF THE INVENTION

[0008]

[0005] The technical problem or task posed by this invention is to create a new, effective, simple and reliable solution for identifying FL.

[0009]

[0006] The technical result achieved by performing the above-mentioned task is an increase in the accuracy of identification of individuals in the register of individual profiles.

[0010]

[0007] The specified technical result is achieved by implementing a method for identifying an individual in a register of individual profiles, performed by at least one computing device, containing the steps of:

[0011] - extract information about the FL, containing the FL parameters, from an external database (DB);

[0012] - form, on the basis of the extracted information, a primary list of individuals intended for identifying individuals in the register of individual profiles;

[0013] - extract a set of values ​​of significant parameters of the FL, where the significant parameters are pre-set by the developer;

[0014] - access the register of FL profiles to search for FL profiles whose significant parameter values ​​match the values ​​of the corresponding significant parameters of the FL;

[0015] - determine, using a pre-trained neural network, the values ​​of the weights that should be assigned when the values ​​of the FL parameters match or do not match;

[0016] - compare the values ​​of all parameters of the FL profile with the FL parameters from the primary list to determine the degree of similarity of the found FL profiles with the FL from the primary list, taking into account the mentioned weight values;

[0017] - based on the values ​​of the degree of similarity of the found FL profiles, an FL profile similar to the FL from the primary list is determined.

[0008] In one of the particular examples of implementing the method, the coincidence of significant parameters is also determined in the case where the parameter value is missing.

[0018]

[0009] In another particular example of implementing the method, an additional step is performed to determine the presence of original profiles among the found FL profiles, wherein the FL profile similar to the FL from the primary list is determined only for the original profiles.

[0019]

[0010] In another particular example of implementing the method, an additional step is performed in which:

[0020] - determine the absence of original profiles among the found FL profiles;

[0021] - for all found FL profiles that are duplicates, a search is performed for the corresponding originals;

[0022] - assign a degree of similarity to the found originals equal to the degree of similarity of their duplicates, and the profile of the individual similar to the FL from the primary list is determined only for the original profiles.

[0023]

[0011] In another particular example of the method implementation, the step of determining the profile of an FL similar to the FL from the primary list comprises the steps of:

[0024] - select from the found profiles the FL profile with the highest degree of similarity;

[0025] - compare the degree of similarity of the selected FL profile with the similarity thresholds; - determine that the degree of similarity of the selected profile is equal to or greater than the upper similarity threshold;

[0026] - form a link to the selected individual profile in the information about the individual in the external database.

[0027]

[0012] In another particular example of the implementation of the method, the step of determining the profile of an FL similar to the FL from the primary list comprises the steps of:

[0028] - determine the presence of original profiles among the found FL profiles;

[0029] - select from the found original profiles the FL profile with the highest degree of similarity;

[0030] - compare the degree of similarity of the selected FL profile with the similarity thresholds; - determine that the degree of similarity of the selected profile is less than the upper similarity threshold;

[0031] z- for all found FL profiles that are duplicates, a search is performed for the corresponding originals;

[0032] - assign a degree of similarity to the found originals equal to the degree of similarity of their duplicates;

[0033] - select a profile with the highest degree of similarity among the previously selected FL profile and the found originals;

[0034] - determine that the degree of similarity of the selected profile is equal to or greater than the upper similarity threshold;

[0035] - form a link to the selected individual profile in the information about the individual in the external database.

[0036]

[0013] In another particular example of implementing the method, additional steps are performed in which:

[0037] - form from the primary list a list of individuals for whom identification has not been performed;

[0038] - determine the degree of similarity between individuals, and if the degree of similarity between individuals is equal to or higher than the upper threshold of similarity, then such individuals are considered to be the same individual;

[0039] - groups are formed from FLs that have been defined as the same FL; - in each group, the main FL is selected, and the rest are FLs that complement it, and the FL that is not included in any of the mentioned groups is defined as the main one without FLs that complement it;

[0040] - for each primary individual, a new profile is created in the registry based on the data of the primary individual;

[0041] - enrich the created profiles with data from individuals that complement them;

[0042] - send a command to the database to generate a link to the profile created for him in the information about the main individual and additional individuals.

[0043]

[0014] In another particular example of implementing the method, an additional step of adjusting the significant parameters is performed in the event that no FL profile was found for the FL from the primary list.

[0044]

[0015] In another preferred embodiment of the claimed solution, a FL identification system is presented, containing at least one computing device and at least one memory containing machine-readable instructions, which, when executed by at least one computing device, perform the above-mentioned method. BRIEF DESCRIPTION OF THE DRAWINGS

[0045]

[0016] The features and advantages of the present invention will become apparent from the following detailed description of the invention and the accompanying drawings, in which:

[0046]

[0017] Fig. 1 shows a general diagram of the interaction of elements of the FL identification system.

[0047]

[0018] Fig. 2 shows an example of a general view of a computing device.

[0048] IMPLEMENTATION OF THE INVENTION

[0049]

[0019] Below, the concepts and terms necessary for understanding this technical solution will be described.

[0050]

[0020] In this technical solution, the term “system” means, among other things, a computer system, a computer (electronic computer), a numerical control (CNC), a PLC (programmable logic controller), computerized control systems and any other devices capable of performing a given, clearly defined sequence of operations (actions, instructions).

[0051]

[0021] A command processing unit is an electronic unit or integrated circuit (microprocessor) that executes machine instructions (programs).

[0052]

[0022] The command processing unit reads and executes machine instructions (programs) from one or more data storage devices. The data storage devices may include, but are not limited to, hard disk drives (HDD), flash memory, ROM (read-only memory), solid-state drives (SSD), and optical drives.

[0053]

[0023] A program is a sequence of instructions intended for execution by a computer control device or a command processing device.

[0054]

[0024] Database (DB) - a collection of data organized in accordance with a conceptual structure that describes the characteristics of this data and the relationships between them, and such a collection of data that supports one or more application areas (ISO / IEC 2382:2015, 2121423 "database").

[0055]

[0025] A signal is a material embodiment of a message for use in transmitting, processing, and storing information.

[0026] A logic element is an element that implements certain logical relationships between input and output signals. Logical elements are commonly used to construct logical circuits for computers and discrete automatic control and management circuits. All types of logic elements, regardless of their physical nature, are characterized by discrete values ​​of input and output signals.

[0056]

[0027] In accordance with the diagram shown in Fig. 1, the claimed system 100 for identifying individuals comprises: a primary DB 1 and a person registration device 2. The said elements of the system 100 can be connected by means of widely known wired or wireless communication means that provide the ability to receive and transmit information by generating corresponding signals.

[0057]

[0028] The person registration device 2 may be implemented on the basis of at least one computing device and is equipped with: a register 10, a primary data processing module 20, a data comparison module 30, a decision-making module 40, a weight determination module 50 and a profile creation module 60. The said modules may be implemented on the basis of logical elements and known means for generating and converting signals that provide the possibility of exchanging data between the modules. The weight determination module 50 may be equipped with at least one neural network consisting of input, output and other layers, as well as an encoder and a coder that allows information to be converted into a format, for example, a vector format, for feeding it to the input layer and converting the information obtained from the output layer into a machine-readable format.

[0058]

[0029] At the first stage, in accordance with a given algorithm or a command received from the operator of the device 2 for recording persons, the primary data processing module 20 accesses the primary database 1, in which information about individuals (IP) is stored, using known methods, to extract the values ​​of the IP parameters and form a primary list of IP.

[0059]

[0030] The generated primary list of individuals may contain the parameter values ​​of at least one individual that allow for the identification of the individual, such as: full name, date of birth, INP (unique taxpayer identifier), DUL data (type, series and number of the identity document), SNILS, INN in the Russian Federation, INN in the country of citizenship, address in the Russian Federation, address in a foreign country (Address in a foreign country), telephone number (work, personal), email address, place of work, position, bank card number, etc. Information about the individual may be added to the primary DB 1 by loading data on the individual's transactions from external source systems. For example, source systems may include payment systems, tax accounting systems, loyalty management systems, social networks, as well as any other systems that contain the personal data of the individual, etc.Source systems can also be given a priority, in particular payment systems can be given the highest priority, while social networks can be given the lowest priority.

[0060]

[0031] The generated primary list of individuals is sent by module 20 to data comparison module 30, which extracts a set of values ​​of significant parameters of the individual for each individual from the primary list. The list of significant parameters may be specified by the developer of module 30 and include the following parameters: full name, date of birth, individual taxpayer identification number (INP), insurance number of individuals (SNILS), taxpayer identification number (INN) in the Russian Federation, and taxpayer identification number (INN) in the country of citizenship.

[0061]

[0032] Next, for each individual from the primary list, module 30 accesses register 10, which contains current individual profiles containing information about the individual, to search for individual profiles in which the values ​​of significant parameters match the values ​​of the corresponding significant parameters of the individual. A match of significant parameters is also determined by module 30 in the case where the value of the parameter is missing, for example, when the SNILS is not specified (see Table 1). Moreover, since an individual profile cannot be completely empty, since mandatory parameters that allow for the identification of the individual are always required to be entered during profile registration, individual profiles will always contain parameters with the exception of those parameters whose indication is optional during profile registration.

[0062]

[0033] If module 30 finds at least one FL profile, module 30 proceeds to determine the degree of similarity of each of the found FL profiles with an individual from the primary list.

[0063]

[0034] To determine the said degree of similarity, module 30 compares the value of each parameter of the individual's profile with the value of the corresponding parameter of the individual from the primary list. In particular, the last name is compared with the last name, the first name with the first name, the middle name with the middle name, etc., to determine the completeness of the parameters and their similarity. Before the comparison, spaces and special characters are removed from the values, and the letter portion of the value is converted to uppercase. In particular, the following parameters are compared: Last Name, First Name, Middle Name, Date of Birth, Taxpayer Identification Number (INP), Legal Entity Number (DUL), Taxpayer Identification Number (INN) in the Russian Federation, Taxpayer Identification Number (INN) in the country of citizenship, SNILS, Address in the Russian Federation, Address in the country of registration.

[0064]

[0035] The degree of similarity is determined by the following formula:

[0065] 5 Similarity level = (£Weight i + WeightNameOtdDR + WeightRazl_DUL_INP + WeightRazl_DUL_INNRF + WeightRazl _DUL_INNIno + WeightNameOtdDR_INP_razlDUL + WeightNon_sovp_DUL_INP_INNRF_INNIno_SNILS_Adr_AdrIno ) / 100, where:

[0066] - WeightNamePatronymic - the weight value assigned when the values ​​of the parameters “name”, “patronymic” and date of birth coincide simultaneously;

[0067] 10 - WeightDiscr_DUL_INP - the weight value assigned when the values ​​of the DUL parameters, for example, the series or passport number, and the INP, do not match at the same time;

[0068] - Weight_DUL_INNRF - the weight value assigned when the values ​​of the DUL and INN parameters in the Russian Federation do not match;

[0069] - WeightRazl_DUL_INNino - the weight value assigned when the values ​​of the DUL and INN parameters in the country of citizenship do not match at the same time;

[0070] WeightNameOttDR_INP_razlDUL - the weight value assigned when the values ​​of the parameters “name”, “patronymic”, date of birth and INP coincide simultaneously and the values ​​of the DUL parameters do not coincide;

[0071] 20 - Weight Does Not Match DUL_INP_INNRF_INNIno_SNILS_Adr_AdrIno - the weight value assigned when the values ​​of the parameters DUL, INP, INNRF, INNino, SNILS, Russian Federation address and address in the country of registration do not match simultaneously; - XWeight i - the sum of the weight values ​​that are assigned as a result of comparing the values ​​of each parameter separately.

[0072]

[0036] Table 1 below provides an example of the values ​​of the parameters of an individual from the primary list (line 1) and the values ​​of the parameters of the FL profile (line 2).

[0073] Table 1 zo

[0074]

[0075]

[0037] Table 2 below provides an example of the weight values ​​that are determined to calculate the degree of similarity of the FLs being compared.

[0076] Table 2

[0077]

[0078]

[0038] Below in Tables 3 and 4, examples of the values ​​of weights that are determined to calculate the degree of similarity of the compared FLs are also presented.

[0079] Table 3

[0080]

[0081] Table 4

[0082]

[0083]

[0039] Accordingly, for the example above, the degree of similarity = (Weight_1_1 + Weight_2_1 + Weight_3_2 + Weight_4_1 + Weight_5_2 + Weight_6_1 + Weight_7_2 + Weight_8_4 + Weight_9_3 + Weight_10_1 + Weight_11 4 + WeightNameOfDR + WeightRazlDUL_INNRF + WeightNameOfDR_INP_razl UL) / 100.

[0084]

[0040] The values ​​of the weights that should be assigned when the values ​​of the above parameters match or do not match can be determined by the weight determination module 50 using a neural network that has been previously trained on a labeled data sample. The neural network must be trained in advance, and the weights obtained as a result of the neural network operation must be entered by the operator through the software interface of the person registration device 2 and, upon the operator's command, saved in DB1 as general system parameters.

[0085]

[0041] Accordingly, for training the neural network, module 50 may be equipped with a memory from which module 50 extracts a labeled data sample or a portion of a labeled data sample containing the number of parameters necessary for the assessment of the degree of similarity performed by module 30. For example, a labeled data sample may contain the following list of parameters - attributes:

[0086] - Surname;

[0087] - Name;

[0088] - Surname;

[0089] - Date of birth;

[0090] - DUL data (type, series and number of the identity document); - INP (unique taxpayer identifier in the source system); - TIN in the Russian Federation;

[0091] - INN INO (TIN in a foreign country);

[0092] - SNILS;

[0093] - Address in the Russian Federation;

[0094] - Address Foreign (Address in a foreign country).

[0095]

[0042] The training algorithm is constructed using a classical supervised learning model and is a modification of one of the standard evolutionary search algorithms (genetic algorithm). Based on the training data set—a set of cases prepared by tax accounting specialists—a model is constructed that predicts the class label (multi-class classification) for an object (a pair of individuals). The object of analysis is a pair of data sets about individuals. The analysis factors are the attribute values ​​of the individuals. The target variable is an integer indicating the degree of coincidence of the two individuals. The higher the number, the greater the degree of coincidence.

[0096]

[0043] Each individual in the learning algorithm is represented by a set of 50 weight values:

[0097] • Four weights for each of the 11 FL attributes:

[0098] o Weight, if the attribute values ​​match

[0099] o Weight if the attribute values ​​are different

[0100] o Weight, if only one of the two individuals being compared has the attribute value filled in

[0101] o Weight, if neither of the two individuals being compared has the attribute value filled in.

[0102] io• Six additional weights for specific combinations of attribute values:

[0103] o The name, patronymic, and date of birth coincided at the same time

[0104] o At the same time, the DUL and INP are different

[0105] o The DUL and INN are different at the same time

[0106] o At the same time, the DUL and INN are different

[0107] o The Name, Patronymic, Date of Birth, and INP coincided at the same time and the DULs were different at the same time

[0108] o At the same time, different DUL, INP, INN, INN Foreign, SNILS, Address in the Russian Federation, Address Foreign

[0109]

[0044] The first step of the training module algorithm is to generate 10,000 mutations of the first “parent” individual, for which the default weight values ​​are set:

[0110] • 10 - if the attribute values ​​(combination of attribute values) match;

[0111] • -5 - if the attribute values ​​(combination of attribute values) are different; • 0 - if the attribute value is not filled in for one or both individuals.

[0112]

[0045] Each mutation is produced by changing three randomly selected weights of the "parent" individual. The new weight value is also chosen randomly within the range specified in the algorithm parameters.

[0113]

[0046] For each mutation obtained, the sum of the weights for each case in the training set is calculated.

[0114]

[0047] Next, the result of the comparison of the values ​​of the FL attributes is determined, depending on the range of values ​​into which the resulting sum of weights falls:

[0115] • TRUE - if the sum of the weights is more than 50, then the individuals are considered identical.

[0116] • UNKNOWN - if the sum of the weights lies in the range from -50 to 50, then the result is considered uncertain and is subject to manual review.

[0117] FALSE - if the sum of the weights is less than -50, the individuals are considered different.

[0048] The fitness function is applied to the obtained results. This function is used to indicate which discrepancy is less critical and which is more so in the event of a discrepancy between the case result and the desired result:

[0118] • If the obtained and desired results coincide, then the obtained result is assigned the maximum degree of conformity.

[0119] • If the obtained and desired results do not match, then the degree of conformity is calculated using a formula. The formula uses increasing and decreasing coefficients to “shift” the degree of conformity upward or downward:

[0120] o A, B, C, D, k1, k2, k3, k4 are positive integers

[0121] o A = MAX (A, B, C, D)

[0122] o Summ - the sum of weights calculated in the previous step

[0123] • Formulas for the “shift” of the degree of conformity for different combinations of the obtained and desired results are given in Table 5.

[0124] • In this way we can configure the algorithm so that, for example, if we get TRUE with the expected result FALSE, this is much more critical (less desirable result) than if we expected FALSE but got UNKNOWN.

[0125] • Thanks to this function, the weight selection algorithm arrives at the desired result faster.

[0126] Table 5. Formulas for the "bias" of the degree of correspondence

[0127]

[0128] PCI7RU2025 / 000046

[0129]

[0130]

[0049] After applying the objective function, the fitness level of each individual in the population is calculated, which is equal to the sum of the degrees of correspondence between the obtained results and the desired results for all cases. 5

[0050] The next step is to sort the individuals by fitness level, and select the number of the fittest individuals specified in the algorithm parameters. Then, for each of the selected fittest individuals, two mutations are made by changing three weights of the "parent" individual.

[0131] 10

[0051] The selected individuals with maximum fitness and their mutations form the next generation, for which the selection of the most fit individuals is again carried out in accordance with the previously mentioned sorting of individuals by degree of fitness.

[0132]

[0052] The algorithm for generating new generations and selecting the best individuals in them is performed until an individual (mutation) or several individuals (mutations) are obtained, for which the obtained result (TRUE, FALSE or UNKNOWN) coincides with the desired one for all cases.

[0133]

[0053] If a match across all cases cannot be achieved due to inconsistencies in the training data, the search for the 20 best mutations is terminated when, over several consecutive generations, the number of cases for which the obtained and desired results match has not increased. The number of such generations is specified in the algorithm parameters.

[0134]

[0054] After the ML algorithm has selected suitable weight values, these weights are applied to identify individuals. If new training cases emerge, the neural network training algorithm can be re-run to obtain more accurate weight values ​​that take these new cases into account.

[0135]

[0055] After the data comparison module 30 has received the said weight values ​​from the weight determination module 50, the said module 30 carries out for each previously found FL profile a determination of its degree PC17RU2025 / 000046

[0136] Similarity to the individual from the primary list using the formula mentioned above. The list of profiles and the resulting similarity values ​​are then sent to decision module 40, which makes a decision on the similarity of the individual's profile to the individual from the primary list.

[0137]

[0056] After receiving the aforementioned list of profiles, module 40 analyzes the profiles for the presence of original profiles. If the list of profiles contains at least one profile that is an original, then decision-making module 40 selects from the aforementioned list all profiles that are original. The original profile of an individual is the profile that does not contain information indicating that this profile is a duplicate of another individual profile. Information that the individual profile is a duplicate may be added by the operator during the process of editing information in register 10, before the start of the individual identification process.

[0138]

[0057] If there are no originals in the list of profiles, then for all profiles in the list that are duplicates, module 40 searches for the corresponding originals in the register 10. Then module 40 assigns a degree of similarity to the found originals equal to the degree of similarity of their duplicates.

[0139]

[0058] Next, module 40 selects from the originals the profile with the highest degree of similarity. If there are several profiles with the highest degree of similarity, module 40 selects from them the profile with the data received from the source system with the highest priority. If there are several profiles with the highest priority of the source system, module 40 selects from them the profile with the highest profile entry identifier in registry 10, i.e., the most recently saved profile entry (the entry identifier is the profile's serial number in registry 10, which is assigned to the profile at the time of its creation).

[0140]

[0059] Next, module 40 proceeds to the stage of comparing the degree of similarity of the selected profiles with similarity thresholds. The similarity thresholds can be specified by the developer of module 40 using known methods, and the following values ​​can be specified as similarity thresholds: upper similarity threshold - "0.5"; lower similarity threshold - "-0.5"

[0141]

[0060] If module 40 determines that the similarity level of the selected profile is equal to or greater than the upper similarity threshold (0.5), this means that the selected profile and the individual from the primary list are the same individual. Decision module 40 sends a command to BD1 to generate a link to the selected profile in the information about the individual. After this, the individual is removed from the primary list of individuals, and identification of the individual is considered complete.

[0142]

[0061] If module 40 determines that the similarity level of the selected FL profile is less than the upper similarity threshold (0.5), then module 40 searches for the corresponding originals in registry 10 in the manner described above for all profiles in the list that are duplicates. Module 40 then assigns a similarity level to the found originals equal to the similarity level of their duplicates, after which it selects the profile with the maximum similarity level among the previously selected FL profile and the found originals. Module 40 compares the similarity level of the selected FL profile with the upper similarity threshold in the manner described above.

[0143]

[0062] If module 40 determines that the similarity level of the selected FL profile is also less than the upper similarity threshold, then module 40 compares the similarity level of the selected FL profile with the lower similarity threshold.

[0144]

[0063] If module 40 determines that the similarity level of the selected profile is less than the lower similarity threshold (-0.5), this means that the selected profile and the FL from the primary list are different FLs. The identification of the FL at this stage is considered not completed, and the FL is left in the primary list for further processing.

[0145]

[0064] If module 40 determines that the similarity level of the selected profile falls within the range from -0.5 (the lower similarity threshold) to 0.5 (the upper similarity threshold), this means that the selected profile and the individual from the primary list are not similar enough to each other (an indeterminate result). At this stage, the identification of the individual is considered not completed, and the individual is left in the primary list for further processing. Additionally, decision module 40 sends a command to BD1 to save information about the fact that similar records for the individual have been found in registry 10, which may relate to the same individual. Later, the operator can analyze similar profiles and, if the profiles contain data about the same individual, manually designate one of the profiles as the original and the others as its duplicates.

[0146]

[0065] Additionally, the data comparison module 30 may be configured to correct significant parameters by excluding at least one parameter from the significant parameters specified by the developer or searching for FL profiles that have the same value of at least one of the following parameters: SNILS, INP, INN in the Russian Federation, INN in the country PCI7RU2025 / 000046

[0147] Citizenship, DUL; or both the full name and date of birth parameters match. This function can be activated if, in the previous step, identification was not completed for all individuals on the primary list, and the similarities between the profiles of individuals found using the specified significant parameters and those from the primary list are determined in the manner described previously.

[0148]

[0066] If, after completion of the work of modules 30 and 40 for all individuals of the primary list, there remain individuals in this list for which identification has not been performed (i.e., a connection with the profile in the registry 10 has not been established), then the decision-making module 40 transmits the list of remaining individuals to the profile creation module 60.

[0149]

[0067] Module 60 determines the degree of similarity between all individuals from the list of remaining individuals using the formula for calculating the degree of similarity mentioned above. If the degree of similarity between two individuals from the list of remaining individuals is equal to or higher than the upper similarity threshold (0.5), then such individuals are considered by Module 60 to be the same individual.

[0150]

[0068] Module 60 then forms groups of FLs that have been identified as the same FL. Each group may contain two or more FLs, or there may be no such groups at all.

[0151]

[0069] In each group, module 60 selects one person as the primary person, and the rest as complementary PLs. The primary person is selected randomly. A PL from the list of remaining PLs that is not included in any of the aforementioned groups is determined by module 40 to be the primary person without complementary PLs.

[0152]

[0070] For each primary FL, profile creation module 60 sends a command to BD1 to create a new profile in registry 10 based on the primary FL's data. If the primary FL has complementary FLs, the profile is enriched with the data of the complementary FLs. For example, the FL profile stores the DULs of not only the primary FL, but also the DULs of the complementary FLs.

[0153]

[0071] Module 60 then sends a command to BD1 to generate a link to the profile created for the primary individual in the information about the primary individual. If the primary individual has additional individuals, then a link to the profile created for the primary individual is also added to the information about the additional individuals.

[0154]

[0072] Thus, due to the fact that the search for FL profiles in the registry 10 is carried out taking into account the significant parameters, and the similarity of the FL profiles with the FL from the primary list is determined taking into account the values ​​of the weights that should be assigned when the values ​​of the FL parameters determined with the help of a trained neural network match or do not match, the accuracy of the identification of the FL in the registry of FL profiles is increased.

[0073] In general terms (see Fig. 2), the computing device contains one or more processors (201), memory means such as RAM (202) and ROM (203), input / output interfaces (204), input / output devices (205), and a device for network interaction (206), united by a common information exchange bus.

[0155]

[0074] The processor (201) (or several processors, a multi-core processor, etc.) can be selected from a range of devices that are widely used at present, for example, from manufacturers such as: Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. Under the processor or one of the processors used in the device (200), it is also necessary to take into account a graphic processor, for example, an NVIDIA or Graphcore GPU, the type of which is also suitable for the full or partial implementation of the method, and can also be used for training and applying machine learning models in various information systems.

[0156]

[0075] RAM (202) is a random access memory and is intended for storing machine-readable instructions executable by the processor (201) for performing the necessary operations for logical data processing. RAM (202), as a rule, contains executable instructions of the operating system and the corresponding software components (applications, software modules, etc.). In this case, the available memory capacity of a graphics card or graphics processor may serve as RAM (202).

[0157]

[0076] ROM (203) represents one or more permanent data storage devices, such as a hard disk drive (HDD), a solid state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.

[0158]

[0077] To organize the operation of the components of the device (200) and to organize the operation of external connected devices, various types of I / O interfaces (204) are used. The selection of the corresponding interfaces depends on the specific design of the computing device, which may be, without limitation: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.

[0159]

[0078] To ensure user interaction with the device (200), various I / O information means (205) are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality means, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification means (a retinal scanner, a fingerprint scanner, a voice recognition module), etc.

[0160]

[0079] The network interaction means (206) ensures the transmission of data via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (206) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.

[0161]

[0080] Additionally, satellite navigation tools may also be used as part of the device (200), for example, GPS, GLONASS, BeiDou, Galileo. The specific selection of elements of the device (200) for the implementation of various software and hardware architectural solutions may vary while maintaining the required functionality provided.

[0162]

[0081] Modifications and improvements to the above-described embodiments of the present technical solution will be apparent to those skilled in the art. The preceding description is provided only as an example and does not carry any limitations. Therefore, the scope of the present technical solution is limited only by the scope of the appended claims.

Claims

CLAUSES OF THE INVENTION 1. A method for identifying individuals (IE) in a register of IE profiles, performed by at least one computing device, comprising the steps of: - extract information about the FL, containing the FL parameters, from an external database (DB); - form, on the basis of the extracted information, a primary list of individuals intended for identifying individuals in the register of individual profiles; - extract a set of values ​​of significant parameters of the FL, where the significant parameters are pre-set by the developer; - access the register of FL profiles to search for FL profiles whose significant parameter values ​​match the values ​​of the corresponding significant parameters of the FL; - determine, using a pre-trained neural network, the values ​​of the weights that should be assigned when the values ​​of the FL parameters match or do not match; - compare the values ​​of all parameters of the FL profile with the FL parameters from the primary list to determine the degree of similarity of the found FL profiles with the FL from the primary list, taking into account the mentioned weight values; - based on the values ​​of the degree of similarity of the found FL profiles, the FL profile similar to the FL from the primary list is determined.

2. The method according to paragraph 1, characterized in that the coincidence of significant parameters is also determined in the case where the value of the parameter is absent.

3. The method according to paragraph 1, characterized in that the step of determining the presence of original profiles among the found FL profiles is additionally performed, wherein the FL profile similar to the FL from the primary list is determined only for the original profiles.

4. The method according to paragraph 1, characterized in that the following stage is additionally performed: - determine the absence of original profiles among the found FL profiles; - for all found profiles of individuals that are duplicates, a search is performed for the corresponding originals; - a similarity level is assigned to the found originals equal to the similarity level of their duplicates, and an individual profile similar to the profile from the primary list is determined only for profiles that are originals.

5. The method according to paragraph 1, characterized in that the step of determining the profile of an individual similar to an individual from the primary list comprises the steps of: - select from the found profiles the FL profile with the highest degree of similarity; - compare the degree of similarity of the selected FL profile with the similarity thresholds; - determine that the degree of similarity of the selected profile is equal to or greater than the upper similarity threshold; - form a link to the selected individual profile in the information about the individual in the external database.

6. The method according to paragraph 1, characterized in that the step of determining the profile of an individual similar to an individual from the primary list comprises the steps of: - determine the presence of original profiles among the found FL profiles; - select from the found original profiles the FL profile with the highest degree of similarity; - compare the degree of similarity of the selected FL profile with the similarity thresholds; - determine that the degree of similarity of the selected profile is less than the upper similarity threshold; - for all found FL profiles that are duplicates, a search is performed for the corresponding originals; - assign a degree of similarity to the found originals equal to the degree of similarity of their duplicates; - select a profile with the highest degree of similarity among the previously selected FL profile and the found originals; - determine that the degree of similarity of the selected profile is equal to or greater than the upper similarity threshold; - form a link to the selected individual profile in the information about the individual in the external database.

7. The method according to paragraph 1, characterized in that the following steps are additionally performed: - form a list of individuals from the primary list for whom identification has not been performed; - determine the degree of similarity of individuals among themselves, and if the degree of similarity of individuals is equal to or higher than the upper threshold of similarity, then such persons are considered to be the same individual; - groups are formed from FLs that have been defined as the same FL; - in each group, the main FL is selected, and the rest are FLs that complement it, and the FL that is not included in any of the mentioned groups is defined as the main one without FLs that complement it; - for each primary individual, a new profile is created in the registry based on the data of the primary individual; - enrich the created profiles with data from individuals that complement them; - send a command to the database to generate a link to the profile created for him in the information about the main individual and additional individuals.

8. The method according to paragraph 1, characterized in that an additional step of adjusting significant parameters is performed in the event that no FL profile was found for the FL from the primary list.

9. An individual identification system comprising at least one computing device and at least one memory containing machine-readable instructions which, when executed by at least one computing device, perform the method according to any one of paragraphs 1-8.