Repeated user determination method and device, equipment, medium and product

By obtaining the work photos, identity information and organizational information of registered users, using convolutional neural network to extract face feature vectors, and using the Chinese Whispers algorithm for clustering, the problem of low accuracy of face recognition technology when screening repeated users is solved, and the accuracy of repeated users is achieved is achieved.

CN120448840APending Publication Date: 2025-08-08CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510590322.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing facial recognition technology has low accuracy when screening repeated users, especially when facing changes in lighting, expressions, angles, occlusions, etc.

Method used

By obtaining the work photos, identity information and organizational information of registered users, the face feature vector is extracted using a convolutional neural network, and the Chinese Whispers algorithm is used to cluster this information, and the duplicate users are determined based on the clustering results.

Benefits of technology

The accuracy of repeated user determination is improved, and through the combination of multi-dimensional information, false recognition is reduced, and the accuracy of screening is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448840A_ABST
    Figure CN120448840A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for determining repeated users, equipment, a medium and a product, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring user information of registered users of at least one institution; the user information comprises a work photo, identity information and mechanism information; determining a face feature vector of the registered user according to the work photo of the registered user; clustering each piece of information in the target information of the registered user to obtain a clustering result corresponding to each piece of information in the registered user and the target information; wherein the target information comprises a face feature vector, identity information and mechanism information; and according to a clustering result, determining repeated users in the registered users. According to the embodiment of the invention, the work photos, the identity information and the mechanism information of the registered users are respectively clustered, and the repeated users in the registered users are determined according to the clustering results corresponding to the information, so that the repeated users can be judged in combination with multi-dimensional information, and the determination accuracy of the repeated users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and apparatus, device, medium, and product for determining repeat users. Background Art

[0002] At present, in order to prevent a user from using different identities to register for an industry, facial recognition technology is usually used for screening.

[0003] Facial recognition technology is a biometric identification technology that uses computer vision and machine learning to automatically identify and verify the identity of people in images or videos. By analyzing and matching facial features, the technology can identify faces from images and compare them with known faces in a database to confirm individual identities.

[0004] However, facial recognition technology often relies on precise feature matching, which can easily lead to misidentification when faced with changes in lighting, expression, angle, occlusion, etc. Therefore, the accuracy of repeated user screening using facial recognition technology is relatively low. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to provide a method and apparatus, device, medium and product for determining repeat users, aiming to improve the accuracy of determining repeat users.

[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for determining duplicate users, the method comprising:

[0007] Obtain user information of a registered user of at least one organization; the user information includes work photos, identity information, and organization information;

[0008] Determining a facial feature vector of the registered user based on the work photo of the registered user;

[0009] Clustering each piece of information in the target information of the registered user to obtain a clustering result corresponding to the registered user and each piece of information in the target information; wherein the target information includes the facial feature vector, the identity information, and the organization information; the clustering result corresponding to each piece of information respectively includes at least one user set, and the similarity of the information of the registered user in each user set of the clustering result corresponding to each piece of information is greater than a threshold value corresponding to the information;

[0010] According to the clustering result, duplicate users among the registered users are determined.

[0011] In some embodiments, determining duplicate users among the registered users based on the clustering results includes:

[0012] Obtaining n clustering relationships of the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1;

[0013] determining a target similarity between the first registered user and the second registered user based on the n clustering relationships;

[0014] Based on the target similarity, the identification results of the first registered user and the second registered user are determined; wherein, when the target similarity is greater than or equal to a preset similarity threshold, the identification results of the first registered user and the second registered user are duplicate users; when the target similarity is less than the preset similarity threshold, the identification results of the first registered user and the second registered user are non-duplicate users.

[0015] In some embodiments, the user information also includes introducer information;

[0016] After obtaining user information of registered users of at least one organization, obtaining n clustering relationships between the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1, the determining method further includes:

[0017] selecting a first registered user from among the registered users of at least one organization;

[0018] The registered user corresponding to the introducer information of the first registered user is determined as the second registered user.

[0019] In some embodiments, determining the target similarity between the first registered user and the second registered user based on the n clustering relationships includes:

[0020] Determining, based on the n clustering relationships, n initial similarities between the first registered user and the second registered user corresponding one-to-one to the n information;

[0021] The n initial similarities are weighted and summed to obtain a target similarity between the first registered user and the second registered user.

[0022] In some embodiments, determining, based on the n clustering relationships, n initial similarities between the first registered user and the second registered user corresponding to the n pieces of information includes:

[0023] For each piece of information among the n pieces of information, if the cluster relationship corresponding to the first registered user and the second registered user with the information is that they are located in the same user set in the clustering result corresponding to the information, determining the initial similarity between the first registered user and the second registered user with respect to the information as a first similarity;

[0024] If the clustering relationship between the first registered user and the second registered user and the information is different user sets in the clustering result corresponding to the information, determining the initial similarity between the first registered user and the second registered user and the information as the second similarity;

[0025] The first similarity is greater than the second similarity.

[0026] In some embodiments, determining the target similarity between the first registered user and the second registered user based on the n clustering relationships includes:

[0027] Obtaining the number of target clustering relationships among the n clustering relationships, wherein the target clustering relationship satisfies: the first registered user and the second registered user are in the same user set;

[0028] Determine the target similarity between the first registered user and the second registered user based on the number; wherein the target similarity is positively correlated with the number.

[0029] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a device for determining a duplicate user, the device comprising:

[0030] An acquisition module, configured to acquire user information of a registered user of at least one organization; the user information includes work photos, identity information, and organization information;

[0031] A first determination module is used to determine a facial feature vector of a registered user based on a work photo of the registered user;

[0032] A clustering module is configured to cluster each piece of information in the target information of the registered user to obtain clustering results corresponding to the registered user and each piece of information in the target information; wherein the target information includes a facial feature vector, identity information, and organization information; the clustering results corresponding to each piece of information respectively include at least one user set, and the similarity of the registered user's information in each user set of the clustering results corresponding to each piece of information is greater than a threshold value corresponding to the information;

[0033] The second determining module is used to determine duplicate users among the registered users according to the clustering results.

[0034] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0035] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0036] To achieve the above-mentioned purpose, the fifth aspect of the embodiments of the present application proposes a computer program product. When the instructions in the computer program product are executed by the processor of an electronic device, the electronic device executes the method described in the first aspect above.

[0037] The present application proposes a duplicate user identification method, apparatus, device, medium, and product. The method obtains user information of registered users of at least one organization; the user information includes work photos, identity information, and organization information. Based on the registered user's work photos, the registered user's facial feature vector is determined. The method then clusters each piece of target information from the registered user to obtain clustering results corresponding to the registered user and each piece of information in the target information. The target information includes the facial feature vector, identity information, and organization information. The clustering results corresponding to each piece of information each include at least one user set, and the similarity of the registered user's information within each user set in the clustering results corresponding to each piece of information exceeds a threshold corresponding to the information. Finally, based on the clustering results, duplicate users are identified among the registered users. By clustering the registered user's work photos, identity information, and organization information separately and identifying duplicate users among the registered users based on the clustering results corresponding to each piece of information, duplicate user identification can be combined with multi-dimensional information to improve the accuracy of duplicate user identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flowchart of a method for determining duplicate users provided in an embodiment of the present application;

[0039] Figure 2 yes Figure 1 Flowchart of step S104 in FIG.

[0040] Figure 3 yes Figure 2 One of the flowcharts of step S202 in FIG.

[0041] Figure 4 yes Figure 3 Flowchart of step S301 in FIG.

[0042] Figure 5 yes Figure 2Flowchart 2 of step S202;

[0043] Figure 6 Schematic diagram of the structure of a device for determining duplicate users provided in an embodiment of the present application;

[0044] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0046] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0048] First, let’s analyze some of the terms used in this application:

[0049] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0050] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.

[0051] Information Extraction: A text processing technology that extracts specified types of entity, relationship, event, and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data is all text information extraction. Of course, the information extracted by text information extraction technology can be of various types.

[0052] Image caption generates a natural language description for an image and uses the generated description to help applications understand the semantics expressed in the image's visual scene. For example, image captioning can convert image retrieval into text retrieval, which can be used to classify images and improve image retrieval results. People can usually describe the details of an image's visual scene with just a quick glance, but automatically adding descriptions to images is a comprehensive and arduous computer vision task that requires converting the complex information contained in the image into a natural language description. Compared to ordinary computer vision tasks, image captioning not only requires identifying objects from images, but also requires associating the identified objects with natural semantics and describing them in natural language. Therefore, image captioning requires people to extract deep features of the image, associate them with semantic features, and convert them to generate descriptions.

[0053] Based on this, embodiments of the present application provide a method and apparatus, device, medium, and product for determining duplicate users, aiming to improve the accuracy of duplicate user determination.

[0054] The recommended method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the recommended method in the embodiments of the present application is described.

[0055] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0056] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0057] The method for determining duplicate users provided in the embodiment of the present application relates to the field of artificial intelligence technology. The method for determining duplicate users provided in the embodiment of the present application can be applied in a terminal, can be applied in a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the method for determining duplicate users, etc., but is not limited to the above forms.

[0058] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0059] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0060] Figure 1 This is an optional flowchart of the method for determining duplicate users provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S104.

[0061] Step S101, obtaining user information of a registered user of at least one organization; the user information includes work photos, identity information and organization information.

[0062] In the embodiments of this application, an "institution" can be understood as the place where a user registers. In one example, the institution can be an insurance company. User information may include, but is not limited to, the registered user's work photo, identity information, and institution information. The registered user's identity information may include, but is not limited to, the registered user's ID photo or ID number. Institution information may include, but is not limited to, the name, location, and type of the registered institution.

[0063] In an embodiment of the present application, the user information of the registered user obtained may be user information within a preset time period, and the preset time period may be a time period from the target historical moment to the current moment. In one example, user information of registered users of at least one organization within three years may be obtained.

[0064] In an embodiment of the present application, by obtaining user information of registered users of at least one institution, it is possible to screen whether the same user has registered in the same institution or different institutions, thereby improving the accuracy of determining duplicate users.

[0065] Step S102: determining a facial feature vector of the registered user based on the work photo of the registered user.

[0066] In an embodiment of the present application, after obtaining a work photo of a registered user, a facial feature vector of the registered user is extracted from the work photo of the registered user. In one example, when extracting the facial feature vector of the registered user, the work photo of the registered user can be input into a pre-trained convolutional neural network (CNN). Convolutional neural networks can effectively extract useful features from facial images and perform classification and recognition, etc. Facial feature extraction is performed using the convolutional neural network to obtain the facial feature vector of the registered user.

[0067] In an embodiment of the present application, after determining the facial feature vector of the registered user, the facial feature vector may be normalized. The normalization process may eliminate errors that may be caused by scale differences.

[0068] Step S103: cluster the various information in the target information of the registered user to obtain clustering results corresponding to the registered user and the various information in the target information; wherein, the target information includes the facial feature vector, the identity information and the organization information; the clustering results corresponding to each information respectively include at least one user set, and the similarity of the information of the registered user in each user set of the clustering results corresponding to each information is greater than the threshold corresponding to the information.

[0069] In this embodiment of the present application, after facial feature extraction is performed on the work photos of registered users, target information can be clustered separately. The target information can include facial feature vectors, identity information, and organization information. Furthermore, the target information can also include the registered user's introducer information.

[0070] By clustering each piece of information in the target information of the registered user, at least one user set corresponding to each piece of information can be obtained.

[0071] In one example, clustering facial feature vectors can yield at least one user set whose facial feature vector similarity exceeds a facial feature vector similarity threshold. Clustering identity information can yield at least one user set whose identity information similarity exceeds an identity information similarity threshold. Clustering organization information can yield at least one user set whose organization information similarity exceeds an organization information similarity threshold. Clustering referrer information can yield at least one user set whose referrer information similarity exceeds an referrer information similarity threshold.

[0072] In an embodiment of the present application, the Chinese Whispers algorithm can be used when clustering. The Chinese Whispers algorithm is a graph-based unsupervised clustering algorithm, which is commonly used for community detection and graph cluster analysis. The goal of the Chinese Whispers algorithm is to group or cluster the nodes in the graph so that connected nodes have similar labels. The basic idea is to assign similar nodes to the same class through information propagation between nodes. The Chinese Whispers algorithm has the characteristics of simple algorithm, and the accuracy of clustering results can be improved by adopting the Chinese Whispers algorithm.

[0073] Step S104: determining duplicate users among the registered users based on the clustering result.

[0074] In the embodiment of the present application, the clustering result may include the clustering result corresponding to each information in the target information. After the clustering result is determined, duplicate users among the registered users may be determined according to the clustering result.

[0075] In an embodiment of the present application, for two registered users, a clustering relationship corresponding to each information in the target information is obtained. In one example, the two registered users are the first registered user and the second registered user. The target information includes a facial feature vector, identity information, and organization information. The clustering relationship between the facial feature vectors of the first registered user and the second registered user, the clustering relationship between the identity information of the first registered user and the second registered user, and the clustering relationship between the organization information of the first registered user and the second registered user are obtained. The target similarity between the first registered user and the second registered user is determined based on the clustering relationship between the facial feature vectors of the first registered user and the second registered user, the clustering relationship between the identity information, and the clustering relationship between the organization information.

[0076] In an embodiment of the present application, after determining the target similarity between a first registered user and a second registered user, an identification result for the first registered user and the second registered user may be determined based on the target similarity and a preset similarity threshold. For example, if the target similarity is greater than or equal to the preset similarity threshold, the identification result for the first registered user and the second registered user is that they are duplicate users. If the target similarity is less than the preset similarity threshold, the identification result for the first registered user and the second registered user is that they are not duplicate users.

[0077] Steps S101 to S104 shown in the embodiment of the present application are performed by obtaining user information of registered users of at least one organization; the user information includes work photos, identity information and organization information. And based on the work photos of the registered users, the facial feature vectors of the registered users are determined. Then, each information in the target information of the registered users is clustered to obtain clustering results corresponding to the registered users and each information in the target information; wherein, the target information includes facial feature vectors, identity information and organization information; the clustering results corresponding to each information respectively include at least one user set, and the similarity of the information of the registered users in each user set of the clustering results corresponding to each information is greater than the threshold value corresponding to the information. Finally, based on the clustering results, duplicate users among the registered users are determined. By clustering the work photos, identity information and organization information of the registered users respectively, and determining duplicate users among the registered users based on the clustering results corresponding to each information, it is possible to combine multi-dimensional information to determine duplicate users, thereby improving the accuracy of duplicate user determination.

[0078] See also Figure 2 In some embodiments, step S104 may include but is not limited to steps S201 to S203:

[0079] Step S201, obtaining n clustering relationships of the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1;

[0080] Step S202: determining the target similarity between the first registered user and the second registered user based on the n clustering relationships;

[0081] Step S203: Determine the identification results of the first registered user and the second registered user based on the target similarity; wherein, when the target similarity is greater than or equal to a preset similarity threshold, the identification results of the first registered user and the second registered user are duplicate users; and when the target similarity is less than the preset similarity threshold, the identification results of the first registered user and the second registered user are non-duplicate users.

[0082] Specifically, the clustering results may include clustering results corresponding to each piece of information in the target information. When determining duplicate users among registered users based on the clustering results, the first and second registered users among the registered users are first obtained. In one example, the first and second registered users may both be any two registered users from at least one organization. In another example, the first registered user may be any one of the registered users, and the second registered user may be the introducer of the first registered user.

[0083] In this embodiment of the present application, n clustering relationships are obtained for the first and second registered users in the clustering results, corresponding one-to-one with n pieces of information included in the target information. In one example, the target information may include a facial feature vector, identity information, and organization information. In this case, the value of n is 3. Three clustering relationships are then obtained for the first and second registered users, respectively, with the facial feature vector, identity information, and organization information.

[0084] In an embodiment of the present application, the target similarity between the first registered user and the second registered user can be determined based on n cluster relationships. In one example, the target similarity between the first registered user and the second registered user can be obtained by determining n initial similarities corresponding to the n cluster relationships and taking a weighted sum of the n initial similarities. In another example, the target similarity between the first registered user and the second registered user can also be determined based on the number of users in the same user set as the first registered user and the second registered user.

[0085] In this embodiment, by determining the target similarity between the first registered user and the second registered user based on the clustering result of each information in the target information of the registered users, and then determining the duplicate users among the registered users, the accuracy of duplicate user determination can be improved.

[0086] In some embodiments, the user information also includes introducer information;

[0087] After obtaining user information of registered users of at least one organization, obtaining n clustering relationships between the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1, the determining method further includes:

[0088] selecting a first registered user from among the registered users of at least one organization;

[0089] The registered user corresponding to the introducer information of the first registered user is determined as the second registered user.

[0090] Specifically, user information may also include introducer information. Introducer information can be understood as the information of the old registrant who introduced the new user. Introducer information may include, but is not limited to, the introducer's name, location, and photo.

[0091] In the embodiment of the present application, after obtaining the user information of the registered users of at least one organization, it is also necessary to determine the first registered user and the second registered user. The first registered user and the second registered user can be understood as two target users for duplicate user determination.

[0092] In this embodiment of the present application, a registered user can be randomly selected from the registered users of at least one organization as the first registered user. When determining the second registered user, the first registered user's introducer information can be obtained first. The registered user corresponding to the introducer information can then be determined as the second registered user. In other words, the second registered user can be the introducer of the first registered user.

[0093] In this embodiment, by determining the introducer of the first registered user as the second registered user, it is possible to screen whether there is a new registered user who registers with different information from the original registered user, thereby improving the accuracy of determining duplicate users.

[0094] See also Figure 3 In some embodiments, step S202 may include but is not limited to steps S301 to S302:

[0095] Step S301: determining n initial similarities between the first registered user and the second registered user corresponding to the n pieces of information based on the n clustering relationships;

[0096] Step S302: performing weighted summation of the n initial similarities to obtain a target similarity between the first registered user and the second registered user.

[0097] Specifically, after obtaining n clustering relationships, n initial similarities corresponding to the n pieces of information can be determined for the first registered user and the second registered user, respectively. In one example, the n clustering relationships may include a first initial similarity between facial feature vectors of the first registered user and the second registered user, a second initial similarity between identity information of the first registered user and the second registered user, and a third initial similarity between the organization information of the first registered user and the second registered user.

[0098] In this embodiment of the present application, each clustering result corresponding to each piece of information in the target information has a preset weight. The specific weight corresponding to each clustering result can be set according to the actual situation. After obtaining n initial similarities corresponding to n pieces of information for the first registered user and the second registered user, a weighted sum of the n initial similarities can be performed. That is, the product of each initial similarity and its corresponding weight is summed to obtain the target similarity between the first registered user and the second registered user.

[0099] In this embodiment, by performing weighted summation on n initial similarities to determine the target similarity between the first registered user and the second registered user, the accuracy of duplicate user determination can be improved by combining the influence of each clustering relationship on duplicate user determination.

[0100] See also Figure 4In some embodiments, step S301 may include but is not limited to steps S401 to S402:

[0101] Step S401: For each of the n pieces of information, if the cluster relationship corresponding to the first registered user and the second registered user with the information is the same user set in the clustering result corresponding to the information, then determine the initial similarity between the first registered user and the second registered user with respect to the information as a first similarity;

[0102] Step S402: If the cluster relationship between the first registered user and the second registered user and the information is different user sets in the cluster result corresponding to the information, then the initial similarity between the first registered user and the second registered user and the information is determined as a second similarity;

[0103] The first similarity is greater than the second similarity.

[0104] Specifically, when determining the n initial similarities of the first registered user and the second registered user corresponding to n information one by one based on n clustering relationships, it can be determined by whether the clustering relationships corresponding to the first registered user and the second registered user and each information in the target information are the same user set in the clustering results corresponding to the target information.

[0105] In one example, when the clustering relationship between the facial feature vectors of user 1 and user 2 is in the same user set, the initial similarity corresponding to the facial feature vectors of the first registered user and the second registered user is determined as the first similarity. For example, the first similarity may be 1.

[0106] In another example, when the clustering relationship of the facial feature vectors of user 1 and user 2 is not in the same user set, the initial similarity corresponding to the facial feature vectors of the first registered user and the second registered user is determined as the second similarity. For example, the second similarity may be 0.

[0107] In the embodiment of the present application, specific values of the first similarity and the second similarity can be set according to specific circumstances, and the first similarity and the second similarity only need to satisfy that the first similarity is greater than the second similarity.

[0108] In this embodiment, by determining the initial similarities corresponding to each information of the first registered user and the second registered user, the target similarities between the first registered user and the second registered user can be determined.

[0109] See also Figure 5 In some embodiments, step S202 may also include but is not limited to steps S501 to S502:

[0110] Step S501: obtaining the number of target clustering relationships among the n clustering relationships, wherein the target clustering relationship satisfies: the first registered user and the second registered user are in the same user set;

[0111] Step S502: determining the target similarity between the first registered user and the second registered user based on the number; wherein the target similarity is positively correlated with the number.

[0112] Specifically, when determining the target similarity between the first registered user and the second registered user based on n cluster relationships, the target similarity may also be determined by the number of cluster relationships among the n cluster relationships that satisfy that the first registered user and the second registered user are in the same user set.

[0113] In the embodiment of the present application, the target cluster relationship can be understood as a cluster relationship that satisfies the first registered user and the second registered user being in the same user set. The first registered user and the second registered user being in the same user set indicates that the first registered user and the second registered user have a high similarity.

[0114] In an embodiment of the present application, the target similarity between the first registered user and the second registered user is determined by obtaining the number of target cluster relationships among n cluster relationships. When determining the target similarity, it is necessary to satisfy the requirement that the target similarity is positively correlated with the number of target cluster relationships, that is, the greater the number of target cluster relationships, the higher the target similarity. The proportional relationship between the target similarity and the target cluster relationships can be set according to the specific situation.

[0115] In this embodiment, by determining the target similarity between the first registered user and the second registered user according to the number of target clustering relationships and a preset proportional relationship, the target similarity can be determined intuitively.

[0116] In a specific embodiment of the present application, when clustering is performed based on the Chinese Whispers algorithm, the following steps are included:

[0117] Step 1, initialization: Initialize each node to a unique class label (which can be the node ID or a random label).

[0118] Step 2: Propagate labels: Each node selects a label by looking at its neighboring nodes. Each node counts the labels of all its neighboring nodes and updates its own label to the label that appears most frequently among its neighboring nodes.

[0119] Step 3, iterative update: Repeat label propagation until the labels of all nodes are stable, that is, they no longer change.

[0120] Step 4: The termination condition is met: when the label of each node no longer changes or the maximum number of iterations is reached, the algorithm stops.

[0121] Step 5: Output the clustering results: the label of each node indicates the class or cluster to which it belongs, and nodes with the same label are considered to belong to the same class.

[0122] The embodiment of the present application provides a method and apparatus, device, medium and product for determining duplicate users, which obtains user information of registered users of at least one organization; the user information includes work photos, identity information and organization information. And based on the work photos of the registered users, the facial feature vectors of the registered users are determined. Then, each information in the target information of the registered users is clustered to obtain clustering results corresponding to the registered users and each information in the target information; wherein, the target information includes facial feature vectors, identity information and organization information; the clustering results corresponding to each information respectively include at least one user set, and the similarity of the information of the registered users in each user set of the clustering results corresponding to each information is greater than the threshold corresponding to the information. Finally, based on the clustering results, duplicate users among the registered users are determined. By clustering the work photos, identity information and organization information of the registered users respectively, and determining duplicate users among the registered users based on the clustering results corresponding to each information, it is possible to combine multi-dimensional information to determine duplicate users, thereby improving the accuracy of duplicate user determination.

[0123] See also Figure 6 The present application also provides a device for determining duplicate users, which can implement the above-mentioned method for determining duplicate users. The device includes:

[0124] An acquisition module, configured to acquire user information of a registered user of at least one organization; the user information includes work photos, identity information, and organization information;

[0125] A first determination module is used to determine a facial feature vector of a registered user based on a work photo of the registered user;

[0126] A clustering module is configured to cluster each piece of information in the target information of the registered user to obtain clustering results corresponding to the registered user and each piece of information in the target information; wherein the target information includes a facial feature vector, identity information, and organization information; the clustering results corresponding to each piece of information respectively include at least one user set, and the similarity of the registered user's information in each user set of the clustering results corresponding to each piece of information is greater than a threshold value corresponding to the information;

[0127] The second determining module is used to determine duplicate users among the registered users according to the clustering results.

[0128] Specifically, the acquisition module can obtain user information of a registered user of at least one organization. The user information can include a work photo, identity information, and organization information. In one example, the organization can be an insurance company, and the registered user is a registered agent of the insurance company. The acquisition module can then obtain the work photo, identity information, and insurance company information of the registered agent of the insurance company.

[0129] The first determination module can determine a facial feature vector of the registered user based on the registered user's work photo. To determine the facial feature vector of the registered user, the work photo of the registered user can be input into a pre-trained convolutional neural network model. The convolutional neural network model extracts facial features to obtain the facial feature vector of the registered user.

[0130] The clustering module can cluster the various pieces of information in the target information of registered users, obtaining clustering results corresponding to the registered users and the various pieces of information in the target information. The Chinese Whispers algorithm can be used for clustering. Chinese Whispers is an unsupervised graph-based clustering algorithm commonly used for community detection and graph cluster analysis.

[0131] The second determination module may determine duplicate users among the registered users based on the clustering result. Specifically, the second determination module may determine whether the first registered user and the second registered user are duplicate users by determining target similarity between the first registered user and the second registered user, and determining whether the first registered user and the second registered user are duplicate users based on the relationship between the target similarity and a preset similarity threshold.

[0132] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-mentioned method for determining duplicate users. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0133] See also Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0134] The processor 701 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0135] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called by the processor 701 to execute the method for determining repeated users in the embodiments of this application.

[0136] Input / output interface 703, used to implement information input and output;

[0137] Communication interface 704, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0138] Bus 705 , which transmits information between various components of the device (e.g., processor 701 , memory 702 , input / output interface 703 , and communication interface 704 );

[0139] The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .

[0140] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the above-mentioned method for determining duplicate users when executed by a processor.

[0141] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] An embodiment of the present application further provides a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the above-mentioned method for determining duplicate users.

[0143] The embodiments of the present application provide a method, apparatus, device, medium, and product for determining duplicate users. The method obtains user information of registered users from at least one organization; the user information includes work photos, identity information, and organization information. Based on the work photos of the registered users, a facial feature vector of the registered users is determined. The method then clusters each piece of information in the target information of the registered users to obtain clustering results corresponding to the registered users and each piece of information in the target information. The target information includes the facial feature vector, identity information, and organization information. The clustering results corresponding to each piece of information each include at least one user set, and the similarity of the registered user's information in each user set of the clustering results corresponding to each piece of information exceeds a threshold corresponding to the information. Finally, based on the clustering results, duplicate users are determined among the registered users. Clustering is performed separately on the work photos, identity information, and organization information of the registered users. Clustering involves not only the facial feature vector, but also identity information, organization information, and information about former students who serve as introducers. This additional information is used to increase the weight of clustering into the same cluster. The resulting "cluster" is not only an independent cluster but also includes additional information. This allows the formation of a "cluster" with a network of relationships. For related networks, you can find other "categories" related to the duplicate personnel that have been identified, and then focus on tracking those with related relationships.

[0144] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0145] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0147] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0148] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0149] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0150] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0151] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0153] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0154] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for determining repeated users, characterized in that: The determination method includes: Obtain user information of a registered user of at least one organization; the user information includes work photos, identity information, and organization information; Determining a facial feature vector of the registered user based on the work photo of the registered user; Clustering each piece of information in the target information of the registered user to obtain a clustering result corresponding to the registered user and each piece of information in the target information; wherein the target information includes the facial feature vector, the identity information, and the organization information; the clustering result corresponding to each piece of information respectively includes at least one user set, and the similarity of the information of the registered user in each user set of the clustering result corresponding to each piece of information is greater than a threshold value corresponding to the information; According to the clustering result, duplicate users among the registered users are determined.

2. The determination method according to claim 1, characterized in that Determining duplicate users among the registered users according to the clustering result includes: Obtaining n clustering relationships of the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1; determining a target similarity between the first registered user and the second registered user based on the n clustering relationships; Based on the target similarity, the identification results of the first registered user and the second registered user are determined; wherein, when the target similarity is greater than or equal to a preset similarity threshold, the identification results of the first registered user and the second registered user are duplicate users; when the target similarity is less than the preset similarity threshold, the identification results of the first registered user and the second registered user are non-duplicate users.

3. The determination method according to claim 2, characterized in that: The user information also includes introducer information; After obtaining user information of registered users of at least one organization, obtaining n clustering relationships between the first registered user and the second registered user in the clustering result that correspond one-to-one to n pieces of information included in the target information, where n is an integer greater than 1, the determining method further includes: selecting a first registered user from among the registered users of at least one organization; The registered user corresponding to the introducer information of the first registered user is determined as the second registered user.

4. The determination method according to claim 2, characterized in that: The determining, based on the n clustering relationships, the target similarity between the first registered user and the second registered user includes: Determining, based on the n clustering relationships, n initial similarities between the first registered user and the second registered user corresponding one-to-one to the n information; The n initial similarities are weighted and summed to obtain a target similarity between the first registered user and the second registered user.

5. The determination method according to claim 4, characterized in that: Determining, based on the n clustering relationships, n initial similarities between the first registered user and the second registered user corresponding to the n information one-to-one, includes: For each piece of information among the n pieces of information, if the cluster relationship corresponding to the first registered user and the second registered user with the information is that they are located in the same user set in the clustering result corresponding to the information, determining the initial similarity between the first registered user and the second registered user with respect to the information as a first similarity; If the clustering relationship between the first registered user and the second registered user and the information is different user sets in the clustering result corresponding to the information, determining the initial similarity between the first registered user and the second registered user and the information as the second similarity; The first similarity is greater than the second similarity.

6. The determination method according to claim 2, characterized in that: The determining, based on the n clustering relationships, the target similarity between the first registered user and the second registered user includes: Obtaining the number of target clustering relationships among the n clustering relationships, wherein the target clustering relationship satisfies: the first registered user and the second registered user are in the same user set; Determine the target similarity between the first registered user and the second registered user based on the number; wherein the target similarity is positively correlated with the number.

7. A device for determining a repeated user, characterized in that: The determining device comprises: An acquisition module, configured to acquire user information of a registered user of at least one organization; the user information includes a work photo, identity information, and organization information; A first determining module is configured to determine a facial feature vector of the registered user based on a work photo of the registered user; a clustering module, configured to cluster each piece of information in the target information of the registered user to obtain a clustering result corresponding to the registered user and each piece of information in the target information; wherein the target information includes the facial feature vector, the identity information, and the organization information; the clustering result corresponding to each piece of information respectively includes at least one user set, and the similarity of the information of the registered user in each user set of the clustering result corresponding to each piece of information is greater than a threshold value corresponding to the information; The second determining module is configured to determine duplicate users among the registered users according to the clustering result.

8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method for determining a repeated user according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for determining a repeated user according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the method for determining a repeated user according to any one of claims 1 to 6.