Model Training for User Similarity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current face recognition technologies face challenges in distinguishing between users with very similar features, such as twins, leading to risks of account mis-registration and misappropriation due to inefficient manual labeling processes in supervised machine learning methods.
Innovation Solution
A model training method that acquires user data pairs with identical data fields, determines user similarity through biological features like facial images or speech data, and trains a binary classification model using sample data to differentiate between similar and dissimilar users without manual labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is used to construct the identification model, then the model can be trained with supervised machine learning, but the process consumes large amounts of manpower resources and time, making model training inefficient
Solution Approach 1:
The system automatically determines user relationships by comparing user features (such as facial recognition results, device information, location data) without requiring manual labeling. The model self-trains by processing user data and automatically identifying potential twin relationships, eliminating the need for human investigators to manually collect and label data
Solution Approach 2:
The patent replaces the mechanical manual labeling process with automated computational methods. Instead of human investigators conducting surveys and manual observations, the system uses algorithmic comparison of user features, device information, and behavioral data to automatically determine user relationships and construct training datasets
2Reliability
If manual labeling is used to obtain user relationships, then the identification model can be constructed, but it consumes large amounts of time for labeling
Solution Approach 1:
The system collects and pre-processes user data (facial recognition results, device information, location data, behavioral patterns) in advance before model training is needed. This preliminary data collection and automatic relationship determination creates a ready-to-use training dataset, eliminating time-consuming manual labeling when the model needs to be trained
Solution Approach 2:
The system automatically determines user relationships by analyzing user features and data patterns without human intervention. The automated process continuously identifies potential twin relationships and constructs training datasets, eliminating the time loss associated with manual investigator work
3Ease of operation
If face recognition is used for identity verification, then user convenience is improved, but users with very similar looks (such as twins) cannot be effectively distinguished, leading to account mis-registration and misappropriation risks
Solution Approach 1:
The patent segments the identification problem into two parts: standard face recognition for general user verification (maintaining convenience) and a specialized twin detection model for distinguishing users with similar features (improving accuracy). The system uses multiple data dimensions (device information, location data, behavioral patterns) to segment and identify twin relationships that standard face recognition cannot distinguish
4Device complexity
If standard face recognition is used, then the verification process is simple and convenient, but it cannot distinguish between users with very similar features
Solution Approach 1:
The patent extends the identification beyond traditional facial feature comparison by adding multiple new dimensions: device information (hardware identifiers, software versions), location data (GPS coordinates, Wi-Fi networks), behavioral patterns (usage time, operation habits), and social relationship data. This dimensional expansion allows the system to distinguish twins who appear identical in standard face recognition
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Embodiments of the present application disclose a model training method, apparatus, and device, and a data similarity determining method, apparatus, and device. The model training method includes: acquiring a plurality of user data pairs, wherein data fields of two sets of user data in each user data pair have an identical part; acquiring a user similarity corresponding to each user data pair, wherein the user similarity is a similarity between users corresponding to the two sets of user data in each user data pair; determining, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, sample data for training a preset classification model; and training the classification model based on the sample data to obtain a similarity classification model. With the embodiments of the present application, rapid training of a model can be implemented, the model training efficiency can be improved, and the resource consumption can be reduced.