Model Training for User Similarity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current face recognition technologies face challenges in distinguishing between users with very similar features, such as twins, leading to risks of account mis-registration and misappropriation due to inefficient manual labeling processes in supervised machine learning methods.

Innovation Solution

A model training method that acquires user data pairs with identical data fields, determines user similarity through biological features like facial images or speech data, and trains a binary classification model using sample data to differentiate between similar and dissimilar users without manual labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is used to construct the identification model, then the model can be trained with supervised machine learning, but the process consumes large amounts of manpower resources and time, making model training inefficient

Engineering Contradiction:
Improveidentification accuracyVSAvoidmodel training efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system automatically determines user relationships by comparing user features (such as facial recognition results, device information, location data) without requiring manual labeling. The model self-trains by processing user data and automatically identifying potential twin relationships, eliminating the need for human investigators to manually collect and label data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual labeling process with automated computational methods. Instead of human investigators conducting surveys and manual observations, the system uses algorithmic comparison of user features, device information, and behavioral data to automatically determine user relationships and construct training datasets

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual labeling is used to obtain user relationships, then the identification model can be constructed, but it consumes large amounts of time for labeling

Engineering Contradiction:
Improveuser relationship identification reliabilityVSAvoidlabeling time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system collects and pre-processes user data (facial recognition results, device information, location data, behavioral patterns) in advance before model training is needed. This preliminary data collection and automatic relationship determination creates a ready-to-use training dataset, eliminating time-consuming manual labeling when the model needs to be trained

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically determines user relationships by analyzing user features and data patterns without human intervention. The automated process continuously identifies potential twin relationships and constructs training datasets, eliminating the time loss associated with manual investigator work

Inventive Principle:
Principle #25Self-service

3Ease of operation

If face recognition is used for identity verification, then user convenience is improved, but users with very similar looks (such as twins) cannot be effectively distinguished, leading to account mis-registration and misappropriation risks

Engineering Contradiction:
Improveidentity verification convenienceVSAvoiduser distinction accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the identification problem into two parts: standard face recognition for general user verification (maintaining convenience) and a specialized twin detection model for distinguishing users with similar features (improving accuracy). The system uses multiple data dimensions (device information, location data, behavioral patterns) to segment and identify twin relationships that standard face recognition cannot distinguish

Inventive Principle:
Principle #1Segmentation

4Device complexity

If standard face recognition is used, then the verification process is simple and convenient, but it cannot distinguish between users with very similar features

Engineering Contradiction:
Improveverification process simplicityVSAvoidfeature distinction precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extends the identification beyond traditional facial feature comparison by adding multiple new dimensions: device information (hardware identifiers, software versions), location data (GPS coordinates, Wi-Fi networks), behavioral patterns (usage time, operation habits), and social relationship data. This dimensional expansion allows the system to distinguish twins who appear identical in standard face recognition

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3611657B1Model training method and method, apparatus, and device for determining data similarity
Publication Date: 2024.08.21 ADVANCED NEW TECHNOLOGIES CO LTD
  • EP3611657B1 patent drawingFigure 1~2
  • EP3611657B1 patent drawingFigure 3
  • EP3611657B1 patent drawingFigure 4

AI summary

Embodiments of the present application disclose a model training method, apparatus, and device, and a data similarity determining method, apparatus, and device. The model training method includes: acquiring a plurality of user data pairs, wherein data fields of two sets of user data in each user data pair have an identical part; acquiring a user similarity corresponding to each user data pair, wherein the user similarity is a similarity between users corresponding to the two sets of user data in each user data pair; determining, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, sample data for training a preset classification model; and training the classification model based on the sample data to obtain a similarity classification model. With the embodiments of the present application, rapid training of a model can be implemented, the model training efficiency can be improved, and the resource consumption can be reduced.