A homogenous picture detection method and apparatus

CN122842162APending Publication Date: 2026-09-29GUIYANG LONGMA COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010767.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本实施例的目的在于提供一种同源图片检测方法及装置,用于解决同源图片判定误差大问题

Benefits of technology

[0004]本实施例的目的在于提供一种同源图片检测方法及装置,用于解决同源图片判定误差大问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842162A_ABST
    Figure CN122842162A_ABST
Patent Text Reader

Abstract

This invention provides a method for detecting images from the same source. The method involves acquiring a first image and a second image; performing face detection and keypoint extraction on the first and second images respectively; rotating, aligning, cropping, and normalizing the detected face regions; calculating the facial feature vectors of the first and second images and determining whether the difference in facial feature proportions exceeds a first threshold; inputting the first and second images into a face feature comparison model to obtain face similarity probabilities; performing background deface preprocessing on the first and second images respectively, and then inputting them into a background feature comparison model to obtain background similarity probabilities; weighted summing of the facial feature proportion difference judgment result, the face similarity probability, and the background similarity probability to obtain a total score; and outputting the final judgment result based on the comparison of the total score with a second threshold. This method solves the problem of large errors in the determination of images from the same source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, and in particular relates to a method and apparatus for detecting images from the same source. Background Technology

[0002] During the identity verification process for SIM card activation and network access by mobile operators, users need to take a photo of their ID document with their mobile phone and upload it. The same user may submit multiple photos in multiple operations, and these photos may have two relationships: Type A is the same original image that has been edited (such as adding watermarks, black borders at the top, cropping, scaling, flipping, rotating, resolution changes, etc.), which should be judged as having the same source; Type B is two photos taken by the same person (such as the same person, the same scene, but taken at different times, with slight differences in expression or posture), which should be judged as having different sources.

[0003] Existing technologies mainly suffer from the following problems: First, methods based on a single deep learning model cannot simultaneously focus on three dimensions: facial details, facial feature proportions, and background differences. When two images have the same watermark, the model is prone to misjudging the watermark as a strong signal from the same source. Second, methods based on facial key point comparison are greatly affected by random errors in key point detection, and watermarks, black borders, and other coverings can cause changes in the relative position of the face in the image, leading to failure in key point comparison. Third, methods based on pixel-level comparison are sensitive to scaling, cropping, and resolution changes. After adding a watermark, large-area changes in pixels make it impossible to identify them as the same original image. Summary of the Invention

[0004] The purpose of this embodiment is to provide a method and apparatus for detecting images from the same source, in order to solve the problem of large errors in the determination of images from the same source.

[0005] A method for detecting images from the same source includes: Get the first image and the second image; Face detection and key point extraction are performed on the first image and the second image respectively; The detected face regions are rotated, aligned, cropped, and normalized. Calculate the facial feature proportion feature vectors of the first image and the second image, and determine whether the difference in facial feature proportions exceeds a first threshold. The first image and the second image are respectively input into the face feature comparison model to obtain the face similarity probability; The first image and the second image are respectively subjected to background removal face preprocessing, and then input into the background feature comparison model to obtain the background similarity probability; The total score is obtained by weighting and summing the results of the facial feature proportion differences, the face similarity probability, and the background similarity probability, and then outputting the final judgment result based on the comparison between the total score and the second threshold.

[0006] Furthermore, the detected face regions are rotated, aligned, cropped, and normalized, including: The rotation angle is calculated based on the center of the left eye and the center of the right eye. The image is rotated with the midpoint of the center of the left eye and the center of the right eye as the rotation center, so that the direction of the line connecting the two eyes is rotated to the horizontal direction. After rotation, the aligned face region is determined based on the position of the original face detection box, and the cropped area is expanded outward by a ratio of 0.5. The cropped face region is normalized into a square image, and the square image is scaled to 256×256 pixels to obtain an aligned and normalized face image.

[0007] Further, the facial feature proportion feature vectors of the first image and the second image are calculated, and it is determined whether the difference in facial feature proportions exceeds a first threshold, including: Calculate the distance between eyes, distance between eyes and nose, mouth width, distance between mouth and eyes, and distance between eyebrows and eyes from 68 extracted facial key points; Based on the above distance, five proportional features r1~r5 are constructed to form the feature vector R=[r1, r2, r3, r4, r5]; Calculate the facial feature vectors R1 and R2 of the first image and the second image respectively, and calculate the Euclidean distance d_ratio between them; If d_ratio is greater than the first threshold T1, then the first image and the second image are determined to be from different sources; if d_ratio ≤ T1, then proceed to the next step.

[0008] Furthermore, the first threshold T1 = 0.15.

[0009] Furthermore, the face feature comparison model adopts a Siamese network structure, including two feature extraction branches with shared weights; the feature extraction branches use EfficientNet-B0 as the backbone network and output a 1280-dimensional feature vector; the two branches output a first face feature vector f1 and a second face feature vector f2 respectively, and f1 and f2 are concatenated and input into the classifier, and the face similarity probability p_face is output through the Sigmoid activation function, with a value range of [0,1].

[0010] Furthermore, the classifier consists of three fully connected layers: the first layer has a 2560-dimensional input and a 512-dimensional output, using ReLU activation and Dropout with a dropout rate of 0.5; the second layer has a 512-dimensional input and a 128-dimensional output, using ReLU activation and Dropout with a dropout rate of 0.3; and the third layer has a 128-dimensional input and a 1-dimensional output, using the Sigmoid activation function.

[0011] Furthermore, background removal and face removal preprocessing are performed on the first image and the second image respectively, including: For each original image, a face detection algorithm is used to detect face regions, the largest face bounding box is selected, and the face bounding box is expanded outward by 20 pixels. All pixel values ​​within the expanded rectangular area are set to 0.

[0012] Furthermore, the background feature comparison model adopts a Siamese network structure, with the two branches sharing weights; the feature extraction backbone network is MobileNetV3-Small, which outputs a 128-dimensional background feature vector; the two branches output the first background feature vector b1 and the second background feature vector b2 respectively. After concatenating b1 and b2, they are input into a binary classifier to output the background similarity probability p_bg, with a value range of [0,1].

[0013] A homology image detection device, comprising: The image acquisition module is used to acquire the first and second images to be detected. The face detection and key point extraction module is used to perform face detection on the first image and the second image respectively and extract 68 facial key points; The face alignment and normalization module is used to perform rotation alignment, cropping and normalization processing on the detected face regions, and output the first face image and the second face image after alignment and normalization. The facial feature ratio calculation and preliminary judgment module is used to calculate the facial feature ratio feature vector based on the 68 facial key points, and calculate the Euclidean distance d_ratio between the feature vectors of two images. If d_ratio is greater than the first threshold T1, it is determined that they are from different sources. The face feature comparison module is used to input the aligned and normalized first face image and the second face image into the face feature comparison model, and output the face similarity probability p_face; The background feature comparison module is used to perform background deface preprocessing on the first image and the second image, and then input the background feature comparison model to output the background similarity probability p_bg. The weighted scoring and judgment module is used to calculate the total score based on the judgment result of the preliminary judgment module of facial feature proportion, the face similarity probability p_face and the background similarity probability p_bg, and compare the total score with the second threshold T2 to output the final judgment result; The output module is used to output the determination result of "same source" or "different source".

[0014] Furthermore, the background feature comparison module includes: The background face removal preprocessing submodule is used to detect the largest face bounding box in the image, expand it outward by 20 pixels, and set the pixels of the expanded area to 0. The feature extraction submodule uses the MobileNetV3-Small Siamese network, with the two branches sharing weights, and outputs 128-dimensional background feature vectors b1 and b2 respectively; The classification submodule is used to concatenate b1 and b2 and input them into a binary classifier, outputting the background similarity probability p_bg.

[0015] This invention provides a method and apparatus for detecting images from the same source, which solves the problem of large errors in determining images from the same source.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] Figure 1 : A flowchart illustrating the steps of the method for detecting homologous images provided in an embodiment of the present invention; Figure 2: A structural diagram of the homogeneous image detection device provided in an embodiment of the present invention; Figure 3: A schematic diagram of the positive sample construction of the homologous image detection method provided in the embodiment of the present invention; Figure 4: A schematic diagram of negative sample construction for the homologous image detection method provided in an embodiment of the present invention; Figure 5: A structural diagram of the background comparison model for the same-source image detection method provided in the embodiment of the present invention. Detailed Implementation

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] This embodiment provides a method for detecting images from the same source. This method is applied to a SIM card activation and network access authentication scenario to determine whether two images to be detected originate from the same original image. The method includes the following steps: Step S1: Obtain the first image and the second image: Acquire a first image and a second image to be detected. The images are color images or grayscale images, preferably ID photos or identity verification photos containing a face. The images can be acquired by taking pictures in real time with a camera, reading them from a storage medium, or receiving them via network transmission.

[0020] Step S2: Perform face detection and key point extraction on the first and second images respectively. Face detection is performed on the first and second images respectively to determine whether the images contain human faces. If no face is detected in either image, they are directly identified as different sources, and the detection process ends; if faces are detected in both images, facial landmarks are extracted from each image.

[0021] Specifically, a regression tree-based facial landmark detection algorithm (such as the 68-point face shape predictor from the dlib library) is used to detect faces and locate key points in the image. The detected face regions are represented by rectangular boxes, denoted as (top, right, bottom, left), where top, right, bottom, and left are the coordinates of the top, right, bottom, and left boundaries of the face box, respectively. Simultaneously, 68 facial landmarks are extracted, including the coordinates of the left eye, right eye, tip of the nose, bridge of the nose, upper lip, lower lip, and chin.

[0022] When there are multiple faces in an image, only the largest face region and its corresponding key points are selected for subsequent processing to avoid interference from non-target faces.

[0023] Step S3: Rotate, align, crop, and normalize the face region: To ensure that facial proportion calculations are unaffected by head pose and image rotation, the detected face regions are rotated and aligned. The specific steps are as follows: Calculate the rotation angle using the centers of the left and right eyes as references. Let the coordinates of the left eye center be (lx, ly) and the coordinates of the right eye center be (rx, ry), then: dx = rx - lx dy = ry - ly angle = arctan2(dy, dx) Using the midpoint between the center of the left eye and the center of the right eye ((lx+rx) / 2, (ly+ry) / 2) as the rotation center, rotate the entire image so that the line connecting the two eyes is rotated to the horizontal direction (i.e., the angle is 0).

[0024] After rotation, the aligned face region is determined based on the position of the original face detection bounding box. To preserve the contextual information around the face, the cropping region is expanded outward by a factor of 0.5. That is, the width of the cropping box is 1.5 times the width of the original face bounding box, and the height is 1.5 times the height of the original face bounding box, while the center point remains unchanged. The cropping box is then confined within the image boundaries.

[0025] The cropped face region is normalized into a square image. Specifically, the longer side of the cropped region is taken as the side length of the square, the cropped image is placed in the center of the square, the missing areas are filled with zero pixels, and then the square image is scaled to 256×256 pixels to obtain an aligned and normalized face image.

[0026] Step S4: Calculate the facial feature proportion feature vectors of the first image and the second image, and determine whether the difference in facial feature proportions exceeds the first threshold. From the 68 facial key points extracted in step S2, the following facial feature proportions are calculated: (1) Eye distance d_eye: Euclidean distance between the center of the left eye and the center of the right eye. The center of the left eye is the mean coordinate of all key points of the left eye, and the center of the right eye is the mean coordinate of all key points of the right eye.

[0027] (2) Eye-nose distance d_eye_nose: The average of the distance from the center of the left eye to the tip of the nose and the distance from the center of the right eye to the tip of the nose. The tip of the nose is taken as the last point of the key point of the tip of the nose (if it exists) or the last point of the key point of the bridge of the nose.

[0028] (3) Mouth width d_mouth: Euclidean distance between the left key point of the upper lip and the right key point of the upper lip.

[0029] (4) Mouth-eye distance d_mouth_eye: The vertical distance from the center of the upper lip to the line connecting the centers of the two eyes. The center of the upper lip is the average of all key points on the upper lip.

[0030] (5) Eyebrow-eye distance: The average of the distance from the center of the left eye to the center of the left eyebrow and the distance from the center of the right eye to the center of the right eyebrow. The center of the left eyebrow is the average of all key points of the left eyebrow, and the center of the right eyebrow is the average of all key points of the right eyebrow.

[0031] Based on the above distances, five proportional features are constructed: r1 = d_eye_nose / d_eye r2 = d_mouth / d_eye r3 = d_mouth_eye / d_eye r4 = d_left_eye_brow / d_eye r5 = d_right_eye_brow / d_eye Where d_left_eye_brow is the distance between the left eyebrow and eye, and d_right_eye_brow is the distance between the right eyebrow and eye.

[0032] The above five proportional features are combined to form a feature vector R = [r1, r2, r3, r4, r5].

[0033] For the first and second images, calculate their facial feature vectors R1 and R2 respectively, and then calculate the Euclidean distance between them: d_ratio = ||R1 - R2||2 If d_ratio is greater than the first threshold T1 (T1=0.15 in this embodiment), the first image and the second image are determined to be from different sources, and the "different source" result is directly output, ending the detection process; if d_ratio ≤ T1, then proceed to the next step.

[0034] Step S5: Input the first image and the second image into the face feature comparison model respectively to obtain the face similarity probability; The aligned and normalized first and second face images obtained in step S3 are input into the face feature comparison model, respectively. The face feature comparison model adopts a Siamese network structure, including two feature extraction branches with shared weights.

[0035] The feature extraction branch uses EfficientNet-B0 as the backbone network. EfficientNet-B0 consists of: a standard convolutional layer (3×3 kernel, stride 2, output channels 32); followed by 7 moving inverted bottleneck convolutional modules (MBConv), with the following configurations: Module 1: 1 layer, expansion factor 1, 3×3 kernel; Module 2: 2 layers, expansion factor 6, 3×3 kernel; Module 3: 2 layers, expansion factor 6, 5×5 kernel; Module 4: 3 layers, expansion factor 6, 3×3 kernel; Module 5: 3 layers, expansion factor 6, 5×5 kernel; Module 6: 4 layers, expansion factor 6, 5×5 kernel; Module 7: 1 layer, expansion factor 6, 3×3 kernel; finally, a global average pooling layer and a fully connected layer are connected to output a 1280-dimensional feature vector.

[0036] The two branches output the first face feature vector f1 and the second face feature vector f2, respectively, both with a dimension of 1280. f1 and f2 are concatenated to obtain a 2560-dimensional vector, which is then input into the classifier. The classifier consists of three fully connected layers: the first layer takes a 2560-dimensional input and outputs a 512-dimensional vector, using ReLU activation and Dropout with a dropout rate of 0.5; the second layer takes a 512-dimensional input and outputs a 128-dimensional vector, using ReLU activation and Dropout with a dropout rate of 0.3; the third layer takes a 128-dimensional input and outputs a 1-dimensional vector, using the Sigmoid activation function. The output value is the face similarity probability p_face, ranging from [0,1].

[0037] The training data for the facial feature comparison model is constructed as follows: Positive sample construction (see) Figure 3): Select the original image and randomly apply one or more editing operations to generate a pair of source images. Editing operations include: adding a text watermark (random content, size, color, and position), adding a logo watermark (random size, transparency, and position), adding a top or bottom black border (5%~15% of the image height, containing random text), random cropping (retaining 70%~95% of the original image), random scaling (0.8~1.2x), random horizontal flipping (50% probability), random rotation (-5°~5°), and resolution transformation (scaling and then restoring). One to three combinations of operations are randomly selected and applied each time.

[0038] Negative sample construction (see) Figure 4 ): Includes three types, mixed in a 1:1:1 ratio: (a) pictures of different people; (b) different frames of a video taken continuously by the same person in the same scene (frame interval 5 to 30 frames); (c) two independent photos taken by the same person.

[0039] Training uses a contrastive loss function: L = (1-y)·d² + y·max(0, margin-d)², where y=1 indicates homogeneity, y=0 indicates dissimilarity, d is the Euclidean distance between f1 and f2, and the margin is set to 1.0. The optimizer is Adam, with an initial learning rate of 0.001, a batch size of 32, and 50 training epochs.

[0040] Step S6: Perform background removal face preprocessing on the first and second images respectively, and then input them into the background feature comparison model to obtain the background similarity probability; this step is executed in parallel with step S5 to improve detection efficiency.

[0041] First, preprocessing is performed on the first and second images to remove faces from the background. Specifically, for each original image, the same face detection algorithm as in step S2 is used to detect face regions. If no face is detected, no processing is performed on the image. If a face is detected, the largest face bounding box is selected, and it is expanded outward by 20 pixels. All pixel values ​​within the expanded rectangular area are set to 0 (pure black) to completely remove face information.

[0042] The first and second images, after face removal preprocessing, are respectively input into the background feature comparison model (see...). Figure 5The background feature comparison model also employs a Siamese network structure, with the two branches sharing weights. The feature extraction backbone network is MobileNetV3-Small, whose structure includes: an input layer (224×224×3); a standard convolutional layer (3×3, stride 2, output 16 channels); followed by multiple depthwise separable convolutional bottleneck modules, each containing depthwise convolution, compressed activation attention, pointwise convolution, and a hard Swish activation function; finally, after global average pooling, a fully connected layer is connected to output a 128-dimensional background feature vector.

[0043] The two branches output the first background feature vector b1 and the second background feature vector b2 (dimension 128), respectively. After concatenating b1 and b2, they are input into a binary classifier (fully connected layer + Softmax) to output the background similarity probability p_bg, with a value range of [0,1].

[0044] The training data for the background feature comparison model also requires face removal preprocessing. Positive samples: Editing operations (watermarking, black borders, cropping, scaling, resolution changes, but excluding flipping and rotation) that do not affect the background are applied to the original images to generate homologous pairs. Negative samples include three sources, mixed in a 1:1:1 ratio: (a) images of different people (different backgrounds after face removal); (b) image pairs from the same person's video, spaced 30-100 frames apart (slightly different backgrounds after face removal); (c) image pairs from two separate shots of the same person (similar but different backgrounds after face removal). The training loss function, optimizer, batch size, etc., are the same as for the face model.

[0045] Step S7: Based on the judgment result of Step S4, the face similarity probability of Step S5 and the background similarity probability of Step S6, perform a weighted sum to obtain the total score, and output the final judgment result based on the comparison between the total score and the second threshold. This step involves a weighted scoring decision. Since step S4 has already passed the first threshold, if this step is entered, the difference in facial proportions does not exceed the first threshold, therefore a base score S1 = 2 points is assigned.

[0046] Let p_face be the face similarity probability obtained in step S5, then the face feature score is S2 = p_face ×6.

[0047] Let the background similarity probability obtained in step S6 be p_bg, then the background feature score is S3 = p_bg × 4.

[0048] Calculate the total score: Score = S1 + S2 + S3 = 2 + p_face × 6 + p_bg × 4, with a maximum score of 12 points.

[0049] The total score is compared with the second threshold T2 (T2=7.0 in this embodiment): if the score ≥ T2, the first image and the second image are determined to be from the same source; if the score < T2, they are determined to be from different sources.

[0050] Step S8: Output the determination result, outputting the determination result of "same source" or "different source" for use by the identity verification system.

[0051] Step S9 (optional): Online hard sample mining and model update; In practical applications, misclassified image pairs are collected online. Image pairs that are actually from different sources but are misclassified as belonging to the same source are labeled as difficult negative samples; image pairs that are actually from the same source but are misclassified as belonging to different sources are labeled as difficult positive samples. When the number of collected difficult samples reaches a preset value (e.g., 1000 pairs), incremental model training is triggered. The difficult samples are mixed with the original training data at a 1:3 ratio, and the model weights are fine-tuned based on the original model weights. The learning rate is set to one-tenth of the initial learning rate (0.0001), and training is performed for 10 epochs. The updated model is used for subsequent detections to continuously adapt to the real data distribution.

[0052] This invention also provides a device for detecting homologous images (see...). Figure 2 This device is used to implement the method described in Embodiment 1. The device includes the following modules: 1. Image acquisition module: This module is used to acquire a first image and a second image to be detected. It may include a camera interface, a network communication interface, or a storage medium reading interface for acquiring image data from external sources.

[0053] 2. Face detection and key point extraction module: This module performs face detection on the first and second images respectively, determining whether faces are present and extracting 68 facial landmarks. If no face is detected in an image, the judgment module is triggered to directly output the different source result. Specifically, this module uses a regression tree-based facial landmark detection algorithm, outputting the coordinates of the face bounding box and 68 landmarks. When multiple faces are present, only the face with the largest area is retained.

[0054] 3. Face alignment and normalization module: This module is used to perform rotation alignment, cropping, and normalization on detected face regions. First, it calculates the rotation angle based on the centers of the left and right eyes, rotating the entire image so that the line connecting the two eyes is horizontal. Then, it expands outward by a factor of 0.5 from the original face bounding box, cropping out the expanded face region. Finally, it normalizes the cropped image into a 256×256 pixel square image. The module outputs the aligned and normalized first and second face images.

[0055] 4. Facial feature proportion calculation and preliminary judgment module: This is used to calculate the facial feature vector based on the 68 facial key points extracted in step 2, and to calculate the Euclidean distance d_ratio between the feature vectors of the two images. If d_ratio > the first threshold T1, it is determined that they are from different sources and the result is output directly; otherwise, the result where d_ratio ≤ T1 (i.e., the facial feature proportions are consistent) is passed to the weighted scoring module as the base score S1=2.

[0056] This module specifically calculates the distance between the eyes, the distance between the eyes and nose, the width of the mouth, the distance between the mouth and eyes, and the distance between the eyebrows and eyes, and constructs five proportional features r1~r5 to form a feature vector R.

[0057] 5. Facial Feature Comparison Module: This module is used to input aligned and normalized first and second face images into a face feature comparison model, extract feature vectors, and calculate the face similarity probability p_face. This module includes: - Feature extraction submodule: The EfficientNet-B0 Siamese network is used, with two branches sharing weights, and outputting 1280-dimensional feature vectors f1 and f2 respectively.

[0058] - Classification submodule: concatenate f1 and f2 and input them into a three-layer fully connected classifier, which outputs p_face via Sigmoid.

[0059] The construction of training data and the training method of this module are consistent with those described in step S5 of Example 1.

[0060] 6. Background Feature Comparison Module: This is used to perform background removal and face preprocessing on the first and second images, and then the preprocessed images are input into the background feature comparison model (see...). Figure 5 This module extracts background feature vectors and calculates the background similarity probability p_bg. - Background face removal preprocessing submodule: Detects the largest face bounding box in the image, expands it outward by 20 pixels, and sets the expanded area pixels to 0 (pure black).

[0061] - Feature extraction submodule: The MobileNetV3-Small Siamese network is used, with two branches sharing weights, and outputting 128-dimensional background feature vectors b1 and b2 respectively.

[0062] - Classification submodule: Concatenates b1 and b2 and inputs them into a binary classifier, outputting p_bg.

[0063] The training data for this module needs to undergo face removal preprocessing, and the training method is consistent with step S6 of Example 1.

[0064] 7. Weighted scoring judgment module: The system receives the base score S1 (S1=2 when d_ratio ≤ T1) from the facial feature ratio preliminary determination module, as well as the p_face output from the face feature comparison module and the p_bg output from the background feature comparison module. The total score Score = S1 + p_face×6 + p_bg×4 is calculated and compared with the second threshold T2 (7.0). If Score≥ T2, the system determines that the two systems are from the same source; otherwise, they are from different sources. The final determination result is then output.

[0065] 8. Output module: Used to output the determination result of "same origin" or "different origin".

[0066] 9. Online feedback and model update module (optional): This method is used to collect online misclassification cases. When the number of difficult samples reaches a preset value, it triggers incremental model training and replaces the original model with updated model weights, achieving continuous model optimization. No further deep learning inference is required. The background deface training method enables the background model to accurately extract background features even when watermarks or other overlays are added.

[0067] This invention provides a method and apparatus for detecting images from the same source, which solves the problem of large errors in determining images from the same source.

[0068] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting images from the same source, characterized in that, include: Get the first image and the second image; Face detection and key point extraction are performed on the first image and the second image respectively; The detected face regions are rotated, aligned, cropped, and normalized. Calculate the facial feature proportion feature vectors of the first image and the second image, and determine whether the difference in facial feature proportions exceeds a first threshold. The first image and the second image are respectively input into the face feature comparison model to obtain the face similarity probability; The first image and the second image are respectively subjected to background removal face preprocessing, and then input into the background feature comparison model to obtain the background similarity probability; The total score is obtained by weighting and summing the results of the facial feature proportion differences, the face similarity probability, and the background similarity probability, and then outputting the final judgment result based on the comparison between the total score and the second threshold.

2. The method for detecting homologous images according to claim 1, characterized in that, The process of rotating, aligning, cropping, and normalizing the detected face region includes: Calculate the rotation angle based on the center of the left eye and the center of the right eye, and rotate the image using the midpoint of the center of the left eye and the center of the right eye as the rotation center, so that the direction of the line connecting the two eyes is rotated to the horizontal direction; After rotation, the aligned face region is determined based on the position of the original face detection box, and the cropped area is expanded outward by a ratio of 0.

5. The cropped face region is normalized into a square image, and the square image is scaled to 256×256 pixels to obtain an aligned and normalized face image.

3. The method for detecting homologous images according to claim 1, characterized in that, The step of calculating the facial feature proportion feature vectors of the first image and the second image, and determining whether the difference in facial feature proportions exceeds a first threshold, includes: Calculate the distance between eyes, distance between eyes and nose, mouth width, distance between mouth and eyes, and distance between eyebrows and eyes from 68 extracted facial key points; Based on the above distance, five proportional features r1~r5 are constructed to form the feature vector R=[r1, r2, r3, r4, r5]; Calculate the facial feature vectors R1 and R2 of the first image and the second image respectively, and calculate the Euclidean distance d_ratio between them; If d_ratio is greater than the first threshold T1, then the first image and the second image are determined to be from different sources; if d_ratio ≤ T1, then proceed to the next step.

4. The method for detecting homologous images according to claim 3, characterized in that, The first threshold T1 = 0.

15.

5. The method for detecting homologous images according to claim 1, characterized in that, The facial feature comparison model adopts a twin network structure, including two feature extraction branches with shared weights; The feature extraction branch uses EfficientNet-B0 as the backbone network and outputs a 1280-dimensional feature vector. The two branches output the first face feature vector f1 and the second face feature vector f2 respectively. f1 and f2 are concatenated and input into the classifier. After passing through the Sigmoid activation function, the face similarity probability p_face is output, with a value range of [0,1].

6. The method for detecting homologous images according to claim 5, characterized in that, The classifier consists of three fully connected layers: the first layer has a 2560-dimensional input and a 512-dimensional output, using ReLU activation and Dropout with a dropout rate of 0.5; the second layer has a 512-dimensional input and a 128-dimensional output, using ReLU activation and Dropout with a dropout rate of 0.3; and the third layer has a 128-dimensional input and a 1-dimensional output, using the Sigmoid activation function.

7. The method for detecting homologous images according to claim 1, characterized in that, The preprocessing of background removal for the first image and the second image includes: For each original image, a face detection algorithm is used to detect face regions, the largest face bounding box is selected, and the face bounding box is expanded outward by 20 pixels. All pixel values ​​within the expanded rectangular area are set to 0.

8. The method for detecting homologous images according to claim 1, characterized in that, The background feature comparison model adopts a Siamese network structure with two branches sharing weights; the feature extraction backbone network is MobileNetV3-Small, which outputs a 128-dimensional background feature vector; the two branches output the first background feature vector b1 and the second background feature vector b2 respectively. After concatenating b1 and b2, they are input into a binary classifier to output the background similarity probability p_bg, with a value range of [0,1].

9. A device for detecting identical images, characterized in that, include: The image acquisition module is used to acquire the first and second images to be detected. The face detection and key point extraction module is used to perform face detection on the first image and the second image respectively and extract 68 facial key points; The face alignment and normalization module is used to perform rotation alignment, cropping and normalization processing on the detected face regions, and output the first face image and the second face image after alignment and normalization. The facial feature ratio calculation and preliminary judgment module is used to calculate the facial feature ratio feature vector based on the 68 facial key points, and calculate the Euclidean distance d_ratio between the feature vectors of two images. If d_ratio is greater than the first threshold T1, it is determined that they are from different sources. The face feature comparison module is used to input the aligned and normalized first face image and the second face image into the face feature comparison model, and output the face similarity probability p_face; The background feature comparison module is used to perform background deface preprocessing on the first image and the second image, and then input the background feature comparison model to output the background similarity probability p_bg. The weighted scoring and judgment module is used to calculate the total score based on the judgment result of the preliminary judgment module of facial feature proportion, the face similarity probability p_face and the background similarity probability p_bg, and compare the total score with the second threshold T2 to output the final judgment result; The output module is used to output the determination result of "same source" or "different source".

10. The image detection device according to claim 9, characterized in that, The background feature comparison module includes: The background face removal preprocessing submodule is used to detect the largest face bounding box in the image, expand it outward by 20 pixels, and set the pixels of the expanded area to 0. The feature extraction submodule uses the MobileNetV3-Small Siamese network, with the two branches sharing weights, and outputs 128-dimensional background feature vectors b1 and b2 respectively; The classification submodule is used to concatenate b1 and b2 and input them into a binary classifier, outputting the background similarity probability p_bg.