A same-model detection method and system based on a deep learning network
Patent Information
- Application Number
- CN202610141118.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-02-02
AI Technical Summary
[0003]然而,将通用模型直接应用于特定领域(如商品检索)时,面临严峻挑战:
首位命中与排名提升:难负例挖掘强化边界学习,多视图与属性增强提升前列排序稳定性。
Smart Images

Figure CN121614634B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep learning technology, and in particular relates to a detection method and system based on deep learning networks. Background Technology
[0002] Content-based image retrieval (CBIR) technology is the core of multimedia information retrieval. With the development of deep learning, especially the rise of vision-language pre-trained models, joint understanding of images and text has become possible. The SigLIP model (Sigmoid Loss for Language-Image Pre-training, proposed by Google Research in 2023, is an advanced vision-language pre-training model. It aims to address some fundamental challenges of its predecessor, the CLIP model, in terms of training efficiency and robustness to large-scale noisy data. Its core innovation lies in adopting a simpler and more efficient training paradigm) achieves powerful general visual semantic representation capabilities through contrastive learning pre-training on large-scale image-text pairs.
[0003] However, applying general models directly to specific domains (such as product retrieval) presents significant challenges: Domain distribution offset: E-commerce images are often accompanied by marketing copy, complex backgrounds, transparent packaging, and differences in subject size, which leads to noise and instability in the feature embedding extracted by the general model.
[0004] Difficulty in fine-grained differentiation: Products of the same category may have subtle differences in attributes such as style, material, and pattern. General models lack domain knowledge and it is difficult to construct clear fine-grained semantic boundaries.
[0005] Inefficient negative sample sampling: Random or simple negative samples cannot provide effective decision boundary information, the model is prone to overfitting simple samples, and has insufficient ability to distinguish difficult samples that are "similar but different".
[0006] Insufficient robustness of retrieval: Retrieval of a single image view is greatly affected by shooting conditions and background interference, resulting in fluctuating retrieval results and a lack of consistency.
[0007] System scalability bottleneck: The real-time retrieval requirements under massive image data place extremely high engineering demands on feature extraction, index construction, and query throughput.
[0008] Existing technologies either focus on optimizing a single model or lack end-to-end systems engineering design, making it difficult to improve broad recall and system robustness while ensuring high Top-K accuracy. Therefore, there is an urgent need for a high-precision image retrieval solution that optimizes both the algorithm and the system.
[0009] Therefore, there is an urgent need for a detection method and system based on deep learning networks. Summary of the Invention
[0010] To achieve the objectives of this invention, the following technical solution is adopted: Specifically, this application provides a duplicate detection system based on deep learning networks, which includes: The preprocessing module preprocesses the raw unstructured image data, transforming it into standardized multi-view image data suitable for model learning and retrieval; The embedding module constructs training images based on the multi-view image data. In the embedding space of the initial or pre-trained SigLIP1, it retrieves the highest number of similar but different product or label negative examples for each training image, maintains a nearest neighbor cache or temporary index, and dynamically selects difficult negative examples within the batch for training according to the current model embedding. It maps the attribute short text generated from the training images to text feature vectors and converts them into text feature vectors through the SigLIP text encoder. Multimodal contrastive learning is achieved based on a loss function covering multiple dimensions. The retrieval fusion module constructs a loss function using contrastive learning loss and hard-to-bear example weights, loads SigLIP pre-trained weights, fine-tunes the parameters of the visual encoder and text encoder, and uses the fine-tuned SigLIP to extract feature vectors from the three views of all images in the image library. It also constructs a master index for the original image view and simultaneously sends the feature vectors of the three views corresponding to the original image view to the vector database. It queries the top-N nearest neighbors for each view and obtains a retrieval list of images with the same features as the user's searched image based on the dynamic weights and the fusion calculation results of the vector database.
[0011] The beneficial effects of this invention are as follows: First-place hit and ranking improvement: Difficult negative example mining enhances boundary learning, and multi-view and attribute enhancement improve the stability of top ranking.
[0012] Recall and interpretability enhancements: Fine-grained attribute constraints improve long-tail and complex scenarios; multi-view fusion enhances robustness.
[0013] The project is scalable: it supports both batch and streaming processing, high-concurrency downloading and GPU batch embedding, and is easy to integrate.
[0014] Furthermore, it is transformed into standardized multi-view image data suitable for model learning and retrieval, specifically including: First, invalid URLs are filtered out. Then, the original unstructured images in the URLs are decoded, and the original unstructured images with alpha channels are uniformly converted into JPG format with a white background. Finally, the pre-trained subject detection model is invoked to obtain the subject bounding box and generate two derived views: a white-background subject image and a cropped subject image (retaining only the area within the bounding box). The original image and the two derived images together constitute a multi-view image.
[0015] Furthermore, for each training image, retrieve the top number of similar but different product or tag-based negative examples, specifically including: Initial embedding vectors are generated for all training images using a SigLIP model pre-trained on a public dataset. For each training image in the training samples, the pre-trained SigLIP is used to extract features for all training images. For each training image, the number of target samples with the highest similarity but different product IDs are retrieved in the feature space, and their similarity set is denoted as {s_i}. Based on the negative examples of the training images, select a target number of negative examples from high to low to obtain the negative examples of the training images.
[0016] Furthermore, training is performed by dynamically selecting difficult negative examples within the batch based on the current model embedding, specifically including: Phase 1, during the first iteration step interval: use random negative examples, sample negative examples from the offline library with a 10% probability, and let the model initially adapt to the domain data distribution; Phase Two, within the second iteration step interval: Based on the sampling strategy and batch adjustment and identification method, determine the input batch of negative samples and the newly input negative samples in different input batches. Update the features of all training samples using the current model. When inputting negative samples in each batch, recalculate the difficult negative samples. Gradually increase the input ratio of difficult negative samples to the target input ratio threshold, and gradually increase the input number of negative samples with higher similarity. Phase 3, within the third iteration step interval: maintain the target proportion of difficult negative sample sampling, introduce the most difficult negative sample, and add gradient clipping to prevent training oscillations, further sharpening the model boundary.
[0017] Furthermore, a search list of similar products matching the user's search results is obtained, specifically including: Using a finely tuned SigLIP image encoder, feature vectors are extracted from the three views of all images in the library. During the indexing phase, a master index is built only for the original view V_orig of each image. Obtain the user's search image, generate three views of the search image, send the feature vectors of the three views to the vector database simultaneously, and query the Top-K nearest neighbors of each view; For each view, the Top-K score list is converted into an ordinal ranking. The ranking is then converted into a standardized score. Different image recognition weight coefficients are determined based on the cosine similarity between image features in each view. The product of the weight coefficients and the standardized scores is summed and deduplicated to obtain a comprehensive similarity coefficient. The final search list is obtained based on the comprehensive similarity coefficient in descending order.
[0018] Secondly, this application provides a method for detecting identical items based on deep learning networks, applied to the aforementioned system for detecting identical items based on deep learning networks, specifically including: S1 uses the products corresponding to the training images as a basis to determine the listing data of the products corresponding to different training images. Based on the listing data, the target quantity is determined. According to the target quantity and the recognition results of the negative sample of different training images, the similarity interval of the negative sample of different training images is determined. S2 determines the sampling strategy for negative samples based on the similarity between the similarity interval and the target similarity interval, and in combination with the composition data of the negative samples. Based on the sampling strategy, S3 embeds hard-to-resolve examples within the dynamically selected batches for training in batches, and determines the adjustment and identification method for the batches according to the usage data of hard-to-resolve examples in different batches and the sampling strategy.
[0019] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0021] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.
[0022] Figure 1 This is a framework diagram of a similar detection system based on deep learning networks; Figure 2 It is a flowchart for converting standardized multi-view image data into data suitable for model learning and retrieval; Figure 3 This is a flowchart for retrieving the top number of similar but different product or tag-based negative examples for each training image; Figure 4 This is a flowchart of training based on the current model embedding and dynamic selection of difficult negative examples within a batch; Figure 5 This is a flowchart of a similar detection method based on deep learning networks. Detailed Implementation
[0023] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0024] Example 1 like Figure 1 As shown, this application provides a similarity detection system based on deep learning networks, specifically including: The preprocessing module preprocesses the raw unstructured image data, transforming it into standardized multi-view image data suitable for model learning and retrieval; The embedding module constructs training images based on the multi-view image data. In the embedding space of the initial or pre-trained SigLIP1, it retrieves the highest number of similar but different product or label negative examples for each training image, maintains a nearest neighbor cache or temporary index, and dynamically selects difficult negative examples within the batch for training according to the current model embedding. It maps the attribute short text generated from the training images to text feature vectors and converts them into text feature vectors through the SigLIP text encoder. Multimodal contrastive learning is achieved based on a loss function covering multiple dimensions. The retrieval fusion module constructs a loss function using contrastive learning loss and hard-to-bear example weights, loads SigLIP pre-trained weights, fine-tunes the parameters of the visual encoder and text encoder, and uses the fine-tuned SigLIP to extract feature vectors from the three views of all images in the image library. It also constructs a master index for the original image view and simultaneously sends the feature vectors of the three views corresponding to the original image view to the vector database. It queries the top-N nearest neighbors for each view and obtains a retrieval list of images with the same features as the user's searched image based on the dynamic weights and the fusion calculation results of the vector database.
[0025] Furthermore, such as Figure 2 As shown, the data is transformed into standardized multi-view image data suitable for model learning and retrieval, specifically including: First, invalid URLs are filtered out. Then, the original unstructured images in the URLs are decoded, and the original unstructured images with alpha channels are uniformly converted into JPG format with a white background. Finally, the pre-trained subject detection model is invoked to obtain the subject bounding box and generate two derived views: a white-background subject image and a cropped subject image (retaining only the area within the bounding box). The original image and the two derived images together constitute a multi-view image.
[0026] It should be noted that the white background main view is the detected subject pasted onto a pure white background, and the subject cropping image is the image that retains only the area inside the bounding box.
[0027] Specifically, the following implementation method is adopted: URL filtering and downloading: The system verifies the validity of the URL and downloads the image data. Assume the original image is a PNG image of a sneaker with a transparent background.
[0028] Format unification and background processing: Decode the image, detect that its color mode is RGBA (including the alpha channel), and create a pure white (RGB: 255,255,255) background canvas of the same size as the original image.
[0029] The RGB channel data of the original image are composited onto a white canvas based on the transparency of its alpha channel to generate a new RGB image.
[0030] The image is then uniformly scaled to a fixed resolution (e.g., 224x224) and converted to JPG format for storage. This step eliminates the distraction caused by the transparent background, allowing the model to focus on the subject.
[0031] Subject detection and multi-view generation: The unified image is input into the subject detection model. Assume the model returns bounding box coordinates [x1, y1, x2, y2].
[0032] White background main image generation: Using this bounding box, the main area is cropped from the original unified image, and then centered and pasted into a new pure white background (224x224) image.
[0033] Main body cropping image generation: The region is directly cropped from the original unified image based on the bounding box coordinates and scaled to 224x224.
[0034] At this point, for a single original input image, we have obtained three view files: shoe_original.jpg, shoe_white.jpg, and shoe_crop.jpg. These views will describe the same product from different perspectives (global context, clean subject, and subject details).
[0035] It should be noted that the subject detection model is constructed using the YOLOv8 model.
[0036] Furthermore, such as Figure 3As shown, for each training image, the top number of similar but different product or tag-based negative samples are retrieved, specifically including: Initial embedding vectors are generated for all training images using a SigLIP model pre-trained on a public dataset. For each training image in the training samples, the pre-trained SigLIP is used to extract features for all training images. For each training image, the number of target samples with the highest similarity but different product IDs are retrieved in the feature space, and their similarity set is denoted as {s_i}. Based on the negative examples of the training images, select a target number of negative examples from high to low to obtain the negative examples of the training images.
[0037] Furthermore, such as Figure 4 As shown, training is performed by dynamically selecting difficult negative examples within a batch based on the current model embedding, specifically including: Phase 1, during the first iteration step interval: use random negative examples, sample negative examples from the offline library with a 10% probability, and let the model initially adapt to the domain data distribution; Phase Two, within the second iteration step interval: Based on the sampling strategy and batch adjustment and identification method, determine the input batch of negative samples and the newly input negative samples in different input batches. Update the features of all training samples using the current model. When inputting negative samples in each batch, recalculate the difficult negative samples. Gradually increase the input ratio of difficult negative samples to the target input ratio threshold, and gradually increase the input number of negative samples with higher similarity. Phase 3, within the third iteration step interval: maintain the target proportion of difficult negative sample sampling, introduce the most difficult negative sample, and add gradient clipping to prevent training oscillations, further sharpening the model boundary.
[0038] Specifically, the first iteration step range is between 0 and 30% of the maximum iteration steps during the current model's training process; the second iteration step range is between 30% and 80% of the maximum iteration steps during the previous model's training process; and the third iteration step range is between 80% and the maximum iteration steps.
[0039] Specifically, the target investment ratio threshold is 40%, and the target ratio is 50%.
[0040] In another embodiment, Phase 1 (warm-up period, 0-30% of total steps): mainly uses random negative examples, sampling negative examples from the offline library with a 10% probability, to allow the model to initially adapt to the domain data distribution.
[0041] Phase Two (Reinforcement Phase, 30%-80% of Total Steps): Every N training steps (e.g., 10 steps), the features of all training samples are updated using the current model, and the hard negative examples are recalculated. The sampling probability of hard negative examples is gradually increased to 40%, and samples with higher similarity are selected (e.g., t_low is increased by 0.1). This forces the model to learn more refined distinctions.
[0042] Phase 3 (convergence period, 80%-100% of total steps): Maintain a high proportion (50%) of difficult negative examples and introduce "most difficult negative examples" (highest similarity and different categories). At the same time, gradient clipping (norm=1.0) is added to prevent training oscillations and further sharpen the model boundary.
[0043] Furthermore, multimodal contrastive learning is achieved based on a loss function that covers multiple dimensions, specifically including: Attribute text generation: Using visual language models (such as GPT-4V, Qwen-VL), generate multiple short attribute texts for each training image; Attribute cleaning and structuring: A domain dictionary is established, and the attribute short texts generated by the visual language model are parsed and normalized through regular expressions and dictionary matching to generate N uniform and conflict-free short text descriptions for each training image, which serve as the text feature vectors of the training images. In a training batch, for a training image, its positive sample pairs are expanded to: image positive examples, all text feature vectors of the training image, and a loss function is constructed that covers the contrast loss in three dimensions: image-image, image-text, and text-image, to ensure the alignment of visual features with semantic attributes.
[0044] Specifically, the steps include: Attribute generation: Using an open-source VLM (such as BLIP-2), generate short attribute text for each training image. Design structured prompts: "Describe the product's category, color, material, pattern, and style. Output in 'attribute:value' format;" Attribute cleaning and vectorization: A domain attribute dictionary is established, short attribute texts are parsed, and synonyms are mapped to standard values (e.g., "dark blue" and "sapphire blue" are mapped to "blue"). Each cleaned attribute text (e.g., "color: blue; material: cotton") is converted into a text feature vector T_j by the SigLIP text encoder.
[0045] Multimodal contrastive learning: In a training batch, for an image I_a, its positive sample pairs are expanded to: 1) positive image I_p; 2) all its own attribute text features {T_j}. The loss function is calculated to cover the contrastive loss in three dimensions: image-image, image-text, and text-image, as shown in the following formula: L_total = λ1 * L_i2i + λ2 * L_i2t + λ3 * L_t2i Where L_i2i is the image-to-image contrast loss, focusing on the hard negative examples; L_i2t and L_t2i are the image-to-text contrast loss, ensuring the alignment of visual features with semantic attributes; and λ1, λ2, and λ3 are weight coefficients.
[0046] Furthermore, fine-tuning of the parameters of the visual encoder and text encoder is performed, specifically including: Load SigLIP pre-trained weights and fine-tune only some parameters of the visual encoder and text encoder in batches to balance training efficiency and adaptation effect. The AdamW optimizer was used, with an initial learning rate of 5e-6 and cosine annealing scheduling. The batch size was set to 128 based on GPU memory. During training, the system evaluates on the validation set at regular intervals and saves the best-performing checkpoints based on the loss function, enabling fine-tuning of the parameters of the visual encoder and text encoder.
[0047] This embodiment details the specific implementation steps, hyperparameter configuration, and training process management for domain-adaptive fine-tuning of a pre-trained SigLIP model. This process aims to efficiently transfer the powerful representational capabilities of a general visual language model to a specific vertical domain (such as product images), while avoiding overfitting and training instability.
[0048] 1. Model initialization and parameter freezing strategies; First, the pre-trained SigLIP model weights are loaded. This invention preferably uses the SigLIP-ViT-B / 16-256 variant, with a Vision Transformer Base visual encoder, an image input resolution of 256x256, and a Transformer-based text encoder. To achieve the core goal of "efficient adaptation," we employ a layer-wise progressive unfreezing strategy instead of full parameter fine-tuning.
[0049] The specific implementation is as follows: Initial Freeze Phase: Immediately after loading the model, all parameters of the visual encoder and text encoder are frozen. Subsequently, only the parameters of the last two Transformer blocks of each encoder are unfrozen. Specifically, for a visual encoder containing 12 Transformer blocks, blocks 11 and 12 are unfrozen; a similar operation is performed for the text encoder. In this phase, most of the model's parameters are frozen, allowing only the highest-level, most abstract feature representations to be fine-tuned to adapt to the semantic concepts of the new domain. This results in extremely high training efficiency and effectively prevents catastrophic forgetting.
[0050] Mid-stage unfreezing: This stage begins when the model's loss on the validation set stops decreasing for two consecutive evaluation epochs. We further unfreeze layers 9 and 10 of the visual encoder. At this point, the model begins to learn how to adjust the combination of domain-relevant intermediate layer features. Simultaneously, the base learning rate is halved to fine-tune these newly unfrozen parameters.
[0051] Late-stage fine-tuning: After the model performance has stabilized, as an optional step, all layers can be unfrozen, but a large weight decay and a small learning rate can be introduced to "polish" the overall model to achieve optimal performance. During this stage, validation set performance needs to be carefully monitored to avoid overfitting.
[0052] 2. Optimizer, learning rate scheduling, and batch configuration: Optimized configuration during training is crucial to the final model performance. Specific settings are as follows: Optimizer: The AdamW optimizer is used, which decouples weight decay from gradient updates, helping to improve the model's generalization ability. Parameters are set as follows: beta1=0.9, beta2=0.999, and weight decay rate (weight_decay) is set to 0.05 to effectively control model complexity.
[0053] Learning rate strategy: Use cosine annealing with warmup.
[0054] Warm-up phase: In the first 1,000 training steps, the learning rate increases linearly from 0 to an initial learning rate of 5e-6. This phase helps stabilize the unstable gradients in the early stages of training.
[0055] Annealing Phase: After warm-up, the learning rate gradually decays from 5e-6 to the preset minimum learning rate of 1e-7 according to the cosine function. The entire annealing cycle covers the remaining total training steps. This smooth decay helps the model to perform more refined optimization in the loss plateau region.
[0056] Batch and Gradient Accumulation: On the GPU, the batch size (per_device_batch_size) is set to 16. To obtain a larger effective batch size of 128 to stabilize the optimization direction, we set the gradient accumulation steps (gradient_accumulation_steps) to 1 (i.e., after calculating the gradients of 16 samples per GPU, we synchronize once, resulting in an effective batch size of 128). The total number of training steps is set to 100,000.
[0057] 3. Training monitoring, evaluation, and checkpoint management; To ensure the training process is controlled and the best model is preserved, a complete monitoring and saving mechanism has been established: Evaluation metrics and frequency: A full evaluation is performed every 2,000 training steps, i.e., on the validation set (containing approximately 5,000 unseen query-candidate pairs). Core monitoring metrics include: Recall@10: This measures the model's recall capability in a wide range of searches and is the main goal of optimization.
[0058] Precision@1: Measures the accuracy of the model’s most relevant results, ensuring that the first-result hit rate does not decrease.
[0059] Validation loss: Monitors how well the model fits on unseen data.
[0060] Checkpoint strategy: Implement the "best checkpoint saving" strategy. Continuously track the core metric Recall@10, and only when this metric exceeds the historical best value will the current model weights, optimizer state, and training steps be completely saved as a checkpoint file.
[0061] Early stopping mechanism: The patience value is set to 5. This means that if the core validation metric (Recall@10) does not improve after 5 consecutive evaluations (i.e., 10,000 consecutive steps), early stopping is triggered, training is terminated, and the system rolls back to the saved best checkpoint. This effectively avoids unnecessary computational resource consumption and potential overfitting.
[0062] Through the above-mentioned refined training configuration and management process, this invention can systematically and efficiently complete the deep adaptation of the SigLIP model in a specific domain, laying a solid model foundation for subsequent high-precision image retrieval.
[0063] Furthermore, a search list of similar products matching the user's search results is obtained, specifically including: Using a finely tuned SigLIP image encoder, feature vectors are extracted from the three views of all images in the library. During the indexing phase, a master index is built only for the original view V_orig of each image. Obtain the user's search image, generate three views of the search image, send the feature vectors of the three views to the vector database simultaneously, and query the Top-K nearest neighbors of each view; For each view, the Top-K score list is converted into an ordinal ranking. The ranking is then converted into a standardized score. Different image recognition weight coefficients are determined based on the cosine similarity between image features in each view. The product of the weight coefficients and the standardized scores is summed and deduplicated to obtain a comprehensive similarity coefficient. The final search list is obtained based on the comprehensive similarity coefficient in descending order.
[0064] Specifically, feature extraction and index construction: Using a fine-tuned SigLIP image encoder, feature vectors are extracted from the three views of all images in the library. To save storage, during the indexing phase, only the original view V_orig of each image is constructed as the main index, since its features have been sufficiently enhanced through multi-view and multimodal learning during training.
[0065] HNSW (Hierarchical Navigable Small World) was chosen as the indexing algorithm due to its good balance between recall and query speed. Parameters were set as follows: number of construction levels M=16, dynamic candidate set size efConstruction=200. When inserting hundreds of millions of vectors, a sharding strategy was employed, horizontally partitioning the vector set and storing it across multiple physical nodes.
[0066] Dynamic multi-view fusion search process: When a user submits a query for image Q, the online service performs the following atomic operations: Query preprocessing: Perform step one on Q to generate its three views Q_orig, Q_white, and Q_crop.
[0067] Parallel vector retrieval: The feature vectors of the three views are simultaneously sent to the vector database, and the top-200 nearest neighbors of each view are retrieved. This step is a parallel I / O operation, and its time consumption mainly depends on network and database performance.
[0068] Adaptive score calibration and fusion: Score Distribution Normalization: Due to subtle differences in the feature spaces of different views, the original cosine similarity score distributions returned will differ. Order-based normalization is employed: For the Top-K score list returned for each view, it is transformed into an ordinal rank, and then the rank is converted into a normalized score: S_norm = 1.0 - (rank - 1) / K. This method is insensitive to the absolute scale of the scores and is more robust.
[0069] Dynamic weight calculation: Instead of using fixed weights, the weights are dynamically calculated based on the characteristics of the query image itself. The cosine similarity sim_w_c between the features of the Q_white and Q_crop views is calculated. If sim_w_c is high (>0.9), it indicates a simple background and a prominent subject, so Q_orig is given a lower weight (e.g., 0.2), while Q_white and Q_crop are given higher weights (0.4 each). If sim_w_c is low, it indicates a complex background or inaccurate subject detection, so the weight of Q_orig is increased (e.g., 0.5), and the weights of other views are decreased. The weight set {w_orig, w_white, w_crop} is normalized to a sum of 1.
[0070] Weighted summation and deduplication: For each candidate image ID, its final score is: S_final = w_orig * S_norm_orig + w_white * S_norm_white + w_crop * S_norm_crop. All candidates are sorted in descending order by S_final to obtain the final search list.
[0071] Furthermore, this also includes deploying SigLIP models using batch processing pipelines or streaming services.
[0072] Type 1: Batch processing pipeline (V1) - suitable for creating full / incremental indexes for offline libraries.
[0073] Phase 1 (Download and Filter): Multiple Workers concurrently pull image URLs from the data source and perform filtering.
[0074] Phase 2 (Preprocessing and Embedding): The downloaded images are placed in a queue and preprocessed and feature extracted in batches by the GPU server.
[0075] Phase 3 (Index Building): Feature vectors are written to the vector database in batches, triggering index building or updates.
[0076] The state is persisted across stages, supporting breakpoint resumption. The monitoring panel allows you to view the length of each queue, processing speed, and error logs.
[0077] Type 2: Streaming Service (V2) - Suitable for near real-time online image import and retrieval services.
[0078] Asynchronous data ingestion: Receive new image URLs or binary data via a message queue.
[0079] Concurrent processing: The service cluster concurrently executes download, preprocessing, and embedding operations. The embedding step utilizes GPUs for real-time computation in small batches (e.g., 32) to ensure low latency.
[0080] Incremental index: The extracted vectors are immediately written to the incremental buffer of the vector database, and the database backend periodically merges the buffer data into the main index.
[0081] Search API: Provides a high-concurrency, low-latency RESTful API that receives query images and returns the fused search results.
[0082] Resource monitoring: Real-time monitoring of API latency, QPS, GPU utilization, and memory consumption, and setting threshold alarms (e.g., triggering when P99 latency > 200ms).
[0083] Example 2 Secondly, such as Figure 5 As shown, this application provides a method for detecting identical items based on deep learning networks, applied to the aforementioned system for detecting identical items based on deep learning networks, specifically including: S1 uses the products corresponding to the training images as a basis to determine the listing data of the products corresponding to different training images. Based on the listing data, the target quantity is determined. According to the target quantity and the recognition results of the negative sample of different training images, the similarity interval of the negative sample of different training images is determined. Furthermore, the product corresponding to the training image is determined based on the product matched with the training image.
[0084] Furthermore, the product listing data corresponding to the training images is determined based on the listing volume of the product type on different platforms.
[0085] Furthermore, the product types are categorized based on the customer groups and styles of the products.
[0086] Specifically, the method for determining the target quantity is as follows: In training an image retrieval model for apparel products, the intensity of hard negative example mining needs to be adjusted differentiated based on the popularity and competitiveness of the apparel category in the market. To this end, this invention designs an adaptive mechanism based on apparel product listing data to dynamically determine the "target number" (K value) of hard negative examples to mine for each anchor point sample in each training batch. The implementation of this method is specifically divided into the following three stages: S11 Based on the product listing data, determine the number of products listed for the corresponding product type in different platforms; S12 determines the type of demand for identifying the same product in the training images based on the number of images uploaded; This phase begins with fine-grained category segmentation and market analysis of all apparel products in the training data. The system connects to public product data interfaces of major e-commerce platforms and social e-commerce platforms to calculate the total number of newly listed apparel products in each category over the past 30 days.
[0087] High-demand apparel category: This refers to categories with huge market inventory, rapid style iteration, and extremely fierce competition due to homogeneity. The criterion is: more than 50,000 pieces available for sale. Typical categories include: Basic T-shirts / shirts: White cotton T-shirts, striped shirts, etc., with numerous similar-looking items. Classic jeans: Blue straight-leg jeans, black skinny jeans, etc. Trending dresses / sneakers: Recently popular styles on social media, leading to a surge in imitations.
[0088] Category II, Apparel in High Demand: This category refers to products with a certain market size and variety of styles. The criteria for this category are: 5,000 to 50,000 items listed. Typical categories include: Designer Tops: Styles with unique cuts or niche design elements. Functional Outerwear: Windbreakers, denim jackets in specific styles, etc. Stylized Accessories: Bags, hats, etc. in specific styles.
[0089] Three categories of low-demand apparel: These refer to relatively niche, long-tail, or new product testing categories. The criterion is: fewer than 5,000 pieces available. Typical categories include: High-end custom apparel: expensive items with extremely low production volumes; Extremely niche style clothing: such as Lolita, cyberpunk, and other culturally specific apparel; Brand new product launches: products that have recently been released and have not yet been widely distributed.
[0090] S13 determines the number of targets based on the same recognition requirement type of different training images.
[0091] It is understood that the type of demand for identifying similar items in the training images is determined based on the number of items uploaded, specifically based on the range in which the number of items uploaded is located.
[0092] Specifically, the same type of identification needs includes three types: Type 1, Type 2, and Type 3, where Type 1 needs are greater than Type 2 needs, and Type 2 needs are greater than Type 3 needs.
[0093] Specifically, the number of targets is determined based on the same recognition requirement type of different training images, including: Case 1: If there is no requirement type among the same recognition requirements of different training images, then the target quantity is determined to be the preset quantity; Strategy A (Baseline Model, K=3): This strategy is activated when no images of any high-demand clothing category appear in the batch. For example, a batch consisting entirely of images of "high-end custom cheongsams" and "niche designer jewelry." In this case, the model's learning focus is on establishing basic category differentiation capabilities, which can be satisfied with a small number of low-duration examples (3).
[0094] Scenario 2: If there are training images of the same recognition requirement type, then obtain the number of training images of the same requirement type. When the proportion of the number of training images of the same requirement type in the training images is greater than the preset proportion threshold, then determine the target number as the second preset number. Strategy B (Reinforcement Mode, K=8): This strategy is activated when the number of images of a high-demand clothing category in a batch exceeds 30%. For example, a batch containing a large number of images of "white basic T-shirts" and "blue straight-leg jeans". This indicates that the batch is filled with easily confused products, and it is necessary to "force" the model to learn extremely subtle differences (such as neckline ribbing, wash marks, and print details) by mining more difficult examples (8).
[0095] Scenario 3: If the proportion of training images of the aforementioned requirement type in the training images is not greater than a preset proportion threshold, then the recognition requirement weight value of the training images is determined based on the same recognition requirement type of the training images, and the target quantity is determined based on the sum of the recognition requirement weight values of different training images.
[0096] The system will assign a dynamic weight to each image in the batch and determine the K value based on the total weight. Single image weight assignment: Category 1 high-demand clothing images, base weight W_high = 3.0; Category 2 medium-demand clothing images, base weight W_medium = 1.5; Category 3 low-demand clothing images, base weight W_low = 1.0. Batch total weight calculation and threshold comparison: Calculate the batch total weight Total_Weight = Σ(weight of each image), and preset a dynamic threshold Threshold = Batch_Size * 2.2 (i.e., Batch_Size is the number of training images). If Total_Weight > Batch_Size * 2.2, then the overall recognition difficulty of this batch is determined to be high, and K=8 is used. If Total_Weight ≤ Batch_Size * 2.2, then K=3 is used.
[0097] It is understood that when the sum of the recognition requirement weight values of different training images is greater than the preset weight threshold, the target number is determined to be the second preset number; when the sum of the recognition requirement weight values of different training images is not greater than the preset weight threshold, the target number is determined to be the preset number.
[0098] It should be noted that the preset quantity is less than the second preset quantity.
[0099] Furthermore, the similarity interval of the negative samples of the training images is determined by selecting a target number of negative samples from the negative samples of the training images, from high to low, to form the similarity interval.
[0100] S2 determines the sampling strategy for negative samples based on the similarity between the similarity interval and the target similarity interval, and in combination with the composition data of the negative samples. Specifically, the target similarity interval is the interval where the similarity is greater than a preset similarity threshold.
[0101] Specifically, the method for determining the sampling strategy for negative samples is as follows: In training clothing image retrieval models, the timing of introducing difficult negative examples into the training process is crucial to the final model performance. This invention proposes a dynamic sampling strategy based on feature space clustering analysis. Its core lies in real-time evaluation of the concentration and competitive landscape of "extremely difficult samples" in the current training batch, thereby intelligently determining the intensity of introducing new difficult negative examples in the next training cycle and reducing the impact on the model when training is incomplete.
[0102] S21 Based on the similarity between the similarity interval and the target similarity interval, determine the training image whose similarity interval falls into the target similarity interval, and use it as the target training image; Define the "target similarity interval": Set the cosine similarity threshold T_hard = 0.78. For any anchor sample, the interval covered by all negative examples with a similarity greater than 0.78 among its 10 hard negative examples is defined as its "target similarity interval". For example, if the similarity of the 10 hard negative examples of a "black slim-fit suit" is [0.82, 0.80, 0.79, 0.76, 0.74, 0.72, 0.70, 0.68, 0.65, 0.63], then its target similarity interval is similarity > 0.78, including the first 3 samples.
[0103] Identifying "Target Training Images": The system scans all samples within a batch. If, among a sample's 10 difficult negative examples, there is at least one sample with a similarity greater than 0.78, that sample is labeled as a "target training image." This indicates that the current model believes the product has highly confusing "nearby competitors" in the market. For example, a basic "white cotton T-shirt" and a trendy "floral dress" are very likely to be labeled as such.
[0104] S22 uses the composition data of the negative sample to determine the composition ratio of the negative sample in the training sample; S23 determines the sampling processing strategy for the negative samples based on the target training image data and the composition ratio of the negative samples.
[0105] It is understandable that, based on the target training image data and the composition ratio of negative samples, the sampling processing strategy for the negative samples is determined, specifically including: In a possible embodiment, the composition ratio of the negative sample is obtained, and it is determined whether the composition ratio of the negative sample is greater than a preset composition ratio threshold. If so, the sampling processing strategy of the negative sample is determined by using the preset ratio; otherwise, proceed to the next step. Level 1: Global negative sample health check. Calculation: Statistical analysis of the proportion of negative samples in the training samples. Decision: If the proportion is >90% (indicating that training is overly reliant on negative samples), to prevent the model from "indigesting" negative samples, the system immediately activates a conservative strategy: In the next round of training, if the original batch input ratio is 20%, then the next round will use a preset ratio plus 20%, i.e., 21%, to determine the proportion of negative samples, in order to ensure training stability.
[0106] It should be noted that, based on the target training image data, it is determined whether a target training image exists. If so, proceed to the next step; otherwise, the second preset ratio is used to determine the sampling processing strategy for the negative sample. Extremely difficult sample existence check: Condition: If the overall proportion is healthy (≤40%), check if the batch contains a "target training image". Decision: If no target training image exists, it indicates that the product differentiation in this batch is generally good (e.g., the batch mainly consists of "down jackets" and "beach shorts"), and there is no extreme challenge. To maintain stability, an aggressive strategy is also adopted. In the next round of training, if the original batch's input proportion is 20%, then the next time a second preset proportion plus 20%, i.e., 22%, is used to determine the proportion of negative samples to ensure training stability.
[0107] Based on the composition data of the target training image, determine the composition ratio of the target training image in the training image, and determine whether the composition ratio of the target training image in the training image is greater than a preset composition ratio threshold. If yes, then use the preset ratio to determine the sampling processing strategy for the negative sample. If no, proceed to the next step. Concentration check of difficult samples: Condition: If target training images exist, calculate their proportion in the training images. Decision: If the proportion is >25%, it means that there are many "difficult" samples. In order to avoid training impact, the system continues to adopt a conservative strategy.
[0108] Based on the similarity between different target training images, target training images with similarity greater than a preset similarity value are identified. Based on the target training image data with similarity greater than the preset similarity value, the sampling processing strategy for the negative sample is determined.
[0109] Specifically, based on target training image data containing target training images with a similarity greater than a preset similarity value, a sampling strategy for the negative sample is determined, including: Target training images with similarity greater than a preset similarity value are considered as overlapping training images. An overlap factor is determined based on the proportion of overlapping training images in the target training images and the average proportion of target training images in the training images. A sampling strategy for negative samples is then determined based on the overlap factor.
[0110] It is understood that when the overlap factor is greater than the preset overlap factor threshold, the sampling processing strategy for the negative sample is determined to be to use a preset ratio to determine the sampling processing strategy for the negative sample; and when the overlap factor is not greater than the preset overlap factor threshold, the sampling processing strategy for the negative sample is determined to be to use a second preset ratio to determine the sampling processing strategy for the negative sample.
[0111] Analysis of the competitive landscape within difficult samples (core step): Scenario: If difficult samples exist but their proportion does not exceed the limit (≤25%), then a detailed analysis is performed. The system calculates the feature similarity between all pairs of "target training images".
[0112] Identify high-cohesion clusters: Find those "target training images" with a similarity greater than 0.85 and label them as "overlapping training images". For example, you might find that "same style hoodies in different colors of the same brand" form such a high-cohesion cluster.
[0113] Calculate the "overlap factor": Calculate the overlap cluster percentage = number of overlapping training images / total number of target training images, calculate the target image percentage = number of target training images / number of training images, and calculate the overlap factor = (overlap cluster percentage + target image percentage) / 2.
[0114] Final decision: If the overlap factor is greater than a preset threshold (e.g., 0.15), it indicates that difficult samples not only exist but are also highly clustered and similar to each other, forming a local "red ocean competition zone." Introducing a large number of new difficult negative examples at this point can easily cause the model to collapse in this region. Therefore, a conservative strategy is adopted.
[0115] If the overlap factor is ≤ 0.15, it indicates that the difficult samples are relatively dispersed and the competition is not concentrated. The system can then activate an aggressive strategy: according to the second preset ratio, to accelerate the sharpening of the model boundary.
[0116] It should be noted that the preset ratio is less than the second preset ratio. That is, when negative sample input is updated each time, the preset ratio or the second preset ratio of negative sample is increased. By reducing the number of negative sample inputs when the similarity is high, the technical problem of a high impact on the model before it is fully trained is avoided when the similarity of negative sample inputs is high and the number of inputs is large.
[0117] Based on the sampling strategy, S3 embeds hard-to-resolve examples within the dynamically selected batches for training in batches, and determines the adjustment and identification method for the batches according to the usage data of hard-to-resolve examples in different batches and the sampling strategy.
[0118] Specifically, the difficult negative examples are negative examples whose similarity to the model's existing training samples is greater than a preset similarity threshold and do not meet the requirement.
[0119] Specifically, the method for determining the batch adjustment identification method is as follows: In the iterative training of clothing image retrieval models, the "discovery" and "input" of difficult examples are dynamically changing. To ensure that the model can learn the increasing number of difficult samples at the most appropriate pace, this invention designs a training batch adaptive adjustment mechanism. This mechanism dynamically determines whether to expand the batch size to accommodate more samples and provide more stable gradient estimates by monitoring the growth trend of difficult examples in multiple consecutive training batches in real time, thereby coping with the high-difficulty learning stage.
[0120] Based on the data on the use of difficult cases in different batches, determine the proportion of new difficult cases in different batches. The proportion of newly added difficult-to-bearing examples: refers to the proportion of newly mined difficult-to-bearing examples that are actually used in the loss function calculation in a training batch, relative to the total number of newly mined difficult-to-bearing examples for all anchor points in that batch. Since the samples are random when they are added, it is related to the sampling strategy.
[0121] The number of newly identified difficult examples: This refers to the total number of newly identified samples that meet the "difficult example" criteria (similarity > threshold T_hard) for all anchor samples in the training set under the current training state. It reflects the objective increase in the number of "difficult opponents" in the model's field of vision and is a direct manifestation of changes in the model's discriminative power.
[0122] Based on the updated identification results of difficult negative examples in different batches, determine the number of newly identified difficult negative examples in different batches; The identification adjustment method for each batch is determined based on the proportion of new inputs and the number of new identifications for difficult-to-handle cases in different batches, combined with the sampling processing strategy.
[0123] It should be noted that the newly added proportion of difficult examples refers to the proportion of newly added difficult examples in the existing training samples of the model.
[0124] Specifically, the number of newly identified difficult-to-bear examples is the number of newly identified difficult-to-bear examples based on the existing training samples.
[0125] Specifically, in the above steps, if the sampling processing strategy is the second ratio, as long as there is a batch where the proportion of newly added difficult examples or the number of newly identified examples does not meet the requirements, it indicates that the situation of adding difficult examples in the current training mode is relatively serious. Therefore, in order to learn difficult examples more comprehensively, it is determined to perform batch adjustment processing, that is, to adjust the number of batches according to the product of the preset factor and the original batch number.
[0126] The system maintains a sliding window to record the aforementioned metrics for the most recent M (e.g., M=5) consecutive training batches. The decision-making process is as follows: Level 1: Rapid Expansion Judgment under Aggressive Strategy. Condition: If the currently effective sampling strategy is "Second Ratio" (aggressive strategy). Decision Logic: An aggressive strategy inherently implies an active increase in training difficulty. At this point, if any batch of "new input ratio" or "new identification number" exceeds its respective warning threshold (e.g., new input ratio greater than 0.5% or new identification number greater than 0.3%) within the monitoring window, it indicates that the growth of difficult examples has been very rapid, and the model faces intensive and continuous high-intensity challenges.
[0127] Adjustment Action: To prevent the model from facing a "big storm" within a "small pond" (original batch size), the system immediately decides to increase the training capacity, adjusting the number of batches for subsequent training to: the original planned batch size × a preset factor (e.g., factor = 1.5). Increasing the batch size helps to obtain a more stable gradient direction and smooth out oscillations.
[0128] It should be noted that the preset factor is greater than 1.
[0129] Furthermore, if the sampling strategy does not belong to the second proportion, then the identification adjustment method for the batch is determined based on the proportion of new inputs of difficult-to-handle cases and the number of new identifications in different batches.
[0130] Furthermore, based on the proportion of newly added difficult-to-handle examples and the number of newly identified examples in different batches, the identification adjustment method for the batch is determined, specifically including: Case 1: When the proportion of new inputs of difficult cases in different batches is greater than the preset input proportion threshold, the number of batches will be adjusted according to the product of the preset factor and the original number of batches. Scenario 1: The intensity of input is generally high. Condition: Within the monitoring window, the "new input ratio" of each batch is greater than a high preset input ratio threshold (such as 0.5%). This indicates that although the strategy is nominally "conservative", due to the accumulation in the early stage or the characteristics of the sample, the intensity of new difficult cases input in each batch is still very high. Adjustment action: Trigger batch expansion as well, and execute according to the original batch number × preset factor (1.5).
[0131] Scenario 2: When the proportion of newly added difficult examples in different batches is not greater than the preset input ratio threshold, the number of newly identified difficult examples in different batches is obtained. Based on the number of newly identified difficult examples in different batches, the ratio of the number of newly identified examples to the number in the existing training samples of the model is determined and used as the new identification ratio. When the new identification ratio in different batches is less than the preset ratio threshold, the batch adjustment process is not performed for the time being. Specifically, to identify sluggish growth in existing samples, the condition is that within the window, not all batches have an excessively high input ratio. In this case, calculate the "new recognition ratio" for each batch = number of new recognitions in this batch / total number of training samples in the model. If this ratio for all batches is less than a preset threshold (e.g., 0.3%), it indicates that the growth rate of newly emerging "hard nuts to crack" in the model's field of vision is very slow, and the learning pressure has not increased significantly. The adjustment action is to temporarily not adjust the batch size and maintain the current pace.
[0132] Scenario 3: When the proportion of newly identified products in different batches is not less than the preset proportion threshold, if the ratio of the number of batches not less than the preset proportion threshold to the number of existing input batches is greater than the preset batch proportion threshold, the number of batches will be adjusted according to the product of the preset factor and the number of original batches. In other cases, if the proportion of newly identified products in the most recent preset number of batches is not less than the preset proportion threshold, the number of batches will be adjusted according to the product of the second preset factor and the number of original batches.
[0133] To identify sudden increases in existing cases or localized outbreaks, the condition is: within a window, at least one batch has a "newly identified proportion" of not less than a threshold (≥0.3%). This indicates the occurrence of a localized outbreak of difficult-to-handle cases.
[0134] Further decision: Calculate the proportion of "high-growth batches" to the total number of batches in the window. If this proportion > the preset batch proportion threshold (e.g., 60%), it means that the increase in difficult-to-recover examples is not accidental, but a trend. The system determines that training stability needs to be significantly enhanced by expanding the batch size by a preset factor (e.g., factor = 1.5) to a larger extent.
[0135] If this proportion is ≤ 60%, but the recognition proportions of the most recent L consecutive batches (e.g., L=3) are all ≥ the threshold: this indicates that although it is not a global trend, the model is continuously encountering strong new challenge clusters. To be on the safe side, a moderate expansion is performed, adjusted by the original batch number × the preset factor (1.2).
[0136] Furthermore, the preset factor is greater than the second preset factor.
[0137] Assuming the model is being trained using a conservative strategy, the original plan was to run 100 batches in the next cycle. Monitoring the last 5 batches yielded the following: The proportion of new inputs in batches 1-5: [0.1%, 0.3%, 0.2%, 0.2%, 0.1%] (all <0.5%, not meeting condition 1) The percentage of newly identified individuals in batches 1-5 is: [0.3%, 0.8%, 0.4%, 1.2%, 0.5%] (all ≥0.3%) Adjustment: Using a preset factor (1.5), the number of batches in the next training cycle is adjusted to 100 × 1.5 = 150. This means that although the number of new difficult examples introduced in each batch is not large (a conservative strategy), by running more batches, the model can gradually digest these newly emerging difficult samples in a richer data mixture and a smoother optimization process, achieving a robust performance improvement.
[0138] In summary, the batch adjustment mechanism in this embodiment forms a perfect synergistic closed loop with the aforementioned sampling strategy: the sampling strategy controls how many units each batch consumes (fine-grained rhythm), while the batch adjustment mechanism determines "how long to consume" or "the scale of the next meal" based on the digestion situation (coarse-grained planning). Together, they achieve global, adaptive optimization of the clothing image retrieval model training process, ensuring that the model can both boldly challenge high-difficulty boundaries and maintain a steady learning pace in an increasingly complex feature space.
[0139] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0140] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0141] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A duplicate detection system based on deep learning networks, characterized in that, Specifically, it includes: The preprocessing module transforms the raw unstructured image into a standardized multi-view image; The embedding module constructs training images using multi-view images. In the initial or pre-training SigLIP embedding space, it retrieves the target number of negative examples with the highest similarity but belonging to different products or labels for each training image, maintains a nearest neighbor cache or temporary index, and dynamically selects difficult negative examples within the batch for training according to the model embedding. The attribute text generated by each cleaned training image is converted into a text feature vector by the SigLIP text encoder. Based on the contrast loss function covering three dimensions of image-image, image-text, and text-image, as well as the weight of difficult negative examples, a total loss function is constructed in each training epoch to update the network parameters. After training, the fine-tuned visual encoder and text encoder are output. The retrieval and fusion module uses the SigLIP visual encoder, which has been fine-tuned by the embedding module, to extract feature vectors from the three views of all images in the image library. In the indexing stage, a master index is built only for the original view of each image, generating the three views of the user's retrieved image. The feature vectors of the three views are sent to the vector database simultaneously, and the top-N nearest neighbors are queried for each view. The top-N score list returned for each view is converted into an ordered ranking and a standardized score. Based on the cosine similarity between the white background main image feature and the main cropped image feature of the retrieved image, the weight coefficient of each view feature is determined. The comprehensive similarity coefficient is obtained by summing and deduplicating the product of the weight coefficient and the standardized score. Based on the comprehensive similarity coefficient, the retrieval list of the same style as the user's retrieved image is obtained in descending order. The three views include a white background main image, an original image, and a cropped main image; Training is performed by dynamically selecting difficult negative examples within a batch based on model embedding, specifically including: During the first iteration step interval, random negative examples are used, and negative examples are sampled from the offline library with a 10% probability. Within the second iteration step interval, based on the sampling processing strategy and batch size adjustment method, the input batch of negative samples and the newly input negative samples in each input batch are determined. The features of all training samples are updated using the current model. When inputting negative samples in each input batch, the difficult negative samples are recalculated. The input ratio of difficult negative samples is gradually increased to the target input ratio threshold, and the similarity threshold is gradually increased to increase the number of input negative samples with higher similarity. The sampling strategy is based on the proportion of negative examples in the training samples, the distribution concentration of target training images in the current batch, and the clustering degree of overlapping training images in the target training images. It determines the proportion of new difficult negative examples to be introduced in the next batch using a first increment or a second increment, where the second increment is greater than the first increment. The target training images are training images whose similarity falls within a preset target similarity interval, and the overlapping training images are training images in the target training images whose similarity to each other is greater than a preset similarity threshold. The batch size adjustment method involves acquiring the proportion of newly added hard-to-handle examples and the proportion of newly identified examples in multiple consecutive training batches in real time; and determining whether to trigger batch size adjustment and the adjustment range based on the type of the currently effective sampling processing strategy, the comparison results of the newly added proportion and the newly identified example with their respective preset thresholds, and the proportion of batches with the newly identified example exceeding the threshold in the multiple consecutive training batches. The hard negative examples are negative examples whose similarity is greater than the threshold T_hard. In the third iteration step interval, the target proportion of difficult negative examples is maintained for sampling, and the most difficult negative examples are introduced, while gradient clipping is added.
2. The same-item detection system based on deep learning networks as described in claim 1, characterized in that, Converting into standardized multi-view images specifically includes: First, invalid URLs are filtered out. Then, the original unstructured images in the URLs are decoded, and the original unstructured images with alpha channels are uniformly converted into JPG format with a white background. Finally, the pre-trained subject detection model is called to obtain the subject bounding box and generate two derived views: a white background subject image and a subject cropped image; the original image and the two derived views together constitute a multi-view image.
3. The same-item detection system based on deep learning networks as described in claim 2, characterized in that, URL filtering and downloading specifically includes: system verification of URL validity and downloading image data.
4. The same-item detection system based on deep learning networks as described in claim 2, characterized in that, The subject detection model is constructed using the YOLOv8 model.
5. The same-item detection system based on deep learning networks as described in claim 1, characterized in that, The first iteration step range is between 0 and 30% of the maximum iteration step during the current model's training process; the second iteration step range is between 30% and 80% of the maximum iteration step during the current model's training process; and the third iteration step range is between 80% and the maximum iteration step.
6. A method for detecting identical items based on deep learning networks, applied to the system for detecting identical items based on deep learning networks as described in any one of claims 1-5, characterized in that, Specifically, it includes: Based on the products corresponding to the training images, the listing data of the products corresponding to different training images is determined. Three demand types are divided according to the listing volume. The target quantity K is determined according to the proportion and / or weight of each demand type image in the same training batch. For each training image, K negative sample samples with the highest similarity and belonging to different products are retrieved in the pre-training embedding space. The similarity interval of the negative sample samples of the training image is determined according to the similarity value of the retrieved K negative sample samples. Based on the inclusion relationship between the similarity interval and the target similarity interval, and combined with the composition ratio of the negative sample samples, the sampling processing strategy of the negative sample is determined. The sampling strategy is as follows: based on the proportion of negative examples in the training samples, the distribution concentration of target training images in the current batch, and the clustering degree of overlapping training images in the target training images, the proportion of new difficult negative examples to be introduced in the next batch is determined by either a first increment or a second increment, wherein the second increment is greater than the first increment; wherein, the target training images are training images whose similarity falls within a preset target similarity interval, and the overlapping training images are training images in the target training images whose similarity to each other is greater than a preset similarity threshold; Based on the aforementioned sampling strategy, training is performed by embedding dynamically selected difficult examples within each batch in batches. The batch size adjustment method is determined based on the usage data of difficult examples in different batches and the sampling strategy. The batch size adjustment method involves: real-time acquisition of the proportion of newly added difficult examples and the proportion of newly identified examples in multiple consecutive training batches; and determining whether to trigger batch size adjustment and the adjustment magnitude based on the type of the currently effective sampling strategy, the comparison results of the newly added and newly identified examples with their respective preset thresholds, and the proportion of batches with newly identified examples exceeding the threshold in the multiple consecutive training batches.
7. The method for detecting identical items based on deep learning networks as described in claim 6, characterized in that, The product listing data corresponding to the training images is determined based on the number of listings of the product type on different platforms.
Citation Information
Patent Citations
Clothing retrieval technology based on deep metric learning
CN111914109A
Feature extraction method and system based on difficult sample mining and multi-granularity division
CN118823371A