Foreign trade product description deduplication batch processing method based on single-pass clustering
Through the batch processing method of foreign trade product data based on single-pass clustering, the problems of high computational overhead and insufficient parallel processing capabilities in traditional methods are solved, efficient deduplication of foreign trade product descriptions is achieved, and the real-time performance and processing capabilities of the system are improved.
Patent Information
- Application Number
- CN202510872029.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional foreign trade product description deduplication methods based on single-pass clustering have high computational overhead in large-scale data processing and are difficult to utilize the parallel processing capabilities of modern computing hardware, resulting in a decrease in system response speed, affecting real-time performance and processing efficiency.
A batch processing method for foreign trade product data based on single-pass clustering is adopted. By setting the similarity threshold of product description feature vectors, a full cluster center vector library is created, which is divided into potential new cluster center groups and attributable groups. Single-pass clustering and batch processing are performed, and the search efficiency is improved through the approximate nearest neighbor search index structure.
It improves the efficiency of deduplication of foreign trade product descriptions, reduces computational complexity, supports efficient processing of large-scale streaming data, and ensures the real-time performance and processing capabilities of the system.
Smart Images

Figure CN120763137A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular to a method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering. Background Art
[0002] In the foreign trade sector, deduplication of product descriptions is crucial for ensuring data quality and improving business efficiency. While traditional deduplication methods based on single-pass clustering are theoretically suitable for processing continuously updated data streams, they face numerous efficiency challenges in practice. These methods require a full-dimensional similarity comparison of each new product description with all existing clusters. As data scale increases, computational overhead increases significantly, leading to a noticeable decrease in system responsiveness.
[0003] In terms of hardware resource utilization, traditional serial processing models struggle to fully exploit the parallel processing capabilities of modern computing hardware. Because algorithms force data processing sequentially, even when faced with batches of similar descriptions, they cannot effectively leverage the parallel computing advantages of multi-core processors or GPUs. Furthermore, frequent random memory accesses lead to inefficient caches, further reducing computing performance. This efficiency bottleneck is particularly pronounced in business scenarios requiring real-time processing, severely impacting overall system throughput.
[0004] Faced with the growing data processing demands of cross-border trade, traditional methods are no longer able to meet the real-time and efficient processing requirements of the business. Insufficient system processing capacity is particularly prominent during peak traffic periods, not only affecting deduplication effectiveness but also causing delays in subsequent data analysis and business decision-making. Summary of the Invention
[0005] In order to improve the efficiency of deduplication of foreign trade product descriptions, this application provides a method for batch processing of foreign trade product data based on single-pass clustering.
[0006] This application provides a method for batch processing of foreign trade product data based on single-pass clustering, which adopts the following technical solutions: A method for batch processing of foreign trade product data based on single-pass clustering includes the following steps: Set the similarity threshold of product description feature vectors and create a full cluster center vector library for storing product description feature vectors; Extracting the feature vector corresponding to the product description in the first foreign trade detailed order as the cluster center and storing it in the full cluster center vector library; Obtain descriptions of newly added batches of foreign trade products and feature vectors corresponding to the newly added batches of foreign trade products, and calculate the similarity between the feature vectors of the newly added batches of product descriptions and the existing cluster center vectors in the full cluster center vector library; Grouping the newly added batches of foreign trade product descriptions according to the similarity and the product description feature similarity threshold into potential new cluster center groups and attributable groups; Performing single-pass clustering on each potential new cluster center in the potential new cluster center group according to the original order of appearance to generate a temporary new cluster center vector library; Merging the temporary new cluster center vector library into the full cluster center vector library; Batch processing the attributable groups, and batch associating product descriptions of the attributable groups to the nearest cluster centers; After clustering is completed, a deduplication threshold is set according to business needs, and product descriptions above the deduplication threshold in each cluster obtained by clustering are deduplicated.
[0007] By adopting this technical solution, we can effectively cluster newly added batches of foreign trade product descriptions by setting a similarity threshold for product description feature vectors and creating a comprehensive database of cluster center vectors. We then divide product descriptions into potential new cluster center groups and assignable groups, then perform both single-pass clustering and batch processing on each description. Finally, we apply a deduplication threshold to remove duplicate product descriptions. This method can efficiently process large amounts of foreign trade product description data and improve the efficiency of deduplication.
[0008] Preferably, the full cluster center vector library is configured with an index structure that supports approximate nearest neighbor search, and uses a memory database or a distributed database to store cluster center information.
[0009] By implementing this technical solution, we can significantly improve search efficiency by configuring an index structure that supports approximate nearest neighbor search within the full cluster center vector library and using an in-memory or distributed database to store cluster center information. Approximate nearest neighbor search can quickly find the feature vector most similar to the target vector in large-scale data, reducing computational effort and search time.
[0010] Preferably, the index structure supporting fast approximate nearest neighbor search specifically calculates the cosine distance metric between product description feature vectors to obtain the similarity between the product description feature vector of the newly added batch of foreign trade details and the existing cluster center vector in the full cluster center vector library.
[0011] By employing this technical solution, cosine distance effectively measures the directional similarity between vectors and is highly adaptable to data such as text feature vectors. Using the cosine distance metric, we can more accurately assess the similarity between the feature vectors describing newly added batches of foreign trade products and the existing cluster center vectors in the full database, enabling more accurate grouping and subsequent clustering.
[0012] Preferably, the newly added batches of foreign trade product descriptions are grouped according to the similarity and the product description feature similarity threshold, and divided into potential new cluster center groups and attributable groups, including the following: If the similarity of the foreign trade product description is greater than or equal to the product description feature vector similarity threshold, then the foreign trade product description with a similarity higher than the product description feature vector similarity threshold is determined to belong to the group that can be assigned; If the similarity of the foreign trade product description is less than the similarity threshold of the product description feature vector, it is determined that the foreign trade product description with the similarity lower than the similarity threshold of the product description feature vector belongs to the potential new cluster center group.
[0013] By employing the above technical solution, we refine the specific rules for grouping based on similarity and product description feature similarity thresholds. This clear grouping rule makes the grouping process clearer, more specific, and easier to implement and operate. It accurately categorizes newly added batches of foreign trade product descriptions into potential new cluster centers and assignable groups, ensuring that each product description is appropriately categorized and providing accurate input data for subsequent single-pass clustering and batch processing.
[0014] Preferably, performing strict single-pass clustering on each potential new cluster center in the potential new cluster center group in the original order of appearance to generate a temporary new cluster center vector library comprises the following steps: Create an empty temporary new cluster center vector library; Performing single-pass clustering on each potential new cluster center in the potential new cluster center group according to the original order of appearance; Taking the feature vector corresponding to the first product description in the potential new cluster center group as the first vector of the temporary cluster center vector library, and storing it in the temporary new cluster center vector library; Calculate the similarity between the product description feature vector in each subsequent potential new cluster center group and all existing temporary cluster center vectors in the temporary new cluster center vector library; if the similarity of the calculated product description feature vector is less than the product description feature vector similarity threshold, then the calculated product description feature vector is used as the new temporary cluster center vector and inserted into the temporary new cluster center vector library; if the similarity of the calculated product description feature vector is greater than or equal to the product description feature vector similarity threshold, then the calculated product description feature vector is assigned to the nearest neighbor temporary cluster center.
[0015] By adopting this technical solution, a strict single-pass clustering method ensures that each potential new cluster center is accurately processed, avoiding duplication or omission. Based on the similarity threshold, it determines whether to use the new potential new cluster representative point as a new temporary cluster center vector or assign it to the nearest neighbor temporary cluster center. This effectively generates an accurate temporary new cluster center vector library. This provides a high-quality data foundation for the subsequent merging of the temporary new cluster center vector library into the full cluster center vector library.
[0016] Preferably, merging the temporary new cluster center vector library into the full cluster center vector library comprises the following steps: Inserting the temporary new cluster center vector library into the full cluster center vector library in batches; Record the number mapping relationship before and after the temporary cluster center is inserted into the full cluster center vector library.
[0017] By adopting the above technical solution, the specific steps of merging the temporary new cluster center vector library into the full cluster center vector library are described, including batch insertion and record number mapping. This merging method can efficiently integrate the newly generated temporary cluster center vector library into the full cluster center vector library, avoiding the tedious operation of inserting one by one and improving the efficiency of data merging. At the same time, the record number mapping relationship can ensure that the new number of the temporary cluster center in the full cluster center vector library can be accurately found in the subsequent processing process, providing accurate cluster center information for subsequent batch processing and deduplication operations.
[0018] Preferably, batch processing of the attributable groups and batch associating product descriptions of the attributable groups to the nearest cluster center comprises the following steps: Calculate and obtain the first maximum similarity between each product description in the group that can be assigned and the temporary cluster center vector library, Calculating a second maximum similarity between each product description in the assignable group and the existing cluster center vectors in the unmerged full cluster center vector library; When the first maximum similarity is less than or equal to the second maximum similarity, the most similar cluster center of the product descriptions in the group that can be assigned is located in the full cluster center vector library. When the first maximum similarity is greater than the second maximum similarity alone, the most similar cluster center of the product description in the group that can be assigned is located at the temporary cluster center; and the new number of the temporary cluster center after being inserted into the full cluster center vector library is obtained. The new number is obtained through the recorded mapping relationship between the temporary cluster center and the full cluster center, and the product description is associated with the new number of the mapped full library cluster center.
[0019] By employing this technical solution, we calculate the maximum similarity with cluster center vectors in both the temporary cluster center vector library and the unmerged full cluster center vector library, and determine the most similar cluster center location based on the similarity. This ensures that each product description that can be grouped is accurately associated with the most appropriate cluster center. Furthermore, by obtaining the new number of the temporary cluster center in the full cluster center vector library through the recorded mapping relationship, the accuracy of the association operation is guaranteed.
[0020] Preferably, after clustering is completed, the deduplication threshold is higher than the product description feature vector threshold.
[0021] By adopting the above technical solution, due to the higher deduplication threshold, product descriptions are considered duplicates only when the similarity between them is very high. This can avoid misjudgment caused by setting the similarity threshold too low.
[0022] In summary, this application includes at least one of the following beneficial technical effects: 1. Through a single-pass clustering and batch merging mechanism, the system can efficiently process newly added batches of foreign trade product data, avoiding the resource consumption of full recalculation required by traditional methods. The use of an approximate nearest neighbor search (ANN) index structure reduces processing time complexity and improves processing efficiency. 2. Using a fixed similarity threshold for product description feature vectors and a dynamically updated full-database cluster center vector library, we strictly adhere to the threshold when introducing new clusters to prevent semantic drift of cluster centers due to incremental updates. 3. Approximate nearest neighbor search indexing technology based on a vector database reduces the complexity of similarity calculations, ensuring that the speed of comparing new foreign trade invoices with cluster centers does not decrease significantly as the number of centers increases, thereby supporting the efficient processing of large-scale streaming data. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is an overall flow chart of an embodiment of the present application; Figure 2 This is a flowchart of step S4 in the embodiment of the present application. DETAILED DESCRIPTION
[0024] The following combination Figure 1-Figure 2 This application is described in further detail.
[0025] The embodiment of the present application discloses a method for batch processing of foreign trade product data based on single-pass clustering.
[0026] Reference Figure 1 A method for batch processing of foreign trade product data based on single-pass clustering includes the following steps: S0: Set the similarity threshold of the product description feature vector and create a full cluster center vector library for storing the product description feature vector; S1: Extract the product description feature vector from the first foreign trade order as the cluster center and store it in the full cluster center vector library; S2: Obtain the feature vectors corresponding to the foreign trade details of the newly added batch and the foreign trade products of the newly added batch, and calculate the similarity between the feature vectors of the description of the newly added batch products and the existing cluster center vectors in the full cluster center vector library; S3: Structural grouping of newly added batches of foreign trade details is performed based on similarity and product description feature vector similarity thresholds, into potential new cluster center groups and attributable groups; S4: Perform single-pass clustering on each potential new cluster center in the potential new cluster center group according to the original order of appearance to generate a temporary new cluster center vector library; S5: Merge the temporary new cluster center vector library into the full cluster center vector library; S6: batch processing the attributable groups, batch associating the product descriptions of the attributable groups to the nearest cluster centers; S7: After clustering is completed, a deduplication threshold is set according to business needs, and duplicate product descriptions above the deduplication threshold are removed from each cluster obtained by clustering; S8: After the clustering of the newly added batch of foreign trade detailed orders is completed, the process jumps to step S2 and waits for the entry of the foreign trade detailed orders of other newly added batches.
[0027] Specifically, step S0: setting a similarity threshold for product description feature vectors and creating a full cluster center vector library for storing product description feature vectors, including the following: By setting the similarity threshold of the product description feature vector, it is possible to determine whether the product description in the foreign trade detailed order belongs to an existing cluster or needs to create a new cluster.
[0028] The full cluster center vector library is used to store cluster center information. It can adopt various storage architectures, such as an in-memory database, which offers fast read and write speeds and improves the efficiency of subsequent similarity calculations; or a distributed database, suitable for storing and managing large-scale data. When storing data, it is also necessary to select an appropriate vector indexing method based on the application scenario, such as a KD tree index, which can quickly find similar vectors, or an inverted file index, which is suitable for searching high-dimensional vectors. In this embodiment, the full cluster center vector library supports approximate nearest neighbor search (ANN), facilitating subsequent steps.
[0029] Step S1: Extract the product description feature vector from the first foreign trade order as the cluster center and store it in the full cluster center vector library, including the following: Specialized machine learning models can be used to extract feature vectors for product descriptions in foreign trade invoices. For example, multiplication neural networks are commonly used for feature extraction from image-based foreign trade invoices. They can extract representative feature vectors of product descriptions from images. Feature extraction libraries, such as the Python Scikit-learn library, can also be used, which provides a variety of feature extraction methods.
[0030] Step S2: Obtain the feature vectors corresponding to the foreign trade details of the newly added batch and the foreign trade products of the newly added batch, and calculate the similarity between the feature vectors of the description of the new batch of products and the existing cluster center vectors in the full cluster center vector library, including the following: When a new batch of product features from a foreign trade order (for example, 256 product descriptions) enters the system, the algorithm efficiently calculates the distance between each product description feature vector and all cluster centers using an approximate nearest neighbor search (ANN) indexing structure (such as HNSW, IVF, and LSH) within the vector database. This ensures that computational complexity does not scale linearly with the number of cluster centers, maintaining a near-constant query speed. Similarity can be calculated using distance metrics such as cosine distance, combined with ANN technology to accelerate the retrieval of large cluster centers. Cosine distance measures similarity by calculating the cosine of the angle between two vectors, prioritizing vector direction rather than length. These distance metrics accurately reflect the degree of similarity between foreign trade orders and cluster centers.
[0031] Step S3: Structural grouping of newly added batches of foreign trade details is performed based on similarity and product description feature vector similarity thresholds, into potential new cluster center groups and attributable groups, including the following: Compare the description feature vectors of multiple newly added batches of products with the preset product description feature vector similarity threshold. If the similarity of the product description feature vectors is higher than the preset product description feature vector similarity threshold, then these product description feature vectors belong to the group that can be assigned, indicating that they are highly consistent with the existing cluster structure and can be directly classified into the corresponding cluster without significantly affecting the cluster structure; if the similarity of the foreign trade detailed orders is lower than the preset product description feature vector similarity threshold, then these foreign trade detailed order product descriptions belong to the potential new cluster center group. These foreign trade detailed order product descriptions are quite different from the existing clusters and may represent new data distribution patterns or anomalies.
[0032] In this embodiment, the preset product description feature vector similarity threshold is 0.8. Among the 256 product description feature vectors, the similarity of 200 product description feature vectors is higher than the preset product description feature vector similarity threshold, and 56 product description feature vectors are lower than the preset product description feature vector similarity threshold.
[0033] refer to Figure 2Step S4: Single-pass clustering for each potential new cluster center in the potential new cluster center group in original order to generate a temporary new cluster center vector library, further comprising the following sub-steps: S41: Create an empty temporary new cluster center vector library, which is a separate storage space specially used for temporarily storing new cluster centers discovered in the current batch processing. The temporary new cluster center vector library is completely isolated from the main cluster library, ensuring that the evaluation process of new clusters is not disturbed by the existing clustering structure, while avoiding the performance overhead caused by frequent operations on the main library. The temporary library uses the same data structure as the main library, supporting fast similarity query and dynamic insertion operation.
[0034] S42: Single-pass clustering for each potential new cluster center in the potential new cluster center group in original order.
[0035] S43: The product description feature vector in the first potential new cluster center group is taken as the temporary cluster center vector and stored in the temporary new cluster center vector library; since the temporary new cluster center vector library is initially empty, the product description feature vector of the first potential new cluster center group in the current batch automatically becomes the temporary cluster center, and the product description feature vector of the foreign trade detailed list is directly inserted into the temporary new cluster center vector library.
[0036] S44: Calculate the similarity of the product description feature vector in each subsequent potential new cluster center group with all existing temporary cluster center vectors in the temporary new cluster center vector library, and compare the calculated similarity with the product description feature vector similarity threshold to obtain a new temporary cluster center vector or belong to the nearest neighbor temporary cluster center.
[0037] The calculation method of similarity is consistent with the main clustering process, ensuring uniform evaluation criteria. If the calculated similarity of the product description feature vector is less than the product description feature vector similarity threshold, the calculated product description feature vector is taken as a new temporary cluster center vector and inserted into the temporary new cluster center vector library; if the calculated similarity of the product description feature vector is greater than or equal to the product description feature vector similarity threshold, the calculated product description feature vector is attributed to the nearest neighbor temporary cluster center.
[0038] S45: When all product description feature vectors in the potential new cluster center group are processed, the process exits normally; otherwise, return to S43 to continue processing the next foreign trade detailed list. The size change of the temporary new cluster center vector library will be monitored in real time during the loop process, and an early warning will be triggered when an abnormal increase in the number of temporary clusters is found, which may indicate that the original threshold setting is unreasonable or the data distribution has changed abruptly. After completing all processing, the new cluster center vector library stores all new cluster centers discovered in the current batch.
[0039] Reference Figure 1, step S5: merge the temporary new cluster center vector library into the full cluster center vector library. The system will batch insert all cluster center vectors in the temporary new cluster center vector library into the full cluster center vector library. This batch operation uses an efficient data writing mechanism (such as batch insertion of the database or expansion of the memory list) to optimize performance and reduce the overhead of inserting a single record. Subsequently, the system will record the number mapping relationship of the temporary cluster center before and after the merger, such as mapping the temporary number to the globally unique number in the full library. This mapping relationship is stored in the form of a dictionary or mapping table to ensure that subsequent operations can correctly reference these clusters. Finally, the system will update the associated data according to the recorded mapping relationship, such as adjusting the product description in the "attributable group" that originally pointed to the temporary number to point to the official number in the full library, thereby ensuring data consistency. The entire process must meet atomicity requirements to avoid data inconsistency caused by partial insertion. At the same time, the persistent storage of the mapping relationship supports subsequent backtracking or incremental update requirements.
[0040] Through this series of steps, the full cluster center vector library can be dynamically expanded to adapt to changes in new data while maintaining efficient query and maintainability.
[0041] Step S6: batch processing the attributable groups, batch associating the product descriptions of the attributable groups to the nearest cluster center, including the following steps: Calculate the first maximum similarity between each product description in the attributable group and the temporary cluster center library, and obtain the second maximum similarity between each product description in the attributable group and the existing cluster center vector in the unmerged full cluster center library, compare the first maximum similarity and the second maximum similarity and select the larger similarity as the basis.
[0042] When the first maximum similarity is individually less than the second maximum similarity, the most similar cluster center of the product description in the attributable group is located in the full cluster center library, so this product description is directly associated with the corresponding cluster in the full cluster center library. When the first maximum similarity is individually greater than the second maximum similarity, the most similar cluster center of the product description in the attributable group is located in the temporary cluster center; at this time, the system will use the recorded number mapping relationship to query the new number assigned to the most similar cluster center after it is merged into the full library, and associate the product description with the mapped new number of the full library cluster center.
[0043] This mechanism not only effectively solves the problem of matching group-attributable product descriptions with the nearest cluster center, but also ensures data consistency and traceability through number mapping, while balancing processing efficiency and accuracy. The entire process relies on the previously established mapping relationship between temporary and full cluster centers, allowing dynamic incremental clustering results to be seamlessly integrated into the global cluster system.
[0044] Preferably, only the cluster identification information or key summary information of the product descriptions that can be assigned to the group is stored, rather than the complete original description data, so as to reduce data storage redundancy.
[0045] S7: After clustering is completed, a deduplication threshold is set according to business requirements, and duplicate product descriptions above the deduplication threshold are removed from each cluster obtained by clustering, including the following: After clustering is complete, the system performs refined deduplication based on actual business needs. This process first requires setting a deduplication threshold, such as 0.95, which is typically higher than the product description feature vector similarity threshold used during initial clustering. Setting a higher deduplication threshold allows us to remove only nearly identical descriptions while retaining a reasonable level of product diversity.
[0046] In specific implementation, the system will traverse all product descriptions in each cluster and calculate their similarity with the center vector of the cluster to which they belong. When the similarity between a product description and the cluster center reaches or exceeds the set deduplication threshold, the product description will be judged to be substantially duplicated with the cluster center description and will be marked as a candidate for deduplication. For example, in a cluster containing 10 product descriptions, there may be 3 descriptions with a similarity of 0.96 to the center vector. These descriptions will be identified and removed by the system, while descriptions with a similarity between 0.8-0.94 will be retained as valid variants.
[0047] The advantages of this hierarchical threshold design are: it ensures inclusiveness in the initial clustering, grouping related products into the same category; it also eliminates highly repetitive descriptions through more rigorous post-processing, ultimately improving data quality while maintaining product catalog diversity. The entire process is based entirely on a quantitative assessment of vector similarity, ensuring objectivity and consistency in deduplication decisions.
[0048] Step S8: After the clustering of the newly added batch of foreign trade details is completed, the system jumps to step S2 and waits for the entry of other newly added batches of foreign trade details. After the processing of the current batch of foreign trade details is completed, the system will perform a series of cleanup and preparation tasks. This includes releasing temporary computing resources, persisting the latest cluster structure, updating monitoring indicators, etc. The process then returns to step S2 and waits for the input of the next batch of data. This loop design allows the algorithm to continuously process streaming data while maintaining the evolving clustering structure. Before jumping, the system will check resource usage and make dynamic adjustments when necessary, such as expanding the vector index capacity or adjusting the processing scale of subsequent batches to ensure long-term stability. Through this incremental processing mechanism, the entire process achieves continuous learning and adaptation of data distribution, balancing the requirements of computing efficiency and clustering quality.
[0049] The implementation principle of the method for batch processing of foreign trade product data based on single-pass clustering in the embodiment of the present application is as follows: in terms of computational efficiency, through the innovative batch processing mechanism, the system can process a large amount of foreign trade detailed order data at the same time. Compared with the traditional point-by-point processing method, the computing speed is significantly improved. At the same time, through the optimized clustering process design, it is ensured that the processing results have the same accuracy as the traditional method. In terms of algorithm stability, a dual guarantee mechanism is adopted: on the one hand, a strict product description feature vector similarity threshold is set as the judgment standard, and on the other hand, a dynamically updated central library maintenance system is established. These two measures work together to effectively prevent semantic deviation problems that may occur in the clustering process. In terms of system performance, by integrating advanced approximate nearest neighbor search technology, an efficient vector index structure is constructed, so that the system can still maintain a stable response speed when processing massive data, and is not affected by the expansion of data scale.
[0050] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering, characterized by: The following steps are involved: Set the similarity threshold of product description feature vectors and create a full cluster center vector library for storing product description feature vectors; Extracting the feature vector corresponding to the product description in the first foreign trade detailed order as the cluster center and storing it in the full cluster center vector library; Obtain descriptions of newly added batches of foreign trade products and feature vectors corresponding to the descriptions of the newly added batches of foreign trade products, and calculate the similarity between the feature vectors of the descriptions of the newly added batches of products and the existing cluster center vectors in the full cluster center vector library; Grouping the newly added batches of foreign trade product descriptions according to the similarity and the product description feature similarity threshold into potential new cluster center groups and attributable groups; Performing single-pass clustering on each potential new cluster center in the potential new cluster center group according to the original order of appearance to generate a temporary new cluster center vector library; Merging the temporary new cluster center vector library into the full cluster center vector library; Batch processing the attributable groups, batch associating the product descriptions of the attributable groups to the nearest cluster center; After clustering is completed, a deduplication threshold is set according to business needs, and product descriptions above the deduplication threshold in each cluster obtained by clustering are deduplicated.
2. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 1, characterized in that: The full cluster center vector library is configured with an index structure that supports approximate nearest neighbor search, and uses an in-memory database or a distributed database to store cluster center information.
3. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 2, characterized in that: The index structure supporting fast approximate nearest neighbor search specifically calculates the cosine distance metric between product description feature vectors to obtain the similarity between the product description feature vector of the newly added batch of foreign trade details and the existing cluster center vectors in the full cluster center vector library.
4. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 1, characterized in that: The newly added batches of foreign trade product descriptions are grouped according to the similarity and the product description feature similarity threshold into potential new cluster center groups and attributable groups, including the following: If the similarity of the foreign trade product description is greater than or equal to the product description feature vector similarity threshold, then the foreign trade product description with a similarity higher than the product description feature vector similarity threshold is determined to belong to the group that can be assigned; If the similarity of the foreign trade product description is less than the similarity threshold of the product description feature vector, it is determined that the foreign trade product description with the similarity lower than the similarity threshold of the product description feature vector belongs to the potential new cluster center group.
5. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 1, characterized in that: Performing strict single-pass clustering on each potential new cluster center in the potential new cluster center group in the original order of appearance to generate a temporary new cluster center vector library includes the following steps: Create an empty temporary new cluster center vector library; Performing single-pass clustering on each potential new cluster center in the potential new cluster center group according to the original order of appearance; Taking the feature vector corresponding to the first product description in the potential new cluster center group as the first vector of a temporary cluster center vector library, and storing it in the temporary new cluster center vector library; Calculate the similarity between the product description feature vector in each subsequent potential new cluster center group and all existing temporary cluster center vectors in the temporary new cluster center vector library; if the similarity of the calculated product description feature vector is less than the product description feature vector similarity threshold, then the calculated product description feature vector is used as the new temporary cluster center vector and inserted into the temporary new cluster center vector library; if the similarity of the calculated product description feature vector is greater than or equal to the product description feature vector similarity threshold, then the calculated product description feature vector is assigned to the nearest neighbor temporary cluster center.
6. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 3, characterized in that: Merging the temporary new cluster center vector library into the full cluster center vector library includes the following steps: Inserting the temporary new cluster center vector library into the full cluster center vector library in batches; Record the number mapping relationship before and after the temporary cluster center is inserted into the full cluster center vector library.
7. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 6, characterized in that: Batch processing the attributable groups and batch associating product descriptions of the attributable groups to the nearest cluster center includes the following steps: Calculate and obtain the first maximum similarity between each product description in the group that can be assigned and the temporary cluster center library, Calculating and obtaining the second maximum similarity between each product description in the attributable group and the existing cluster center vectors in the unmerged full cluster center library; When the first maximum similarity is less than or equal to the second maximum similarity, the most similar cluster center of the product descriptions in the group that can be assigned is located in the full cluster center vector library. When the first maximum similarity is greater than the second maximum similarity alone, the most similar cluster center of the product description in the group that can be assigned is located at the temporary cluster center; and the new number of the temporary cluster center after being inserted into the full cluster center vector library is obtained. The new number is obtained through the recorded mapping relationship between the temporary cluster center and the full cluster center, and the product description is associated with the new number of the mapped full library cluster center.
8. The method for batch processing of deduplication of foreign trade product descriptions based on single-pass clustering according to claim 1, characterized in that: The deduplication threshold is higher than the product description feature vector threshold.