A keyword mining method and system based on multi-factor and double semantic channels

By employing a keyword mining method based on multi-factor and dual semantic channels, the problems of chaotic data structure and static thesaurus on e-commerce platforms are solved. This enables accurate keyword mining and updating, improves business efficiency and optimizes traffic structure, and avoids duplication and homogenization. It is suitable for apparel companies in cross-border e-commerce.

CN121961701BActive Publication Date: 2026-07-31ZHEJIANG ZIBUYU ELECTRONIC COMMERCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG ZIBUYU ELECTRONIC COMMERCE CO LTD
Filing Date
2026-03-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, when merchants use search term reports from e-commerce platforms, the data structure is chaotic and lacks structured classification, resulting in a mixture of high-traffic terms, long-tail terms, and brand terms, making it difficult to assess the breadth and depth of coverage, and manual screening is inefficient; the term library is static and lacks semantic expansion, making it impossible to actively discover potential new terms, and the keyword construction results for the same type of products are prone to duplication.

Method used

A keyword mining method based on multi-factor and dual semantic channels is adopted. The data reconstruction module divides the products into different combinations to generate a long table structure. The index building module performs vectorization encoding, and the intelligent recommendation module performs matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors. Combined with the MPS multi-factor scoring model and differentiated dynamic damping attenuation technology, the accurate mining and updating of keywords are achieved.

Benefits of technology

It solves the problems of "internal competition" and "homogenized recommendations" for apparel companies when optimizing keywords for a massive number of products. It dynamically selects the most suitable keyword mining range to ensure that each product mines accurate traffic and forms a differentiated positioning, thereby improving the overall recall rate and commercial effectiveness of the system and avoiding keyword duplication and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961701B_ABST
    Figure CN121961701B_ABST
Patent Text Reader

Abstract

This invention provides a keyword mining method and system based on multi-factor and dual semantic channels, belonging to the field of data processing technology. Specifically, it includes: a file acquisition module responsible for intelligently filtering target worksheet sets containing valid data; a slicing module responsible for partitioning the worksheet content to obtain multiple data slices; inputting the header and sample data of each data slice into a large language model workflow to intelligently identify the data's dimension type and structural features; an execution module responsible for generating management strategies and prompt words, calling the large language model to dynamically generate targeted data processing function code, executing the generated function code through a function execution engine, standardizing the data slices, and merging all successfully processed data for export as a standardized spreadsheet, thus improving the efficiency of spreadsheet standardization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, and in particular relates to a keyword mining method and system based on multi-factor and dual semantic channels. Background Technology

[0002] With the development of cross-border e-commerce, accurate keyword targeting is crucial for merchants to acquire traffic. However, current technologies often present the following problems when merchants utilize search term reports (such as ABA reports) provided by e-commerce platforms: Disorganized data structure and lack of classification: Raw data usually exists in the form of "wide tables" and lacks structured classification, resulting in a mixture of high-traffic keywords, long-tail keywords and brand keywords, making it difficult to assess the breadth and depth of coverage, and making manual screening inefficient.

[0003] The thesaurus is static and lacks semantic expansion: existing report analysis is mostly based on static review of historical data (existing thesaurus), which cannot actively discover potential new words that are "semantically related" to high-performing words but do not appear in the report (incremental word expansion).

[0004] To address the aforementioned technical issues, the invention patent application CN202411505725.6, "Method, System, and Computing Device for Optimizing Product Titles," utilizes a large language model based on product information to select optimized traffic keywords from third-party candidate traffic keywords and adds these optimized traffic keywords to the original title, thereby generating a new title. This approach better helps merchants select traffic-driving keywords, more conveniently optimize product titles, and enhance product competitiveness. However, the above technical solution has the following technical problems: When using the same reference product's search and conversion data to construct keywords for the same type of products, the keyword construction results may have a high degree of duplication. Therefore, how to determine the keyword mining strategy based on the reference product data of the products in the combination and the composition data of the same type of products, so as to reduce the risk of duplication from the root, has become an urgent technical problem to be solved.

[0005] Therefore, there is an urgent need for a keyword mining method and system based on multi-factor and dual semantic channels. Summary of the Invention

[0006] To achieve the objectives of this invention, the following technical solution is adopted: Specifically, this application provides a keyword mining system based on multi-factor and dual semantic channels, which includes: The data reconstruction module divides products into different combinations based on their type. Using similar product data and product data within each combination, it determines a keyword mining strategy for the products in that combination. Based on this strategy, it obtains reference products for each product. The reference product data from the e-commerce platform is then converted into a long table structure with "search term-product" as the granularity. This table is then linked to product brand and type metadata. Based on multiple dimensions, a comprehensive matching score is calculated for each "search term-product" pair. The index building module performs vectorized encoding on the search terms and product titles in the long table, respectively, to generate semantic vectors for the search terms and semantic vectors for the product titles, and stores them in the retrieval engine along with the matching scores. The intelligent recommendation module, in response to the user's input query terms, performs parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors through the retrieval engine to obtain recall results. After deduplication of the recall results, they are reordered according to the matching scores to output the final keyword recommendation list. The module also uses the keyword mining and processing method of the products in the combination to determine the update method of the keyword mining and processing results of the products.

[0007] The beneficial effects of this invention are as follows: By using similar product data and product data within a product portfolio to determine keyword mining strategies for that portfolio, apparel companies face the challenges of "internal competition" and "homogenized recommendations" when optimizing massive amounts of product keywords. Through intelligent analysis of the density of similar products and the richness of the internal product line, the most suitable keyword mining reference range is dynamically selected. This ensures that keywords are mined for each apparel product that can capture precise traffic while also creating a differentiated positioning. It avoids the technical problem of overly similar keyword mining results caused by using the same similar products, ultimately achieving optimal overall store traffic structure and maximized sales conversion.

[0008] Through its original MPS multi-factor scoring model, the characteristics of four orthogonal dimensions—conversion efficiency, position confidence, ranking long-tail effect, and market size—are non-linearly coupled, overcoming the one-sidedness of single-indicator evaluation and constructing a quantitative system that more closely approximates real business value.

[0009] By constructing independent vector spaces for both search terms and product titles, and establishing a hybrid retrieval architecture, a leap from "literal matching" to "semantic and intent matching" has been achieved. This enables the discovery of potential keywords that differ in form but are highly semantically related, effectively solving the challenges of cold start and incremental keyword expansion.

[0010] The scoring model innovatively applies differentiated dynamic damping attenuation technology, using strong attenuation for product position and very weak attenuation for search term ranking. This effectively suppresses "traffic hegemony" while ensuring the value of top keywords, accurately discovering high-conversion long-tail keywords that are ignored by traditional methods, and improving the overall recall rate and commercial utility of the system.

[0011] By utilizing keyword mining and processing methods for products in a combination, a method for updating the keyword mining and processing results of products is determined. This ensures the timeliness of the update processing for a small number of "updated products" with weak reference product data foundations and requiring high-frequency monitoring when reference keywords are updated. Furthermore, by combining the overlap between "non-updated products" with a large amount of reference product data and updated products, the reliability of keyword update processing when updated products respond to changes in reference product keywords is determined. The timeliness of the update processing is used to determine the keyword update processing strategy for non-updated products, thereby achieving absolute tilt in resource allocation and maintaining the basic health of the keyword database for the same type of product at extremely low cost.

[0012] Furthermore, the dimensions include product performance factors, product position factors, search term ranking factors, and search term volume factors.

[0013] Secondly, this application provides a keyword mining method based on multi-factor and dual-semantic channels, applied to the aforementioned keyword mining system based on multi-factor and dual-semantic channels, specifically including: S1 divides products into different groups based on their type. Using similar product data and product data in the group, it determines the keyword mining strategy for the products in the group. When the mining strategy is an optimized mining strategy, it proceeds to the next step. S2 determines the reference product of the product based on the analysis results of the similarity between the product image of the product and similar products. Based on the reference product data of the product and the similarity between the reference product and the reference products of the products in the combination of the product, the product in the product is selected for keyword mining processing using all reference products, and it is taken as the target product. Based on the composition data of the target product in the combination and the similarity between the product and the reference products of the target product, the keyword mining processing method of the product is determined. S3 uses the keyword mining and processing method to mine and process the keywords of the product to obtain the keyword mining and processing results, and uses the keyword mining and processing method of the product in the combination to determine the method for updating the keyword mining and processing results of the product.

[0014] Furthermore, the products are divided into different groups, specifically including: Group products of the same type into the same category.

[0015] Furthermore, similar products are those that belong to the same type as the product.

[0016] Furthermore, the method for determining the keyword mining strategy for the products in the combination is as follows: Based on the similar product data of the product, determine the quantity of similar products of the product; Based on the product data in the combination, determine the quantity of products in the combination; Based on the number of similar products in the combination and the number of products in the combination, a keyword mining strategy for the products in the combination is determined.

[0017] Furthermore, the method for determining the keyword mining and processing method for the product is as follows: Based on the composition data of the target products in the combination, determine the proportion of the target products in the combination; Based on the degree of similarity between the product and the reference product of the target product, determine the number of overlaps between the product and the reference product of the target product; Based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product, a keyword mining and processing method for the product is determined.

[0018] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0020] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0021] Figure 1 This is a framework diagram of a keyword mining system based on multi-factor and dual semantic channels; Figure 2 This is a flowchart of a keyword mining method based on multi-factor and dual semantic channels; Figure 3 This is a flowchart illustrating the method for determining the keyword mining strategy for products in a combination; Figure 4 This is a flowchart illustrating the method for determining the target product. Detailed Implementation

[0022] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0023] Example 1 like Figure 1 As shown, this application provides a keyword mining system based on multi-factor and dual semantic channels, specifically including: The data reconstruction module divides products into different combinations based on their type. Using similar product data and product data within each combination, it determines a keyword mining strategy for the products in that combination. Based on this strategy, it obtains reference products for each product. The reference product data from the e-commerce platform is then converted into a long table structure with "search term-product" as the granularity. This table is then linked to product brand and type metadata. Based on multiple dimensions, a comprehensive matching score is calculated for each "search term-product" pair. The index building module performs vectorized encoding on the search terms and product titles in the long table, respectively, to generate semantic vectors for the search terms and semantic vectors for the product titles, and stores them in the retrieval engine along with the matching scores. The intelligent recommendation module, in response to the user's input query terms, performs parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors through the retrieval engine to obtain recall results. After deduplication of the recall results, they are reordered according to the matching scores to output the final keyword recommendation list. The module also uses the keyword mining and processing method of the products in the combination to determine the update method of the keyword mining and processing results of the products.

[0024] The beneficial effects of this invention are as follows: By using similar product data and product data within a product portfolio to determine keyword mining strategies for that portfolio, apparel companies face the challenges of "internal competition" and "homogenized recommendations" when optimizing massive amounts of product keywords. Through intelligent analysis of the density of similar products and the richness of the internal product line, the most suitable keyword mining reference range is dynamically selected. This ensures that keywords are mined for each apparel product that can capture precise traffic while also creating a differentiated positioning. It avoids the technical problem of overly similar keyword mining results caused by using the same similar products, ultimately achieving optimal overall store traffic structure and maximized sales conversion.

[0025] Through its original MPS multi-factor scoring model, the characteristics of four orthogonal dimensions—conversion efficiency, position confidence, ranking long-tail effect, and market size—are non-linearly coupled, overcoming the one-sidedness of single-indicator evaluation and constructing a quantitative system that more closely approximates real business value.

[0026] By constructing independent vector spaces for both search terms and product titles, and establishing a hybrid retrieval architecture, a leap from "literal matching" to "semantic and intent matching" has been achieved. This enables the discovery of potential keywords that differ in form but are highly semantically related, effectively solving the challenges of cold start and incremental keyword expansion.

[0027] The scoring model innovatively applies differentiated dynamic damping attenuation technology, using strong attenuation for product position and very weak attenuation for search term ranking. This effectively suppresses "traffic hegemony" while ensuring the value of top keywords, accurately discovering high-conversion long-tail keywords that are ignored by traditional methods, and improving the overall recall rate and commercial utility of the system.

[0028] By utilizing keyword mining and processing methods for products in a combination, a method for updating the keyword mining and processing results of products is determined. This ensures the timeliness of the update processing for a small number of "updated products" with weak reference product data foundations and requiring high-frequency monitoring when reference keywords are updated. Furthermore, by combining the overlap between "non-updated products" with a large amount of reference product data and updated products, the reliability of keyword update processing when updated products respond to changes in reference product keywords is determined. The timeliness of the update processing is used to determine the keyword update processing strategy for non-updated products, thereby achieving absolute tilt in resource allocation and maintaining the basic health of the keyword database for the same type of product at extremely low cost.

[0029] Furthermore, the dimensions include product performance factors, product position factors, search term ranking factors, and search term volume factors.

[0030] In one possible specific implementation, an original wide table containing the search term, the top N ranked products (top_product_asin_1 to N), and their corresponding click share and conversion share are obtained from an e-commerce platform. Through a traversal transformation algorithm, each row of the wide table data is broken down into N independent records, forming a long table with (search term, product ASIN) as the primary key. During this process, each record is clearly associated with product brand (top_brand) and category (top_category) information, resolving the ambiguity in the original data.

[0031] Click share refers to the percentage of clicks received by a specific product (usually your product or its competitors) in the search results page after a consumer searches for a specific search term within a specific statistical period and e-commerce platform (such as the Amazon Brand Analytics reporting period), out of the total number of clicks received by all products under that search term. It directly measures the product's exposure conversion attractiveness or click competitiveness under a specific search term.

[0032] Conversion share refers to the percentage of actual purchases (or order conversions) achieved by a specific product through that search term within the same search term and statistical period, out of the total number of purchases achieved by all products under that search term. It measures the product's ultimate sales monetization ability or conversion efficiency under a specific search term.

[0033] This metric is designed for deeper quantification from the perspective of "commercial value realization." Clicks are important, but traffic that converts clicks into actual purchases is high-value traffic. Conversion share reveals a product's ability to meet users' genuine purchase intent in a specific search context and successfully complete a transaction in a competitive environment. It is directly related to sales revenue and is a key indicator for measuring the purity of keyword commercial value. Similar to click share, using the "share" format also ensures the comparability of product conversion efficiency across search terms with different traffic volumes.

[0034] MPS score calculation: A matching score (MPS) is calculated for each record in the long table, which is defined as the product of four factors: MPS = PPF × PDF × RDF × VF.

[0035] Product Performance Factor (PPF): PPF = (click_share × w_click) + (conversion_share × w_conv), where w_conv (e.g., set to 3) is greater than w_click (e.g., set to 1) to give higher weight to conversion behavior and filter high-net-worth intent keywords.

[0036] Product Position Factor (PDF): The exponential decay model PDF = exp(-λ_pos × (product_position - 1)) is adopted, where the decay coefficient λ_pos can be 0.3, simulating the non-linear decay of user attention as the ranking position decreases.

[0037] Search Rank Factor (RDF): The RDF model is adopted with a small decay, RDF = exp(-λ_rank × (search_rank- 1)), where the decay coefficient λ_rank takes a minimum value such as 0.001, so as to smoothly retain long-tail ranking terms and avoid precipitous cut-off of their value.

[0038] Search term volume factor (VF): Incorporating external monthly search volume data, VF = log(1 + search_volume). Utilizing the compression properties of the logarithmic function, it balances the numerical span between top-ranking keywords and long-tail keywords, ensuring the model's numerical stability and comparability across different traffic volumes.

[0039] In another embodiment, an index is created whose mapping includes the following core fields: the search_term and product_title fields for traditional text retrieval; the search_term_vector and product_title_vector fields for vector retrieval; and metadata fields such as MPS score, brand, and category for filtering and display.

[0040] Hybrid retrieval and ranking: Multi-path recall: When the user enters the query term Q, the system executes two-path retrieval in parallel.

[0041] Path A (Text Recall): Use the BM25 algorithm to perform full-text matching recall on the search_term field.

[0042] Path B (Semantic Recall): After vectorizing Q, perform approximate nearest neighbor search on the search_term_vector and product_title_vector fields respectively to recall semantically similar search terms and related products.

[0043] Deduplication of results: Merge the two recall results, utilize the field folding function of the search engine, use search_term as the folding key, and retain only the record with the highest MPS score in each group to achieve keyword-level deduplication.

[0044] Fine-ranking output: The top K candidate results after deduplication are subjected to a second fine-ranking process. The final ranking score is dominated by the offline calculated MPS score, with a slight combination of the relevance score from real-time retrieval, for example: Final_Score = MPS × (1 + ε × Relevance_Score). The recommended keyword list is output in descending order of the final score.

[0045] In one possible implementation, the messy original data is tidied up—the "wide table" is transformed into a "long table": The original "wide table" was constructed based on reference products. The search term was "study light". The top-ranked product was Product A (60% click-through rate, 20% conversion rate), the second-ranked product was Product B (30% click-through rate, 10% conversion rate), and the third-ranked product was Product C (10% click-through rate, 5% conversion rate).

[0046] The converted "long table" looks like a clear Excel spreadsheet: Table 1. Converted Table

[0047] Step Two: Scientific Scoring – “MPS Multifactor Scoring Model” Patented technology: Calculate a composite score MPS = PPF × PDF × RDF × VF for each row (a combination) in the table.

[0048] Using the data line "study light - Product A" for calculation: Product performance factors: mainly focus on conversion value.

[0049] PPF = Click Share * 1 + Conversion Share * 3 (assuming the conversion weight is 3 times that of the click); Calculate: 60%*1 + 20%*3 = 0.6 + 0.6 = 1.2; Product position factor: Considering attention decay. The products ranked 1st and 3rd have different chances of being noticed by users.

[0050] PDF = exp(-0.3 * (ranking - 1)), calculate (ranking 1st): exp(-0.3 * 0) = 1.0 (ranking first, full weight); Search term ranking factors: Consider long-tail value. The term "study light" may rank 5000th on the platform's overall ranking, but as long as it is effective, it should not be discarded.

[0051] RDF = exp(-0.001 * (rank - 1)), calculated (assuming the total ranking of the search term is 5000): exp(-0.001*4999) ≈ exp(-5) ≈ 0.0067 (the decay is very small, and the value is preserved); Search term volume factor: Consider market size. The monthly search volume for "study light" may be very large, requiring logarithmic smoothing.

[0052] VF = log(1 + monthly search volume), calculated (assuming monthly search volume is 100,000): log(1+100,000) ≈ log(100,001) ≈ 11.5.

[0053] Final MPS score: MPS = 1.2 × 1.0 × 0.0067 × 11.5 ≈ 0.092.

[0054] In layman's terms: This score combines four factors: "high conversion rate (high PPF)," "ranked first (high PDF)," "useful despite being a niche keyword (RDF not zero)," and "significant market size (high VF)." The system calculates this for all keywords in this way to fairly compare top-performing keywords and potential long-tail keywords.

[0055] Step 3: Intelligent Association – “Dual-Channel Semantic Vector Retrieval” This is the key to achieving intelligent word expansion, solving the problem of "how to associate 'study light' with 'reading lamp'".

[0056] Traditional method (literal matching): If you search for "study light", the system can only find words that completely contain these words, such as "good study light".

[0057] Patent method (semantic matching): Channel 1: Word-to-word association: The system converts "study light" into a mathematical vector (which can be understood as "semantic coordinates"), and then searches for the nearest word to it in the vector space, such as "reading lamp".

[0058] Channel Two: Keyword and Product Association: The system will also match the vector for "study light" with the vectors for all product titles. If a product titled "LED Desk Lamp for Students" has a very similar vector, and this product is associated with the search term "desk light," then "desk light" will also be recalled.

[0059] Effect: In this way, the system can "fill in" and mine keywords with different shapes but highly related meanings from massive amounts of data, greatly expanding the range of words to be selected.

[0060] Step 4: Deduplication and Sorting – “Hybrid Search and Final Recommendation” When you enter the seed keyword "desk lamp" into the system for a search, the system will launch two engines simultaneously: Text search engine: Quickly find all words that literally contain "desk lamp", such as "led desklamp".

[0061] Semantic search engine: Through the above two channels, find words with similar meanings, such as "study light", "reading lamp", and "workspace light".

[0062] Next comes the "finishing": Deduplication and merging: The term "study light" may be recalled multiple times from multiple paths, and the system will merge them into one.

[0063] Sort by MPS: This is the decisive step. The system no longer simply sorts by relevance or search volume, but instead sorts by the MPS score, which we calculated earlier and which best represents business value.

[0064] Final output: You will get a recommendation list similar to the table below, which directly tells you which related keywords have the highest value and are most worth prioritizing: Table 2 Keyword Recommendation Results

[0065] Example 2 Specifically, such as Figure 2 As shown, this application provides a keyword mining method based on multi-factor and dual-semantic channels, applied to the aforementioned keyword mining system based on multi-factor and dual-semantic channels, specifically including: S1 divides products into different groups based on their type. Using similar product data and product data in the group, it determines the keyword mining strategy for the products in the group. When the mining strategy is an optimized mining strategy, it proceeds to the next step. S2 determines the reference product of the product based on the analysis results of the similarity between the product image of the product and similar products. Based on the reference product data of the product and the similarity between the reference product and the reference products of the products in the combination of the product, the product in the product is selected for keyword mining processing using all reference products, and it is taken as the target product. Based on the composition data of the target product in the combination and the similarity between the product and the reference products of the target product, the keyword mining processing method of the product is determined. S3 uses the keyword mining and processing method to mine and process the keywords of the product to obtain the keyword mining and processing results, and uses the mining and processing results and the keyword mining and processing method to determine the method for updating the keyword mining and processing results of the product.

[0066] Furthermore, the products are divided into different groups, specifically including: Group products of the same type into the same category.

[0067] Furthermore, similar products are those that belong to the same type as the product.

[0068] Furthermore, such as Figure 3 As shown, the method for determining the keyword mining strategy for the products in the combination is as follows: The core decision-making objective of this embodiment is to solve the challenges of "internal competition" and "homogenized recommendations" faced by apparel companies when optimizing massive amounts of product keywords. Its core logic lies in dynamically selecting the most suitable keyword mining reference range by intelligently analyzing the density of external market competitors and the richness of internal product lines. This ensures that keywords are mined for each apparel product that can both capture precise traffic and create a differentiated positioning, ultimately achieving the optimal overall traffic structure and maximized sales conversion for the store.

[0069] This method follows a closed-loop decision-making process of "market insight → self-assessment → strategy selection → execution mining". First, assess the level of competition the target product faces on the platform; second, assess the SKU depth of similar products within the company; finally, based on the two-dimensional analysis of "external competition and internal abundance", intelligent decision-making determines whether to adopt a "broad-spectrum reference" strategy for rapid coverage or to initiate a "refined grouping" strategy for in-depth mining to drive the precise allocation of marketing resources.

[0070] S11 Based on the similar product data of the product, determine the quantity of similar products of the product; S12 determines the quantity of goods in the combination based on the goods data in the combination; S13 determines the keyword mining strategy for the products in the combination based on the number of similar products in the combination and the number of products in the combination.

[0071] It is understandable that if the number of similar products of the product is less than the preset threshold for the number of similar products, then the number of similar products of the product is small. If the reference products for keyword mining of different products are classified in a secondary manner, the reference value will be low due to the small number. Therefore, the keyword mining strategy for the products in the combination is to use all similar products for keyword mining.

[0072] Additionally, it should be noted that if the number of similar products is not less than a preset threshold for the number of similar products, then the number of products in the combination is obtained, and it is determined whether the number of products in the combination is greater than the preset threshold for the number of products. If yes, then the keyword mining strategy for the products in the combination is determined to be an optimized mining strategy; otherwise, then the keyword mining strategy for the products in the combination is determined to be keyword mining processing using all similar products.

[0073] In one possible embodiment, a cross-border e-commerce apparel company that sells "yoga pants" is used as an example to illustrate the entire decision-making process.

[0074] Product bundle definition: All "yoga pants" category products are grouped into the same bundle.

[0075] Similar products definition: On e-commerce platforms, all other brand products that belong to the same "yoga pants" category as your own products.

[0076] Step S1: Assess the external market competition environment (number of similar products). The system automatically calls the platform API or category data to count the total number of active products for sale under the "Yoga Pants" category (excluding its own products).

[0077] Example result: A search revealed approximately 50,000 yoga pants available for sale on the platform. This enormous number indicates extremely fierce market competition.

[0078] Step S2: Inventory the scale of the internal product line (number of products in the bundle) and count the number of all "yoga pants" SKUs in the company's own product library.

[0079] Example result: The company has a wide product line with 120 different yoga pants (covering high-waisted, cropped, printed, solid color, and different materials).

[0080] Step S3: Strategy decision based on two-dimensional analysis: The system compares the two sets of data with preset thresholds and automatically generates a data mining strategy. Assume the preset thresholds are: similar product quantity threshold = 10,000, and product quantity threshold = 50.

[0081] Scenario A: Blue ocean competition or single product line → Adopt a "broad-spectrum reference strategy" Judgment criteria: The number of similar products is less than 10,000 or the number of products in the combination is less than 50.

[0082] Strategic logic: When external competition is not intense, you can boldly refer to all competitor keywords; or when you have few products, you don't need to worry about internal competition, and the goal is to quickly acquire mainstream market traffic.

[0083] Example: If a company only sells 5 basic black yoga pants, it can directly refer to 50,000 competitors across the entire network, dig out high-traffic generic keywords such as "yoga pants" and "workout leggings" to quickly establish brand awareness.

[0084] Scenario B: Intense competition and a rich product line → Activate the "Optimization and Discovery Strategy": Judgment criteria: The number of similar products is ≥ 10,000 and the number of products in the combination is ≥ 50.

[0085] Strategic Logic: This is precisely the situation faced by the example company (50,000 competing products, 120 self-operated products). If all 120 products compete for the broad term "yoga pants," it will lead to severe internal traffic cannibalization and wasted advertising budget. Refined segmentation must be implemented.

[0086] Decision: Since both conditions are met, the system automatically determines to "optimize the mining strategy".

[0087] Furthermore, when the mining strategy is an optimized mining strategy, different reference products for keyword mining processing are determined, and keyword mining processing is carried out in a targeted manner.

[0088] Furthermore, the reference product is a currently listed product whose similarity to the product image meets the requirements.

[0089] Specifically, such as Figure 4 As shown, the method for determining the target product is as follows: This embodiment aims to address the starting point selection problem in product keyword mining: within a product portfolio, which products should be prioritized for in-depth, global keyword mining, referencing all their visual competitors, to most efficiently build a unique and comprehensive keyword base? The core decision-making objective is to quantitatively assess the "scarcity" of each product's external visual competitors and its "network independence" from other internal products, selecting those products with a simple competitive environment and low risk of internal keyword homogenization as "target products." This ensures that the initial keyword assets generated possess high value and low redundancy.

[0090] This method follows a screening process of "assessing competitor scarcity → analyzing internal network overlap → hierarchical filtering to determine targets". First, products are sorted according to the number of visual competitors to identify unique styles in the market. Second, the overlap between each product's competitors and those of other products in the portfolio is analyzed to assess its independence. Finally, a four-level decision funnel is used to progressively filter products based on four dimensions: "absolute scarcity", "segment independence", "global share control", and "network influence", ultimately determining the most suitable target product list as the benchmark for global keyword mining.

[0091] S21. Using the reference product data of the product, determine the quantity of the reference product of the product, and use the quantity of the reference product to determine the sorting result of the product in the combination; Quantify and rank the scarcity of competing products: "The number of reference products for the product" refers to the total number of other brand competitors that are visually highly similar to the target product in appearance, style, and design, and are currently available for sale, as identified through image search or image similarity algorithms on e-commerce platforms. This number directly reflects the intensity of direct and homogeneous competition that the product faces in the open market.

[0092] Visual similarity is a key factor influencing consumer perception and search comparison. The fewer visual competitors a product has, the less reference data is available for keyword mining. Therefore, when mining keywords, it's crucial to comprehensively consider the needs of the reference products. Targeting such products for in-depth analysis helps to gain a competitive edge in a niche market, and the keywords discovered are more likely to include blue ocean terms, avoiding competition with a vast number of generic terms repeated by competitors.

[0093] This step establishes preliminary screening criteria, transforming "design uniqueness" into a quantifiable "competitive scarcity" indicator, and ranking all products within the portfolio to provide an objective and comparable data foundation for subsequent screening.

[0094] Specific examples: The system scans through an image retrieval interface to obtain the number of visual competitors for each product: Product P (Pioneer Model): 65 reference products.

[0095] Product Q (Niche Style): 120 reference products.

[0096] Product R (Basic): 2200 reference products.

[0097] The products are sorted from least to most in quantity as follows: Product P (65) -> Product Q (120) -> Product R (2200) -> ...

[0098] S22 determines the number of reference products that overlap between the product and the reference products of different products in the combination based on the similarity between the product's reference product and the product's reference product in the combination. Analyze the overlap of internal competitor networks: "The overlap between a product and the reference products of different products in the portfolio" refers to the size of the intersection between the visual competitor set of the current product and the visual competitor set of another product in the portfolio. It accurately measures the degree of overlap in the visual competitor customer groups covered and competed for by the two products in the external market.

[0099] Simply having a small number of competitors is insufficient to determine whether a product is suitable as a "target" for comprehensive keyword mining. If a product's few competitors also highly overlap with several of its own products, then keywords mined based on that product may severely intersect with the organic keyword flow of other products, leading to internal traffic competition. The purpose of analyzing overlap is to assess the "potential intrusiveness" or "unique contribution" of the keyword assets generated by this product if it were selected as a target to other members of the portfolio.

[0100] This step shifts from external competition analysis to internal ecosystem synergy diagnosis. It constructs a "competitive overlap network" among products to identify those products that are not only unique externally but also relatively independent internally. This is a crucial step in avoiding internal homogenization of keyword assets.

[0101] Specific examples (continued from S21): The system further analyzes the overlap between the competitor set of product P and the competitor sets of other products in the portfolio: Of the 65 competing products of product P, 5 overlap with the number of competing products of product Q.

[0102] Of the 65 competing products of product P, 10 overlap with the number of competing products of product R.

[0103] S23 determines whether the product is the target product based on the sorting result of the product in the combination and the number of times the product overlaps with the reference product of different products in the combination.

[0104] It should be noted that the target product is the product for which keyword mining was performed using all reference products.

[0105] It is understood that the sorting result of the goods in the combination is obtained by sorting from the fewest to the most based on the number of reference goods in the combination to which the goods belong.

[0106] Specifically, determining whether the product is the target product includes: S231 Based on the sorting result of the goods in the combination, determine whether the sorting result of the goods in the combination is before the preset position. If yes, then determine the goods as the target goods. If no, proceed to step S232. Preliminary Selection – Capturing Absolutely Scarce Items: "The sorting result of the product in the combination" refers to the order of the products in ascending order based on the number of reference products obtained in step S21. "Before the preset position" is an absolute order threshold, such as the top 3. Products with an extremely small number of reference products have a small amount of reference data when extracting keywords. Therefore, if secondary segmentation is performed, the keyword extraction results will be less comprehensive and accurate.

[0107] The preset positions are the top 2. Product P (ranked 1st) is directly selected as the target product because it has very few competitors (65). Product Q (ranked 2nd, 120 products) is also selected.

[0108] S232 determines reference products that do not overlap with the products in the combination based on the number of overlaps between the product and the reference products of different products in the combination, and treats them as independent products. It then determines whether the proportion of independent products in the reference products of the product is greater than a preset independent product proportion threshold. If yes, the product is determined to be the target product; otherwise, it proceeds to step S233. Multiple selection – Identifying products in independent product categories: "Independent products" refer to those competing products in the reference product set of the current product that do not overlap with the reference product set of any other product in the portfolio. "Independent product ratio" is the percentage of independent competing products out of the total number of competing products.

[0109] Some products may not have the fewest total number of competitors, but their competitor pool has extremely low overlap with other products within the same category. This means they are actually carving out a completely independent competitive track. Setting these products as targets ensures that the resulting keyword pool has minimal overlap with the keyword pools of other existing target products, thereby maximizing the overall coverage of the keyword library.

[0110] This rule supplements S231, expanding from "scarcity" to "independent relationship," enabling the discovery of "track pioneer" products that, despite being in a competitive environment, have highly differentiated positioning.

[0111] For product R not selected in S231, calculate its independent product ratio. Of its 2200 competitors, 2000 overlap with other internal products such as product P or product Q, leaving only 200 independent competitors, a ratio of 9.1%. Assuming the preset independent ratio threshold is 60%, then 9.1% < 60%, which does not meet the condition, so proceed to S233.

[0112] S233 Obtain the proportion of the number of target products in the combination, and determine whether the proportion of the number of target products in the combination is greater than the preset target product proportion threshold. If yes, determine that the product does not belong to the target product. If no, proceed to step S234. Regulation – Controlling the overall scale of the target commodity: "Percentage of target products in the portfolio" refers to the percentage of the total number of products in the portfolio that have been identified as targets.

[0113] A company's operational resources (such as the cost of in-depth data analysis and the focus of advertising budgets) are limited. When there are too many target products, it is already possible to ensure that a large number of products in the portfolio can undergo keyword mining using relatively comprehensive reference products. The mining results are already reliable enough. Therefore, if the reliability is sufficient, products with few independent products can be directly determined not to be target products.

[0114] Specific decisions: Currently, two products, P and Q, have been identified as targets (2 products). Assuming the total number of products in the combination is 10, the target percentage is 20%. The preset target product percentage threshold is set at 25%. Since 20% < 25%, the upper limit has not been triggered, therefore product R is still eligible for final evaluation (S234).

[0115] S234 determines the influence coefficient of the product on the products in the combination based on the number of overlaps with the reference products in the combination and the proportion of the number of reference products in the product. Based on the influence coefficient of the product on the products in the combination, it determines whether the product is a target product.

[0116] It is understood that when the average value of the influence coefficient of the product on the products in the combination is less than the preset influence coefficient threshold, the product is determined to be the target product.

[0117] Final Review – Assessing Internal Network Influence: The "Influence Coefficient of a Product on Products in the Combination" is a comprehensive metric used to quantify the "dilution risk" or "homogenization influence" that a product may cause to the keyword uniqueness of other products due to the high overlap between its competitors and those of other internal products. Its calculation is typically based on the proportion of the number of reference products the product overlaps with to the total number of reference products for those other products, and then the average value is taken. The lower the average value, the "purer" the competitor pool of the product, and the smaller the threat to the keyword uniqueness of other products.

[0118] For products that have passed the first three rounds of screening, this is the final and most meticulous balancing act. Even if a product has many competitors and a low independent ratio, if its competitor pool has little overlap with its peers (i.e., a low influence coefficient), it means the additional information it brings is still relatively unique. Conversely, if the influence coefficient is high, it means its competitors severely overlap with others. Setting it as a target for global keyword mining will result in keywords that largely overlap with existing target products, limiting its value. Therefore, the final screening is conducted based on the core objective of maintaining the uniqueness of the overall keyword asset. This ensures that each newly added target product brings significant, non-repetitive incremental value to the keyword library, minimizing internal redundancy.

[0119] Specific decisions: Calculate the influence coefficient of product R. Its significant overlap with products P and Q results in an average influence coefficient as high as 0.55. Assuming a preset influence coefficient threshold of 0.2, 0.55 > 0.2. Therefore, product R is ultimately determined not to be a target product.

[0120] This embodiment constructs a progressive intelligent selection system for target products, moving from the surface to the core, and from static indicators to dynamic network relationships. It innovatively transforms the visual competitive landscape of products into calculable scarcity indicators and network relationship data. Through a rigorous four-tiered filtering logic—"absolute scarcity priority → independent track supplementation → global scale control → final review of network influence"—it accurately identifies key products from a massive pool of goods that represent specific market segments without polluting the internal keyword ecosystem, serving as the starting point for in-depth analysis.

[0121] This application lays a solid foundation for unique keyword assets: by selecting target products with scarce competitors and low internal overlap for the first round of mining, it can ensure that the generated core keyword library is inherently highly unique and has low internal conflict, laying a high-quality foundation for the entire product's keyword system and effectively preventing internal traffic competition during operation: from the source, it avoids the problem of homogenization of subsequent keyword recommendations caused by selecting highly overlapping products as targets, thereby preventing potential internal consumption of advertising bidding between products and mutual influence of traffic.

[0122] Specifically, such as Figure 4 As shown, the method for determining the keyword mining and processing method for the product is as follows: This embodiment aims to address the core contradiction of keyword strategy in multi-product operations: how to ensure, through systematic rules, the keyword sets mined for different products are highly unique, minimizing overlap. Its core objective is not to actively seek differentiated markets, but rather to forcefully divert data from the source through a "defensive" or "isolated" reference product screening mechanism. This ensures that the keyword mining process for different products is based on as many different competitor samples as possible, thereby automatically generating keyword assets with inherent differences in the results and completely avoiding keyword homogenization caused by overlapping reference sources.

[0123] This method follows a control process of "global monitoring → source isolation diagnosis → representative allocation within the cluster". First, it monitors the density of "target products" that already have priority in mining within the monitoring portfolio. Second, for each product to be processed, it analyzes the overlap between its external competitor pool and the target product, and assesses the size of its independent competitor pool. Finally, for clusters with highly overlapping competitors among ordinary products, a "single representative" mining system is adopted. The core operation of the entire process is the intelligent allocation and isolation of the "reference product" data source, rather than post-processing deduplication of the mining results.

[0124] S31 uses the composition data of the target products in the combination to determine the proportion of the target products in the combination; Keyword explanation: "The proportion of target products in the portfolio" refers to the percentage of products that have been designated by the system for global, in-depth keyword mining out of the total number of products in the portfolio. Keyword mining for these target products will be unrestricted, using all their reference products.

[0125] This is the master switch for activating different control levels. If the target product's proportion is too high (e.g., exceeding 50%), it means that most products within the portfolio already have independent, prioritized keyword mining channels. At this point, to prevent the remaining few ordinary products from competing with these "protagonists" for reference resources, the system will enforce the strictest "reference source isolation" rule: ordinary products must not only completely avoid the target product's reference pool, but also, when highly similar to each other, can only be mined by one of them, eliminating any possible overlap and duplication from the source. Establishing isolation awareness from the top level of the portfolio prevents unnecessary internal data source competition in a resource-biased ecosystem, ensuring the depth and purity of keyword mining for the target products.

[0126] The combination contains 5 products, with 2 being target products A and B. The target product ratio = 2 / 5 = 40%. Assuming the preset target product ratio threshold is 50%, 40% is not greater than the threshold. Therefore, the highest level of mandatory isolation mode is not triggered, and the personalized diagnostic process for product C begins.

[0127] S32 determines the number of reference products that overlap between the product and the target product based on the degree of similarity between the product and the reference product of the target product; Conflict between quantification and the reference source for the "target product": "The overlap between the reference products of the current product and the target product" refers to the size of the intersection between the external competitor set of the current ordinary product and the external competitor set of a certain target product. It directly measures how much of the potential data source for keyword mining is common between the two.

[0128] This is the first crucial step in achieving "reference source isolation." If a regular product shares a large number of identical competitor references with a target product, the keywords mined from these common competitors will inevitably be highly similar, thus undermining uniqueness. Therefore, it is essential to identify this overlap and provide data support for subsequent "avoidance" operations.

[0129] The abstract goal of "avoiding keyword overlap" is transformed into the quantifiable and actionable specific task of "reducing the intersection of reference product sets".

[0130] Specific example (continued from S31): Product C has 2000 reference products (competitors). Analyzing the overlap with the target product: the number of reference products overlapping with target product A is 120, and the number of reference products overlapping with target product B is 30. These overlapping reference products are the risk source of keyword homogenization.

[0131] S33 determines the keyword mining and processing method for the product based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product.

[0132] It should be noted that if the proportion of the target product in the combination is greater than the preset target product proportion threshold, then the keyword mining processing method for the product is to no longer consider the reference products of the target product, and only consider the reference products that belong to multiple products excluding the target product when performing keyword mining processing on the product with the fewest reference products.

[0133] Furthermore, if the proportion of the target product in the combination is not greater than a preset target product proportion threshold, the following is also included: S331 Obtain the number of reference products that overlap between the product and the target product. Based on the proportion of the number of reference products that overlap between the target product and the product in the number of reference products of the product, determine the influence coefficient of the target product on the product. Determine whether there is a target product whose influence coefficient is greater than a preset influence coefficient threshold. If so, all reference products except the reference products of the target product need to be considered when the product is processed for keyword mining. If not, proceed to step S332. Diagnosis – Is it "covered" by a specific target product? The "Influence Coefficient of the Target Product on the General Product" refers to the proportion of the number of reference products that overlap between a single target product and a current general product, relative to the total number of reference products for that general product. A higher coefficient indicates that the target product's competitor pool "covers" a large portion of the general product's market share.

[0134] If a target product has an extremely high influence coefficient (e.g., exceeding 30%), it means that most of the keywords for ordinary products likely originate from the same competitor pool as the target product. To ensure uniqueness, the most direct method is to instruct the ordinary product to completely avoid all reference products of the target product during keyword mining, using all remaining reference products as data sources. This is the most lenient isolation measure, providing protective isolation for ordinary products overshadowed by a strong target product. It does not require consideration of overlap with reference products of products other than the target product, making it the most effective means of ensuring keyword uniqueness.

[0135] Calculate the influence coefficient of target product A on product C: 120 / 2000 = 0.06 (6%). Assuming the preset influence coefficient threshold is 0.3 (30%), 0.06 < 0.3. This isolation strategy is not triggered; proceed to the next step.

[0136] S332 removes reference products that overlap with the target product as available reference products, and determines whether the proportion of the available reference products in the reference products of the product is less than a preset available proportion threshold. If so, all reference products except the reference products of the target product need to be considered when the product is processed for keyword mining. If not, proceed to step S333. Diagnosis – Is there sufficient independent data source? "Available reference products" refers to the set of completely independent competing products remaining after removing all overlaps with any target product from the total reference products of ordinary products. "Available reference product percentage" is the size percentage of this independent set.

[0137] Even if a product isn't covered by a single target product, if the independent competitor pool of a typical product is too small (e.g., <50%), it indicates a limited number of "clean" data sources available for differentiated keyword mining. Forcing it to use only this portion could lead to insufficient keyword mining. In this case, as a trade-off, allowing it to consider all reference products that overlap with the target product during mining, thereby expanding the scope of reference products, without needing to consider overlap with reference products other than the target product, is the most comprehensive way to ensure keyword uniqueness. This strikes a balance between ensuring uniqueness and guaranteeing the feasibility of mining, avoiding data source depletion due to excessive isolation.

[0138] Specific decisions: The number of available reference products for product C is 2000 - 120 - 30 = 1850, representing 1850 / 2000 = 92.5%. Assuming a preset available percentage threshold of 50%, 92.5% > 50%, indicating that product C has extremely rich independent data sources. This step will not be triggered; proceed to the final diagnosis.

[0139] S333 Based on the number of reference products that overlap with different products, determine the similarity coefficient between the product and the reference products of different products, and determine whether there are products with a similarity coefficient greater than a preset similarity coefficient threshold. If yes, proceed to step S334. If no, remove the reference products other than the reference products of the target product. All reference products need to be considered when the product is processed for keyword mining. S334 groups products with a similarity coefficient greater than a preset similarity coefficient threshold into the same group, and determines whether the number of products in the group is greater than a preset product quantity threshold. If not, the reference products other than the target product's reference products are removed. The reference products are only considered when keyword mining is performed on the product with the fewest preset number of reference products. If so, the reference products other than the target product's reference products are removed. The reference products are only considered when keyword mining is performed on the product with the fewest reference products.

[0140] Diagnosis and Decision Making – Handling Internal Overlaps Between Common Goods: The "similarity coefficient" measures the similarity between the reference product sets of two ordinary products. Specifically, it is determined by the ratio of the number of overlapping reference products in the combination of the two products to the number of reference products in the product with the fewest reference products among the two products. Products with a "similarity coefficient greater than a preset threshold" constitute a highly cohesive cluster, meaning that their external competitors' perspectives highly overlap.

[0141] This is crucial for preventing keyword homogenization among ordinary products. If multiple ordinary products reference almost the same competitor pool, even if they all avoid the target product, the keywords they generate will still be largely duplicated. To address this issue, the system groups these products into a "cluster" and implements a "representative system" within the cluster: only one product in the cluster (usually the one with the fewest reference products, as its perspective may be the most unique) is designated to use all available reference products for keyword mining, while other products in the cluster directly share its mining results or only use a subset assigned to them. In this way, the entire cluster generates only one set of keywords based on a common reference source, fundamentally avoiding internal duplication.

[0142] The principle of "reference source isolation" is extended from "target product - ordinary product" to "ordinary product - ordinary product". Through clustered management and representative mining, reference source merging and isolation are also achieved at the ordinary product level, ensuring that the keyword mining process of any two products will not be based on highly overlapping data sources.

[0143] In one possible embodiment, S333 (cluster identification): Suppose the system finds that the similarity coefficients between products C, D, and E are all 0.8 (threshold 0.7), and the three are classified into the same cluster.

[0144] S334 (representing specification): This combination has 3 members (C, D, E). Assuming the preset combination size threshold is 3, where 3 is not greater than 3, the overlap risk is low. The system specifies that for each reference product, only the two products in the combination with the fewest number of reference products containing the reference product need to be considered when mining keywords.

[0145] For example, for a certain reference product F, if the reference products of products C, D, and E all contain F, then since the number of reference products of C and D is the smallest, reference product F only needs to be considered when keyword mining is performed on products C and D, that is, it is included as part of the keyword mining process.

[0146] This embodiment constructs a keyword uniqueness assurance system centered on "intelligent isolation and allocation of reference product resource pools." Through rigorous three-level diagnostic logic (coverage diagnosis, sufficiency diagnosis, and cluster diagnosis), it dynamically allocates suitable, and as non-overlapping as possible, reference product data sources to each product. Its core idea is to solve the problem from the "input end" rather than the "output end," naturally guaranteeing the uniqueness of the "finished product" (keywords) by controlling the exclusivity of the mined "raw materials" (reference products).

[0147] Specifically, the method for determining the update method of the keyword mining and processing results for the product is as follows: This embodiment targets the apparel industry's business scenario, characterized by numerous styles, rapid iteration, and a significant long-tail effect. It aims to establish a "core-driven, radiating update" keyword mining and updating mechanism. Its core objective is to prioritize a small number of "updated products" with weak reference product data foundations and requiring frequent monitoring, ensuring the extreme timeliness of their keywords. Updates for other "non-updated products" with solid data foundations are designed as incidental actions triggered by update events of the updated products. This means updates are only performed when the reliability of keyword updates for updated products and reference products is poor. This ensures absolute resource allocation bias, maintaining the basic health of the overall product keyword database at extremely low cost.

[0148] This method follows a unidirectional, "strong core, weak radiation" process: monitoring and forcibly updating core products → assessing the impact of core updates on radiation products → updating radiation products only when the impact is significant and the evidence is conclusive. System resources are continuously and proactively invested in monitoring and updating "updated products"; while for "non-updated products," the system adopts a passive and cautious approach, only initiating updates for the latter when the updates from "updating merchants" cannot accurately and comprehensively reflect market changes. The entire process ensures that updates begin and end at the core, with radiation updates merely a cautious extension of core updates.

[0149] S41 Based on the keyword mining and processing method, determine the proportion of the reference products considered in the keyword mining and processing of the product among the reference products of the product, and use the proportion of the reference products considered in the keyword mining and processing of the product among the reference products of the product as a reference proportion. "Reference ratio" refers to the percentage of reference products actually adopted and used for analysis when executing a given keyword mining strategy, out of the total number of theoretically obtainable reference products. For example, a dress may have 200 visual or category competitors, but due to the strategy excluding competitors, only 150 are mined, resulting in a reference ratio of 75%.

[0150] The reference ratio is a core quantitative indicator for measuring the "data completeness" and "decision confidence" of keyword mining results. A lower ratio means that the sample data upon which the conclusions are based is insufficient, and the keyword library may not be comprehensive enough. Therefore, products with low reference ratios have the highest risk of keyword incompleteness and require the highest priority and most frequent system maintenance.

[0151] This step provides objective and quantifiable screening criteria for the entire update system. It transforms the strategic decision of "which products need to be prioritized for maintenance" into an automated judgment based on "data completeness" scores, ensuring the scientific and consistent allocation of resources and laying the foundation for building a "core-radial" structure.

[0152] The system calculates the reference ratio for each dress in the store. Product F (a collaboration tie-dye dress) has a unique design, but only 60 reference products are available, some of which are used for data mining, resulting in a reference ratio of 60%, less than 70%. Due to weak data foundation, the system determines it should be marked as an "updated product" (core). Product X (the classic little black dress), on the other hand, has 1800 reference products and a reference ratio greater than 70%, thus it is classified as a "non-updated product" (outlier).

[0153] S42 determines, based on the reference ratio, the products whose keyword mining results need to be updated, and uses them as the updated products; In this method, "updating products" specifically refers to products whose reference ratio is below a preset threshold (e.g., 70%). The system assigns them the highest update priority; given that the generated keyword results are not accurate or comprehensive enough, updating them is a mandatory and primary task of the system.

[0154] Given resource constraints, it is neither possible nor necessary to perform deep updates on all products at the same frequency. Defining products with a low reference ratio as "updated products" acknowledges the fragility of their data foundation and the comprehensiveness bias of their keywords. Through continuous, high-frequency monitoring and updates, the system can reliably update the keywords for these products.

[0155] Product F (the tie-dye collaboration) has been officially marked as a core "updated product" by the system. The system's backend monitoring thread will continuously scan for keyword changes in its 36 reference products and be ready to trigger the update process at any time. Meanwhile, Products X and Y are marked as "non-updated products," and the system will not actively perform periodic scans on them.

[0156] S43 determines the update method for keyword mining and processing of the product based on the reference ratio and the updated product data in the combination.

[0157] Furthermore, the updated product is a product whose reference ratio is less than a preset reference ratio threshold.

[0158] Furthermore, based on the reference ratio and the updated product data in the combination, the update method for keyword mining and processing of the product is determined, specifically including: Scenario 1: If the product is an updated product, keyword mining will be performed whenever the keywords of the reference product being considered change. Once any keyword change is detected in any reference product of the core "updated product" (product F), the system immediately and unconditionally triggers a complete keyword re-mining process for product F.

[0159] This directly reflects the "strong core" principle. Since the core product data foundation is weak, the comprehensiveness of the keyword database is improved through dynamic updates. Immediate response is the bottom line requirement for maintaining the usability and competitiveness of its keyword database.

[0160] Of the 36 reference products for Product F, 15 updated their keywords (for example, adding trending terms such as "summer oil painting feel" and "new Chinese tie-dye"). The system detected this change and automatically recalculated all keywords for Product F within one hour, generating a new keyword library containing the latest trending terms.

[0161] Case 2: If the product is not an updated product, determine whether the proportion of updated products in the combination is greater than the preset updated product proportion threshold. If yes, when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than the preset change proportion threshold, keyword mining processing is performed. If no, proceed to the next step. "The percentage of updated products in the portfolio" is an evaluation metric used to quantify whether the keywords of the products in the portfolio can be reliably and comprehensively updated when the keywords of the reference products change. It represents the reliability and timeliness of the keyword update process for the products in the portfolio.

[0162] Specific example (continuing from S43 case 1): The updated product in the combination is F, and there are a total of 3 products. The proportion of the updated products in the combination is 0.33, which is less than 0.5. When the keywords of the reference product change, the timeliness of the update is poor, so it is necessary to proceed to the next step.

[0163] The reference products that both the product and the updated product need to consider when performing keyword mining are used as the product's associated reference products. It is determined whether the proportion of associated reference products among the reference products that need to be considered when performing keyword mining is greater than a preset proportion threshold. If so, keyword mining is performed when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than a preset change proportion threshold. If not, keyword mining is performed when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than a second preset change proportion threshold.

[0164] It should be noted that the preset change percentage threshold is greater than the second preset change percentage threshold.

[0165] Additional update decision based on representativeness: The system presets a low "preset change percentage threshold" (e.g., 2.0%). The calculated radiation impact is then compared with this threshold.

[0166] If the proportion of related reference products considered during keyword mining for a product exceeds a threshold (20%), it is considered that the core update event can comprehensively represent the market changes that the product needs to address. In this case, the update of the updated product can be considered to reflect the keyword update requirements when reference products change. For the non-updated product, due to its large reference proportion, its keyword extraction results are sufficiently reliable and comprehensive. Only when the proportion of reference products with changed keywords among the considered reference products exceeds a preset change proportion threshold is it necessary to fully utilize the reference products considered during keyword mining for the product to perform keyword mining.

[0167] If the proportion of related reference products considered during keyword mining for a product is less than the threshold (20%), it is considered that the core update event cannot fully and accurately reflect the keyword update needs when reference products change, i.e., it cannot effectively and comprehensively reflect changes in the market. The update needs when reference products change are not met. In this case, the system needs to further check the total change ratio of the reference product's own pool of reference products. As long as its own change ratio exceeds another lower "second preset change ratio threshold" (e.g., 1%), a separate, complete update (accompanying update) will be triggered.

[0168] This rule strictly adheres to the "weak radiation" principle. Only when the impact of changes to the reference product is sufficiently broad (high impact), and the updated product's keyword mining fails to fully reflect the impact of these changes, will the system initiate keyword mining and updating for that product under fixed conditions. This minimizes intervention in non-updated products, ensuring the long-term stability of the keyword strategy for non-updated products, while simultaneously not overlooking truly significant market changes that have a global impact.

[0169] Specific decision-making (following up from S43, scenario 2): For product Y (the proportion of related reference products considered during keyword mining should not exceed 20%): the representativeness of the updated product is insufficient. The system then checks the total change rate of product Y's 500 reference products. If more than 1% (5) of its reference products have changed, reaching the "passive update threshold," an independent update for product Y is triggered. Otherwise, it is completely ignored.

[0170] This embodiment successfully constructed a keyword update intelligent scheduling system characterized by "core-driven, cautiously radiating" approach. It scientifically defines core products through a "reference ratio" and builds a complete evaluation and decision-making chain around update events for these core products. The system prioritizes resources to ensure the timely update of products with fewer reference products during keyword mining. Updates of non-updating products are only executed as a cautious extension of the core update process when the updated product fails to update its keywords according to the changed reference products—that is, when the updated product's keywords do not accurately reflect the changes in the reference product's keywords.

[0171] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0172] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0173] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A keyword mining system based on multi-factor and dual semantic channels, characterized in that, Specifically, it includes: The data reconstruction module divides its own products into different combinations based on their types. Using the product data of its own products and similar products in the combination, it determines the keyword mining strategy for its own products in the combination. The mining strategy is used to obtain reference products for its own products. The reference product data of the e-commerce platform is converted into a long table structure with "search term-product" as the granularity, and associated with product brand and type metadata. Based on multiple dimensions, a comprehensive matching score is calculated for each "search term-product" pair. Dividing one's own products into different groups specifically includes: grouping products of the same type into the same group; the similar products of one's own products are competing products of the same type as one's own products; The method for determining the keyword mining strategy for the products within the aforementioned combination is as follows: Based on the similar product data of the product itself, determine the quantity of similar products of the product itself; Based on the product data of the products in the combination, determine the quantity of each product in the combination. Based on the number of its own products and the number of similar products in the combination, a keyword mining strategy for its own products in the combination is determined; if the number of similar products of its own products is less than a preset threshold for the number of similar products, the keyword mining strategy for its own products in the combination is determined to be to use all similar products for keyword mining. The reference product for the product itself is a visual competitor that is currently listed and whose product image is similar to the product itself to meet the requirements. The index building module performs vectorized encoding on the search terms and product titles in the long table, respectively, to generate semantic vectors for the search terms and semantic vectors for the product titles, and stores them in the retrieval engine along with the matching scores. The intelligent recommendation module responds to the query terms input by the user, and uses the retrieval engine to perform parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors to obtain recall results. After deduplication of the recall results, they are re-ranked according to the matching scores, and the final keyword recommendation list is output. The module also uses the keyword mining and processing method of its own products in the combination to determine the update method of the keyword mining and processing results of its own products. The similarity retrieval based on dual-channel semantic vectors combines word and word association channels. That is, the query word is converted into a mathematical vector, and then the word closest to it is found in the vector space. The word and product association channel: the system matches the mathematical vector of the query word with the product vectors of all product titles. The method for determining the update method of the keyword mining and processing results of the product itself is as follows: Based on the keyword mining and processing method, the proportion of the reference products considered in the keyword mining and processing of the product itself among the reference products of the product itself is determined, and the proportion of the reference products considered in the keyword mining and processing of the product itself among the reference products of the product itself is used as a reference proportion. Based on the reference ratio, determine the products whose keyword mining results in the products themselves need to be updated, and use them as the updated products; Based on the reference ratio and the updated product data in the combination, determine the update method for mining and processing the keywords of the product itself.

2. A keyword mining method based on multi-factor and dual-semantic channel, applied to the keyword mining system based on multi-factor and dual-semantic channel as described in claim 1, specifically comprising: Based on the type of goods, the goods are divided into different groups. Using the similar goods data and goods data of the goods in each group, the keyword mining strategy of the goods in each group is determined. When the mining strategy is the optimized mining strategy, proceed to the next step. Based on the analysis results of the similarity between the product image of the product and similar products, the reference product of the product is determined. Based on the reference product data of the product and the similarity between the reference product and the reference product of the product in the combination of the product, the product in the product is selected for keyword mining using all reference products, and it is taken as the target product. Based on the composition data of the target product in the combination and the similarity between the product and the reference product of the target product, the keyword mining method of the product is determined. The keyword mining and processing method described above is used to mine and process the keywords of the product itself to obtain the keyword mining and processing results. The keyword mining and processing method of the product itself in the combination is used to determine the method for updating the keyword mining and processing results of the product itself.

3. The keyword mining method based on multi-factor and dual semantic channels as described in claim 2, characterized in that, The method for determining the keyword mining and processing method for the product itself is as follows: Based on the composition data of the target products in the combination, determine the proportion of the target products in the combination; Based on the similarity between the product itself and the reference products of the target product, determine the number of reference products that overlap between the product itself and the target product. Based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product, a keyword mining and processing method for the product itself is determined.

4. The keyword mining method based on multi-factor and dual semantic channels as described in claim 2, characterized in that, The updated product is the product itself whose reference ratio is less than the preset reference ratio threshold.